Multimodal AI / AI Đa phương thức

✍️

📖 Multimodal AI

Definition (English):

Multimodal AI refers to AI systems that can process, understand, and generate information across multiple types of modalities — text, images, audio, video, and sensor data. These models can handle diverse input types simultaneously and generate outputs in different formats. Examples include GPT-4V (text + images), Gemini (text + image + audio + video), and DALL-E (text → image).


📖 AI Đa phương thức

Định nghĩa (Tiếng Việt):

AI Đa phương thức đề cập đến hệ thống AI có thể xử lý, hiểu và tạo thông tin qua nhiều loại phương thức — văn bản, hình ảnh, âm thanh, video và dữ liệu cảm biến. Các mô hình này có thể xử lý đồng thời nhiều loại đầu vào đa dạng và tạo đầu ra ở các định dạng khác nhau. Ví dụ bao gồm GPT-4V (văn bản + hình ảnh), Gemini (văn bản + hình ảnh + âm thanh + video) và DALL-E (văn bản → hình ảnh).


📂 Phân loại: Advanced Topics

Hỏi AI Edu?