Video Diffusion Model / Mô hình Khuếch tán Video

✍️
Từ điển và thuật ngữ trí tuệ nhân tạo

📖 Video Diffusion Model

Definition (English):

Video diffusion models extend image diffusion to temporal sequences, generating coherent video from text prompts or reference images. They add temporal attention layers to spatial diffusion architectures. Key models: Runway Gen-3, Pika Labs, Kling, and OpenAI Sora (reported). Challenges: maintaining temporal consistency, motion realism, long-horizon coherence, and computational cost.


📖 Mô hình Khuếch tán Video

Định nghĩa (Tiếng Việt):

Mô hình khuếch tán video mở rộng khuếch tán hình ảnh sang chuỗi thời gian, tạo video mạch lạc từ prompt văn bản hoặc hình ảnh tham chiếu. Chúng thêm lớp attention thời gian vào kiến trúc khuếch tán không gian. Mô hình chính: Runway Gen-3, Pika Labs, Kling và OpenAI Sora (đồn). Thách thức: duy trì nhất quán thời gian, tính thực tế chuyển động, nhất quán dài hạn và chi phí tính toán.


📂 Phân loại: Video AI

HỆ SINH THÁI CiCC

Đi tiếp cùng nội dung này

Hỏi AI Edu?