Video Diffusion Model / Mô hình Khuếch tán Video
📖 Video Diffusion Model
Definition (English):
Video diffusion models extend image diffusion to temporal sequences, generating coherent video from text prompts or reference images. They add temporal attention layers to spatial diffusion architectures. Key models: Runway Gen-3, Pika Labs, Kling, and OpenAI Sora (reported). Challenges: maintaining temporal consistency, motion realism, long-horizon coherence, and computational cost.
📖 Mô hình Khuếch tán Video
Định nghĩa (Tiếng Việt):
Mô hình khuếch tán video mở rộng khuếch tán hình ảnh sang chuỗi thời gian, tạo video mạch lạc từ prompt văn bản hoặc hình ảnh tham chiếu. Chúng thêm lớp attention thời gian vào kiến trúc khuếch tán không gian. Mô hình chính: Runway Gen-3, Pika Labs, Kling và OpenAI Sora (đồn). Thách thức: duy trì nhất quán thời gian, tính thực tế chuyển động, nhất quán dài hạn và chi phí tính toán.
📂 Phân loại: Video AI