KV-Cache / Bộ nhớ đệm KV

✍️
Từ điển và thuật ngữ trí tuệ nhân tạo

📖 KV-Cache

Definition (English):

KV-Cache (Key-Value Cache) stores previously computed key and value tensors during autoregressive generation to avoid redundant recomputation. Without it, each new token would require reprocessing the entire sequence. KV-Cache grows linearly with sequence length and batch size, consuming significant GPU memory for long sequences. PagedAttention (vLLM) manages KV-Cache memory like virtual memory pages.


📖 Bộ nhớ đệm KV

Định nghĩa (Tiếng Việt):

KV-Cache lưu trữ tensor key và value đã tính trước đó trong quá trình tạo autoregressive để tránh tính toán lại thừa. Nếu không, mỗi token mới sẽ cần xử lý lại toàn bộ chuỗi. KV-Cache tăng tuyến tính theo độ dài chuỗi và batch size, tiêu tốn bộ nhớ GPU đáng kể cho chuỗi dài. PagedAttention (vLLM) quản lý bộ nhớ KV-Cache như trang bộ nhớ ảo.


📂 Phân loại: Inference

HỆ SINH THÁI CiCC

Đi tiếp cùng nội dung này

Hỏi AI Edu?