KV-Cache / Bộ nhớ đệm KV
📖 KV-Cache
Definition (English):
KV-Cache (Key-Value Cache) stores previously computed key and value tensors during autoregressive generation to avoid redundant recomputation. Without it, each new token would require reprocessing the entire sequence. KV-Cache grows linearly with sequence length and batch size, consuming significant GPU memory for long sequences. PagedAttention (vLLM) manages KV-Cache memory like virtual memory pages.
📖 Bộ nhớ đệm KV
Định nghĩa (Tiếng Việt):
KV-Cache lưu trữ tensor key và value đã tính trước đó trong quá trình tạo autoregressive để tránh tính toán lại thừa. Nếu không, mỗi token mới sẽ cần xử lý lại toàn bộ chuỗi. KV-Cache tăng tuyến tính theo độ dài chuỗi và batch size, tiêu tốn bộ nhớ GPU đáng kể cho chuỗi dài. PagedAttention (vLLM) quản lý bộ nhớ KV-Cache như trang bộ nhớ ảo.
📂 Phân loại: Inference
HỆ SINH THÁI CiCC