Flash Attention / Flash Attention
📖 Flash Attention
Definition (English):
Flash Attention is an IO-aware exact attention algorithm that reduces memory usage and speeds up transformer training/inference by computing attention in tiles and minimizing GPU memory (HBM) accesses. It achieves 2-4x speedups and reduces memory from O(N²) to O(N). Flash Attention 2 and 3 further optimize for modern GPU architectures (A100, H100). It has become standard in LLM training frameworks.
📖 Flash Attention
Định nghĩa (Tiếng Việt):
Flash Attention là thuật toán attention chính xác nhận biết IO giảm sử dụng bộ nhớ và tăng tốc huấn luyện/suy luận transformer bằng cách tính attention theo tile và tối thiểu hóa truy cập bộ nhớ GPU (HBM). Nó đạt tốc độ nhanh hơn 2-4x và giảm bộ nhớ từ O(N²) xuống O(N). Flash Attention 2 và 3 tối ưu hơn cho kiến trúc GPU hiện đại (A100, H100). Nó đã trở thành tiêu chuẩn trong khung huấn luyện LLM.
📂 Phân loại: DL Techniques
HỆ SINH THÁI CiCC