Mixed Precision Training / Huấn luyện Độ chính xác Hỗn hợp
📖 Mixed Precision Training
Definition (English):
Mixed precision training uses lower numerical precision (FP16 or BF16) for most computations while maintaining a master copy of weights in FP32. This reduces memory usage by ~50% and speeds up training on GPUs with tensor cores (NVIDIA V100+, A100). Loss scaling prevents gradient underflow. BF16 is preferred over FP16 for LLM training due to its larger dynamic range.
📖 Huấn luyện Độ chính xác Hỗn hợp
Định nghĩa (Tiếng Việt):
Huấn luyện độ chính xác hỗn hợp sử dụng độ chính xác số thấp hơn (FP16 hoặc BF16) cho hầu hết phép tính trong khi duy trì bản chính weights ở FP32. Điều này giảm sử dụng bộ nhớ ~50% và tăng tốc huấn luyện trên GPU có tensor core (NVIDIA V100+, A100). Loss scaling ngăn gradient underflow. BF16 được ưa chuộng hơn FP16 cho huấn luyện LLM nhờ dải động lớn hơn.
📂 Phân loại: Training
HỆ SINH THÁI CiCC