Distributed Training / Huấn luyện Phân tán
📖 Distributed Training
Definition (English):
Distributed training is the technique of splitting the training process across multiple computing devices (GPUs, TPUs, nodes) to speed up model training and handle datasets that are too large for a single device. Key strategies include data parallelism, model parallelism, pipeline parallelism, and tensor parallelism. Frameworks like DeepSpeed, FSDP, and Megatron-LM enable efficient distributed training.
📖 Huấn luyện Phân tán
Định nghĩa (Tiếng Việt):
Huấn luyện phân tán là kỹ thuật chia quá trình huấn luyện trên nhiều thiết bị tính toán (GPU, TPU, nút) để tăng tốc huấn luyện mô hình và xử lý tập dữ liệu quá lớn cho một thiết bị. Các chiến lược chính bao gồm song song dữ liệu, song song mô hình, song song pipeline và song song tensor. Các khung như DeepSpeed, FSDP và Megatron-LM cho phép huấn luyện phân tán hiệu quả.
📂 Phân loại: Infrastructure