Data Parallelism / Song song Dữ liệu
📖 Data Parallelism
Definition (English):
Data parallelism splits the training dataset across multiple GPUs, where each GPU holds a complete copy of the model and processes a different data batch. Gradients are synchronized across GPUs (typically via all-reduce) after each step. It is the simplest form of distributed training but requires each GPU to fit the full model. Combined with model parallelism for very large models.
📖 Song song Dữ liệu
Định nghĩa (Tiếng Việt):
Song song dữ liệu chia tập dữ liệu huấn luyện trên nhiều GPU, trong đó mỗi GPU giữ bản sao đầy đủ của mô hình và xử lý batch dữ liệu khác nhau. Gradient được đồng bộ hóa qua GPU (thường qua all-reduce) sau mỗi bước. Đây là hình thức đơn giản nhất của huấn luyện phân tán nhưng yêu cầu mỗi GPU vừa mô hình đầy đủ. Kết hợp với song song mô hình cho mô hình rất lớn.
📂 Phân loại: Training
HỆ SINH THÁI CiCC