Model Parallelism at Inference / Song song Mô hình khi Suy luận
📖 Model Parallelism at Inference
Definition (English):
At inference time, model parallelism distributes a large model across multiple GPUs when it does not fit on a single GPU. Tensor parallelism splits individual layers (attention heads, MLP columns) across GPUs. Pipeline parallelism assigns different layers to different GPUs. For production serving, tensor parallelism within a node and pipeline parallelism across nodes is common.
📖 Song song Mô hình khi Suy luận
Định nghĩa (Tiếng Việt):
Tại thời điểm suy luận, song song mô hình phân phối mô hình lớn trên nhiều GPU khi nó không vừa một GPU. Tensor parallelism chia các lớp riêng lẻ (attention heads, cột MLP) trên GPU. Pipeline parallelism gán các lớp khác nhau cho GPU khác nhau. Trong phục vụ sản xuất, tensor parallelism trong một node và pipeline parallelism giữa các node là phổ biến.
📂 Phân loại: Inference
HỆ SINH THÁI CiCC