Mixture of Experts (MoE) / Hỗn hợp Chuyên gia
📖 Mixture of Experts (MoE)
Definition (English):
MoE is an architecture where the input is routed to a subset of "expert" sub-networks, with a gating network determining which experts to activate. Only a fraction of total parameters are used per input, enabling models to be much larger without proportionally increasing compute. Switch Transformer, Mixtral, and GPT-4 (rumored) use MoE. Key challenges include load balancing and training stability.
📖 Hỗn hợp Chuyên gia
Định nghĩa (Tiếng Việt):
MoE là kiến trúc trong đó đầu vào được định tuyến đến một tập con mạng con "chuyên gia", với mạng cổng xác định chuyên gia nào được kích hoạt. Chỉ một phần nhỏ tham số tổng được sử dụng mỗi đầu vào, cho phép mô hình lớn hơn nhiều mà không tăng tính toán tương ứng. Switch Transformer, Mixtral và GPT-4 (đồn đại) sử dụng MoE. Thách thức chính bao gồm cân bằng tải và ổn định huấn luyện.
📂 Phân loại: DL Techniques
HỆ SINH THÁI CiCC