RLHF / RLHF

✍️
Từ điển và thuật ngữ trí tuệ nhân tạo

📖 RLHF

Definition (English):

RLHF (Reinforcement Learning from Human Feedback) fine-tunes language models using human preference data. Human rankers compare model outputs, training a reward model. The LLM is then optimized via PPO to maximize the reward model score while staying close to the original model (KL penalty). RLHF is how ChatGPT and Claude are trained to be helpful, harmless, and honest.


📖 RLHF

Định nghĩa (Tiếng Việt):

RLHF (Reinforcement Learning from Human Feedback) tinh chỉnh mô hình ngôn ngữ bằng dữ liệu ưu tiên con người. Human rankers so sánh đầu ra mô hình, huấn luyện reward model. LLM sau đó được tối ưu hóa qua PPO để maximize điểm reward model trong khi gần mô hình gốc (hình phạt KL). RLHF là cách ChatGPT và Claude được huấn luyện để hữu ích, vô hại và trung thực.


📂 Phân loại: RL Tech

HỆ SINH THÁI CiCC

Đi tiếp cùng nội dung này

Hỏi AI Edu?