Inverse Reinforcement Learning / Học Phần thưởng Đảo
📖 Inverse Reinforcement Learning
Definition (English):
Inverse RL (IRL) recovers the reward function from observed expert behavior, rather than learning a policy from a given reward. Given demonstrations of desired behavior, IRL infers what the expert is optimizing. Applications: learning driving behavior from human drivers, imitating surgical procedures, understanding human preferences. GAIL (Generative Adversarial Imitation Learning) combines IRL with adversarial training.
📖 Học Phần thưởng Đảo
Định nghĩa (Tiếng Việt):
IRL (Học Phần thưởng Đảo) khôi phục hàm phần thưởng từ hành vi chuyên gia quan sát được, thay vì học chính sách từ phần thưởng đã cho. Với demonstrations hành vi mong muốn, IRL suy ra điều gì chuyên gia đang tối ưu hóa. Ứng dụng: học hành vi lái xe từ tài xế, bắt chước quy trình phẫu thuật, hiểu ưu tiên con người. GAIL (Generative Adversarial Imitation Learning) kết hợp IRL với huấn luyện đối kháng.
📂 Phân loại: RL Tech