Actor-Critic Methods / Phương pháp Actor-Critic
📖 Actor-Critic Methods
Definition (English):
Actor-Critic methods combine policy gradient (actor) with value function estimation (critic). The actor decides actions; the critic evaluates them by estimating value functions. A2C (Advantage Actor-Critic) and A3C (Asynchronous Advantage Actor-Critic) are key variants. SAC (Soft Actor-Critic) adds entropy maximization for exploration. Actor-critic is the foundation of modern deep RL including PPO.
📖 Phương pháp Actor-Critic
Định nghĩa (Tiếng Việt):
Phương pháp Actor-Critic kết hợp gradient chính sách (actor) với ước tính hàm giá trị (critic). Actor quyết định hành động; critic đánh giá chúng bằng cách ước tính hàm giá trị. A2C (Advantage Actor-Critic) và A3C (Asynchronous Advantage Actor-Critic) là biến thể chính. SAC (Soft Actor-Critic) thêm maximize entropy cho khám phá. Actor-critic là nền tảng của RL sâu hiện đại bao gồm PPO.
📂 Phân loại: RL Tech
HỆ SINH THÁI CiCC