Agent Evaluation (AgentBench, SWE-Bench) / Đánh giá Agent

✍️

📖 Agent Evaluation (AgentBench, SWE-Bench)

Definition (English):

Agent evaluation benchmarks measure autonomous agent capabilities on realistic tasks. AgentBench: 8 environments (OS, DB, web, coding, etc.) testing planning, tool use, reasoning. SWE-Bench: 2,294 real GitHub issues from 12 Python repos; agents must patch code to pass tests. WebShop: e-commerce navigation. Mind2Web: web interaction. τ-bench: tool use. These benchmarks drive agent research and expose limitations (hallucination, planning failures, tool misuse).


📖 Đánh giá Agent

Định nghĩa (Tiếng Việt):

Benchmark đánh giá agent đo lường khả năng agent tự chủ trên tác vụ thực tế. AgentBench: 8 môi trường (OS, DB, web, coding, etc.) kiểm tra lập kế hoạch, sử dụng công cụ, lý luận. SWE-Bench: 2,294 issue GitHub thực từ 12 repo Python; agent phải patch code để pass test. WebShop: điều hướng e-commerce. Mind2Web: tương tác web. τ-bench: sử dụng công cụ. Các benchmark này thúc đẩy nghiên cứu agent và phơi bày hạn chế (ảo giác, lỗi lập kế hoạch, sai công cụ).


📂 Phân loại: AI Agents

Hỏi AI Edu?