HumanEval / HumanEval

✍️

📖 HumanEval

Definition (English):

HumanEval is a benchmark for evaluating code generation capabilities of LLMs. It consists of 164 Python programming problems with test cases. Models receive a function signature and docstring, then generate the implementation pass@k metric measures the probability of generating at least one correct solution in k attempts. GPT-4 achieves ~90% pass@1. MBPP and SWE-Bench extend code evaluation.


📖 HumanEval

Định nghĩa (Tiếng Việt):

HumanEval là benchmark đánh giá khả năng tạo mã của LLM. Nó bao gồm 164 bài toán lập trình Python với test case. Mô hình nhận function signature và docstring, sau đó tạo triển khai. Chỉ số pass@k đo xác suất tạo ít nhất một giải pháp đúng trong k lần thử. GPT-4 đạt ~90% pass@1. MBPP và SWE-Bench mở rộng đánh giá mã.


📂 Phân loại: Benchmarks

Hỏi AI Edu?