HumanEval is OpenAIs code generation benchmark consisting of 164 hand-written Python programming problems. Format: Each problem has function signature, docstring with specification, and unit tests. Model generates function body. Evaluation metric: pass@k - probability that at least one of k generated solutions passes all tests. Typically report pass@1, pass@10, pass@100. Problem types: String manipulation, math, algorithms, data structures. Roughly interview-level difficulty. Scoring: Functional correctness only - if unit tests pass, solution is correct. Limitations: Small dataset (164 problems), Python only, tests may be incomplete, some problems have ambiguity. Extensions: HumanEval+ (more tests), MultiPL-E (multiple languages), variants with harder problems. Baseline scores: GPT-4: around 67% pass@1, Claude 3 Opus: similar range. Top models now approach 90%+ with scaffolding. Use cases: Compare code models, track progress, evaluate prompting strategies. Concerns: Possible data contamination, narrow coverage of programming skills. Standard first benchmark for code generation evaluation.
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.