BIG-bench (Beyond the Imitation Game Benchmark) is a collaborative benchmark consisting of 200+ diverse tasks designed to probe the capabilities and limitations of large language models — created by hundreds of researchers submitting "tasks where humans excel but LLMs fail".
Diversity
- Tasks: Emoji movie guessing, chess state tracking, irony detection, swahili translation, biology, physics.
- Hard: Specifically designed to be "future-proof" — many tasks were near 0% performance for GPT-3.
- Lite: BIG-bench Lite is a distinct subset of roughly 24 tasks used for cheaper evaluation.
Why It Matters
- Broadness: Moving away from just "GLUE" (NLU) to "Everything" (Reasoning, Humor, Coding).
- Emergence: Used to study "Emergent Abilities" — skills that suddenly appear only at scale (10B+ params).
- Canary: Uses a "canary string" to prevent the benchmark data from leaking into future training sets.
BIG-bench is the gauntlet — a massive, community-driven suite of weird and hard checks to find the breaking points of Large Language Models.
big-benchevaluation
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.