mmlu (massive multitask language understanding)
MMLU (Massive Multitask Language Understanding) is a comprehensive evaluation benchmark that tests language models across 57 academic subjects spanning STEM, humanities, social sciences, and professional domains, measuring both breadth and depth of knowledge far beyond what earlier benchmarks like GLUE assessed. Introduced by Hendrycks et al. in 2021, MMLU evaluates whether language models have acquired broad world knowledge and can apply it to answer multiple-choice exam questions at difficulty levels ranging from elementary to advanced professional. The 57 subjects include: STEM (abstract algebra, anatomy, astronomy, college biology, college chemistry, college computer science, college mathematics, college physics, electrical engineering, machine learning, etc.), humanities (formal logic, high school European history, jurisprudence, moral disputes, philosophy, prehistory, world religions, etc.), social sciences (econometrics, high school geography, high school government, macroeconomics, marketing, professional psychology, sociology, etc.), and professional/applied (clinical knowledge, global facts, management, medical genetics, nutrition, professional accounting, professional law, professional medicine, etc.). Each question has four answer choices with one correct answer. MMLU contains approximately 15,900 questions, with training, validation, and test splits. Evaluation typically reports accuracy per subject, averaged within each domain, and overall average accuracy. MMLU has become the most widely reported benchmark for comparing foundation models: GPT-4 achieves ~86.4%, Gemini Ultra achieved 90.0% (first to surpass human expert average of ~89.8%), Claude 3 Opus achieves ~86.8%. MMLU's significance lies in testing knowledge that requires genuine understanding rather than pattern matching — questions often require multi-step reasoning, numerical computation, or integrating knowledge across domains. Variants include MMLU-Pro (harder questions with more answer choices) and multilingual MMLU for cross-lingual evaluation.