Home Knowledge Base MMLU (Massive Multitask Language Understanding)

MMLU (Massive Multitask Language Understanding) is the benchmark of 57 academic and professional subjects — from elementary mathematics to medical licensing exams — that became the de facto standard for measuring LLM knowledge depth and breadth — first exposing the massive gap between early language models and human expert performance, then tracking the rapid progress that brought AI to near-expert levels within three years.

What Is MMLU?

The 57 Subjects

STEM:

Humanities:

Social Sciences:

Professional / Applied:

Why MMLU Became the Standard

Performance Timeline

ModelYearMMLU Score
GPT-3 175B202043.9%
InstructGPT202252.0%
GPT-3.5202270.0%
GPT-4202386.4%
Claude 3 Opus202488.2%
Gemini Ultra202490.0%
Expert Human~89.8%

MMLU Variants and Extensions

Limitations

Evaluation Best Practices

MMLU is the comprehensive IQ test for language models — measuring not just what a model has memorized but whether it can integrate knowledge across 57 disciplines to correctly answer questions that require the depth of a medical professional, lawyer, or scientist.

mmlummluevaluation

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.