gsm8k

GSM8K (Grade School Math 8K) is a benchmark of 8,500 high-quality, linguistically diverse grade school math word problems requiring 2-8 step reasoning to solve, designed to evaluate the mathematical reasoning capabilities of language models. Introduced by Cobbe et al. (2021) at OpenAI, GSM8K tests whether models can perform multi-step arithmetic reasoning — a capability that requires understanding problem structure, setting up equations, performing calculations, and maintaining state across reasoning steps. Each problem is a natural language word problem solvable by a sequence of basic arithmetic operations (addition, subtraction, multiplication, division), with answer values being positive integers. Problems span everyday scenarios: shopping calculations, cooking measurements, distance and time problems, work rate scenarios, and simple probability. The dataset includes detailed step-by-step solutions that show the intermediate reasoning and calculations, making it valuable for training and evaluating chain-of-thought reasoning. The training set contains approximately 7,500 problems, and the test set contains approximately 1,000 problems. GSM8K has become a crucial benchmark for measuring reasoning progress because: it requires genuine multi-step inference (not just pattern matching or memorization), the problems are novel enough that memorization of training data is insufficient, it tests a well-defined capability (basic arithmetic reasoning) that can be objectively verified, and performance correlates with broader reasoning capabilities. Early models struggled significantly — GPT-3 achieved only about 20% accuracy. Modern models have made dramatic progress: GPT-4 achieves ~92%, Claude 3 Opus achieves ~95%, and Gemini Ultra achieves ~94.4%, though perfect performance remains elusive. The gap between model accuracy and human accuracy (~95-100%) reveals persistent challenges in reliable multi-step reasoning. GSM8K spawned related benchmarks like MATH (competition-level mathematics) for testing more advanced mathematical capability.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account