GSM8K
Also known as Grade School Math 8K
GSM8K is 8,500 grade-school maths word problems, each solvable in a few elementary steps, released by OpenAI in 2021 to diagnose why language models failed at multi-step arithmetic. Frontier models now score above 99% on it. Its lasting value is not the score but the two follow-up studies that used it to show how much of a benchmark result can be memorisation.
| What it measures | Multi-step grade-school arithmetic word-problem reasoning. |
|---|---|
| Built by | OpenAI (Cobbe et al.), 2021 |
| Format | 8,500 grade-school word problems (7,500 train, 1,000 test), each solvable in a few elementary steps |
| Scoring metric | Exact-match accuracy on the final numeric answer |
| Status | Saturated |
| Representative top score | ~99.6% · Frontier models broadly · read 2026-05 |
| Official leaderboard | llm-stats.com/benchmarks/gsm8k |
How GSM8K works
Each problem is a short natural-language scenario requiring between two and eight elementary arithmetic steps, written and checked by human problem writers so the language is varied rather than templated. The split is 7,500 training and 1,000 test problems. Scoring is exact match on the final numeric answer, which means the benchmark says nothing about whether the reasoning that produced the number was sound, only whether the number is right. The original paper's contribution was as much methodological as it was the dataset: it introduced training a verifier to rank multiple sampled solutions, which prefigured much of the later work on test-time compute.
History and current status
Cobbe and colleagues at OpenAI published GSM8K in October 2021 (arXiv 2110.14168) under the title Training Verifiers to Solve Math Word Problems. It became the default arithmetic reasoning row on every model card for about three years, and the benchmark chain-of-thought prompting was demonstrated on. Scores climbed from roughly a third of problems to above 90% and then past 99%, and the benchmark passed into use as a smoke test. Two studies then reopened it: Scale AI commissioned GSM1k in 2024 and Apple published GSM-Symbolic the same year, both designed to test whether the scores were real.
What the score does not tell you
The follow-up work is the criticism. Scale AI built GSM1k, 1,000 new problems written entirely by human annotators with no model assistance and matched to GSM8K on human solve rate, step count and answer magnitude. Evaluating leading models on it produced accuracy drops of up to 8%, with several model families showing systematic overfitting across almost all sizes, and a positive relationship (Spearman's r squared of 0.36) between how likely a model was to generate a GSM8K example verbatim and how far its score fell. Apple's GSM-Symbolic rebuilt the problems as templates with swappable names and numbers and found drops from 0.3% for the strongest model to 9.2% for a small one, with performance on the original wording sitting in the right tail of the template distribution, which is hard to explain without contamination.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, why benchmark saturation makes a high score meaningless.
GSM8K is not one of the 15 benchmarks scored in the benchmark trust ranking, which rates the most-quoted tests on those same four axes and sorts them most trustworthy first.
What GSM8K scores actually mean
There is effectively no informative band left on GSM8K. Frontier models sit near 99.6%, and the spread between the best and the twentieth-best model is smaller than the label-noise rate in the dataset itself, so differences at the top are not measuring capability. The meaningful number is the delta, not the level: how far a model falls when the same problems are rewritten. A model that scores 95% on GSM8K and 87% on GSM1k has told you something real, which is that roughly eight points of its score were the benchmark rather than the arithmetic. That framing is the useful inheritance from GSM8K, and it generalises to every saturated set in this directory.
Who reports GSM8K, and how to read it
Hardly anyone now, and that is the correct outcome. GSM8K disappeared from frontier model cards once the ceiling was reached, surviving mainly as a regression check in training pipelines and as a row in evaluations of small and open-weight models where it still separates. When you do see it quoted in 2026, it is usually in marketing for a small model, where a high-nineties figure sounds impressive and means very little. Ask what the same model scores on GSM1k or GSM-Symbolic instead.
When to weight GSM8K in a model choice
Do not use GSM8K to select a model in 2026. Use it in two narrower ways. As a cheap regression test, it catches a training run that has broken basic arithmetic, which is what it is still good for. As a contamination probe, run a model on GSM8K and then on GSM1k or a GSM-Symbolic template set and read the gap, which is one of the few cheap, reproducible ways to detect memorisation in a model you did not train. If you need a live maths signal for model selection, move to AIME, FrontierMath or PutnamBench, all of which still separate the frontier.
Benchmarks to read alongside this one
MATH
Step-by-step solving of high-school competition mathematics across algebra, geometry, number theory, probability and precalculus.
AIME 2025
Olympiad-track competition mathematics at the level of the American Invitational Mathematics Examination, used as a high-difficulty LLM eval.
MMLU
Broad academic and professional knowledge across 57 subjects via four-choice multiple-choice questions.
GSM8K: frequently asked questions
- What is GSM8K?
- GSM8K is a 2021 OpenAI benchmark of 8,500 grade-school maths word problems, split 7,500 training and 1,000 test, each solvable in a few elementary arithmetic steps. It is scored by exact match on the final numeric answer.
- Is GSM8K still a useful benchmark?
- Not for choosing a model. Frontier scores sit near 99.6% and no longer separate anything. It survives as a regression test and as a contamination probe when paired with GSM1k or GSM-Symbolic.
- Did models memorise GSM8K?
- Partly. Scale AI built GSM1k, 1,000 fresh human-written problems matched to GSM8K, and found accuracy drops of up to 8% with several model families showing systematic overfitting. Apple GSM-Symbolic found drops of 0.3% to 9.2% when names and numbers were swapped.
Sources
- Cobbe, Kosaraju, Bavarian, Chen, Jun, Kaiser et al. (2021). Training Verifiers to Solve Math Word Problems. arXiv preprint 2110.14168. arxiv.org/abs/2110.14168
- Zhang, Da, Lee, Robinson, Wu, Song et al. (2024). A Careful Examination of Large Language Model Performance on Grade School Arithmetic (GSM1k). arXiv preprint 2405.00332. arxiv.org/abs/2405.00332
- Mirzadeh, Alizadeh, Shahrokhi, Tuzel, Bengio, Farajtabar (2024). GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models. Apple Machine Learning Research. machinelearning.apple.com/research/gsm-symbolic
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.
← All 117 benchmarks in the directoryWhich benchmarks are worth trusting →