Capital & Compute

TruthfulQA

Benchmark· Safety, hallucination & factuality· Checked 2026-06-29

TruthfulQA tests whether a model repeats widespread human misconceptions or answers truthfully. The questions were written specifically so that an imitative model, one that reproduces what people commonly say, gives a false answer. Its most influential finding was counterintuitive: larger models were often less truthful, because they imitated human text more faithfully.

Key facts about the TruthfulQA benchmark
What it measuresWhether a model avoids repeating common human misconceptions when answering questions, rather than imitating popular falsehoods.
Built byLin, Hilton, Evans (Oxford and OpenAI), 2021
Format817 questions across 38 categories, designed so a naive imitator gives a false answer; generation and multiple-choice formats
Scoring metric% truthful (and % truthful-and-informative)
StatusActive
Representative top scoreNot independently confirmed; see the leaderboard

How TruthfulQA works

The benchmark contains 817 questions across 38 categories such as health, law, finance and politics, each constructed so that a plausible but false answer is the one most represented in human writing. Both generation and multiple-choice formats exist. Scoring reports the percentage of truthful answers, and separately the percentage that are both truthful and informative, since refusing to answer is truthful but useless.

History and current status

Lin, Hilton and Evans introduced it in 2021 through work at Oxford and OpenAI. At release the best model was truthful on 58% of questions against 94% for humans. It became the standard honesty row in model cards during the early alignment era. Instruction tuning and reinforcement learning from human feedback substantially improved scores, which is part of why it is now aging.

What the score does not tell you

The truthful-and-informative split matters because a model can raise its truthfulness score by hedging or refusing, which is not the behaviour anyone wants. Alignment training has partly optimised directly against this benchmark, so improvement may reflect targeted tuning rather than general honesty. With 817 questions and no single live leaderboard, current comparisons are also hard to source.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

What TruthfulQA scores actually mean

TruthfulQA was built adversarially: its 817 questions across 38 categories were selected because humans commonly answer them wrongly, so a model that imitates its training data confidently gets them wrong too. At release the best model was truthful on 58% of questions against 94% for humans, and the paper's most cited finding was inverse scaling, where larger models were less truthful because they imitated human misconceptions more fluently. RLHF-tuned models have since largely closed that gap, which means a high score in 2026 partly measures refusal-and-hedging training rather than knowledge. The truthful-and-informative variant is the more honest metric, since a model that declines to answer scores well on truthfulness alone.

Who reports TruthfulQA, and how to read it

It appears less often in 2026 frontier model cards than it did in 2023, having been displaced for factuality by SimpleQA, which measures whether a model knows what it does not know, and by hallucination-rate leaderboards. TruthfulQA remains a useful concept and a weakening measurement. Its inverse-scaling result remains the more durable contribution, and it is cited far more than the benchmark is now run.

When to weight TruthfulQA in a model choice

Treat TruthfulQA as a historical result worth knowing rather than a current selection input, and read the truthful-and-informative figure if you use it at all, because plain truthfulness rewards evasion. Its enduring value is conceptual: it is the cleanest demonstration that scale alone does not produce accuracy, and that a benchmark can be gamed by teaching a model to hedge. For current factuality work, prefer domain-specific evaluations and retrieval-grounded testing on your own corpus, since generic factuality scores transfer poorly to specific subject matter.

2021
First released
Lin, Hilton, Evans (Oxford and OpenAI)
Active
Status today
As of September 4, 2026
See board
Representative top score
Not independently confirmed

Benchmarks to read alongside this one

TruthfulQA: frequently asked questions

What is TruthfulQA?
TruthfulQA is a 2021 benchmark of 817 questions across 38 categories, each written so that the answer most common in human writing is false. It measures whether a model repeats popular misconceptions, reporting both the percentage truthful and the percentage truthful and informative.
What was the inverse scaling finding?
At release, larger models were often less truthful than smaller ones. The explanation is that a bigger model imitates human text more faithfully, and human text contains the misconceptions the benchmark targets. Scaling alone did not produce honesty.
Is TruthfulQA still used?
Less than it was. Alignment training improved scores substantially, partly by optimising against the benchmark itself, and there is no single live leaderboard. SimpleQA and hallucination-rate leaderboards have largely taken over the factuality role in current model cards.

Sources

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory