The AI model value leaderboard
Most rankings tell you which model scores highest. This one also tells you which model is worth the money. Each LLM is rated two ways: its independent benchmark score, and its value, the points it buys you per dollar of tokens. The two orders are not the same.
Which AI model is the best value right now?
On benchmark scores alone, the strongest LLM you can buy here is GPT-5.6 Sol (80.0 of 100 on the coding composite). But cost flips the ranking: GPT-5.6 Luna delivers about 166.7 coding points per dollar of blended token price, roughly 16.7x the value of GPT-5.6 Sol. The cheaper models, mostly from Chinese labs, win on value; the priciest US flagships win on raw capability. Which one is "best" depends entirely on whether you are buying ability or buying ability per dollar.
The value ladder, in one chart
Coding points per dollar of blended token price, for the models you can buy today. The ranking is almost the inverse of the raw-score ranking: the cheapest capable models sit at the top because they score within range of the frontier at a fraction of the price.
| Item | Value |
|---|---|
| GPT-5.6 Luna | 166.7 |
| MiniMax M3 | 111.6 |
| Qwen3.7 Plus | 99.8 |
| Nemotron 3 Ultra | 36.5 |
| Kimi K2.7 Code | 35.5 |
| Devstral 2 | 34.8 |
| Llama 4 Maverick | 31.8 |
| Grok 4.5 | 25.3 |
| Grok 4.3 | 22.5 |
| GLM-5.2 | 21.6 |
| Qwen3.7 Max | 17.6 |
| GPT-5.6 Terra | 17.1 |
| Gemini 3.1 Pro | 15.3 |
| GPT-5.6 Sol | 10.0 |
| Claude Opus 4.8 | 7.4 |
| Claude Fable 5 | 3.8 |
The same picture, on general intelligence
Coding has no score for every model, but the broader intelligence composite does, so this ladder includes all the buyable models, DeepSeek, the GPT-5 and GPT-5.6 tiers, Grok, Sonnet and Haiku among them. The story holds: the cheap models lead on value, and the priciest flagships (GPT-5.2 Pro, Claude Fable 5, Claude Opus 4.8) fall to the bottom, where a top score cannot outrun a high token price.
| Item | Value |
|---|---|
| GPT-5.6 Luna | 115.6 |
| MiniMax M3 | 84.6 |
| DeepSeek V4 Flash | 78.8 |
| Qwen3.7 Plus | 69.6 |
| Nemotron 3 Ultra | 28.0 |
| Llama 4 Maverick | 27.9 |
| Kimi K2.7 Code | 24.5 |
| DeepSeek V4 Pro | 22.7 |
| Devstral 2 | 21.3 |
| Qwen3.8 Max | 19.3 |
| Grok 4.5 | 18.7 |
| Inkling | 16.3 |
| GLM-5.2 | 16.2 |
| Grok 4.3 | 15.9 |
| GPT-5.6 Terra | 12.7 |
| Qwen3.7 Max | 12.3 |
| Claude Haiku 4.5 | 11.9 |
| Gemini 3.1 Pro | 10.3 |
| Gemini 3.5 Flash | 10.3 |
| Kimi K3 | 10.0 |
| GPT-5.3 Codex | 9.2 |
| GPT-5.6 Sol | 7.6 |
| Gemini 3 Pro Preview | 7.4 |
| Claude Sonnet 4.6 | 5.7 |
| Claude Opus 4.8 | 5.6 |
| Grok 4 | 5.6 |
| GPT-5.2 | 5.4 |
| Claude Fable 5 | 3.0 |
| GPT-5.2 Pro | 0.7 |
Rank it yourself
Switch the benchmark between coding and general intelligence, and switch the sort between value and raw score. The default view is coding, ranked by value.
Best LLM by coding, ranked by value
Composite of code generation, understanding, and problem-solving (0-100). Value is coding points per dollar of blended token price.
| # | Model | Coding | Input $/Mtok | Output $/Mtok | Value (pts/$) |
|---|---|---|---|---|---|
| 1 | GPT-5.6 Luna OpenAI✓ | 75.0 | $0.20 | $1.20 | 166.7 |
| 2 | MiniMax M3 MiniMax | 58.6 | $0.30 | $1.20 | 111.6 |
| 3 | Qwen3.7 Plus Alibaba | 55.9 | $0.32 | $1.28 | 99.8 |
| 4 | Nemotron 3 Ultra Nvidia | 49.3 | $0.60 | $3.60 | 36.5 |
| 5 | Kimi K2.7 Code Moonshot✓ | 60.8 | $0.95 | $4 | 35.5 |
| 6 | Devstral 2 Mistral | 31.3 | $0.90 | $0.90 | 34.8 |
| 7 | Llama 4 Maverick Meta | 16.3 | $0.35 | $1 | 31.8 |
| 8 | Grok 4.5 xAI✓ | 76.0 | $2 | $6 | 25.3 |
| 9 | Grok 4.3 xAI | 35.2 | $1.25 | $2.50 | 22.5 |
| 10 | GLM-5.2 Zhipu✓ | 46.5 | $1.40 | $4.40 | 21.6 |
| 11 | Qwen3.7 Max Alibaba✓ | 66.0 | $2.50 | $7.50 | 17.6 |
| 12 | GPT-5.6 Terra OpenAI✓ | 77.0 | $2 | $12 | 17.1 |
| 13 | Gemini 3.1 Pro Google✓ | 68.8 | $2 | $12 | 15.3 |
| 14 | GPT-5.6 Sol OpenAI✓ | 80.0 | $4 | $20 | 10.0 |
| 15 | Claude Opus 4.8 Anthropic✓ | 74.3 | $5 | $25 | 7.4 |
| 16 | Claude Fable 5 Anthropic✓ | 76.5 | $10 | $50 | 3.8 |
Switch the benchmark or sort to re-rank. Scores are 0-100 composites from the source dataset. A ✓ next to the provider marks a price reconciled with our verified registry against the provider source; the rest are the author-direct or endpoint rate reported by the source dataset, not yet independently re-verified.
How to read this leaderboard
Two grounded inputs, one derived number. The benchmark scores are composite Coding and Intelligence scores read from the Price Per Token dataset, which aggregates independent benchmarks and cites Artificial Analysis, the HuggingFace Open LLM Leaderboard, and LayerLens. They are composites on a 0-100 scale, not a single named test. A few of the newest frontier models (the GPT-5.6 tiers and Grok 4.5) are not yet in that dataset, so their scores are read directly from Artificial Analysis and marked per row. The prices carry a check when they are reconciled with this site's own verified registry, the same numbers behind the model release tracker; the rest are the author-direct or endpoint rate reported by the source dataset, labeled per row and not yet independently re-verified. Each links to its provider source.
From those two, the leaderboard computes value: the benchmark score divided by a blended token price, where blended price weights input and output tokens 3 to 1, ((3 × input) + output) ÷ 4. The 3:1 mix mirrors how an agentic coding session actually bills: it reads far more context than it writes. Value is a cost-efficiency measure, not a verdict on quality. A model can top the value ranking and still be the wrong choice for work that needs the highest absolute score. It is the same lesson as the price reversal in per-task cost: the headline number and the number that matters are rarely the same.
The independence rule
This leaderboard sells no placement. Ranking is never for sale, and no model is promoted for payment. The entire point of a value table is to be a neutral referee of what each model actually costs to use; the day a vendor could pay to look cheaper, it would be worthless. The benchmark scores come from an independent third party; the prices come from primary provider sources.
What is not here, and why
A model is listed only once an independent composite score exists for it, so the table is not padded with vendor-reported numbers. That leaves a few notable models tracked but not yet ranked:
- GPT-6 Astra. OpenAI's new flagship, released 2026-09-03 at $10/$50 per Mtok with a $1 cache read, 2.5x the GPT-5.6 Sol rate this board already scores. Artificial Analysis puts it at Intelligence Index 61 at max effort (v4.1.1, read 2026-09-04), which ties Sol at max and trails Claude Fable 5.1 at 66, Claude Opus 5 at 63 and Muse Spark 1.3 at 62, and its AA Coding Agent Index of 67 trails Fable 5.1 in Claude Code at 70. Price Per Token, the base source this board uses, has published no composite for it, and a row scored from a different index would not be comparable with the scored rows. Note also that this board's blended price formula ((3*input + output)/4) puts Astra at $20, level with Fable 5.1 and Fable 5, which hides the thing that actually separates them: on Terminal-Bench 4.0 Astra solves a task for $17.02 against Fable 5.1's $32.69 because it burns 1.53B tokens to Fable 5.1's 2.75B. That effect is modeled on /coding-agent-cost-per-task/ instead. Promote once Price Per Token publishes.
- Claude Fable 5.1. Anthropic's newest frontier model, released September 1, 2026 at $10/$50 per Mtok, the same sticker as the Fable 5 it extends. Its one price change is the cache read, cut from $1.00 to $0.25 per Mtok, which this board's blended price formula ((3*input + output)/4) does not see at all: on blended price Fable 5.1 and Fable 5 are identical at $20, so a scored row would rank it exactly where Fable 5 already sits and hide the only thing that changed. Price Per Token, the base source this board uses, lists only anthropic-claude-fable-5 as of this check and has published no composite for 5.1. Left off pending that composite, the same treatment as Claude Opus 5 and Claude Sonnet 5; the cache-read effect is modeled on /coding-agent-cost-per-task/ instead, where cache tokens are billed separately.
- GLM-5.3-Flash. Z.ai's official reveal of the Ox Alpha stealth model, released 2026-08-26 at $0.15/$0.50 per Mtok standard rate ($0.03 cache read), a 320B-A18B mixture-of-experts model, MIT-licensed. Artificial Analysis scores it Intelligence Index 57, read 2026-08-27, level with Claude Opus 4.8 (also 57, same read date) and just behind Qwen3.8 Max (58). Price Per Token, the base source this board uses, has not published a composite for it yet (only z-ai-glm-5.3 and z-ai-glm-5.2 exist in that dataset as of this check). Left off the scored rows pending that composite, the same treatment as Claude Opus 5 and Claude Sonnet 5 below; GLM-5.2 continues to represent Zhipu on the scored board for now.
- Claude Opus 5. Anthropic's new flagship, released July 24, 2026 at $5/$25 per Mtok (the same rate as Opus 4.8, which it replaces). Artificial Analysis rates it Intelligence Index 63 at max effort, re-read 2026-08-07 (up from 61), still the highest score on the index and just ahead of Fable 5 at 62 in its fallback configuration. It is also the most expensive model to evaluate: $3,836.05 to run the full Intelligence Index. Price Per Token, the base source this board uses, has not yet published a coding/intelligence composite for it. It is left off pending that composite to keep the Anthropic rows on one source; Opus 4.8 represents the tier below for now.
- Claude Sonnet 5. Released June 30, 2026. Price Per Token has not yet published an independent coding/intelligence composite for it. Artificial Analysis scores it at Intelligence Index 55 at max effort, re-read 2026-08-07 (up from 53), and Anthropic's own launch benchmarks (SWE-bench Pro 63.2%, Terminal-Bench 2.1 80.4%) are self-reported; it is left off pending a Price Per Token composite to keep the Anthropic rows on one source.
- GPT-5.5. OpenAI's prior flagship, now superseded by the GPT-5.6 family (Sol, Terra, Luna), which is scored above. Price Per Token has not published a composite for GPT-5.5 and it is no longer OpenAI's current tier, so it is not scored here.
- Meta Muse Spark 1.1. Meta's first paid model, released July 9, 2026 at $1.25/$4.25 per Mtok. Artificial Analysis scores it at Intelligence Index 53 in the Xhigh Effort configuration, read 2026-08-07. That is 10 points above the 43 this note previously carried, which is far larger than the index-wide drift of the same period; the earlier read was recorded without a variant label, so it was most likely a lower-effort configuration rather than a change in the model. Treat the 43 as unattributable and the 53 as the Xhigh figure specifically. Artificial Analysis has since added a Meta Muse Spark 1.2 at 57 in the same configuration, which this board does not yet track. Price Per Token, the base source this board uses, has not published a composite for either. Muse Spark leads on tool-use and agentic tests but trails the frontier on coding. Meta Llama 4 Maverick represents Meta on the board for now.
- Cohere North Mini Code. Free on hosted endpoints and open-weight, so a per-token value score is undefined. It posts a 33.4 Artificial Analysis Coding Index, which was not re-read on 2026-08-07 because the Coding Index is not exposed on the public models leaderboard; its Intelligence Index reads 20 there, a reminder that the two indices are different scales and must not be swapped for one another. The real cost is self-hosted compute, not a token rate.
- Smaller and older variants. Models below roughly 10B parameters, superseded 2024-era releases (Claude 3.5, GPT-4 Turbo, o1), and narrowly tracked or unpriced entries are left off to keep the board to current, recognizable, buyable models.
- Google Gemini 3.8 Flash. Released 2026-09-02 at $0.75/$3.75 per Mtok, the same rate Gemini 3.6 Flash and 3.7 Flash carry, and introductory through 2026-12-31 before doubling on 2027-01-01. Artificial Analysis scores it 59 on the Intelligence Index at high effort (read 2026-09-03). Price Per Token, the base source this board uses, has published no composite for it, and a board score built from a different index would not be comparable with the scored rows. Promote once Price Per Token publishes.
- Meta Muse Spark 1.3. Released 2026-09-02 at $1.25/$4.25 per Mtok, unchanged from Muse Spark 1.2. Artificial Analysis scores the generally available xhigh configuration 61 on the Intelligence Index and the partner-only max configuration 62 (read 2026-09-03), against 57 for 1.2 at xhigh. Price Per Token has published no composite for any Muse Spark generation, so Meta is still represented on the board by Llama 4 Maverick. Promote once Price Per Token publishes.
- Qwen3.8-Max-0902. Qwen's in-place flagship refresh, dated 2026-09-02 on the QwenCloud changelog, at the same $2/$6 per Mtok as the August Qwen3.8-Max base. Qwen's own table claims Terminal-Bench 3.0 doubled from 11.3 to 29.0, vendor-reported and unreproduced, and the comparison predates Claude Fable 5.1. Price Per Token has published no composite for the snapshot, and a board score built from a different index would not be comparable with the scored rows; Qwen3.8 Max continues to represent Alibaba on the scored board for now. Promote once Price Per Token publishes.
For per-token rates and release dates across every model the site follows, see the AI model release tracker, and for the dated record of what each model charged when it launched, AI model releases by month. To turn these rates into the cost of a real job, use the cost-per-task calculator or put two models head to head. To pay nothing at all, see which AI models are free to use and good enough to ship with.
Frequently asked questions
- What is the best LLM for coding in 2026?
- Among models you can actually buy, GPT-5.6 Sol posts the highest coding composite in this dataset, 80.0 of 100, just ahead of Claude Opus 4.8 and the restored Claude Fable 5. On a value basis, points bought per dollar, cheaper models such as GPT-5.6 Luna lead instead, because they score within range of the frontier at a fraction of the token price.
- What is the best value AI model?
- On the coding composite, GPT-5.6 Luna is the value leader at about 166.7 points per dollar of blended price, roughly 16.7 times GPT-5.6 Sol. Other low-cost models (Qwen3.7 Plus, Kimi K2.7) cluster near the top too, and on the intelligence composite DeepSeek V4 Flash leads outright at $0.14/$0.28 per Mtok. Value rewards low price, so it favors capable cheap models over the most expensive flagships: GPT-5.2 Pro, at $21/$168 per Mtok, lands last on value despite a top-tier score.
- How is the value score calculated?
- Value equals the benchmark composite score divided by a blended token price. The blended price weights input and output tokens 3 to 1: (3 times input + output) divided by 4, in dollars per million tokens. The 3:1 mix reflects agentic coding, which reads far more context than it writes. A higher value means more measured ability per dollar; it is a cost-efficiency measure, not a quality ranking on its own.
- Where do the benchmark scores come from?
- The composite Coding and Intelligence scores are read from the Price Per Token dataset, which aggregates independent benchmarks and cites Artificial Analysis, the HuggingFace Open LLM Leaderboard, and LayerLens. They are composite scores on a 0-100 scale, not a single named test. Scores for the newest frontier models the source dataset has not composited yet (the GPT-5.6 tiers and Grok 4.5) are read directly from Artificial Analysis, the same independent benchmark family it aggregates, and flagged per row. Prices marked with a check are reconciled with this site's own verified registry against the provider source; the rest are the author-direct or endpoint rate reported by the source dataset, labeled per row and not yet independently re-verified. This leaderboard sells no placement: ranking is never for sale.
- Are Grok 4.5 and GPT-5.6 on the leaderboard?
- Yes. GPT-5.6 (Sol, Terra, Luna) went generally available July 9, 2026 and Grok 4.5 released July 8, and both are scored here now. Because the Price Per Token dataset this board uses as its base has not published composites for them yet, their scores are read directly from Artificial Analysis (the same independent benchmark family Price Per Token aggregates) and flagged per row. GPT-5.6 Sol posts the top coding composite in the set; GPT-5.6 Luna is the best-value US model on coding; Grok 4.5 lands mid-pack on value at $2/$6 per Mtok.
Sources
- Price Per Token. LLM API Pricing and Benchmarks dataset (composite Coding and Intelligence scores). Scores read 2026-07-13. https://pricepertoken.com/
- Artificial Analysis. Independent LLM benchmarks and intelligence index (cited by the source dataset as a benchmark origin). https://artificialanalysis.ai/
- Capital & Compute. AI model registry (verified per-token API prices, each linked to a provider source). /ai-models/
Machine-readable data: /ai-model-leaderboard.json.