Capital & Compute

The AI model value leaderboard

Leaderboard· Updated September 23, 2026

Most rankings tell you which model scores highest. This one also tells you which model is worth the money. Each LLM is rated two ways: its independent benchmark score, and its value, the points it buys you per dollar of tokens. The two orders are not the same.

Which AI model is the best value right now?

On benchmark scores alone, the strongest LLM you can buy here is Claude Fable 5.1 (81.6 of 100 on the coding composite). But cost flips the ranking: GLM-5.3-Flash delivers about 301.1 coding points per dollar of blended token price, roughly 73.4x the value of Claude Fable 5.1. The cheaper models, mostly from Chinese labs, win on value; the priciest US flagships win on raw capability. Which one is "best" depends entirely on whether you are buying ability or buying ability per dollar.

GLM-5.3-Flash
Best value, coding
301.1 coding points per dollar (blended)
Claude Fable 5.1
Top coding score (buyable)
81.6 of 100 on the coding composite
$0.20
Cheapest to run
GLM-5.3-Flash, blended (3:1) price per Mtok
36
Models tracked
27 scored on coding, 36 on intelligence

The value ladder, in one chart

Coding points per dollar of blended token price, for the models you can buy today. The ranking is almost the inverse of the raw-score ranking: the cheapest capable models sit at the top because they score within range of the frontier at a fraction of the price.

AI model coding value: composite coding score per dollar of blended token priceA lollipop chart ranking buyable models by coding value, defined as coding composite score divided by blended token price. The cheapest models rank highest, led by the open-weight GLM-5.3-Flash well ahead of MiniMax M3 and Qwen3.7 Plus; the most expensive flagships rank lowest.0.0100.0200.0300.0400.0GLM-5.3-Flash301.1MiniMax M3111.6Qwen3.7 Plus99.8GPT-5.6 Luna87.3Gemini 3.8 Flash50.9Nemotron 3 Ultra36.5Kimi K2.7 Code35.5GLM-5.334.8Devstral 234.8Llama 4 Maverick31.8Grok 4.524.1Qwen3.8 Max23.9Grok 4.322.5Grok 4.622.1GLM-5.221.6Inkling20.3Qwen3.7 Max17.6Claude Sonnet 516.6Gemini 3.1 Pro15.3Kimi K312.0GPT-5.6 Terra11.6GPT-5.6 Sol8.1Claude Opus 57.8Claude Opus 4.87.4Claude Fable 5.14.1Claude Fable 53.8GPT-6 Astra3.8
AI model coding value: composite coding score per dollar of blended token price
ItemValue
GLM-5.3-Flash301.1
MiniMax M3111.6
Qwen3.7 Plus99.8
GPT-5.6 Luna87.3
Gemini 3.8 Flash50.9
Nemotron 3 Ultra36.5
Kimi K2.7 Code35.5
GLM-5.334.8
Devstral 234.8
Llama 4 Maverick31.8
Grok 4.524.1
Qwen3.8 Max23.9
Grok 4.322.5
Grok 4.622.1
GLM-5.221.6
Inkling20.3
Qwen3.7 Max17.6
Claude Sonnet 516.6
Gemini 3.1 Pro15.3
Kimi K312.0
GPT-5.6 Terra11.6
GPT-5.6 Sol8.1
Claude Opus 57.8
Claude Opus 4.87.4
Claude Fable 5.14.1
Claude Fable 53.8
GPT-6 Astra3.8
Coding value (composite score divided by blended token price) for buyable models. Higher is more coding ability per dollar. Blended price weights input to output 3 to 1.Source: Price Per Token (composite scores) and Capital & Compute verified pricing

The same picture, on general intelligence

Coding has no score for every model, but the broader intelligence composite does, so this ladder includes all the buyable models, DeepSeek V4 Pro, the GPT-5 and GPT-5.6 tiers, Grok, Sonnet and Haiku among them. The story holds: the cheap models lead on value, and the priciest flagships (GPT-5.2 Pro, Claude Fable 5.1, Claude Fable 5) fall to the bottom, where a top score cannot outrun a high token price.

AI model intelligence value: composite intelligence score per dollar of blended token priceA lollipop chart ranking every buyable model by intelligence value, defined as intelligence composite score divided by blended token price. Cheap models such as GLM-5.3-Flash, MiniMax M3 and Qwen3.7 Plus rank highest; expensive flagships such as GPT-5.2 Pro rank lowest.0.050.0100.0150.0200.0GLM-5.3-Flash176.4MiniMax M356.4Qwen3.7 Plus46.1GPT-5.6 Luna37.3Gemini 3.8 Flash27.5GLM-5.320.9Llama 4 Maverick18.1Nemotron 3 Ultra17.3Kimi K2.7 Code15.4Qwen3.8 Max13.4Grok 4.513.0Grok 4.611.8DeepSeek V4 Pro10.5GLM-5.210.4Devstral 210.4Inkling9.9Grok 4.39.3Qwen3.7 Max8.0Claude Haiku 4.57.7Claude Sonnet 57.2Gemini 3.5 Flash7.1Gemini 3.1 Pro6.8GPT-5.3 Codex6.8Kimi K35.8Claude Opus 55.1GPT-5.6 Terra5.0Gemini 3 Pro Preview5.0Claude Opus 4.84.2Claude Sonnet 4.63.9Grok 43.8GPT-5.6 Sol3.5GPT-5.23.5Claude Fable 5.12.7Claude Fable 52.5GPT-6 Astra2.3GPT-5.2 Pro0.5
AI model intelligence value: composite intelligence score per dollar of blended token price
ItemValue
GLM-5.3-Flash176.4
MiniMax M356.4
Qwen3.7 Plus46.1
GPT-5.6 Luna37.3
Gemini 3.8 Flash27.5
GLM-5.320.9
Llama 4 Maverick18.1
Nemotron 3 Ultra17.3
Kimi K2.7 Code15.4
Qwen3.8 Max13.4
Grok 4.513.0
Grok 4.611.8
DeepSeek V4 Pro10.5
GLM-5.210.4
Devstral 210.4
Inkling9.9
Grok 4.39.3
Qwen3.7 Max8.0
Claude Haiku 4.57.7
Claude Sonnet 57.2
Gemini 3.5 Flash7.1
Gemini 3.1 Pro6.8
GPT-5.3 Codex6.8
Kimi K35.8
Claude Opus 55.1
GPT-5.6 Terra5.0
Gemini 3 Pro Preview5.0
Claude Opus 4.84.2
Claude Sonnet 4.63.9
Grok 43.8
GPT-5.6 Sol3.5
GPT-5.23.5
Claude Fable 5.12.7
Claude Fable 52.5
GPT-6 Astra2.3
GPT-5.2 Pro0.5
Intelligence value (composite score divided by blended token price) for every buyable model. Higher is more measured reasoning per dollar. The expensive flagships rank low here despite strong raw scores.Source: Price Per Token (composite scores) and Capital & Compute / source-dataset pricing

Rank it yourself

Switch the benchmark between coding and general intelligence, and switch the sort between value and raw score. The default view is coding, ranked by value.

Best LLM by coding, ranked by value

Composite of code generation, understanding, and problem-solving (0-100). Value is coding points per dollar of blended token price.

Live ranking
Benchmark

Coding: code-writing and problem-solving. Intelligence: general reasoning. Both are 0-100 composite scores.

Sort by

Value: points per dollar (best bang for the buck). Raw score: highest benchmark, price aside.

What is value? It is how many coding points you get per dollar of tokens: coding score ÷ blended price, where blended price = (3 × input + output) ÷ 4 per million tokens (input is weighted higher because coding agents read far more than they write). A higher value means more measured ability per dollar; it is a cost-efficiency measure, not a verdict on which model is best.
#ModelCodingInput $/MtokOutput $/MtokValue (pts/$)
1GLM-5.3-Flash Zhipu✓71.5$0.15$0.50301.1 Best value
2MiniMax M3 MiniMax✓58.6$0.30$1.20111.6
3Qwen3.7 Plus Alibaba✓55.9$0.32$1.2899.8
4GPT-5.6 Luna OpenAI✓39.3$0.20$1.2087.3
5Gemini 3.8 Flash Google✓76.3$0.75$3.7550.9
6Nemotron 3 Ultra Nvidia49.3$0.60$3.6036.5
7Kimi K2.7 Code Moonshot✓60.8$0.95$435.5
8GLM-5.3 Zhipu74.8$1.40$4.4034.8
9Devstral 2 Mistral✓31.3$0.90$0.9034.8
10Llama 4 Maverick Meta16.3$0.35$131.8
11Grok 4.5 xAI✓72.4$2$624.1
12Qwen3.8 Max Alibaba✓71.8$2$623.9
13Grok 4.3 xAI✓35.2$1.25$2.5022.5
14Grok 4.6 xAI✓66.3$2$622.1
15GLM-5.2 Zhipu✓46.5$1.40$4.4021.6
16Inkling Thinking Machines✓52.1$1.87$4.6820.3
17Qwen3.7 Max Alibaba✓66.0$2.50$7.5017.6
18Claude Sonnet 5 Anthropic✓66.4$2$1016.6
19Gemini 3.1 Pro Google✓68.8$2$1215.3
20Kimi K3 Moonshot✓72.0$3$1512.0
21GPT-5.6 Terra OpenAI✓52.3$2$1211.6
22GPT-5.6 Sol OpenAI✓65.1$4$208.1
23Claude Opus 5 Anthropic✓78.0$5$257.8
24Claude Opus 4.8 Anthropic✓74.3$5$257.4
25Claude Fable 5.1 Anthropic✓81.6$10$504.1
26Claude Fable 5 Anthropic✓76.5$10$503.8
27GPT-6 Astra OpenAI✓76.2$10$503.8

Switch the benchmark or sort to re-rank. Scores are 0-100 composites from the source dataset. A ✓ next to the provider marks a price reconciled with our verified registry against the provider source; the rest are the author-direct or endpoint rate reported by the source dataset, not yet independently re-verified.

How to read this leaderboard

Two grounded inputs, one derived number. The benchmark scores are composite Coding and Intelligence scores read from the Price Per Token dataset, which aggregates independent benchmarks and cites Artificial Analysis, the HuggingFace Open LLM Leaderboard, and LayerLens. They are composites on a 0-100 scale, not a single named test. Every row is read from that one source. Until the 2026-09-11 refresh, seven rows carried Artificial Analysis Intelligence Index figures instead, because the base dataset had no composite for them. Those two scales are not comparable, so the fallback has been dropped and every row re-read, which moved several rankings sharply. A model with no composite is now listed as unscored rather than ranked on a number from a different scale. The prices carry a check when they are reconciled with this site's own verified registry, the same numbers behind the model release tracker; the rest are a named host rate or a launch rate, labeled per row. Each links to its provider source.

From those two, the leaderboard computes value: the benchmark score divided by a blended token price, where blended price weights input and output tokens 3 to 1, ((3 × input) + output) ÷ 4. The 3:1 mix mirrors how an agentic coding session actually bills: it reads far more context than it writes. Value is a cost-efficiency measure, not a verdict on quality. A model can top the value ranking and still be the wrong choice for work that needs the highest absolute score. It is the same lesson as the price reversal in per-task cost: the headline number and the number that matters are rarely the same.

The independence rule

This leaderboard sells no placement. Ranking is never for sale, and no model is promoted for payment. The entire point of a value table is to be a neutral referee of what each model actually costs to use; the day a vendor could pay to look cheaper, it would be worthless. The benchmark scores come from an independent third party; the prices come from primary provider sources.

What is not here, and why

A model is listed only once an independent composite score exists for it, so the table is not padded with vendor-reported numbers. That leaves a few notable models tracked but not yet ranked:

  • Gemini 4 Argon. Announced September 30, 2026 at an introductory $2 and $10 per million tokens, rising to $4 and $20 after an introductory period Google has not dated. It is rolling out only to cyber defenders in Google's Fairwind Program, is not on the Gemini API pricing page, and Price Per Token has published no composite for it. Artificial Analysis scores it 53 at high effort, level with GPT-6 Astra at max and five behind Claude Opus 5.5, but that is a different scale from this board's composite and is not substituted in.
  • Claude Sonnet 5.5. Released September 28, 2026 at Sonnet 5's unchanged $2 and $10 per million tokens. Price Per Token has published no composite for it yet. Artificial Analysis scores it 56 at max effort, third on its index and two points behind Claude Opus 5.5, but at $7.60 per index task against $5.98 for Opus 5.5 and $5.09 for Sonnet 5: it is very verbose on that workload, so a lower rate card does not make it the cheaper model to run.
  • GPT-6.1 Sol. Released September 29, 2026 at DevDay at GPT-6 Sol's $2 and $10 per million tokens with the cached-input rate halved to $0.10. Price Per Token has published no composite for it yet. Artificial Analysis scores it 52 at max effort, one point behind GPT-6 Astra at a fifth of Astra's token price, but that is a different scale from this board's composite and is not substituted in.
  • GPT-6 Sol. Released September 22, 2026 at $2 and $10 per million tokens, exactly half the GPT-5.6 Sol it replaces and exactly the Claude Sonnet 5 rate card. Price Per Token has published no composite for it yet, and the vendor numbers cannot stand in: OpenAI reports 68.8% on DeepSWE v1.1, which is below the 73% the independent DeepSWE board records for the GPT-5.6 Sol it supersedes, and OpenAI does not say which harness produced its figure.
  • GPT-6 Luna. Released September 22, 2026 at $0.10 and $0.50 per million tokens, the cheapest row on this site. Price Per Token has published no composite for it yet. Worth watching rather than ranking: Artificial Analysis scores it 37 at max effort for $0.07 per index task, against Claude Opus 5.5 at 58 for $5.98, so it buys about two thirds of the index score for about a hundredth of the money.
  • Claude Opus 5.5. Released September 22, 2026 as Anthropic's flagship at $4 and $20 per million tokens, 20% under the Opus 5 it replaces, with the cache read cut 60% to $0.20. Price Per Token has published no composite for it yet, and the vendor numbers cannot stand in: Anthropic reports 66.4% on Terminal-Bench 4.0, but the Terminal-Bench board had no entry for the model as of September 23, so there is no independent score on this board's scale to rank it by.
  • Grok 4.7. Released September 21, 2026 at $2 and $6 per million tokens, unchanged from Grok 4.6. Price Per Token has published no composite for it yet, and the vendor numbers cannot stand in: on Terminal-Bench 4.0 alone xAI reports 38.0% from its own Grok Build harness while Artificial Analysis measures 25.8% on a standardized one, so there is no single score to rank it by.
  • DeepSeek V4.1 Flash. Released September 10, 2026 at $0.30 and $1.20 per million tokens at peak, and the only price cut of the month. Price Per Token has published no composite for it yet, and DeepSeek's own reported benchmarks are not independent, so it gets a row here rather than a rank.
  • Meta Muse Spark 1.3. Released September 2, 2026. Not carried in the Price Per Token dataset at all, so there is no composite on this board's scale to rank it by.
  • Qwen3.8-Max-0902. A September 2, 2026 post-training snapshot of Qwen3.8 Max on the same base and the same rate card. No separate composite upstream; the Qwen3.8 Max row stands in for it.
  • Gemini 3.8 Flash Cyber. Restricted to Google's Fairwind Program and carries no published rate, so it can be neither priced nor bought.
  • Ling-3.0-flash-Fin and Sante. Ant Group publishes no first-party per-token rate for either fine-tune, and neither has an independent composite. Every visible price is a third-party host, several of them promotional.
  • GPT-5.5. Scored upstream, but this site holds no verified first-party rate for it, and the value column is meaningless without one.
  • DeepSeek V4 Flash. Retired September 10, 2026. Its model id now serves V4.1 Flash at the V4.1 Flash price, so the row was removed rather than left as a rate nobody can pay.
  • Smaller and older variants. Mini, nano, distilled and superseded point releases are left off unless they change a price or a ranking.

For per-token rates and release dates across every model the site follows, see the AI model release tracker, and for the dated record of what each model charged when it launched, AI model releases by month. To turn these rates into the cost of a real job, use the cost-per-task calculator or put two models head to head. To pay nothing at all, see which AI models are free to use and good enough to ship with.

Frequently asked questions

What is the best LLM for coding in 2026?
Among models you can actually buy, Claude Fable 5.1 posts the highest coding composite in this dataset, 81.6 of 100, ahead of Claude Opus 5 at 78.0 and Gemini 3.8 Flash at 76.3. Anthropic's newer Claude Opus 5.5, released September 22, is not ranked here: the source dataset has published no composite for it yet, so it sits in the unscored list rather than being given a vendor-reported number. On a value basis, points bought per dollar, cheaper models such as GLM-5.3-Flash lead instead, because they score within range of the frontier at a fraction of the token price. Gemini 3.8 Flash is the notable case of both at once: it is the fifth-highest raw score on the board and still cheap enough to rank fifth on value.
What is the best value AI model?
On the coding composite, GLM-5.3-Flash is the value leader at about 301.1 points per dollar of blended price, roughly 73.4 times Claude Fable 5.1. It scores 71.5 on coding, within ten points of the frontier, at a blended $0.24 per million tokens. That gap to the next row is wide enough to be worth checking the rate against Z.ai before committing to it. MiniMax M3 and Qwen3.7 Plus cluster behind it. Value rewards low price, so it favors capable cheap models over the most expensive flagships: GPT-5.2 Pro, at $21/$168 per Mtok, lands last on value.
How is the value score calculated?
Value equals the benchmark composite score divided by a blended token price. The blended price weights input and output tokens 3 to 1: (3 times input + output) divided by 4, in dollars per million tokens. The 3:1 mix reflects agentic coding, which reads far more context than it writes. A higher value means more measured ability per dollar; it is a cost-efficiency measure, not a quality ranking on its own.
Where do the benchmark scores come from?
The composite Coding and Intelligence scores are read from the Price Per Token dataset, which aggregates independent benchmarks and cites Artificial Analysis, the HuggingFace Open LLM Leaderboard, and LayerLens. They are composite scores on a 0-100 scale, not a single named test. Every row is read from that one source. Until the refresh on 2026-09-11, seven rows instead carried Artificial Analysis Intelligence Index figures, because the base dataset had published no composite for them; the two scales are not comparable, and GPT-5.6 Luna scoring 75 on one and 39.3 on the other is the clearest example. Those rows have been re-read and the fallback dropped, which moved several rankings sharply. A model with no composite is listed as unscored rather than given a vendor-reported number. Prices marked with a check are reconciled with this site's own verified registry against the provider source; the rest are a named host rate or a launch rate, labeled per row. This leaderboard sells no placement: ranking is never for sale.
Are the September 2026 models on the leaderboard?
Most of them. Claude Fable 5.1, released September 1, now posts the top coding composite at 81.6; GPT-6 Astra, released September 3, scores 76.2 and Gemini 3.8 Flash 76.3. Claude Opus 5, Claude Sonnet 5, GLM-5.3, GLM-5.3-Flash and Grok 4.6 joined the board on the same refresh, once the source dataset published composites for them. Several September releases are listed as unscored instead: Claude Opus 5.5, Claude Sonnet 5.5, GPT-6 Sol and Luna, Grok 4.7, DeepSeek V4.1 Flash and Meta Muse Spark 1.3 have no independent composite yet, and the Qwen3.8-Max-0902 snapshot shares its base and rate card with the Qwen3.8 Max row. Claude Opus 5.5 is the one worth watching, because it is the first flagship in this record to launch under the price of the model it replaced, at $4 and $20 per million tokens against the $5 and $25 that Opus 5 charges.

Sources

  • Price Per Token. LLM API Pricing and Benchmarks dataset (composite Coding and Intelligence scores). Scores read 2026-09-11. https://pricepertoken.com/
  • Artificial Analysis. Independent LLM benchmarks and intelligence index (cited by the source dataset as a benchmark origin). https://artificialanalysis.ai/
  • Capital & Compute. AI model registry (verified per-token API prices, each linked to a provider source). /ai-models/

Machine-readable data: /ai-model-leaderboard.json.

← Back to Capital & Compute