Vals Index
Also known as Vals AI Index, Vals AI benchmark
The Vals Index is a single score from Vals AI, an independent evaluation company, that averages a model's results on agentic finance, coding, legal and tax tasks and weights each sector by its share of US GDP. It is built to answer how much economically useful work a model can do, not how much it knows, and the weighting means finance tasks decide more than half of the result.
| What it measures | A composite of agentic finance, coding, legal and tax tasks, weighted by each sector's share of US GDP, meant to estimate the economic impact of a model rather than its general intelligence. |
|---|---|
| Built by | Vals AI (independent), 2025 |
| Format | A GDP-weighted average of eight benchmarks in four sectors; v2.1 (25 September 2026) computes (8.0 x Finance + 5.6 x Coding + 1.2 x Legal + 0.5 x Tax) / 15.3, with six private and two public component benchmarks |
| Scoring metric | Weighted accuracy (%), with standard-error bars, plus cost per test and latency; scores are not comparable across index versions |
| Status | Active |
| Representative top score | 68.90% (v2.1) · Gemini 4 Argon (high effort) · read 2026-09 |
| Official leaderboard | www.vals.ai/benchmarks/vals_index |
How Vals Index works
Each sector score is the plain average of its component benchmarks, and the four sectors are then combined with fixed weights taken from Bureau of Economic Analysis value-added-by-industry data, as published on FRED. Version 2.1 uses the formula (8.0 x Finance + 5.6 x Coding + 1.2 x Legal + 0.5 x Tax) / 15.3. Finance is Finance Agent v2 and the Excel Modeling Benchmark. Coding is Terminal-Bench 4.0, Vibe Code Bench and Code Migration. Legal is Legal Research Bench and HLAB. Tax is Tax Agent Bench. Worked through, finance carries 52.3% of the index, coding 36.6%, legal 7.8% and tax 3.3%. Six of the eight components are private test sets that no outside party fully holds; Terminal-Bench 4.0 and Vibe Code Bench are public. Vals runs every model through a fixed harness it controls, reports a standard error on each score, and publishes the cost per test and wall-clock latency of each run alongside the accuracy figure.
History and current status
Vals AI introduced the index on 10 October 2025 as part of a website redesign, describing it as a way to quantify the potential economic impact of AI models on the US economy. It has been revised five times in 2026. On 4 May, Vibe Code Bench was added and the saturated CaseLaw benchmark removed. On 13 May, finance moved to Finance Agent v2. On 27 May, coding moved to Terminal-Bench 2.1. Version 2, on 13 August, alongside the company's $40 million Series A, swapped CorpFin for the Excel Modeling Benchmark, added Code Migration, Legal Research Bench and HLAB, and dropped SWE-bench Verified. Version 2.1, on 25 September 2026, added the Tax sector and moved coding to Terminal-Bench 4.0, which raised the formula denominator from 14.8 to 15.3. Five days later, on 30 September, Gemini 4 Argon became the first Gemini model to top the board.
What the score does not tell you
The weighting is the index's thesis and its main weakness. Because finance is more than half the score, a model tuned for spreadsheet modelling and financial research can lead the index while trailing on general coding or reasoning, and a reader who takes the number as overall capability will be misled. The GDP lens is also a US one: it fixes the sector mix to one economy and leaves out healthcare, manufacturing and most of the service sector entirely. The private test sets protect against contamination but mean the results cannot be reproduced by anyone outside Vals, and the methodology page notes that a larger private validation set is licensed to companies, so the same firm that grades the models sells evaluation data. The version cadence makes scores perishable: Claude Sonnet 5.5's own model page still carries a 69.22% summary figure from an earlier board, above the 67.04% it scores on v2.1. Finally, the leading margins are small relative to the published error bars.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Vals Index is not one of the 15 benchmarks scored in the benchmark trust ranking, which rates the most-quoted tests on those same four axes and sorts them most trustworthy first.
What Vals Index scores actually mean
On v2.1, read 1 October 2026, the top of the board is tight. Gemini 4 Argon scores 68.90% plus or minus 0.97 at $15.68 per test, Claude Sonnet 5.5 67.04% plus or minus 0.92 at $21.34, Claude Opus 5.5 66.97% at $32.14 and Claude Fable 5.1 65.83% at $28.71. The gap between first and second is 1.86 points, about 1.4 combined standard errors, short of the roughly two that would make it a clear lead. Cost and speed separate the leaders more than accuracy does. Vals records a latency of 46 minutes 33 seconds for Argon, against about 78 minutes for each of the three Claude models, but GPT-6 Astra (27 minutes 48 seconds, 63.13%) and Muse Spark 1.3 Max (23 minutes 33 seconds, 58.16%) are faster. GPT-6.1 Sol reaches 61.15% for $3.24 a test, a fifth of Argon's cost, which is the better value read for most buyers.
Who reports Vals Index, and how to read it
Labs cite Vals component benchmarks in launch materials: Google's Gemini 4 Argon announcement table carried a Vals Finance Agent v2 row, and Vals says its results have been cited in model cards from OpenAI, Anthropic, Google, Meta and xAI. The composite index itself travels mostly through Vals' own posts and through community reposts, such as the 30 September 2026 r/ClaudeCode thread claiming Argon tops the index on speed, cost and accuracy. That claim holds for accuracy, holds for cost only among the top seven, and does not hold for speed. When you see a Vals Index figure, ask for the version, the reasoning effort Vals ran the model at, and the error bar.
When to weight Vals Index in a model choice
Use the Vals Index when the work you are buying looks like its sector mix: financial research, spreadsheet modelling, legal research and agentic coding. For that profile it is one of the few boards that grades finished professional work against an expert standard, and its cost-per-test column turns the ranking into a budget. Discount it when your workload is mostly software engineering, research or general reasoning, where the Artificial Analysis Intelligence Index or a single coding board such as Terminal-Bench is closer to the job. Read the sector scores rather than the composite when they are available, treat any lead smaller than about two points as a tie, and record the version, the effort setting and the read date with any number you keep, because the formula has changed five times in 2026.
Benchmarks to read alongside this one
Artificial Analysis Intelligence Index
A composite index of overall model intelligence aggregating performance across reasoning, coding, knowledge, science and agentic tasks.
Terminal-Bench
Whether an AI agent can complete hard, realistic command-line tasks (build, configure, train, debug, secure) end to end inside a real terminal.
LegalBench
Legal reasoning across the specific skills lawyers actually use, as defined by legal professionals.
Vals Index: frequently asked questions
- What is the Vals Index?
- A composite AI benchmark from Vals AI that averages agentic finance, coding, legal and tax tasks, weighting each sector by its share of US GDP. Version 2.1 uses (8.0 x Finance + 5.6 x Coding + 1.2 x Legal + 0.5 x Tax) / 15.3.
- Which model leads the Vals Index?
- Gemini 4 Argon leads version 2.1 at 68.90%, read 1 October 2026, ahead of Claude Sonnet 5.5 at 67.04% and Claude Opus 5.5 at 66.97%. The first-to-second gap is about 1.4 combined standard errors, so it is a narrow lead.
- Is the Vals Index the same as the Artificial Analysis Intelligence Index?
- No. Vals weights sectors by GDP, so finance is 52% of its score. The Artificial Analysis index weights agents, general knowledge, coding and scientific reasoning, and ranks Claude Opus 5.5 first at 58 with Argon at 53.
- Are Vals Index test sets public?
- Mostly not. Six of the eight component benchmarks are private test sets held by Vals AI; Terminal-Bench 4.0 and Vibe Code Bench are public. That limits contamination but also means outsiders cannot reproduce the scores.
Sources
- Vals AI (2026). Vals Index leaderboard and methodology, version 2.1, updated 30 September 2026. www.vals.ai/benchmarks/vals_index
- Vals AI (2026). Methodology: private test sets, fixed harness, error bars, cost and latency. www.vals.ai/methodology
- Vals AI (2026). Gemini 4 Argon benchmarks, cost and capabilities (model page). www.vals.ai/models/google_gemini-4-argon
- Vals AI (2026). Series A: Always a Higher Peak, 13 August 2026. www.vals.ai/blogs/series-a
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.
← All 117 benchmarks in the directoryWhich benchmarks are worth trusting →