Capital & Compute

Vals Index vs Artificial Analysis: Why Rankings Differ

Gemini 4 Argon is first on the Vals Index and tied third on Artificial Analysis. The split comes from weighting: finance is 52% of the Vals score.

· ai· benchmarks· evaluation· By Capital & Compute
Gemini 4 Argon rises from third on Artificial Analysis to first on the Vals Index as Claude Opus 5.5 drops to third.

The Vals Index and the Artificial Analysis Intelligence Index disagree because they weight different work, not because one of them is broken. Vals weights each sector by its share of US GDP, which makes finance tasks 52% of its score. Artificial Analysis weights agents, general knowledge, coding and scientific reasoning. That is why Gemini 4 Argon is first on Vals at 68.90% and tied for third on Artificial Analysis at 53, behind Claude Opus 5.5 at 58. If your work is financial analysis, the Vals ordering is closer to your job. If it is software engineering or research, the Artificial Analysis ordering is.

The question started on Reddit. On September 30, 2026, a r/ClaudeCode thread posted a screenshot captioned “Gemini 4 Argon tops Val AI benchmark on speed, cost and accuracy.” The top replies called it benchmaxxing, and one said “these benchmarks are all fake.” Both reactions skip the useful part, which is what this particular benchmark measures.

What is the Vals Index?

The Vals Index is a composite score from Vals AI, an independent evaluation company. It averages a model’s results on agentic tasks in four sectors and weights each sector by its value added to the US economy, using Bureau of Economic Analysis data. According to the Vals Index methodology page, version 2.1 (September 25, 2026) computes:

(8.0 × Finance + 5.6 × Coding + 1.2 × Legal + 0.5 × Tax) / 15.3

Sector Component benchmarks Weight Share of index Test sets
Finance Finance Agent v2, Excel Modeling Benchmark 8.0 52.3% Both private
Coding Terminal-Bench 4.0, Vibe Code Bench, Code Migration 5.6 36.6% 2 public, 1 private
Legal Legal Research Bench, HLAB 1.2 7.8% Both private
Tax Tax Agent Bench 0.5 3.3% Private

The share column is arithmetic on Vals’ published weights. Two things follow from it. First, a model that is strong at spreadsheet modelling and financial research starts with an advantage that no amount of coding skill can fully offset. Second, about 76% of the index by weight comes from test sets that only Vals holds, which is a strong defence against labs training on the answers. The full entry is in the AI benchmark directory under Vals Index.

Why does Gemini 4 Argon rank first on the Vals Index?

Because it is strongest where the weight is. The Vals model page for Gemini 4 Argon lists it first on Finance Agent v2, which carries half of the 52% finance share, and second on both Vibe Code Bench and Code Migration. Google’s own Gemini 4 Argon announcement made the same case, putting a Vals Finance Agent v2 row in its launch table.

Where Argon is weaker, the Vals weighting barely registers. On Terminal-Bench 4.0, Google’s own table has Argon at 57.4% against 66.4% for Claude Opus 5.5, as transcribed by VentureBeat. That is a vendor-reported number in a harness Google does not name, so treat it as a direction, not a measurement. Terminal-Bench is one of three coding benchmarks, so on Vals it carries about 12% of the index. On Artificial Analysis, Terminal-Bench 4.0 is a 10% component of a score where coding and agent work together carry half the weight.

Vals runs Argon at high reasoning effort and prices it at Google’s $4 and $20 standard rate per million tokens, not the introductory $2 and $10. That makes its $15.68 cost per test the honest long-run figure. The full pricing story is in the Gemini 4 Argon pricing and benchmarks breakdown.

Did Argon really win on speed, cost and accuracy?

On accuracy, narrowly. On cost, only among the leaders. On speed, no. Here is the top of the v2.1 board as read on October 1, 2026:

Vals rank Model Accuracy Cost per test Latency
1 Gemini 4 Argon 68.90% ± 0.97 $15.68 46m 33s
2 Claude Sonnet 5.5 67.04% ± 0.92 $21.34 1h 18m
3 Claude Opus 5.5 66.97% $32.14 1h 19m
4 Claude Fable 5.1 65.83% $28.71 1h 17m
6 GPT-6 Astra 63.13% $18.46 27m 48s
8 GPT-6.1 Sol 61.15% $3.24 43m 10s
9 Muse Spark 1.3 Max 58.16% $3.79 23m 33s

Accuracy. Argon leads Sonnet 5.5 by 1.86 points. Using the standard errors Vals publishes on each model page, that gap is about 1.4 combined standard errors, below the roughly two that would make it a clear statistical lead. It is a real first place and a thin one.

Cost. Argon is the cheapest of the top seven per test, about 27% below Sonnet 5.5 and about half of Opus 5.5. It is not the cheapest model on the board. GPT-6.1 Sol reaches 61.15% for $3.24, about a fifth of Argon’s cost, for 7.75 fewer points.

Speed. Argon is far faster than the three Claude models above and below it, at roughly 47 minutes against 78. But Muse Spark 1.3 Max and GPT-6 Astra both post lower latency, and of the top 15 on the board, eight models are faster than Argon. “Fastest of the top four” is accurate. “Tops on speed” is not.

How is the Vals Index different from the Artificial Analysis Intelligence Index?

The Artificial Analysis Intelligence Index is the other one-number score most launch coverage quotes. Side by side:

Vals Index Artificial Analysis Intelligence Index
Maker Vals AI (independent) Artificial Analysis (independent)
Current version v2.1, September 25, 2026 v4.3, September 7, 2026
Weighting basis Share of US GDP by sector Capability category
Weights Finance 52%, coding 37%, legal 8%, tax 3% Agents 30%, general 30%, coding 20%, science 20%
Components 8 benchmarks 11 evaluations
Private test share About 76% by weight (derived) 45% by weight
Score shape Weighted accuracy, with error bars 0 to 100 composite
Cost reported Dollars per test Cost to run the full index
Leader, October 1, 2026 Gemini 4 Argon, 68.90% Claude Opus 5.5 (max), 58

The Artificial Analysis weights and component count come from its v4.3 announcement. The leader comes from the Artificial Analysis models leaderboard, read between September 25 and October 1, 2026.

The two indexes were built to answer different questions. Artificial Analysis asks how capable a model is across the kinds of tasks frontier models get tested on: knowledge, reasoning, coding and agent work. Vals asks how much of the US economy’s professional work a model could do, and the economy it models is mostly finance. Neither question is wrong. A buyer has to know which one they are asking.

Vals Index rank vs Artificial Analysis rankGemini 4 Argon moves from tied third on Artificial Analysis to first on Vals, while Claude Opus 5.5 drops from first to third. The Claude models hold the next places on both boards.AA Intelligence IndexVals Index v2.11. Claude Opus 5.5 582. Claude Sonnet 5.5 563. Gemini 4 Argon 534. Claude Fable 5.1 535. GPT-6 Astra 536. GPT-6.1 Sol 527. Muse Spark 1.3 488. GPT-6 Sol 489. Grok 4.7 4610. Gemini 3.8 Flash 411. Gemini 4 Argon 68.90%2. Claude Sonnet 5.5 67.04%3. Claude Opus 5.5 66.97%4. Claude Fable 5.1 65.83%5. GPT-6 Astra 63.13%6. GPT-6.1 Sol 61.15%7. Muse Spark 1.3 58.16%8. GPT-6 Sol 57.54%9. Grok 4.7 54.95%10. Gemini 3.8 Flash 54.83%
Vals Index rank vs Artificial Analysis rank
ModelAA Intelligence IndexRank on Artificial AnalysisVals Index v2.1Rank on Vals Index
Gemini 4 Argon53368.90%1
Claude Opus 5.558166.97%3
Claude Sonnet 5.556267.04%2
Claude Fable 5.153465.83%4
GPT-6 Astra53563.13%5
GPT-6.1 Sol52661.15%6
Muse Spark 1.348758.16%7
GPT-6 Sol48857.54%8
Grok 4.746954.95%9
Gemini 3.8 Flash411054.83%10
Rank on the Artificial Analysis Intelligence Index (left) against rank on the Vals Index v2.1 (right), for the ten models scored by both. The gold line ranks higher on Vals than on Artificial Analysis; the red line ranks lower; grey lines hold their place. Ties on Artificial Analysis (three models at 53, two at 48) keep their Vals order. Artificial Analysis values are max effort except Argon and Gemini 3.8 Flash (high) and Grok 4.7 (xhigh); Vals runs its own settings per model. Sources: Artificial Analysis models leaderboard, read 2026-09-25 to 2026-10-01; Vals Index v2.1, read 2026-10-01.

The chart shows how little actually moves. Below the top four, the two orderings agree almost model for model: GPT-6 Astra, GPT-6.1 Sol, Muse Spark 1.3, GPT-6 Sol, Grok 4.7 and Gemini 3.8 Flash sit in the same order on both. The disagreement is concentrated at the top, where Argon and Opus 5.5 trade places and the gaps on both boards are a few points. That is the pattern expected from two honest boards with different weights, not from one board that has been gamed.

Is the Vals Index benchmaxxed?

The Reddit thread’s theory was that a lab pours compute into the benchmark run and then serves a weaker model to everyone else. The Vals design answers part of that and leaves part of it open.

What counts in its favour. Most of the index is private: six of the eight component benchmarks are test sets that Vals says “no other party has full access to,” so a lab cannot train on them directly. Vals runs every model through a fixed harness it controls, so a lab cannot pick a flattering scaffold. And it publishes a standard error for each score, which lets a reader see when a lead is noise. These are the same defences covered in the explainer on benchmark contamination.

What stays open. A private test set cannot be checked by anyone outside Vals, so the results rest on trust in one company. That company also sells evaluation data: the Vals methodology page says a larger private validation set is licensed to companies. Vals announced a $40 million Series A led by a16z at a $400 million valuation on August 13, 2026, and says its results have been cited in model cards from OpenAI, Anthropic, Google, Meta and xAI. None of that is evidence of bias, but it is the commercial setting a reader should know about. Finally, no benchmark can test the model a provider serves next month. A score describes the checkpoint Vals ran on the day it ran it.

So “fake” is the wrong word. “Narrow” is the right one. The Vals Index is a reasonable measure of a specific mix of finance, coding and legal work, and a poor measure of general capability. The broader version of this argument is in are AI benchmarks reliable.

Which index should you use to pick a model?

Pick the index whose weights look like your workload, then check the component benchmark that matches your job.

  • Financial analysis, modelling, legal research. Weight the Vals Index. Argon leads it, and its Finance Agent v2 result is the single most relevant number. Check access first: at launch Argon was open only to a restricted early-access programme.
  • Software engineering and coding agents. Weight Artificial Analysis and a single coding board such as Terminal-Bench. Claude Opus 5.5 leads Artificial Analysis and beats Argon on Terminal-Bench 4.0 in Google’s own table.
  • Budget-constrained general use. Read the cost columns, not the rank. GPT-6.1 Sol is 7.75 points behind Argon on Vals for about a fifth of the cost per test, and one point behind Argon on Artificial Analysis.
  • Any decision. Treat a lead of under two points on either index as a tie, and record the version and the read date with any score you keep.

Frequently asked questions

Why does Gemini 4 Argon rank higher on Vals than on Artificial Analysis?
The Vals Index weights sectors by US GDP, so finance tasks are 52% of its score, and Argon ranks first on Finance Agent v2. Artificial Analysis weights agents, general knowledge, coding and scientific reasoning, where Claude Opus 5.5 scores 58 against Argon at 53.
Is Gemini 4 Argon the fastest model on the Vals Index?
No. Vals records 46 minutes 33 seconds of latency for Argon, faster than the three Claude models near it at about 78 minutes, but GPT-6 Astra and Muse Spark 1.3 Max are faster, and eight of the top 15 models beat Argon on latency.
How much of the Vals Index is coding?
About 37%. Coding has a weight of 5.6 out of 15.3 in version 2.1, split evenly across Terminal-Bench 4.0, Vibe Code Bench and Code Migration. Finance is about 52%, legal about 8% and tax about 3%.
Which is more trustworthy, the Vals Index or Artificial Analysis?
Neither is more trustworthy in general. Vals relies more on private test sets, about 76% of its weight against 45% for Artificial Analysis, which limits contamination but prevents outside checks. Choose by workload: Vals for finance and legal work, Artificial Analysis for coding and general capability.

Sources

Vals AI (2026). Vals Index Leaderboard and Methodology, version 2.1. Vals AI. https://www.vals.ai/benchmarks/vals_index Read 2026-10-01.

Vals AI (2026). Gemini 4 Argon benchmarks, cost and capabilities. Vals AI model page. https://www.vals.ai/models/google_gemini-4-argon Read 2026-10-01.

Vals AI (2026). Claude Sonnet 5.5 benchmarks, cost and capabilities. Vals AI model page. https://www.vals.ai/models/anthropic_claude-sonnet-5-5 Read 2026-10-01.

Vals AI (2026). Methodology. Vals AI documentation. https://www.vals.ai/methodology Read 2026-10-01.

Vals AI (2026). Series A: Always a Higher Peak. Vals AI blog, August 13, 2026. https://www.vals.ai/blogs/series-a

Artificial Analysis (2026). Announcing the Artificial Analysis Intelligence Index v4.3. Artificial Analysis. https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3

Artificial Analysis (2026). LLM Leaderboard: Comparison of models. Artificial Analysis. https://artificialanalysis.ai/leaderboards/models Read 2026-09-25 to 2026-10-01.

Google (2026). Gemini 4 Argon. Google blog, September 30, 2026. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/

Franzen, C. (2026). Google unveils Gemini 4 Argon, retaking benchmark lead over OpenAI and Anthropic, but in limited release. VentureBeat (secondary transcription of Google’s benchmark table). https://venturebeat.com/technology/google-unveils-gemini-4-argon-retaking-benchmark-lead-over-openai-and-anthropic-but-in-limited-release

r/ClaudeCode (2026). Plot twist: Gemini 4 Argon tops Val AI benchmark on speed, cost and accuracy! Reddit thread, community reaction. https://www.reddit.com/r/ClaudeCode/comments/1wuft2u/plot_twist_gemini_4_argon_tops_val_ai_benchmark/

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← Back to Models & benchmarks