Qwen3.8 Max Benchmarks: The First Real Numbers
Artificial Analysis scored Qwen3.8 Max at 58 on its Intelligence Index. The token price fell 20 percent but the cost to run the same eval rose 64 percent.
By Capital & Compute
Qwen3.8 Max has an independent benchmark score for the first time. Artificial Analysis, an independent AI benchmarking service, scores it at 58 on its Intelligence Index, ninth of the 185 models it tracks, read on August 7, 2026. That is 11 points above Qwen3.7 Max and two points below Moonshot’s Kimi K3, both measured the same day.
The score is the headline. The economics underneath it are the story. Alibaba cut the list price 20 percent against its own previous flagship, yet the cost of running the full Intelligence Index on Qwen3.8 Max rose 64 percent, and measured output speed fell to a third of Qwen3.7 Max’s. Cheaper tokens, more expensive work.
What did Qwen3.8 Max actually score?
Artificial Analysis publishes one composite, the Intelligence Index, built from nine separate evaluations: GDPval-AA v2, tau-cubed Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. One number therefore blends reasoning, knowledge, science, agentic terminal work, and long-context retrieval. It is not a coding-specific score, and Artificial Analysis has not published a separate coding composite for Qwen3.8 Max, so anyone claiming a coding rank for this model is inferring one.
Here is the full measured picture as read on August 7, 2026, alongside the models it is most often compared against.
| Model | Intelligence Index | Input / output per Mtok | Blended price | Cost to run the Index | Output tokens/sec |
|---|---|---|---|---|---|
| Qwen3.8 Max | 58 | $2.00 / $6.00 | $1.18 | $1,741.41 | 67.6 |
| Qwen3.7 Max | 47 | $2.50 / $7.50 | $1.60 | $1,063.86 | 201.9 |
| Kimi K3 (max effort) | 60 | $3.00 / $15.00 | $2.31 | $2,425.11 | 38.5 |
| GLM-5.2 (max effort) | 53 | $1.35 / $4.29 | $0.86 | $714.52 | 111.7 |
| Claude Opus 5 (max) | 63 | $5.00 / $25.00 | $3.85 | $3,836.05 | 53.3 |
Blended price is Artificial Analysis’s own 7:2:1 weighting of cache-hit, input, and output tokens, not a rate anyone is billed directly. Cost to run the Index is what Artificial Analysis actually spent evaluating each model, which is the closest public proxy for what a fixed batch of real work costs on each one.
Two specification notes. Qwen3.8 Max carries a 1-million-token context window and a $0.25 per million cache-hit rate, an 88 percent discount on input. And Artificial Analysis publishes no reasoning-effort variant for it, unlike Kimi K3, GLM-5.2, and Claude Opus 5, which each appear at several effort levels. That matters when comparing: 58 is the only configuration published, while 60 for Kimi K3 is its max-effort setting specifically.
The score moved three times in four days
This is the part worth remembering the next time a benchmark number gets quoted at you.
July 19, 2026
Qwen3.8-Max-Preview debuts
Preview access via the Alibaba Token Plan, Qoder, and QoderWork. No price, no benchmarks, no open-weight date.
August 3, 2026
Pricing and an open-weights window announced
A $2/$6 per Mtok rate and a roughly one-week open-weights timeline for Qwen3.8-Max and Qwen3.8-27B, posted on X. Still no independent benchmark.
Around August 5, 2026
A score of 53 appears, then is removed
A first Intelligence Index result of 53 was posted and then pulled. Artificial Analysis attributed the removal to intermittent issues on the endpoint being tested, as reported by OfficeChai.
August 6, 2026
Press reports 56
The Decoder and OfficeChai both report 56, level with Claude Opus 4.8 and one point behind Kimi K3 at 57.
August 7, 2026
The published score reads 58
Artificial Analysis model page and leaderboard both read 58, ninth of 185. Kimi K3 now reads 60 rather than 57, so the gap widened rather than closed.
Note what happened to the comparison, not just the score. On August 6 the press framing was “level with Opus 4.8, one point behind Kimi K3.” By August 7 the numbers underneath that sentence had all moved: Qwen3.8 Max to 58, Kimi K3 to 60, and the whole index up by one to three points across every model checked. The relative claim reversed direction while the story about it stayed online.
That is not evidence of anything dishonest. Recalibration and endpoint re-testing are what a serious evaluator does. It does mean a specific discipline is required, and it is one this site now enforces in its own AI model value leaderboard: index scores read on different dates cannot be compared with each other. Every Artificial Analysis figure on that board was re-read on August 7 for exactly this reason, which moved Kimi K3 from 57 to 60 and Claude Opus 5 from 61 to 63. Had that not happened, adding Qwen3.8 Max at 58 would have made the board claim Qwen beats Kimi, which is the opposite of what a same-day read shows. For more on why benchmark numbers behave this way, see whether AI benchmarks are reliable.
Cheaper tokens, more expensive work
Alibaba cut the list rate from $2.50/$7.50 to $2.00/$6.00, a clean 20 percent off both sides. On Artificial Analysis’s blended measure the drop is larger, 26 percent, from $1.60 to $1.18.
Then the same organization ran the same nine-evaluation suite on both models. Qwen3.7 Max cost $1,063.86 to evaluate. Qwen3.8 Max cost $1,741.41. The bill went up 64 percent while the unit price went down 26 percent, which is only possible if the model emits far more tokens to finish the same work. Artificial Analysis already flagged Qwen3.7 Max as unusually verbose, noting it generated 100 million tokens on the Index against a 70 million median. Qwen3.8 Max is more verbose still.
Dividing one figure by the other gives a rough efficiency number: index points per thousand dollars of measured evaluation work. That calculation is this site’s, not Artificial Analysis’s, and it is arithmetic on the two columns above rather than a published metric.
| Model | Index points | Eval cost | Points per $1,000 of eval work |
|---|---|---|---|
| GLM-5.2 | 53 | $714.52 | 74.2 |
| Qwen3.7 Max | 47 | $1,063.86 | 44.2 |
| Qwen3.8 Max | 58 | $1,741.41 | 33.3 |
| Kimi K3 | 60 | $2,425.11 | 24.7 |
| Claude Opus 5 | 63 | $3,836.05 | 16.4 |
Qwen3.8 Max is less efficient than the model it replaces. It buys 11 more index points, and it buys them at a worse rate per dollar of work than Qwen3.7 Max managed. That is a legitimate trade, and for hard tasks that Qwen3.7 Max simply fails it is the right one, but it is not the trade the price cut advertises. The cost-per-task framing this site applies to coding agents exists precisely because per-token rates hide this.
| Item | Cost to run the Intelligence Index (USD) | Artificial Analysis Intelligence Index |
|---|---|---|
| GLM-5.2 | $0.7k | 53 |
| Qwen3.7 Max | $1.1k | 47 |
| Qwen3.8 Max | $1.7k | 58 |
| Kimi K3 | $2.4k | 60 |
| Claude Opus 5 | $3.8k | 63 |
The frontier reading is fair to Alibaba: nothing on this chart delivers 58 points for less than Qwen3.8 Max does. Kimi K3 charges 39 percent more evaluation spend for two more points, and Claude Opus 5 charges 2.2 times as much for five more. What Qwen3.8 Max does not do is beat GLM-5.2 on efficiency, and Zhipu’s model is open weight under an MIT license while Alibaba’s is not yet.
How fast is it?
Speed is where the regression is starkest, and it is the number least covered in the launch reporting.
| Item | Value |
|---|---|
| Qwen3.7 Max | 201.9 tok/s |
| GLM-5.2 | 111.7 tok/s |
| Qwen3.8 Max | 67.6 tok/s |
| Claude Opus 5 | 53.3 tok/s |
| Kimi K3 | 38.5 tok/s |
Time to first token is 2.60 seconds. Combine that with three times slower generation and a higher token count per task, and the wall-clock effect on an agentic loop compounds: an agent that makes forty model calls to finish a job waits far longer on Qwen3.8 Max than on Qwen3.7 Max, even before the retry behavior of a longer-running loop is considered. For interactive coding work that latency is felt directly. Qwen3.8 Max is still faster per token than Kimi K3 and Claude Opus 5, so it is mid-pack rather than slow in absolute terms. The comparison that stings is against its own predecessor.
Are the Qwen3.8 open weights out?
No. Not as of August 7, 2026.
The August 3 announcement promised weights for both Qwen3.8-Max and a smaller Qwen3.8-27B within about a week, which points at the week of August 10. The Qwen organization on Hugging Face has no Qwen3.8 repository of any kind, and its most recently updated models are unrelated speech and alignment checkpoints. Artificial Analysis classifies Qwen3.8 Max as proprietary. No license has been named for either model.
Alibaba’s own Model Studio pricing documentation still does not list Qwen3.8-Max either, and that page shows a last-updated date of July 15, 2026. So the $2/$6 rate remains sourced to the Qwen team’s X post rather than a published rate card, though the independent measurement adds real weight: Artificial Analysis observed exactly that rate on a live Alibaba Cloud endpoint. The best open-weight models guide counts a model as open only once files and a license exist, and Qwen3.8 has not crossed that line.
This matters for the comparison above. Kimi K3 is now listed as open weights, and GLM-5.2 is MIT licensed. Among the three Chinese frontier models on that chart, the one with the promised-but-absent weights is the Alibaba one.
What Alibaba claimed versus what was measured
The August 3 announcement made four capability claims: autonomous coding sessions running 10 or more days, production-quality deliverables spanning hundreds of professions, more than 500 turns of closed-loop chip-design optimization, and a 365-day e-commerce strategy run. None has been independently reproduced. There is no task specification, no scoring method, and no third-party evaluator for any of the four, and the linked demonstration repository sits under a qwen-code-dev-bot account rather than the official QwenLM organization.
The architecture figures are in a similar state. The 2.4-trillion total parameter count comes from Alibaba. An active count of 95 billion parameters per forward pass has been reported by both The Decoder and OfficeChai, but no model card confirms it, so it is secondary reporting of a vendor figure rather than documentation.
What the Intelligence Index result does establish is narrower and more useful than any of that: on a nine-evaluation suite run by someone with no stake in the outcome, Qwen3.8 Max is a genuine step up from Qwen3.7 Max and a genuine step below the current frontier leaders.
Should you use Qwen3.8 Max?
Three cases, based on the measured numbers rather than the announcement.
Worth switching to from Qwen3.7 Max if capability is the binding constraint. Eleven index points is a large generational gain, and if the tasks that matter are ones Qwen3.7 Max fails outright, the higher token consumption is a price worth paying. Budget for more than a 20 percent price cut suggests: on the one fixed workload with public numbers, total spend rose 64 percent.
Not worth switching to if throughput or unit economics is the binding constraint. A three-times drop in output speed and worse points-per-dollar-of-work than the predecessor are both real regressions. Latency-sensitive and high-volume workloads should stay put or look at GLM-5.2, which does more per dollar of measured work than anything else in this group.
Do not plan an open-weight deployment yet. The weights were promised for the week of August 10 with no license named. Alibaba has a good track record of shipping downloadable models, so the promise is credible, but a deployment plan needs files.
For a same-basis comparison across every model tracked here, the model comparison tool and cost-per-task calculator now include Qwen3.8 Max at its announced rate, and the model tracker carries it at preview status with the open-weights position noted. The AI benchmark directory explains what each of the nine evaluations inside the Intelligence Index actually measures.
Frequently asked questions
- What is Qwen3.8 Max scored on the Artificial Analysis Intelligence Index?
- Artificial Analysis scores Qwen3.8 Max at 58 on its Intelligence Index, ninth of the 185 models it tracks, read on August 7, 2026. That is 11 points above Qwen3.7 Max at 47 and two points below Kimi K3 at 60 in its max-effort configuration.
- Is Qwen3.8 Max better than Kimi K3?
- Not on the Intelligence Index. Kimi K3 scores 60 at max effort against 58 for Qwen3.8 Max on the same day. Qwen3.8 Max is the cheaper of the two to run: Artificial Analysis spent $1,741.41 evaluating it against $2,425.11 for Kimi K3. Kimi K3 is also open weight, and Qwen3.8 Max is not yet.
- Why did the Qwen3.8 Max benchmark score change?
- A first result of 53 was published and then removed, which Artificial Analysis attributed to intermittent issues on the endpoint being tested, as reported by OfficeChai. Press coverage on August 6 reported 56. The published figure read 58 on August 7. Scores across the whole index rose by one to three points over the same window, so comparisons drawn between figures read on different dates do not hold.
- Is Qwen3.8 Max cheaper than Qwen3.7 Max?
- Per token, yes: $2.00 and $6.00 per million input and output tokens against $2.50 and $7.50, a 20 percent cut on both. Per unit of finished work, no. Running the full Intelligence Index cost $1,741.41 on Qwen3.8 Max against $1,063.86 on Qwen3.7 Max, a 64 percent increase, because the newer model emits many more tokens per task.
- Are Qwen3.8 open weights available?
- No. As of August 7, 2026 there is no Qwen3.8 repository on the Qwen Hugging Face organization, no license has been named, and Artificial Analysis classifies the model as proprietary. Alibaba promised weights for Qwen3.8-Max and Qwen3.8-27B within about a week of the August 3 announcement.
- How fast is Qwen3.8 Max?
- Artificial Analysis measures 67.6 output tokens per second with a 2.60 second time to first token. That is roughly a third of the 201.9 tokens per second it measures for Qwen3.7 Max, though still faster than Kimi K3 at 38.5 and Claude Opus 5 at 53.3.
Sources
- Artificial Analysis. (2026). Qwen3.8 Max: Intelligence, Performance and Price Analysis [independent benchmark; Intelligence Index 58 and rank 9 of 185, $2/$6 per Mtok with $0.25 cache hits, blended $1.18, 67.6 output tokens per second, 2.60s time to first token, 1M context, proprietary status, $1,741.41 cost to run the Intelligence Index, and the nine evaluations composing the index]. Verified August 7, 2026.
- Artificial Analysis. (2026). Qwen3.7 Max: Intelligence, Performance and Price Analysis [independent benchmark; Intelligence Index 47, blended $1.60, 201.9 output tokens per second, $1,063.86 cost to run the Intelligence Index, and the note that the model generated 100M tokens against a 70M median]. Verified August 7, 2026.
- Artificial Analysis. (2026). Kimi K3: Intelligence, Performance and Price Analysis [independent benchmark; Intelligence Index 60 at max effort, blended $2.31, 38.5 output tokens per second, $2,425.11 cost to run the Intelligence Index, open-weight status]. Verified August 7, 2026.
- Artificial Analysis. (2026). GLM-5.2: Intelligence, Performance and Price Analysis [independent benchmark; Intelligence Index 53 at max effort, $1.35/$4.29 per Mtok, blended $0.86, 111.7 output tokens per second, $714.52 cost to run the Intelligence Index, MIT license]. Verified August 7, 2026.
- Artificial Analysis. (2026). Claude Opus 5: Intelligence, Performance and Price Analysis [independent benchmark; Intelligence Index 63 at max effort, blended $3.85, 53.3 output tokens per second, $3,836.05 cost to run the Intelligence Index]. Verified August 7, 2026.
- Artificial Analysis. (2026). LLM leaderboard [independent benchmark; the cross-model Intelligence Index values and effort-variant labels used for the drift comparison, including Claude Fable 5 at 62 and GPT-5.6 Sol at 61]. Verified August 7, 2026.
- The Decoder. (2026). Qwen3.8 Max catches Claude Opus 4.8 but Kimi K3 still scores higher for 25 percent less [secondary reporting; the August 6 score of 56, the per-task index cost figures that do not reconcile with Artificial Analysis totals, and the reported 95B active parameter count]. Published August 6, 2026.
- OfficeChai. (2026). Qwen 3.8 Max scores 56 on Artificial Analysis Intelligence Index [secondary reporting; the August 6 score of 56 and the quoted Artificial Analysis explanation for the withdrawn 53, attributed as reported rather than read from Artificial Analysis directly]. Published August 6, 2026.
- Qwen Team. (2026). Qwen3.8-Max announcement [official X post; the $2/$6 rate, the open-weights timeline, the four long-horizon capability claims, and the linked demonstration repository]. Published August 3, 2026.
- Alibaba Cloud. Model Studio model pricing [official pricing documentation; used for the Qwen3.7 Max list rate and to confirm that no Qwen3.8-Max row exists, page last updated July 15, 2026]. Verified August 7, 2026.
- Hugging Face. Qwen organization [official model repository; checked to confirm no Qwen3.8-Max or Qwen3.8-27B weights are published]. Verified August 7, 2026.
- Capital & Compute. Qwen3.8 Max Preview: Release, Access, and Open Weights [this site’s coverage of the July 19, 2026 preview and the original vendor claim].