Capital & Compute

Muse Spark 1.3 Max vs xhigh: Is Max Worth It?

Muse Spark 1.3 max reasoning opened to every developer on September 4 at the same price. It scores 3 points more for 17% more per task. When it pays.

· ai· pricing· benchmarks· By Capital & Compute
Bar chart comparing Muse Spark 1.3 at xhigh and max reasoning on five benchmarks.

Muse Spark 1.3 launched on September 2, 2026 with its strongest setting held back. Meta’s best benchmark numbers came from max reasoning, which it said was still finishing safety testing. Two days later, on September 4, max opened to every developer on the Meta Model API and in Muse Code, as reported by OrcaRouter and confirmed by the “now available” line on Meta’s release post.

It costs the same per token as every other setting. So the question is not what max costs. It is what max buys, and the short answer is: a lot on computer-use work, almost nothing on terminal coding, and three index points overall for about 17% more per task.

Muse Spark 1.3: xhigh against max reasoning on five benchmarksOSWorld 2.0: xhigh 57.2, max 66.9. JobBench: xhigh 61.2, max 64.9. DeepSearchQA: both 89.4. Terminal-Bench 2.1: xhigh 89.2, max 88.8. Artificial Analysis Intelligence Index: xhigh 45, max 48.xhighmaxOSWorld 2.057.266.9JobBench61.264.9DeepSearchQA89.489.4Terminal-Bench 2.189.288.8AA Intelligence Index4548
Muse Spark 1.3: xhigh against max reasoning on five benchmarks
Metricxhighmax
OSWorld 2.057.266.9
JobBench61.264.9
DeepSearchQA89.489.4
Terminal-Bench 2.189.288.8
AA Intelligence Index4548
Muse Spark 1.3 at xhigh (left) against max (right) on five boards. Max wins big on OSWorld 2.0, modestly on JobBench and the Intelligence Index, ties on DeepSearchQA and loses narrowly on Terminal-Bench 2.1. The first four boards are Meta's own figures as reported by VentureBeat; the Intelligence Index row is Artificial Analysis's independent reading.Source: Meta scorecard via VentureBeat (vendor-reported); Artificial Analysis Intelligence Index v4.3.2, read September 29, 2026

How much does Muse Spark 1.3 Max cost?

Muse Spark 1.3 Max costs the same per token as every other Muse Spark 1.3 setting: $1.25 per million input tokens, $0.15 per million cached input tokens and $4.25 per million output tokens. Meta’s pricing and rate limits page lists one standard rate for muse-spark-1.3 and no separate line for any reasoning setting. What changes is how many tokens a request uses.

Max is an effort value on the same model ID, not a separate SKU, and reasoning tokens bill as output, per OrcaRouter’s write-up of the launch. That makes max a volume decision rather than a price decision. The rest of the rate card also carries over:

  • Contributor tier. muse-spark-1.3-contributor lists at $0.10 input, $0.002 cached input and $0.20 output, 21.25 times cheaper on output, in exchange for Meta training on the prompts and completions sent to it.
  • Rate limits. The standard tier allows 3,000 requests and 4 million tokens per minute, counted per team rather than per API key. A max-effort agent that reasons for a long time reaches the token cap sooner.
  • Context. 1M tokens on both tiers.

The launch rate card, the contributor trade-off and how Muse Spark 1.3 compares with Gemini 3.8 Flash are covered in Gemini 3.8 Flash and Muse Spark 1.3 pricing.

What does max cost per task compared with xhigh?

On the independent Artificial Analysis index, read September 29, 2026, max costs $1.60 per task against $1.37 for xhigh, 17% more, and scores 48 against 45. It uses 170 million output tokens to run the index, against 140 million for xhigh, which is 21% more.

Artificial Analysis, index v4.3.2 xhigh max Difference
Intelligence Index 45 48 +3 points
Rank of 216 #28 #19 9 places
Cost per index task $1.37 $1.60 +17%
Output tokens to run the index 140M 170M +21%
Output speed (tokens per second) 178.1 180.9 about even

Two things stand out. First, the premium is small. Max spends more tokens, not several times more, and Meta’s own claim that 1.3 finishes coding work with about 20% fewer tool calls and 25% fewer tokens than 1.2 suggests the model was trained not to wander. By comparison, Anthropic’s Claude Sonnet 5.5 costs 2.8 times as much per index task at max as at xhigh on the same service ($7.60 against $2.74), covered in Claude Sonnet 5.5 pricing and benchmarks.

Second, output speed does not drop. Tokens come out at about the same rate. Max is slower only because it writes more of them before it answers, and Artificial Analysis records a 33.65-second wait to the first token at max. For an interactive tool, that half-minute pause is the real cost of max, more than the 23 cents.

What does max actually improve?

Meta’s scorecard, as reported by VentureBeat, shows where the extra reasoning goes. These are vendor-reported numbers, and Meta’s release post publishes the scorecard as an image rather than text, so they are cited here through that secondary report.

  • OSWorld 2.0, computer use: 57.2 to 66.9. A 9.7-point jump, and the only large one. Long computer-use tasks, where the model plans a sequence of clicks and checks its own state, reward extra thinking more than anything else on the card.
  • JobBench: 61.2 to 64.9. A modest 3.7 points on multi-step professional tasks.
  • GDPval-AA v2: 1,709 to 1,754 Elo. A 45-point rating gain on knowledge-work tasks, left out of the chart because an Elo rating does not share a scale with percentages.
  • DeepSearchQA: 89.4 and 89.4. No change. Search-heavy retrieval is limited by what the tools return, not by how long the model thinks.
  • Terminal-Bench 2.1: 89.2 to 88.8. Max loses by 0.4 points. Terminal coding at xhigh is already near the top of that board, and more thinking adds nothing.

That is a clear pattern. Max pays on long-horizon agent work, where the model has to plan and recover, and does nothing on work that is either tool-bound (search) or already saturated (terminal coding).

Is Muse Spark 1.3 Max good value against other models?

Per token it is one of the cheapest frontier-class models on the market. Per task, the independent index is less kind. Here is Muse Spark 1.3 max next to the models that score about the same, all read from the Artificial Analysis leaderboard on September 29, 2026:

Model and setting Output price per Mtok Intelligence Index Cost per index task
GPT-6 Sol, max $10 48 $1.06
Claude Sonnet 5.5, high $10 47 $1.08
Muse Spark 1.3, max $4.25 48 $1.60
Muse Spark 1.3, xhigh $4.25 45 $1.37

GPT-6 Sol at max matches Muse Spark 1.3 max’s score for about a third less per task, despite an output rate more than twice as high. The reason is tokens: Artificial Analysis records Sol using 77 million output tokens to run the whole index, against 170 million for Muse at max. A cheap rate multiplied by more than twice the tokens is not cheap.

This is the same lesson as the Sonnet 5.5 numbers, from the other direction. A per-token price tells you the cost of a unit of work only if every model uses the same number of units, and they do not. More on this in GPT-6 Sol and Luna pricing.

Muse Spark 1.3 does win on price once you use the contributor tier. Max adds about 30 million output tokens over xhigh on the full index. At the standard $4.25 that is the 23-cent premium per task; at the contributor tier’s $0.20, if the premium is almost all output tokens, as it should be when the tasks are the same, it falls to about one cent. If your prompts can go to Meta for training, max is effectively free.

When should you use max?

Use max when the work is long, agentic and not about terminal coding. Use xhigh otherwise.

  1. Computer-use and browser agents: max. The 9.7-point OSWorld gain is the biggest effect on the card, and the per-task premium is under a quarter.
  2. Terminal coding and repository work in Muse Code: xhigh. Max scored slightly lower on Terminal-Bench 2.1 and waits longer to start.
  3. Search and retrieval pipelines: xhigh. The DeepSearchQA result is identical.
  4. Interactive chat or autocomplete: neither. A 33-second wait to the first token at max is too long for a person watching; drop below xhigh.
  5. Batch jobs on the contributor tier: max, if the data-sharing terms are acceptable. The extra tokens cost almost nothing at $0.20 output.

Current rates for Muse Spark and every other tracked model are in the AI model tracker. The full September release record, including the date max went live, is in new AI models released in September 2026.

Frequently asked questions

When did Muse Spark 1.3 max reasoning become available?
Meta released Muse Spark 1.3 on September 2, 2026 with max reasoning held back for extra safety testing, and opened max to all developers on the Meta Model API and in Muse Code on September 4, 2026.
Does Muse Spark 1.3 Max cost more per token?
No. Max is an effort setting on the same muse-spark-1.3 model, billed at $1.25 per million input tokens, $0.15 cached and $4.25 output. It costs more per task only because it produces more reasoning tokens, which bill as output.
How much better is Muse Spark 1.3 max than xhigh?
On the Artificial Analysis Intelligence Index v4.3.2, max scores 48 and xhigh 45, for $1.60 and $1.37 per index task. On Meta's own scorecard the largest gain is OSWorld 2.0, 57.2 to 66.9; Terminal-Bench 2.1 is slightly lower at max.
Is Muse Spark 1.3 Max in Muse Code?
Yes. Since September 4, 2026, max is a selectable reasoning setting in Muse Code, Meta's coding agent, as well as on the Meta Model API.

Sources

Meta (2026). Introducing Muse Spark 1.3. Meta AI Research release post, including the statement that max reasoning is available on Muse Code and the Meta Model API. https://research.meta.ai/blog/introducing-muse-spark-1-3 Verified 2026-09-29.

Meta (2026). Pricing and rate limits. Meta Model API documentation: standard and contributor rates, rate limits. https://dev.meta.ai/docs/pricing-rate-limits/ Verified 2026-09-29.

Artificial Analysis (2026). Muse Spark 1.3 (max) and Muse Spark 1.3 (xhigh) model pages. Independent benchmarking service: Intelligence Index v4.3.2, cost per index task, output tokens, output speed and time to first token. https://artificialanalysis.ai/models/muse-spark-1-3 Read 2026-09-29.

Artificial Analysis (2026). Models leaderboard. Intelligence Index and cost per task for GPT-6 Sol and Claude Sonnet 5.5 by effort setting. https://artificialanalysis.ai/leaderboards/models Read 2026-09-29.

VentureBeat (2026). Meta says Muse Spark 1.3 has frontier performance, but its best results come from a model developers can’t broadly use yet. Secondary report of Meta’s max and xhigh scorecard. https://venturebeat.com/technology/meta-says-muse-spark-1-3-has-frontier-performance-but-its-best-results-come-from-a-model-developers-cant-broadly-use-yet Read 2026-09-29.

OrcaRouter (2026). Muse Spark 1.3 max reasoning goes live. Secondary report of the September 4 general availability and the effort parameter. https://www.orcarouter.ai/blog/muse-spark-1-3-max-reasoning-ga Read 2026-09-29.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← Back to Models & benchmarks