GPT-6 Astra: Pricing, Benchmarks, Cost
GPT-6 Astra lists $10 and $50 per million tokens, 2.5x GPT-5.6 Sol. On Terminal-Bench it still solves a task for half of what Claude Fable 5.1 costs.
OpenAI released GPT-6 Astra on September 3, 2026 at $10 per million input tokens and $50 per million output. That is two and a half times what GPT-5.6 Sol charges, and it is the first increase in the frontier rate in years.
It is also, on the only public leaderboard that publishes what a run costs, the cheapest way to finish a hard task.
| Model | Accuracy | Cost per solved task | On the cost-efficiency frontier |
|---|---|---|---|
| Astra (max) | 58.18% | $17.02 | Yes |
| Fable 5.1 | 57.88% | $32.69 | No |
| Astra (high) | 57.88% | $11.88 | Yes |
| Astra (medium) | 54.24% | $10.70 | Yes |
| Opus 5 | 51.82% | $34.91 | No |
| Astra (low) | 50.61% | $9.33 | Yes |
| Fable 5 | 44.55% | $49.42 | No |
| GLM-5.3 | 41.82% | $19.77 | No |
| GPT-5.6 Sol | 37.27% | $20.66 | No |
| Opus 4.8 | 23.64% | $83.09 | No |
| GPT-5.6 Terra | 21.52% | $24.42 | No |
| Grok 4.6 | 20.3% | $53.61 | No |
| Gemini 3.8 Flash | 19.09% | $29.03 | No |
| GPT-5.6 Luna | 17.27% | $6.08 | Yes |
| Grok 4.5 | 12.42% | $51.08 | No |
| Sonnet 5 | 12.42% | $234.24 | No |
| Gemini 3.7 Flash | 11.21% | $34.10 | No |
How much does GPT-6 Astra cost?
The rate card, from the OpenAI API pricing documentation rather than the launch post, reads $10 input, $1 cached input, $12.50 cache writes and $50 output per million tokens. Batch and Flex run at half price.
Two surcharges sit on top of that, and both are easy to trip over.
| Line | Standard | Prompt over 272,000 tokens |
|---|---|---|
| Input | $10.00 | $20.00 |
| Cached input | $1.00 | $2.00 |
| Cache writes | $12.50 | $25.00 |
| Output | $50.00 | $75.00 |
The long-context tier applies to the whole request, not to the tokens above the threshold. A 273,000-token prompt is billed at double throughout. Separately, Fast mode is twice the standard rate, and OpenAI notes it is unavailable with EU data residency.
Against GPT-5.6 Sol at $4 and $20, Astra is a 2.5x rise on input and output alike. Two qualifications matter. The Sol rate is itself flagged promotional through at least November 21, 2026, so the comparison is against a discount that expires. And $10 and $50 is not a record: it is precisely what Claude Fable 5.1 started charging two days earlier.
Specifications: a 1,050,000-token context window (922,000 maximum input, 128,000 maximum output), a knowledge cutoff of April 30, 2026, five reasoning-effort levels from low to max, text and image in and text out. It serves Chat Completions, Responses and Batch, with no Realtime, fine-tuning or embeddings endpoint. The model reference carries the full list.
Why the higher rate produces a lower bill
Terminal-Bench 4.0, maintained by Stanford and the Laude Institute, is unusual in publishing the dollar cost of every run next to the score. Sixty-six tasks, five attempts each, 330 trials per entry, and a full-precision total spend.
Divide that spend by the trials a model actually solved and you get the number that decides procurement.
| Model | Agent | Effort | Accuracy | Solved | Total cost | Per solved task |
|---|---|---|---|---|---|---|
| GPT-6 Astra | Codex | max | 58.18% | 192/330 | $3,267.18 | $17.02 |
| Claude Fable 5.1 | Claude Code | max | 57.88% | 191/330 | $6,243.50 | $32.69 |
| GPT-6 Astra | Codex | high | 57.88% | 191/330 | $2,269.42 | $11.88 |
| Claude Opus 5 | Claude Code | max | 51.82% | 171/330 | $5,969.11 | $34.91 |
| GPT-5.6 Sol | Codex | max | 37.27% | 123/330 | $2,541.70 | $20.66 |
Astra at high effort scores 57.88%, the same figure Claude Fable 5.1 posts, and costs $11.88 against $32.69. It beats Claude Opus 5 by more than six accuracy points at roughly a third of the cost per solved task. Against its own predecessor it is 21 points more accurate and cheaper per finished task, despite charging 2.5 times more per token.
The reason is token volume. Astra burned 1.53 billion tokens across the benchmark at max effort. Fable 5.1 burned 2.75 billion and Opus 5 burned 6.53 billion.
| Scenario | Cached input | Uncached input | Output | Cache writes | Total |
|---|---|---|---|---|---|
| GPT-6 Astra (58.18%, 192 solved, 1.53B tokens) | $1,443 | $624 | $1,199 | $0 | $3,267 |
| Claude Fable 5.1 (57.88%, 191 solved, 2.75B tokens) | $617 | $2,160 | $3,156 | $310 | $6,244 |
| Claude Opus 5 (51.82%, 171 solved, 6.53B tokens) | $3,136 | $946 | $1,650 | $237 | $5,969 |
| GPT-5.6 Sol (37.27%, 123 solved, 4.41B tokens) | $1,727 | $280 | $469 | $66 | $2,542 |
That chart contains a check worth running. Astra consumed 1,443,351,702 cached input tokens, 62,437,628 uncached and 23,988,992 output. Priced at OpenAI’s published $1, $10 and $50, that comes to $3,267.18. The leaderboard publishes $3,267.18. The board and the rate card agree to the cent, which is the strongest evidence available that the cost column means what it appears to mean.
The comparison with Fable 5.1 is the one that should give buyers pause, because it runs against the obvious intuition. Anthropic cut the Fable 5.1 cache read to $0.25 per million tokens on September 1, a quarter of what Astra charges, and that cut is real: Fable 5.1 spent only $616.76 on cached input against Astra’s $1,443.35. It still finished the run at twice the cost, because it emitted 63.1 million output tokens to Astra’s 24.0 million, and output is where the expensive rate lives. A better cache rate does not rescue a model that talks more.
What the effort dial actually buys
Astra exposes five reasoning-effort levels, and the leaderboard ran all five.
| Effort | Accuracy | Solved | Total cost | Per solved task |
|---|---|---|---|---|
| max | 58.18% | 192/330 | $3,267.18 | $17.02 |
| xhigh | 57.88% | 191/330 | $2,350.51 | $12.31 |
| high | 57.88% | 191/330 | $2,269.42 | $11.88 |
| medium | 54.24% | 179/330 | $1,914.80 | $10.70 |
| low | 50.61% | 167/330 | $1,557.30 | $9.33 |
Max effort buys 0.30 accuracy points over high effort for 43% more money per solved task. The published confidence interval on the max run is plus or minus 2.8 points, so that 0.30 is not a measurable difference. On this benchmark, high is the setting to run and max is the setting to justify.
Note also that xhigh and high land on identical accuracy, with high the cheaper of the two. The dial is not monotonic in value.
Where Astra does not win
The cost-per-task picture is the flattering one. On the Artificial Analysis Intelligence Index, an independent composite of nine evaluations, Astra scores 61 at max effort on version 4.1.1. That places it behind Claude Fable 5.1 at 66, Claude Opus 5 at 63 and Meta Muse Spark 1.3 at 62, and exactly level with the GPT-5.6 Sol it replaces at 2.5 times the price per token.
The same pattern as Terminal-Bench holds underneath: Artificial Analysis reports the cost of running its whole index at $3,013.30 for Astra against $3,836.05 for Opus 5 and $8,523.16 for Fable 5.1. Lower score, materially lower cost to reach it.
Artificial Analysis also records a regression. In its Astra benchmarking writeup, the model gains roughly 80 Elo points on AA-Briefcase and six points on Humanity’s Last Exam, cuts its AA-Omniscience hallucination rate from 92% to 51%, and loses about 80 Elo points on GDPval-AA v2. Its Coding Agent Index of 67 trails Fable 5.1 in Claude Code at 70.
OpenAI’s own system card adds a caveat that no marketing page carries: Astra “has lower CoT monitorability than GPT-5.6 Sol across most CoT token lengths.” The same document reports Astra as the first model to reach the Critical cybersecurity capability level under OpenAI’s Preparedness Framework, and indirect prompt-injection robustness rising from 96.23% to 99.79%.
The benchmark claim that needs its harness named
OpenAI markets 99.9% on ARC-AGI-3. The ARC Prize Foundation writeup, published the same day, shows what that figure requires.
| ARC-AGI-3, semi-private set | Standard harness | OpenAI Provider Adapter |
|---|---|---|
| max | 62.7% | 98.6% |
| xhigh | 59.3% | 98.4% |
| high | 54.8% | 99.9% |
| medium | 38.6% | 98.4% |
| low | 17.5% | 98.0% |
| none | 35.2% | 96.7% |
The Provider Adapter is OpenAI’s own harness, and it preserves reasoning state between turns. Under it, every effort setting including “none” scores above 96%, which tells you the adapter rather than the model is carrying most of the result. On the standard stateless harness the same model runs from 17.5% to 62.7%.
Both numbers are real. Only one of them is a like-for-like comparison with how other models are scored, and it is the 62.7%. For context, GPT-5.6 Sol manages 7.78% on the same semi-private set, so the underlying jump is large without the adapter. The benchmark tracker here records the standard-harness figure.
One point in OpenAI’s favour on this: it undersold. The announcement claims 57.7% on Terminal-Bench 4.0 and the independent board records 58.18%.
Should you switch to GPT-6 Astra?
If the work is long-horizon agentic coding and you are already on Codex, the case is straightforward. Astra at high effort is cheaper per finished task than every model on the Terminal-Bench board that scores above 20%, and it is the most accurate entry on that board at max. Run high, not max.
If you are on Claude Code with Fable 5.1 or Opus 5, the switching cost is a harness migration and the measured advantage is partly a harness effect, so the honest answer is that the Terminal-Bench gap overstates what you would see. Fable 5.1 also still leads the Artificial Analysis index by five points.
If the work is short prompts, high output volume, or anything where a task finishes in one turn, Astra is simply expensive. The entire argument for it rests on burning fewer tokens across many turns. On a single call, $50 output is $50 output, and GPT-5.6 Luna at $0.20 and $1.20 does the cheap work.
If your prompts run long, watch the 272,000-token threshold. Crossing it doubles the bill for the whole request, and a 1,050,000-token context window makes that easy to do by accident.
Current rates for every tracked model are in the AI model tracker, and the modeled comparison across all of them is in cost per coding task. To price your own workload, use the AI coding cost calculator.
Frequently asked questions
- How much does GPT-6 Astra cost?
- GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens, with cached input at $1 and cache writes at $12.50. Prompts over 272,000 tokens are billed at double for the entire request, with output at $75 per million. Fast mode is twice the standard rate. Batch and Flex are half price. That is 2.5 times the $4 and $20 GPT-5.6 Sol charges, and identical to Claude Fable 5.1.
- When was GPT-6 Astra released?
- OpenAI released GPT-6 Astra on September 3, 2026, replacing GPT-5.6 Sol as its frontier flagship. It carries a 1,050,000-token context window, a 128,000-token maximum output, a knowledge cutoff of April 30, 2026, and five reasoning-effort levels from low to max.
- Is GPT-6 Astra better than Claude Fable 5.1?
- It depends on the measure. On Terminal-Bench 4.0 the two are statistically tied, 58.18 percent for Astra at max effort against 57.88 percent for Fable 5.1, but Astra solves a task for $17.02 against $32.69 because it burns 1.53 billion tokens to Fable 5.1 2.75 billion. On the Artificial Analysis Intelligence Index, Fable 5.1 leads 66 to 61. The Terminal-Bench runs also used different agent harnesses, Codex and Claude Code, so part of the gap belongs to the harness rather than the model.
- Did GPT-6 Astra really score 99.9 percent on ARC-AGI-3?
- Only with OpenAI own Provider Adapter harness, which preserves reasoning state between turns. ARC Prize published both tables on launch day: on the standard harness Astra scores 62.7 percent at max effort on the semi-private set, and 54.8 percent at high effort. Under the Provider Adapter every effort level from none to max scores between 96.7 and 99.9 percent, which indicates the adapter rather than the reasoning setting is producing the result.
- Which GPT-6 Astra reasoning effort should you use?
- High, on the current evidence. On Terminal-Bench 4.0 high effort scores 57.88 percent at $11.88 per solved task while max scores 58.18 percent at $17.02. The 0.30-point gain sits well inside the published plus or minus 2.8 point confidence interval, so max costs 43 percent more per solved task for no measurable accuracy benefit. Low effort still reaches 50.61 percent at $9.33.
- Why is GPT-6 Astra more expensive than GPT-5.6 Sol?
- OpenAI raised the frontier rate from $4 and $20 to $10 and $50 per million tokens. Note that the Sol figures are promotional pricing flagged as running through at least November 21, 2026, so the comparison is against a discounted rate. Despite the increase, Astra costs less per finished task on Terminal-Bench 4.0 than Sol does, $17.02 against $20.66, because it uses about a third of the tokens.
Sources
OpenAI (2026). API pricing. OpenAI developer documentation. https://developers.openai.com/api/docs/pricing Verified 2026-09-04.
OpenAI (2026). GPT-6 Astra. OpenAI developer documentation, model reference. https://developers.openai.com/api/docs/models/gpt-6-astra Verified 2026-09-04.
OpenAI (2026). GPT-6 Astra system card. OpenAI deployment safety documentation. https://deploymentsafety.openai.com/gpt-6-astra Verified 2026-09-04.
Terminal-Bench (2026). Terminal-Bench 4.0 leaderboard. Stanford University and the Laude Institute, independent benchmark. https://www.tbench.ai/leaderboard/terminal-bench/4.0 Verified 2026-09-04.
Artificial Analysis (2026). GPT-6 Astra. Artificial Analysis, independent benchmarking service, Intelligence Index v4.1.1. https://artificialanalysis.ai/models/gpt-6-astra Verified 2026-09-04.
Artificial Analysis (2026). Benchmarking GPT-6 Astra. Artificial Analysis, independent benchmarking service. https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra Verified 2026-09-04.
ARC Prize Foundation (2026). GPT-6 Astra on ARC-AGI-3. ARC Prize Foundation, independent benchmark. https://arcprize.org/blog/astra Verified 2026-09-04.
OpenAI’s launch announcement at openai.com/index/gpt-6-astra/ returned HTTP 403 on every attempt on 2026-09-04, so it was not read for this post. Vendor benchmark figures quoted above are attributed to OpenAI as relayed by secondary coverage, and every rate and specification comes from the OpenAI developer documentation instead.