GPT-5.6 Pricing: $4/$20 and the Cost Per Task
GPT-5.6 Sol lists $4 input and $20 output per million tokens. GPT-6 Sol replaced it at half that on September 22. What each works out to per finished task.
GPT-5.6 Sol bills $4.00 per million input tokens and $20.00 per million output, with cached input at $0.40, a rate OpenAI labelled promotional through November 21, 2026 and then undercut with its own successor two months early. What it costs you is a different number: on those rates a single question runs about $0.018, and the twelve-turn agent loop that actually does the work runs about $0.83, or $3.31 when it takes four attempts to land.
What GPT-5.6 costs right now
All three tiers, per million tokens, from OpenAI’s API pricing documentation as verified on September 12, 2026:
| Tier | Input | Cached input | Output | Batch and Flex | Fast mode |
|---|---|---|---|---|---|
| GPT-5.6 Sol | $4.00 | $0.40 | $20.00 | $2.00 / $0.20 / $10.00 | $8.00 / $40.00 |
| GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | $1.00 / $0.10 / $6.00 | $4.00 / $24.00 |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | $0.10 / $0.01 / $0.60 | $0.40 / $2.40 |
None of those is the launch price. GPT-5.6 was previewed on June 26, 2026 to a gated partner set and became generally available on July 9 at $5.00 / $30.00 for Sol, $2.50 / $15.00 for Terra and $1.00 / $6.00 for Luna. Three cuts followed in six weeks: Terra down 20 percent and Luna down 80 percent on July 30, then Sol down to $4.00 / $20.00 on August 21.
The Sol number carried an asterisk the other two did not. OpenAI labelled the rate promotional and named November 21, 2026 as the date it was committed through, with $5.00 / $30.00 as the only published fallback. That question was settled early, and not by the calendar. On September 22, 2026 OpenAI shipped GPT-6 Sol at $2.00 / $10.00, exactly half the promotional rate, and GPT-6 Luna at $0.10 / $0.50. The GPT-5.6 tiers stay listed and callable for pinned workloads at the rates in the table above.
Sol is also no longer the flagship, and has not been since GPT-6 Astra arrived on September 3, 2026 at $10.00 / $50.00. Live per-token rates for every tracked model sit on the AI model tracker.
What one task actually costs on Sol
The rate card is four numbers. The bill is a different question, and it is the one worth modeling before committing a workload. Every figure in this section is arithmetic on the rates above, with the token assumptions stated, not a published OpenAI benchmark.
| Task | Modeled cost on GPT-5.6 Sol |
|---|---|
| Batch classification (per record, 700 in / 120 out, Batch rate) | $0.0026 |
| One-shot question (1,500 in / 600 out, uncached) | $0.018 |
| Codebase question (80K cached + 2K fresh in / 800 out) | $0.056 |
| Document summary (100K in / 1,200 out, uncached) | $0.42 |
| Agent loop, 12 turns (90% cache hit, 2.5K out per turn) | $0.83 |
| Same loop, 4 passes (three failures before it lands) | $3.31 |
The working, so the numbers can be checked or re-run against a different workload. A one-shot question is (1,500 x $4 + 600 x $20) / 1,000,000, or $0.018. A question against a cached 80,000-token codebase is (80,000 x $0.40 + 2,000 x $4 + 800 x $20) / 1,000,000, or $0.056: the cached context is the cheapest part of it. The twelve-turn loop assumes each turn carries 25,000 input tokens at a 90 percent cache hit rate and emits 2,500 tokens, so 12 x (22,500 x $0.40 + 2,500 x $4 + 2,500 x $20) / 1,000,000, or $0.828.
One split inside that last figure matters more than the total. Of the $0.83, output tokens account for $0.60, or 72 percent. Input is nearly free by comparison once caching is working. That is why loop count and verbosity move the bill and the input rate barely does, and it is why a cut to the output rate is worth roughly three times a cut to input on agentic work. Benchmark-measured rather than modeled numbers, on real coding tasks, are in the breakdown of what Sol costs per solved task on the coding leaderboard.
Why cost per task beats the sticker price
Here is the part most pricing pages miss. The per-token rate is the headline, but it is rarely the bill.
Picture two developers on the same model at the same $4 / $20. One asks a question, gets an answer, moves on: a few thousand tokens, a fraction of a cent. The other runs an autonomous coding agent that reads twelve files, plans, edits, runs tests, reads the failures, and tries again, four times over. Same rate. The second task costs roughly 180 times more, because agentic work multiplies tokens through loops, and output tokens, the expensive ones, stack up with every retry.
That is why a frontier price increase can matter less than it sounds, and why a price cut can matter less too. A model that costs 25 percent more per token but needs one fewer retry is cheaper in practice. The method for turning token rates into a real per-task figure is the same one used in this breakdown of what Claude Code actually costs per task, and it applies to any model: estimate tokens per task, multiply by loop count, weight output heavily. To run your own token counts rather than the assumptions above, the AI coding cost calculator takes them directly.
So the number to extract from a rate card is not $X per million. It is dollars per task you actually run, on your actual workload.
Where the rate stops being $4/$20
Three rules push a GPT-5.6 request above the headline rate, and two pull it below. All five come from the same pricing documentation, and the first one is the easiest to walk into by accident.
Sol carries a 1,050,000-token context window, but any request above 272,000 tokens bills the entire request at 2x input and 1.5x output, or $8.00 / $30.00. This is not a surcharge on the tokens past the threshold. It reprices the whole call, so crossing 272,000 tokens by one token costs the same as crossing it by three hundred thousand.
| Context size | As Sol bills it | If $4 / $20 held |
|---|---|---|
| 64K | $0.34 | $0.34 |
| 128K | $0.59 | $0.59 |
| 272K | $1.17 | $1.17 |
| 273K | $2.30 | $1.17 |
| 400K | $3.32 | $1.68 |
| 700K | $5.72 | $2.88 |
| 1.05M | $8.52 | $4.28 |
The other two multipliers are smaller but easier to trigger. Fast mode bills at exactly twice standard, so Sol Fast is $8.00 / $40.00. Cache writes bill at 1.25x the uncached input rate, or $5.00 per million on Sol, which is the price of putting a codebase into cache before the cheap $0.40 reads start paying it back. A cached context has to be read about four times before it comes out ahead of paying fresh input every turn.
Pulling the other way: Batch and Flex both bill at 50 percent of standard. On Sol that is $2.00 input, $0.20 cached and $10.00 output. Any workload that tolerates latency, and classification, extraction, evaluation and backfills usually do, is paying double by running on the standard tier. That halving is a larger saving than the entire August 21 price cut, and it is available on all three tiers.
What happened instead of November 21, 2026
This section used to model the risk that the $4.00 / $20.00 promotion lapsed and Sol reverted to $5.00 / $30.00, a 43 percent increase on the twelve-turn loop above. That is no longer the scenario to plan for.
OpenAI resolved it two months early and in the other direction. On September 22, 2026 it released GPT-6 Sol at $2.00 / $10.00 and GPT-6 Luna at $0.10 / $0.50, halving both GPT-5.6 tiers rather than letting the promotional window close. Run the same twelve-turn loop on GPT-6 Sol and it costs about $0.414 against the $0.828 modeled above, because every rate in the calculation halves at once.
The budgeting lesson survives the specific forecast, inverted. A rate with a stated expiry is not a durable basis for a plan, and that holds whichever way it moves: the risk here turned out to be over-budgeting against a reversion that never came, not under-budgeting against one that did. The durable practice is to re-read the rate card rather than extrapolate from a vendor’s stated intention in either direction. How Sol stacks up against Claude Opus 4.8 and Fable 5 at their own rates is the wider version of that question.
Frequently asked questions
- How much does GPT-5.6 cost?
- As of September 23, 2026, per million tokens: Sol at $4.00 input, $0.40 cached input and $20.00 output. Terra at $2.00, $0.20 and $12.00. Luna at $0.20, $0.02 and $1.20. The Sol rate was promotional with a stated end date of November 21, 2026, but OpenAI settled that early by releasing GPT-6 Sol at $2.00 and $10.00 on September 22. All three GPT-5.6 tiers remain listed and callable.
- Is GPT-5.6 still the flagship OpenAI model?
- No. GPT-6 Astra replaced it at the top of the lineup on September 3, 2026, at $10.00 input and $50.00 output per million tokens, and GPT-6 Sol replaced the Sol tier itself on September 22 at $2.00 and $10.00. GPT-5.6 remains fully available for pinned workloads, at twice what its GPT-6 counterpart now charges.
- What does one task actually cost on GPT-5.6 Sol?
- On the current rates, a one-shot question of 1,500 input and 600 output tokens costs about $0.018. A twelve-turn agent loop with a 90 percent cache hit rate costs about $0.83, and about $3.31 if it takes four attempts to land. Output tokens are roughly 72 percent of that agent bill, which is why loop count moves the total far more than the headline input rate does.
- Is GPT-5.6 cheaper than GPT-5.5?
- Yes. GPT-5.5 is still listed and unchanged at $5.00 input and $30.00 output, and Sol undercuts it at $4.00 and $20.00. The gap widened again on September 22, 2026, when GPT-6 Sol arrived at $2.00 and $10.00, a third of the GPT-5.5 rate.
- What happened to the GPT-5.6 promotional rate that was due to expire on November 21, 2026?
- OpenAI resolved it two months early and in the opposite direction from the risk. Rather than letting Sol revert to $5.00 and $30.00, it released GPT-6 Sol on September 22, 2026 at $2.00 and $10.00, exactly half the promotional rate. GPT-5.6 Sol remains listed at $4.00 and $20.00 for workloads pinned to it.
- Do Batch and Flex change the GPT-5.6 price?
- Both bill at 50 percent of standard. On Sol that is $2.00 input, $0.20 cached input and $10.00 output per million tokens. On Terra it is $1.00, $0.10 and $6.00, and on Luna $0.10, $0.01 and $0.60. Fast mode runs the opposite way at twice standard, so Sol in Fast mode is $8.00 and $40.00.
- Does a long prompt cost more on GPT-5.6?
- Yes. Sol carries a 1,050,000-token context window, but any request above 272,000 tokens bills the entire request at twice the input rate and one and a half times the output rate, or $8.00 and $30.00. A turn that would cost $1.68 at 400,000 tokens costs $3.32 instead. The threshold reprices the whole call, not just the tokens past it.
Sources
- OpenAI. (2026). API Pricing. Primary; the GPT-5.6 Sol, Terra and Luna rate card, the Batch, Flex and Fast mode multipliers, the cache write rate, and the GPT-6 Astra rates. Verified 2026-09-12. developers.openai.com/api/docs/pricing
- OpenAI. (2026). GPT-5.6 Sol model documentation. Primary; the promotional status and the November 21, 2026 end date, the 272,000-token long-context rule (2x input, 1.5x output), the 1,050,000-token context window and the February 16, 2026 knowledge cutoff. Verified 2026-09-12. developers.openai.com/api/docs/models/gpt-5.6-sol
- OpenAI. (2026). Previewing GPT-5.6 Sol. Primary; the June 26, 2026 gated preview of Sol, Terra and Luna and the launch tier pricing. Verified 2026-06-27. openai.com/index/previewing-gpt-5-6-sol
- OpenAI. (2026). API Pricing. Primary; GPT-5.5 still listed and unchanged at $5.00 / $30.00 with cached input at $0.50. Verified 2026-09-12. openai.com/api/pricing