# Claude Haiku 5.5: Pricing, Benchmarks, Cost

> Claude Haiku 5.5 costs $0.10/$0.50 per million tokens for prompts under 100K, 90% below Haiku 4.5, and scores 43 on Artificial Analysis at $0.21 a task.

Source: https://capitalandcompute.net/blog/claude-haiku-5-5-pricing-benchmarks/
Published: 2026-10-08
Updated: 2026-10-08

Anthropic released Claude Haiku 5.5 on October 7, 2026 at $0.10 per million input tokens and $0.50 per million output, for prompts up to 100,000 tokens. That is a tenth of what Haiku 4.5 charged, and exactly what OpenAI charges for GPT-6 Luna.

The verdict: pick Haiku 5.5 when quality matters and GPT-6 Luna when only the bill does, and keep Haiku prompts under 100,000 tokens, because past that line every token costs five times as much. On the Artificial Analysis Intelligence Index, Haiku 5.5 at max effort scores 43 for $0.21 a task, five points above Luna’s 38 at three times Luna’s $0.07\. At high effort Haiku scores 38 for $0.08, a tie.

| Price per million tokens | Prompt up to 100K tokens | Prompt over 100K tokens |
| ------------------------ | ------------------------ | ----------------------- |
| Input                    | $0.10                    | $0.50                   |
| Output                   | $0.50                    | $2.50                   |
| Cache read               | $0.01                    | $0.05                   |
| Cache write (5 minutes)  | $0.125                   | $0.625                  |
| Batch input / output     | $0.05 / $0.25            | $0.25 / $1.25           |

**Where you can use it:** the Claude API as `claude-haiku-5-5`, Amazon Bedrock, Google Cloud and Microsoft Foundry, Claude Code, and the Claude apps on every plan from Free to Enterprise. Rates are from [Anthropic’s pricing documentation](https://platform.claude.com/docs/en/about-claude/pricing) and availability from the [Claude Haiku product page](https://www.anthropic.com/claude/haiku), both read October 8, 2026.

## How much does Claude Haiku 5.5 cost?

Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens when the prompt is 100,000 tokens or shorter, and $0.50 and $2.50 when it is longer. Cache reads are $0.01 and $0.05, and batch requests are half price. Anthropic’s [launch announcement](https://www.anthropic.com/claude-haiku-5-5) puts the saving against Haiku 4.5 at 90% below the 100,000-token line and 50% above it.

That is the rate card. What a job costs depends on how many tokens the model spends to finish it, and Haiku 5.5 is the first Haiku with an effort setting. The [models overview](https://platform.claude.com/docs/en/about-claude/models/overview) lists its default as medium, with a 1M-token context window, 128K maximum output and a June 2026 knowledge cutoff.

Artificial Analysis, an independent benchmarking service, ran its Intelligence Index at four of those settings. Each step up the dial adds two to four points and costs 1.6 to 1.7 times as much as the step below:

| Effort (Artificial Analysis) | Intelligence Index | Cost per index task |
| ---------------------------- | ------------------ | ------------------- |
| Medium (the API default)     | 34                 | $0.047              |
| High                         | 38                 | $0.079              |
| Xhigh                        | 41                 | $0.124              |
| Max                          | 43                 | $0.213              |

Source: [Artificial Analysis, Claude Haiku 5.5](https://artificialanalysis.ai/models/claude-haiku-5-5), Intelligence Index v4.3.2, read October 8, 2026\. No low-effort run was published.

Max is 4.5 times the cost of medium for nine more points. For classification and routing, where the answer is short and the question is easy, medium is the setting to start from. Extraction from messy documents and anything with several tool calls is where high or xhigh pays.

Anthropic’s own estimate is that work costs “around 75% less to run” than on Haiku 4.5\. The footnote explains why it is less than 90%: about one Haiku 4.5 request in ten was over 100,000 tokens, and Haiku 5.5’s new tokenizer, similar to Sonnet 5.5’s and Opus 5.5’s, uses slightly more tokens for the same text.

## What does the 100,000-token tier mean for your bill?

It means the price of a request is set by its prompt length, and the change is not gradual. Anthropic’s pricing page states that “a prompt of over 100,000 tokens pays higher prices,” and the higher price applies to the whole request, not the tokens past the line. A 101,000-token prompt bills every token at the over-100K rate.

__Claude Haiku 5.5 vs GPT-6 Luna: what one request costs as the prompt grows__
| Context size | Claude Haiku 5.5 | GPT-6 Luna |
| ------------ | ---------------- | ---------- |
| 25K          | $0.00            | $0.00      |
| 50K          | $0.01            | $0.01      |
| 100K         | $0.01            | $0.01      |
| 101K         | $0.06            | $0.01      |
| 200K         | $0.10            | $0.02      |
| 272K         | $0.14            | $0.03      |
| 273K         | $0.14            | $0.06      |
| 500K         | $0.26            | $0.10      |
| 900K         | $0.46            | $0.18      |

Modeled cost of one request with 2,000 output tokens, uncached, as the prompt grows. Below 100,000 tokens Claude Haiku 5.5 and GPT-6 Luna bill identically. At 101,000 tokens the Haiku request jumps from $0.011 to $0.056, five times as much. Luna holds its rate until 272,000 tokens, where its own surcharge (2x input, 1.5x output) applies, and it stays below Haiku all the way out to its 922,000-token input limit.Source: [Modeled from Anthropic's Claude pricing documentation and OpenAI's API pricing documentation, read October 8, 2026](https://platform.claude.com/docs/en/about-claude/pricing)

Two practical consequences follow.

**Short-prompt work gets the headline price.** Classification, routing, summarizing a support ticket and pulling fields out of an invoice all sit far below 100,000 tokens. Anthropic says about 90% of the requests its previous Haiku received did too. For that traffic, $0.10 and $0.50 is the real rate.

**Long-context work does not.** Feed Haiku 5.5 a 300,000-token codebase or a stack of filings and you pay $0.50 and $2.50, five times the headline and half of Haiku 4.5’s old flat rate. It is still cheap in absolute terms. But the 1M-token window comes with a price that changes at 10% of the way in, and an agent that grows its context turn by turn will cross the line without anyone deciding to. If you run Haiku as a subagent, keeping each call under 100,000 tokens, by chunking documents or compacting history, is the single biggest cost lever you have.

## Claude Haiku 5.5 vs GPT-6 Luna: which is cheaper?

Per token, neither: both list at $0.10 input, $0.50 output and a $0.01 cache read below 100,000 tokens, according to [OpenAI’s API pricing](https://developers.openai.com/api/docs/pricing) and Anthropic’s. Per task, Luna is cheaper at the top of the dial and level lower down.

* **At each model’s max effort**, Haiku 5.5 scores 43 for $0.21 a task and GPT-6 Luna scores 38 for $0.07\. Haiku is five points better and three times the cost.
* **At the same score**, Haiku 5.5 at high effort reaches 38 for $0.079, against Luna’s 38 for $0.07\. The difference is about a cent a task.
* **On long prompts**, Luna wins outright. Its surcharge starts at 272,000 tokens rather than 100,000 and is smaller (2x input and 1.5x output, against Haiku’s 5x on both), as the chart above shows.

Anthropic’s launch table compares the two directly, and the gap is widest on agent work. These are vendor-reported numbers from Anthropic’s announcement, run on settings Anthropic chose, and none has been reproduced independently:

| Benchmark (vendor-reported)  | Haiku 5.5 | GPT-6 Luna | Haiku 4.5 | Sonnet 5.5 (reference) |
| ---------------------------- | --------- | ---------- | --------- | ---------------------- |
| Terminal-Bench 4.0           | 39.2%     | 16.4%      | 0.0%      | 70.6%                  |
| OSWorld 2.1 (offline subset) | 72.4%     | 48.9%      | 15.7%     | 83.9%                  |
| FrontierCode 1.1 (Main)      | 46.4%     | 42.4%      | not given | 52.1% (xhigh)          |
| Chartography (no tools)      | 46.4%     | 29.1%      | 6.4%      | 61.6%                  |
| GDPval-AA v2.1 (Elo)         | 1620      | 1437       | 735       | 1840                   |
| AA-Briefcase v1.1 (Elo)      | 1578      | 1336       | 614       | 1824                   |
| Humanity’s Last Exam, tools  | 57.4%     | not given  | 18.7%     | 64.5%                  |

Terminal-Bench and OSWorld are the two to read. Both measure an agent finishing a multi-step job, at a terminal and on a desktop, and Haiku 5.5’s lead there is 23 to 24 points. On the [independent index](https://artificialanalysis.ai/models/gpt-6-luna) the gap is five. Which number matters depends on whether your small model is answering questions or doing work. The full GPT-6 Luna picture, including its own effort settings, is in [GPT-6 Sol and Luna pricing and benchmarks](/blog/gpt-6-sol-luna-pricing-benchmarks/).

## Should you switch from Claude Haiku 4.5?

Yes, for almost any workload. Haiku 4.5 costs $1 input and $5 output on every request, so Haiku 5.5 is cheaper on both sides of the 100,000-token line: 90% cheaper below it, 50% cheaper above it.

The quality gap is larger than the price gap. On Anthropic’s own table Haiku 4.5 scores 0.0% on Terminal-Bench 4.0 and 15.7% on OSWorld 2.1, against 39.2% and 72.4% for Haiku 5.5\. Anthropic’s announcement quotes customers who tested it (vendor-published testimonials, not independent evaluations): Box says Haiku 5.5 scored 11 points higher than Haiku 4.5 at about half the latency, and a document question-answering team reports 0.84 against 0.76 across 400 queries.

Two things to check before you move traffic:

1. **Token counts.** The new tokenizer produces slightly more tokens for the same text. Re-measure the prompt sizes you rely on, especially any close to 100,000 tokens, because a prompt that fit under the line on Haiku 4.5’s tokenizer may not on Haiku 5.5’s.
2. **Cyber refusals.** Anthropic says Haiku 5.5’s cybersecurity safeguards are more restrictive than Haiku 4.5’s, though looser than Sonnet 5.5’s. They block penetration testing. A security pipeline built on Haiku 4.5 should be re-tested.

The models overview says Haiku 5.5 will stay available until at least October 7, 2027.

## Does Haiku 5.5 replace Sonnet 5.5?

No, and Anthropic says so: “Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks.” Its own numbers agree. Sonnet 5.5 leads Haiku 5.5 by 31 points on Terminal-Bench 4.0 and 13 on the Artificial Analysis index (56 against 43).

What changed is the price of the comparison. In the same announcement Anthropic halved Sonnet 5.5’s cache reads from $0.20 to $0.10 per million tokens, which it says means Sonnet 5.5 “now runs around 20% cheaper on most agentic work.” Artificial Analysis now puts Sonnet 5.5 at max effort at $5.46 per index task, down from the $7.60 it measured when [Sonnet 5.5 launched](/blog/claude-sonnet-5-5-pricing-benchmarks/), which now puts it below [Opus 5.5’s $5.98](/blog/claude-opus-5-5-pricing-benchmarks/).

Haiku 5.5 at max is still about 26 times cheaper per index task than Sonnet 5.5 at max. Anthropic’s intended pattern is to use both: a larger model plans the work and hands the narrow pieces (summaries, lookups, compaction) to Haiku 5.5 subagents. At these prices that split is worth building even for a small team.

## Which small model should you use?

* **High-volume classification, routing and extraction under 100K tokens:** Claude Haiku 5.5 at medium or high effort. At high it matches GPT-6 Luna’s best index score for about the same cost per task, and it is far stronger on agent benchmarks.
* **The cheapest possible bill, quality second:** GPT-6 Luna at $0.07 a task at max effort, cheaper than any Haiku 5.5 setting that scores more than 34.
* **Long documents or growing agent contexts over 100K tokens:** GPT-6 Luna, which keeps the base rate to 272,000 tokens. Haiku 5.5 costs 2.5 to 5 times as much in that range.
* **Near-frontier quality on a budget:** Haiku 5.5 at max. It scores 43, one point below Kimi K3 (44) and two below GLM-5.3 (45), at $0.21 a task against their $2.00 and $2.01 on the same index. That is roughly a tenth of the cost for one or two points.
* **Hard agentic coding:** Sonnet 5.5 or Opus 5.5, not a Haiku.

The [AI coding cost calculator](/ai-coding-cost-calculator/) runs any of these rate cards against your own token mix, and the [AI model tracker](/ai-models/) has every current rate in one table.

## How to get started with Claude Haiku 5.5

1. **Switch the model ID.** Call `claude-haiku-5-5` on the Claude API. The same ID works on Google Cloud and Microsoft Foundry; Bedrock uses `anthropic.claude-haiku-5-5`.
2. **Pick the effort setting.** The default is medium. Run your own eval set at medium and high before you pay for xhigh or max, since each step roughly multiplies the cost per task by 1.6 to 1.7.
3. **Measure prompt length.** Log the token count of each request for a day. Anything regularly over 100,000 tokens is billed at five times the headline rate and is a candidate for chunking or compaction.
4. **Cache and batch.** Cache reads are a tenth of the input price on either side of the line, and the Batch API halves everything for work that can wait.
5. **Check API credits if you subscribe.** Anthropic is adding a monthly API credit to its Max and Team plans: $100 on Max 5x, $200 on Max 20x, and for Team $20 per Standard seat and $100 per Premium seat, pooled and capped at $500 a month. Pro gets none. Per the [Claude Help Center article on monthly API credits](https://support.claude.com/en/articles/17154008), the credits cover the Claude API, Batches, Managed Agents and the Agent SDK, not interactive Claude Code. They expire each billing cycle, need seven days on the plan before you can claim them, and do not work on Bedrock, Vertex or Foundry. Plan prices are compared in the [AI coding plan pricing guide](/ai-pricing/).

Haiku 5.5 is one of several October releases. The rest of the month is in [new AI models released in October 2026](/blog/new-ai-models-october-2026/).

## Frequently asked questions

When was Claude Haiku 5.5 released?

Anthropic released Claude Haiku 5.5 on October 7, 2026, the third Claude 5.5 model after Opus 5.5 and Sonnet 5.5\. The API model ID is claude-haiku-5-5.

How much does Claude Haiku 5.5 cost?

For prompts up to 100,000 tokens, $0.10 per million input tokens and $0.50 per million output, with cache reads at $0.01\. For longer prompts the whole request bills at $0.50 input, $2.50 output and $0.05 cache read. Batch requests are half price.

Is Claude Haiku 5.5 better than GPT-6 Luna?

On quality, yes: 43 against 38 on the Artificial Analysis Intelligence Index at each model's max effort, and much higher on Anthropic's agent benchmarks. On cost per task, Luna is cheaper at max ($0.07 against $0.21), the two tie at a score of 38, and Luna is far cheaper on prompts over 100,000 tokens.

Is Claude Haiku 5.5 cheaper than Haiku 4.5?

Yes. It is 90% cheaper per token on prompts up to 100,000 tokens and 50% cheaper on longer ones. Anthropic estimates real workloads cost about 75% less once its slightly larger tokenizer is counted.

What is the Claude Haiku 5.5 context window?

One million tokens, with up to 128,000 output tokens. Unlike Sonnet 5.5 and Opus 5.5, the price is not flat across the window: prompts over 100,000 tokens pay five times the base rate.

## Sources

Anthropic (2026). _Introducing Claude Haiku 5.5_. Anthropic announcement: release, vendor-reported benchmark table against Haiku 4.5, GPT-6 Luna and Sonnet 5.5, the 75% cost estimate and its footnote, safeguards, customer testimonials and the Sonnet 5.5 cache-read cut. <https://www.anthropic.com/claude-haiku-5-5> Verified 2026-10-08.

Anthropic (2026). _Pricing_. Claude Platform documentation: Haiku 5.5 rates by prompt length, cache, batch and long-context rules. <https://platform.claude.com/docs/en/about-claude/pricing> Verified 2026-10-08.

Anthropic (2026). _Models overview_. Claude Platform documentation: model IDs, default effort, context window, maximum output, knowledge cutoff and retirement dates. <https://platform.claude.com/docs/en/about-claude/models/overview> Verified 2026-10-08.

Anthropic (2026). _Claude Haiku_. Anthropic product page: availability by plan and platform. <https://www.anthropic.com/claude/haiku> Verified 2026-10-08.

Anthropic (2026). _Monthly API credits for Max and Team plans_. Claude Help Center article: credit amounts, eligible products and expiry. <https://support.claude.com/en/articles/17154008> Verified 2026-10-08.

Artificial Analysis (2026). _Claude Haiku 5.5_ model page. Independent benchmarking service, Intelligence Index v4.3.2 score and cost per task at medium, high, xhigh and max effort. <https://artificialanalysis.ai/models/claude-haiku-5-5> Read 2026-10-08.

Artificial Analysis (2026). _GPT-6 Luna_, _Claude Sonnet 5.5_, _Claude Opus 5.5_, _Kimi K3_ and _GLM-5.3_ model pages. Intelligence Index score and cost per task. <https://artificialanalysis.ai/models/gpt-6-luna> Read 2026-10-08.

OpenAI (2026). _API pricing_. OpenAI documentation: GPT-6 Luna rates and the 272,000-token surcharge. <https://developers.openai.com/api/docs/pricing> Rates as recorded on this site’s model tracker, verified 2026-09-23.
