Anthropic released Claude Haiku 5.5 on 7 October at $0.10 per million input tokens and $0.50 per million output tokens. That is one tenth of what Haiku 4.5 cost, and it matches the price of OpenAI's GPT-6 Luna. The first independent tests show something the price list hides: Haiku 5.5 has five effort settings, and the setting you choose changes the cost of a finished job by ten times and the wait for an answer from 13 seconds to more than seven minutes. This article covers what Anthropic shipped, what Artificial Analysis measured, how Haiku compares with Luna and with Sonnet 5.5, and which setting to use.
What Anthropic released on 7 October
Haiku is the small, cheap tier of Claude. Anthropic's announcement positions Haiku 5.5 for high-volume work such as summaries, compaction, database queries and classification, and as a helper model that does the bulk of the reading while a larger model such as Opus 5.5 or Sonnet 5.5 makes the decisions.
- Price. $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Prompts over 100,000 tokens pay $0.50 and $2.50. Reading from the prompt cache costs $0.01 per million, and the Batch API halves the price again.
- Effort setting. It is the first Haiku-class model with an adjustable effort level: Low, Medium, High, Xhigh and Max. You trade cost and speed against intelligence on every request.
- Where to get it. Anthropic says it is available now on the Claude Platform as
claude-haiku-5-5, on Amazon Web Services, Google Cloud and Microsoft Azure, in Claude Code, and in claude.ai for Free, Pro, Max, Team and Enterprise users. - Context window. Artificial Analysis lists 1 million tokens. Anthropic's announcement page does not state it.
- The catch Anthropic is open about. It recommends Sonnet 5.5 or Opus 5.5 for complex agentic coding, and says Haiku 5.5 suits narrower jobs. Its cybersecurity safeguards are stricter than Haiku 4.5's.
Anthropic puts the saving at about 75% on average against Haiku 4.5, and 90% for prompts under 100,000 tokens, which it says were about 90% of Haiku 4.5 requests. It adds that the new tokenizer uses slightly more tokens per task and that the savings figures already include that.
The price: the same as GPT-6 Luna, one tenth of Haiku 4.5
On the price list, Haiku 5.5 sits exactly where OpenAI's cheapest model sits. Both list $0.10 in and $0.50 out, and Artificial Analysis lists a 90% cache discount for each. Against Anthropic's own range, the drop is steep. Prices are per million tokens, checked on 8 October 2026.
| Model | Input | Output | Cached input |
|---|---|---|---|
| Claude Haiku 5.5 (up to 100K prompt) | $0.10 | $0.50 | $0.01 |
| Claude Haiku 5.5 (over 100K prompt) | $0.50 | $2.50 | $0.05 |
| GPT-6 Luna | $0.10 | $0.50 | 10% of input |
| Claude Haiku 4.5 | $1.00 | $5.00 | $0.10 |
| Claude Sonnet 5.5 | $2.00 | $10.00 | $0.10 |
Read the 100,000-token line carefully. Anthropic prices by prompt length, and the higher rate applies to the whole prompt once it crosses the line. By our arithmetic from the table, a 99,000-token prompt costs about $0.0099 to read, and a 101,000-token prompt about $0.0505, five times as much for 2% more text. If you feed Haiku long contracts or ledgers, split the work into pieces under 100,000 tokens.
What an independent test says: the effort setting is the price
Artificial Analysis runs every major model through the same ten-test Intelligence Index and records what it cost to finish one task. Haiku 5.5 appears five times, once per effort setting. Cost per task is what Artificial Analysis paid on average, and response time is how long a task took end to end. All figures were read on 8 October 2026.
Going from Low to Max buys 14 more points of score, from 29 to 43, for about ten times the cost and a wait of seven minutes instead of thirteen seconds. Artificial Analysis calls Max "very verbose": it used 440 million output tokens to finish the index, against a median of 100 million for comparable models. Low used 32 million. Every extra token is billed at the output rate, which is why a model with a cheap price per token can still produce an expensive job.
Haiku 5.5's price per token is a floor, not a quote. The effort setting you leave it on decides what you actually pay.
Haiku 5.5 vs GPT-6 Luna: the head-to-head at the same price
Because the list prices are identical, the comparison comes down to how many tokens each model spends to reach a given score. Here are the closest pairs on the same leaderboard.
| Model and setting | Score | Cost per task | Response time |
|---|---|---|---|
| GPT-6 Luna, low | 22 | $0.0045 | 7 s |
| Claude Haiku 5.5, low | 29 | $0.02 | 13 s |
| GPT-6 Luna, medium | 30 | $0.02 | n/a |
| GPT-6 Luna, high | 33 | $0.03 | 26 s |
| Claude Haiku 5.5, medium | 34 | $0.05 | 17 s |
| GPT-6 Luna, xhigh | 35 | $0.04 | 41 s |
| GPT-6 Luna, max | 38 | $0.07 | 115 s |
| Claude Haiku 5.5, high | 38 | $0.08 | 31 s |
| Claude Haiku 5.5, xhigh | 41 | $0.12 | 90 s |
| Claude Haiku 5.5, max | 43 | $0.21 | 434 s |
- At a score of 38 it is a tie on cost. Luna at max is $0.07 and Haiku at high is $0.08. But Haiku gets there in 31 seconds and Luna in 115, so for the same money Haiku is faster.
- Luna is cheaper at the bottom. If a job needs only a score around 22, Luna at low costs $0.0045, less than a quarter of Haiku's cheapest setting. For plain sorting and tagging that is hard to beat.
- Haiku goes higher. Luna stops at 38. Haiku reaches 43 at its top setting, though that setting is the expensive, slow one.
- Anthropic's own comparison. On the tests Anthropic published, Haiku 5.5 scores 1620 on GDPval-AA against Luna's 1437, 72.4% against 48.9% on OSWorld, and 39.2% against 16.4% on Terminal-Bench 4.0. Those are Anthropic's figures, run by Anthropic, and the gaps are larger than the independent index suggests.
Our view: neither model wins outright. Pick Luna when the job is easy and the volume is huge. Pick Haiku when you need the score to go above 38, or when waiting two minutes is not acceptable.
Against Haiku 4.5, Sonnet 5.5 and the rest of the field
The generational jump is the headline. Haiku 4.5 scores 17 on the same index at $0.28 per task. Haiku 5.5 at Low scores 29 for $0.02: a higher score for one fourteenth of the cost, from the same company. If you run Haiku 4.5 today, switching is the easy decision on this page.
Against the bigger Claude models the saving is real but smaller than the price list suggests. Claude Sonnet 5.5 at medium scores 41 at $0.48 per task. Haiku 5.5 at Xhigh scores 41 at $0.12, four times cheaper, but takes 90 seconds against Sonnet's 7. Sonnet exists for a reason. Anthropic's numbers show it ahead of Haiku on every test it published, for example 70.6% against 39.2% on Terminal-Bench 4.0 and 83.9% against 72.4% on OSWorld.
Haiku at Max is the one setting we would avoid. At $0.21 per task it costs the same as GPT-6.1 Sol at medium, which scores 48. Sol at low scores 42 for $0.13, one point below Haiku's Max score at 62% of the price. And Xiaomi's MiMo-V2.6-Pro scores 46 at $0.13 on the same leaderboard. Haiku's sweet spot is Low to High, where it is cheap and quick. Above that, others are better value.
Is it fast enough for a live chat or a phone call?
Anthropic lists live support and browser use among Haiku 5.5's uses, and Asana told Anthropic it saw over 30% lower latency and up to 2.5 times faster inference per agent turn against the model it was using. That is a customer's report to the vendor, on its own workload.
Artificial Analysis measures something stricter: the time until the first answer token arrives, because the model reasons before it speaks. At Low that is 9.9 seconds for Haiku 5.5, against 3.1 seconds for Luna at low. At High it is 28 seconds, and at Max it is 432 seconds. For a text chat where the customer sees "typing", ten seconds is tolerable. For a phone call, it is not. If you build anything live, test the real delay on your own prompts before you commit, and keep Max for overnight batch work.
Who should switch, and who should wait
- Switch now if you run Haiku 4.5 for summaries, tagging, extraction or triage. The same job gets better and costs a fraction. Start at Low, check quality on 50 of your own examples, and move up one setting only if it fails.
- Test it if you run Luna at medium or above. At a score of 38 the cost is the same and Haiku is faster, and Haiku can go higher if you need it.
- Stay with Luna if the work is simple and the volume is large. Its low setting is far cheaper per task.
- Stay with Sonnet 5.5 or Opus 5.5 if the job is coding, multi-step agent work or anything where a wrong answer is expensive. Anthropic says so itself.
- Whatever you pick, cap the spend. Set a maximum output length per request and a monthly limit in the provider's console, and log the cost of every job. Never leave a reasoning model on Max by default.
One thing we cannot tell you yet is how it handles Urdu or Arabic. Anthropic's announcement does not report results for either, so run your own sample before you move a customer-facing workflow.
Your next step: measure your own cost per job
Every launch week brings a new "cheapest" model. The cheapest one for you finishes your jobs in the fewest tokens at a quality you accept. We set out how to measure that in the cheapest AI model for a business, by cost per task, and compare the two mid-tier models Haiku sits beneath in Claude Sonnet 5.5 vs GPT-6.1 Sol. We will refresh this page when Artificial Analysis re-runs the model or Anthropic changes the price.
If you would rather use AI than manage models, that is what our products are for. The AI Chat Assistant answers website visitors, captures leads and hands harder chats to your team over live chat or WhatsApp, and CallSentinel scores your sales and support calls and reads Urdu and Arabic calls back in English. If you are weighing a bigger project and want a second opinion on which model fits, talk to us.