Anthropic launches Claude Haiku 5.5: around 75% cheaper than Haiku 4.5, and much stronger
Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, 90% below Haiku 4.5 for most prompts, while scoring 72.4% on OSWorld. Anthropic also halved Sonnet 5.5 cache reads and added monthly API credits for Max and Team plans.

Anthropic has finished its Claude 5.5 lineup. On 7 October 2026 the company released Claude Haiku 5.5, the smallest and cheapest model of the family, and it did so with an unusually aggressive price tag. According to Anthropic, Haiku 5.5 is "the cheapest, fastest, and most capable small model" it has ever shipped, and on average it costs around 75 percent less to run than Haiku 4.5, the model it replaces. The launch arrives only weeks after Opus 5.5 and Sonnet 5.5, so the whole Claude 5.5 generation is now available.
The announcement was not only about one model. Anthropic also cut the price of cache reads on Sonnet 5.5 in half and introduced monthly API credits for people who already pay for a Claude Max or Team plan. Taken together, the three changes say a lot about where the AI market is heading: the frontier models keep getting smarter, but the real competition for everyday work is moving to price, speed and volume.

What Haiku 5.5 costs
The headline number is the price. For prompts up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. Haiku 4.5 cost $1.00 and $5.00. That is a cut of 90 percent on both sides. Cache reads drop from $0.10 to $0.01 per million tokens, and cache writes from $1.25 to $0.125.
There is a second price tier for long prompts. Above 100,000 tokens, Haiku 5.5 costs $0.50 per million input tokens and $2.50 per million output tokens, which is still 50 percent below Haiku 4.5. Anthropic says that around 90 percent of requests to its previous Haiku model stayed below the 100,000 token line, which is why the cheaper tier matters most in practice.
So why does Anthropic talk about "around 75 percent" and not 90 percent? The company explains this in a footnote. Haiku 5.5 uses an updated tokenizer, similar to the one in Sonnet 5.5 and Opus 5.5, and it needs slightly more tokens to finish the same piece of work. Once you mix the two price tiers and the extra tokens, the average saving lands at about 75 percent. That is still a dramatic drop for a model that, according to the benchmarks, is far more capable than its predecessor.
For comparison, Sonnet 5.5 costs $2.00 per million input tokens and $10.00 per million output tokens. At list price, Haiku 5.5 is therefore 20 times cheaper than Sonnet 5.5 for short and medium prompts.
Faster and much stronger than Haiku 4.5
Cheap models usually come with an obvious trade off. Haiku 5.5 is clearly weaker than Sonnet 5.5, and Anthropic says so openly. But the distance to its own predecessor is huge, and on several tests it beats GPT-6 Luna, which Anthropic included as a reference point.

A few numbers from Anthropic's own table stand out:
- Computer use: on the offline subset of OSWorld 2.1, where an agent has to operate a real computer through long multi step tasks, Haiku 5.5 scores 72.4 percent. Haiku 4.5 managed 15.7 percent and GPT-6 Luna 48.9 percent. Sonnet 5.5 is at 83.9 percent.
- Agentic coding: on Terminal-Bench 4.0, Haiku 5.5 reaches 39.2 percent, up from 0.0 percent for Haiku 4.5. GPT-6 Luna scores 16.4 percent, Sonnet 5.5 70.6 percent.
- Hard reasoning: on Humanity's Last Exam without tools, Haiku 5.5 gets 45.9 percent, compared with 10.2 percent for Haiku 4.5. With tools it reaches 57.4 percent.
- Knowledge work: on GDPval-AA v2.1, a rating from Artificial Analysis that covers real professional tasks in 44 occupations, Haiku 5.5 scores 1620 points. Haiku 4.5 had 735, GPT-6 Luna 1437 and Sonnet 5.5 1840.
- Charts and visuals: on the Chartography test for visual reasoning, Haiku 5.5 reaches 46.4 percent, against 6.4 percent for Haiku 4.5.
Benchmarks published by a model maker should always be read with some care. The company picks the tests, the settings and the comparison models. Still, the size of the jump is hard to ignore. A model that went from almost nothing to almost 40 percent on a demanding terminal benchmark, while getting cheaper, is a different product, not a small refresh.
Anthropic also says Haiku 5.5 is its fastest model to date at standard speed. The only exception is Opus in the special Fast Mode, which runs quicker. Speed matters for the jobs Haiku is built for, because a customer waiting in a chat window or an agent clicking through a website cannot afford long pauses.
The first Haiku with an effort dial
Haiku 5.5 is the first Haiku class model with an adjustable effort setting. Anthropic's larger models already let developers choose how much the model should think before it answers. More effort usually means better answers but more tokens and higher cost. Less effort means faster and cheaper replies.
For a small model this dial is especially useful. A company can run simple classification jobs at low effort and switch to a higher setting for harder cases, without moving to a bigger and more expensive model. Anthropic published charts that show accuracy against cost at each effort level on OSWorld, GDPval-AA and Humanity's Last Exam, and the message is the same in each of them: Haiku 5.5 sits far to the left on cost, while getting close to the much pricier models on accuracy at its higher settings.
What it is meant for
Anthropic is clear that Haiku 5.5 is not a replacement for Sonnet or Opus on complex coding. The company writes that Sonnet 5.5 and Opus 5.5 remain the better choice for long agentic coding tasks like those in Terminal-Bench. Haiku 5.5 is designed for narrow, high volume jobs that used to be too expensive to hand to Claude.
Typical examples from the announcement:
- Summaries and compaction: shortening long chats or documents so that a bigger model can keep working with less context.
- Classification and database queries: sorting tickets, tagging content, pulling a specific value out of a table.
- Subagents: a big model plans a task, and many small Haiku agents do the legwork in parallel, for example reading a long report and extracting one number.
- Live customer support: fast replies where every second of waiting counts.
- Browser use: clicking through websites and filling in forms on behalf of a user.
The subagent idea is probably the most important one. Modern AI coding tools and agent frameworks increasingly split work between a lead model and many helpers. If those helpers become ten times cheaper, a team can run far more of them, or run the same setup for a fraction of the old bill.

To make the price change more concrete, we ran a simple example. Imagine a support tool that answers one million customer questions per month. Each request sends 5,000 tokens of context and gets a 500 token answer back. Without caching, that is 5 billion input tokens and 500 million output tokens. At list prices, the bill would be about $15,000 with Sonnet 5.5, $7,500 with Haiku 4.5 and $750 with Haiku 5.5. Because of the new tokenizer the real Haiku 5.5 figure would be a bit higher, but the order of magnitude stays the same.
What early customers report
Anthropic shared feedback from several companies that tested the model before launch. As with any launch quote, these are hand picked, but they are specific enough to be interesting:
- Asana said latency for task completions in its AI Teammates product dropped by more than 30 percent compared with its current model, with up to 2.5 times faster inference per agent turn.
- HubSpot reported that Haiku 5.5 got the best score it has seen on its CRM test suite for smaller models, 92.8 percent averaged over three runs.
- AlphaSense runs a document question feature with about 8 million calls per week. In a test with 400 queries, Haiku 5.5 scored 0.84 against 0.76 for Haiku 4.5.
- Box said Haiku 5.5 scored 11 points higher than Haiku 4.5 at about half the latency.
- Cognition, the company behind the Devin coding agent, now offers Haiku 5.5 as a helper model next to Opus 5.5 as the lead.
Cheaper Sonnet and free API money for subscribers
Two more changes came with the launch. First, cache reads on Sonnet 5.5 now cost $0.10 per million tokens instead of $0.20. Cache reads are the part of a request that reuses context the model has already seen, and in agent work they make up a large share of all tokens. Anthropic says this cut makes Sonnet 5.5 about 20 percent cheaper on most agentic tasks.
Second, Anthropic is rolling out monthly API credits for Max and Team subscribers this week. The amounts are generous compared with the subscription prices.

According to Anthropic's documentation, Max 5x subscribers get $100 per month, Max 20x subscribers $200. Team plans get $20 per Standard seat and $100 per Premium seat, pooled into one balance and capped at $500 per month. Free, Pro and Enterprise plans are not eligible. New subscribers can claim the credits after seven days on an eligible plan.
The credits come with rules. They only work on the Claude API in the Claude Console, not on AWS, Google Cloud or Microsoft's platforms. They cover the API, Claude Managed Agents, the Claude Agent SDK and the playground, but not Claude Code or extra usage inside the Claude apps. They expire at the end of each billing cycle with no rollover, and you have to link a Claude Console organization to your plan to receive them.
The idea behind this is easy to see. Many people pay for Max to use Claude in the chat app or in Claude Code, but they never touch the API. With $100 or $200 a month of free credits and a model as cheap as Haiku 5.5, it suddenly costs nothing to try building a small agent or app. At Haiku 5.5 prices, $100 buys roughly a billion input tokens. That is a strong hook to turn subscribers into developers.
Safety and availability
Anthropic says Haiku 5.5 shows major improvements on almost all of its alignment tests compared with Haiku 4.5, with far fewer cases of misaligned behavior and a lower willingness to help with misuse. The details are in the Haiku 5.5 system card. On cybersecurity, the safeguards are stricter than for Haiku 4.5, but somewhat looser than for other recent Claude models. They allow a wider range of defensive security tasks than Sonnet 5.5 does, but still block penetration testing and similar attacker techniques. The biology safeguards match those of Sonnet 5, Sonnet 5.5 and Opus 5.
The model is available now on the Claude Platform under the name claude-haiku-5-5, and on Amazon Web Services, Google Cloud and Microsoft Azure. Anthropic also updated its Python and TypeScript SDKs with beta support for computer use and browser use, and it points to Haiku 5.5 as a natural fit for those tasks.
Why it matters
For a long time, the race in AI was mostly about who has the smartest model. That race continues at the top, but launches like this one show a second race that may matter even more for daily use: who can deliver good enough intelligence at the lowest price and the highest speed. A model that handles summaries, support chats and simple agent steps for a tenth of the old price changes what companies are willing to automate.
It also puts pressure on competitors. OpenAI is rolling out GPT-6 in ChatGPT this week, and every big lab now sells a small, fast model tier. Anthropic choosing to compare Haiku 5.5 directly with GPT-6 Luna is a clear signal that the small model segment is now a battlefield of its own.
For regular users, the effect is indirect but real. Cheaper models make it possible for apps to add AI features without charging extra, and they make agents that run many steps in the background affordable. For developers with a Max or Team plan, the new monthly credits are worth claiming today, even if it is just to experiment.
Sources
Source: anthropic.com