Anthropic released Claude Haiku 5.5 on October 7, 2026, a new small model the company calls the cheapest, fastest, and most capable Haiku it has ever shipped. The headline number is the price: for prompts up to 100,000 tokens, Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. That is roughly 90% below Haiku 4.5 pricing, and Anthropic says the model runs about 75% cheaper on average once typical workloads are factored in.

The launch matters because small models do the unglamorous bulk work of the AI economy: classification, extraction, summarization, routing, and the subagent calls that let bigger models delegate. Anthropic is positioning Claude Haiku 5.5 explicitly as a subagent for Opus 5.5 and Sonnet 5.5 on coding work, and as the engine for high-volume agent pipelines where inference cost, not raw capability, is the binding constraint on what gets built.

Pricing follows a two-tier structure that splits at 100,000 prompt tokens. Below the line, input costs $0.10 and output $0.50 per million tokens, with cache reads at $0.01 and five-minute cache writes at $0.125. Above 100,000 tokens, rates step up to $0.50 input and $2.50 output per million, still a 50% cut from Haiku 4.5's flat $1.00 and $5.00 pricing. Anthropic notes that about 90% of Haiku 4.5 requests fell under the 100,000-token threshold, which is exactly where the steepest savings land. Batch processing takes another 50% off the already reduced rates.

The context window also jumps from 200,000 tokens on Haiku 4.5 to 1 million tokens on Claude Haiku 5.5, with up to 128,000 output tokens per response. One wrinkle tempers the headline savings: the new tokenizer counts the same text as roughly 30% more tokens than Haiku 4.5 did, a detail Anthropic discloses in its migration guide, as reported by MarkTechPost. Developers moving workloads over will need to watch token accounting closely.

Benchmarks: punching above its price

Anthropic's published benchmark table puts Claude Haiku 5.5 well ahead of its predecessor and competitive with pricier rivals. The model scores 72.4% on the offline subset of OSWorld 2.1, a computer-use benchmark, versus 15.7% for Haiku 4.5 and 48.9% for OpenAI's GPT-6 Luna, according to Artificial Analysis. On Terminal-Bench 4.0, an agentic coding evaluation where Haiku 4.5 scored 0.0%, the new model reaches 39.2%, ahead of GPT-6 Luna's 16.4%.

On Humanity's Last Exam, a test of expert-level academic knowledge and reasoning, Claude Haiku 5.5 reports 45.9% without tools and 57.4% with them, compared with 10.2% and 18.7% for Haiku 4.5. Artificial Analysis also flags the model's relatively low hallucination rate: its knowledge accuracy trails larger models, as expected for a small model, but it is more willing to admit when it does not know an answer, which keeps its hallucination rate below rivals like Gemini 3.8 Flash and GPT-6 Luna.

Perhaps the most interesting addition is the adjustable effort setting, the first in a Haiku-class model. Developers can dial a request toward lower cost or higher intelligence without switching model tiers, with adaptive thinking on by default and effort defaulting to medium. The model takes text and image input, produces text output, and carries a June 2026 knowledge cutoff.

A price war in the small-model tier

The pricing lands Claude Haiku 5.5 exactly on top of OpenAI's GPT-6 Luna short-context rates, which list identical $0.10 and $0.50 figures. That is a direct challenge in the segment where developers are most price-sensitive. For prompts over 100,000 tokens the comparison gets murkier: Luna's higher tier starts only above 272,000 input tokens at $0.20 and $0.75, making it cheaper on list price for a mid-range 150,000-token prompt, as MarkTechPost's analysis points out.

Anthropic sweetened the broader platform alongside the launch. Cache reads on Sonnet 5.5 drop from $0.20 to $0.10 per million tokens, which the company says makes Sonnet 5.5 roughly 20% cheaper on most agentic work. Paid-plan subscribers also get new monthly API credits: Max 5x users receive $100 per month, Max 20x users get $200, and Team subscribers receive up to $500 pooled across their users, usable on any Anthropic model.

Early customer evidence comes from Asana, where staff software engineer Aaron Vinh says the company saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn. Anthropic is also updating its Python and TypeScript SDKs with beta support for computer use and browser use, areas where it says the mix of speed, capability, and price makes Claude Haiku 5.5 a particularly good fit.

Why it matters

The economics of AI agents live or die on the cost of the small, repeated calls that orchestrate them. Every routing decision, every tool-call summary, and every context compaction is a token spend, and at high volume those fractions of a cent compound into the difference between a demo and a business. By collapsing the price of that layer while expanding context to a million tokens, Anthropic is betting that the next wave of agent adoption will be won on unit economics rather than benchmark leaderboards.

For developers, the practical move is to test Claude Haiku 5.5 as a drop-in subagent for classification, extraction, and summarization workloads currently running on pricier models. For everyone else, the takeaway is simpler: the AI running quietly inside more apps is about to get dramatically cheaper to operate, which usually means there is about to be a lot more of it. Related coverage: TikTok's AI shopping assistant shows where consumer agents are heading, while Cloudflare's 39-millisecond decision models attack the same cost-and-latency problem from the infrastructure side. For the standards layer underneath it all, see the Sierra-Meta personal agent protocol.