Anthropic launched Claude Haiku 5.5 on October 7, 2026, and the headline number is hard to miss: for prompts under 100,000 tokens, API prices fall 90% compared with the previous generation. Input tokens now start at $0.10 per million and output tokens at $0.50 per million, according to VentureBeat. That pricing lands exactly on OpenAI's GPT-6 Luna rates, setting up a direct contest for the cheapest way to run millions of small, repetitive AI tasks.

The launch matters most for one audience: agents. Claude Haiku 5.5 is the smallest model in the Claude family, and Anthropic describes it as built for high-volume work such as summarization, classification, database queries, and the tiny subtasks that larger models delegate to smaller ones. When an agent fleet fires thousands of background jobs every hour, the per-token price becomes the bill. Cutting the cost of Claude Haiku 5.5 by 90% changes which workflows are affordable to automate at all.

Anthropic says the new model costs roughly 75% less to run on average than Haiku 4.5, once a new tokenizer and different token consumption patterns are factored in. That figure is a vendor estimate, not an independent audit, and the details decide how much of the savings any given workload actually captures.

The fine print: 100,000 tokens

The 90% cut comes with a boundary. Above 100,000 prompt tokens, Claude Haiku 5.5 costs $0.50 per million input tokens and $2.50 per million output tokens, reported by Unite.AI. Those rates are still half of Haiku 4.5's prices, but they represent a fivefold jump over the headline tier. Anthropic says roughly 90% of requests to the older model stayed under the threshold, so the cheaper band covers most real traffic — but any long-context workflow needs its own cost calculation.

There is also a quiet change in the Claude Haiku 5.5 token math. The new tokenizer uses more tokens for the same work, with one developer analysis estimating about 30% more. A cheaper token is not automatically a cheaper task, and teams porting over from Haiku 4.5 should re-measure per-task cost rather than per-token price.

Caching gets cheaper too. Cache reads drop to $0.01 per million tokens in the lower tier, and Anthropic separately halved Sonnet 5.5 cache reads from $0.20 to $0.10 per million. The company estimates roughly 20% savings on typical agent workloads that reuse stored context heavily.

Why agents are the real audience

Anthropic positioned Haiku as a support worker for its larger Opus and Sonnet models. In one example shared by financial AI company Rogo, a larger model assembles a presentation while Haiku retrieves the revenue figure needed for a single slide. The applied-AI team member quoted in the announcement said the model is accurate enough to trust for such jobs and fast and cheap enough to run constantly.

Early customer reports point in the same direction. Asana reported task-completion latency falling more than 30% with inference per agent turn accelerating by as much as 2.5 times, according to VentureBeat. Box reported an 11-point improvement over Haiku 4.5 at roughly half the latency. HubSpot cited a 92.8% average across three runs of its CRM evaluation, the strongest result among the models it tested. These are company-supplied testimonials, not comparable independent throughput figures.

The model also introduces adjustable effort levels for the first time in the Haiku line, letting developers trade cost against capability. But the setting changes the bill and the benchmarks: Anthropic's Terminal-Bench score of about 39% was achieved at maximum effort, while the medium default lands near 20%. Developers choosing higher effort tiers should verify that the gains pay for themselves on their own evaluations.

This launch arrives in an ecosystem already racing to hand more work to smaller models. Recent coverage of agent economics has tracked both the opportunity and the risk of cheaper autonomous labor, from acquisitions built around agent-led customer experience to governance failures when firms act on bad AI calls.

A pricing war at the small-model tier

Claude Haiku 5.5 does not stand alone at the bottom of the price ladder. OpenAI's GPT-6 Luna matches all four of Haiku's lower-tier rates, including caching, according to the comparison compiled by VentureBeat. Google's Gemini 3.5 Flash-Lite sits higher at $0.30 per million input tokens, while Grok 4.3 starts at $1.25. The pricing contest is real, but price alone does not determine cost per completed job — token consumption, retries, and accuracy all feed into the final number.

Anthropic's benchmark table shows Claude Haiku 5.5 with gains over Haiku 4.5 and leads over GPT-6 Luna on selected evaluations. On the knowledge-work test GDPval-AA v2.1, Haiku 5.5 scored 1,620 against 735 for its predecessor and 1,437 for GPT-6 Luna. On the computer-use benchmark OSWorld 2.1's offline subset, it posted 72.4% versus 15.7% for Haiku 4.5. Sonnet 5.5 remains ahead on every measure, which supports reserving more complex work for larger models.

As reported by Okay News, the model became available immediately on the Claude platform and through AWS, Google Cloud, and Microsoft Azure under the model ID claude-haiku-5-5. Anthropic also published a migration guide for developers moving from the older model.

The extras: credits, safeguards, and tighter security

The announcement extends beyond the small model. Monthly API credits are rolling out for subscription customers: $100 for Max 5x subscribers, $200 for Max 20x, and up to $500 shared among Team users. The credits can be spent on any model through Anthropic's platform.

On safety, Claude Haiku 5.5 tightens cybersecurity restrictions relative to its predecessor, including blocking penetration testing under standard safeguards. Anthropic notes that defensive cybersecurity work remains permitted, and organizations seeking broader access can apply through its verification programs.

Anthropic's cybersecurity stance is worth watching given recent reports of autonomous agents being turned against infrastructure, and given the wave of agent-led acquisitions reshaping customer experience platforms. Cheaper agents mean more agents, and Claude Haiku 5.5 joining the pricing war means the governance questions get louder, not quieter.

What builders should do this week

First, the practical steps. Developers adopting Claude Haiku 5.5 should check whether their workloads stay under 100,000 tokens per request, since crossing that line multiplies prices by five. Teams with heavy cache reuse stand to gain the most from the Sonnet and Haiku cache-read reductions. And anyone tempted by the maximum-effort benchmark numbers should rerun their own evaluations at the effort level they actually plan to pay for.

The bigger picture is that the cheapest AI labor keeps getting cheaper. A year ago, the smallest Claude model cost $1 per million input tokens — expensive for the thousands of tiny background jobs modern agents run every minute. Now two rival labs price the same work at a tenth of that. For agent builders, the constraint is shifting from token budgets to orchestration: deciding which small decisions deserve a small model, and which need a bigger one. The price war made the small model nearly free. The engineering judgment still costs something.