Anthropic released Claude Sonnet 5.5 on Monday, the second model in its Claude 5.5 family — and it pulled off the rare trick of getting cheaper without getting cheaper. The sticker price didn't move: it still costs $2 per million input tokens and $10 per million output tokens, the same as its predecessor. But Anthropic says the model generates output more than 30% faster and costs up to 30% less per completed task in its testing.

The headline number isn't the speed, though. It's the step count. In early enterprise testing, Claude Sonnet 5.5 needed meaningfully fewer tool calls to finish the same jobs — and for the agentic workloads now spreading through every industry, every skipped tool call is one fewer chance for an autonomous workflow to go off the rails.

The price of tokens was never the real bill

For a chatbot, token price is the whole invoice. For an agent, it's barely the cover charge. An agentic coding session burns tokens, sure — but it also burns time on tool calls, shell executions, retries, and the occasional catastrophic wrong turn that a human has to unwind. That's why the most interesting data in Anthropic's launch isn't the pricing table. It's the enterprise efficiency numbers, as reported by VentureBeat.

At Lovable, co-founder and CTO Fabian Hedin said evaluations found Claude Sonnet 5.5 required about one-third fewer tool calls — and roughly half as many shell executions — to complete coding jobs. For coding agents, fewer tool calls may prove the most important metric of all: each call is a decision point, a latency hit, and a surface for error.

Box's VP of AI Products, Yashodha Bhavnani, told Anthropic the new model was more accurate, ran 2.4 times faster, and used 12% fewer total tokens. Perhaps more tellingly, she said it rechecked source documents and caught errors its predecessor missed — the kind of quiet diligence that matters when agents are handling real documents.

Zendesk tested the model across hundreds of support cases spanning customer replies and escalation decisions. Director of AI Abhinay Kathuria said tickets were processed 20% faster and the model made fewer incorrect decisions than the Claude models Zendesk currently runs in production. At Slack, principal engineer Curtis Allen said Claude Sonnet 5.5 outperformed Sonnet 5 on nearly all of the company's offline Slackbot evaluations without any prompt changes — while requiring fewer steps and roughly 14% fewer output tokens.

And at Base44, across 118 real application builds, the new model reached results comparable to the far pricier Opus 5 in an average of 3.6 iterations per build versus 7.7 — while producing the fewest failed tool calls among the models it compared. As VentureBeat's analysis put it, token prices tell only part of the story: every unnecessary tool call, failed action, and retry adds latency, infrastructure cost, and another opportunity for an autonomous workflow to go off course.

Fewer steps meets the real bottleneck: integration

That framing lands harder when you set it against Anthropic's own research. The company's 2026 State of AI Agents report, published this month, drew on data from more than 500 technical leaders plus Anthropic's own internal deployment metrics. Internally, Anthropic runs roughly 30,000 AI agents simultaneously on research and engineering work, with AI now leading 26% of the company's R&D tasks — up from under 1% in February. Ninety percent of internally produced code is written by AI agents.

The external survey data shows 80% of organizations now report measurable ROI from AI agents. But the number that tells the actual story, in Shawn Livermore's breakdown of the report, is this: 46% of organizations name integration with existing systems as their primary challenge. Not model quality. Not cost. Not governance frameworks or regulatory concerns. Integration.

A separate survey makes the same point from the developer's side. TechCoffeeHouse reported on a September study of more than 800 developers and engineering leaders across Southeast Asia and India: 53% now say AI agents are in production or broad use at their organizations, yet only 38% consider their codebase mostly or fully ready for a fully autonomous agent. Cost has now overtaken reliability as developers' biggest worry about wider adoption — and 86% of developers still review or validate AI-generated outputs most of the time.

This is exactly the tension the agent economy is living with right now: agents are capable and deployed, but the plumbing around them — identity, discovery, review capacity, receipts — hasn't caught up. A companion piece on the agent-to-agent economy's plumbing problem makes the case, where marketplaces raise millions while agents wait 140 hours in review queues. Cheaper, faster, more reliable steps are half the fix. The other half is the infrastructure that watches, verifies, and pays those steps. Claude Sonnet 5.5 improves the first half on a big scale; the second half is still being built.

What shipped, and what's next

The model itself is straightforward. As Unite.AI detailed, Claude Sonnet 5.5 is available on all major platforms — Amazon Web Services, Google Cloud, and Microsoft Azure — listed as Claude Sonnet 5.5 on each provider's model catalog, and it's offered with zero data retention. Anthropic describes it as a faster, lower-cost complement to Claude Opus 5.5, strongest at well-scoped everyday tasks: fixing bugs, producing polished documents, slides, and spreadsheets. The company also highlights a sharper eye for design, from refining user interfaces to producing template-following slides that need minimal editing.

On safety, the new model is the first Sonnet to launch with cyber safeguards and fallbacks similar to those used for Anthropic's most capable models — a notable step, given that Sonnet is the tier most enterprises actually deploy at scale. As 9to5Mac noted, its biology safeguards remain unchanged from Sonnet 5, targeting a narrow set of high-risk requests while leaving routine software development and most life sciences work unaffected. Early testers described it as a better collaborator than Sonnet 5, with clearer writing and speed suited to quick iteration.

The timing fits a broader rollout. Claude Sonnet 5.5 follows the launch of Opus 5.5 last week — the flagship tier at $4 per million input tokens and $20 per million output — and a third model, Haiku 5.5, designed for fast high-volume work, is due soon. Reuters, via Froggyweb, notes Anthropic is expanding its product lineup ahead of a planned IPO. Enterprise customers account for about 80% of the company's business, with clients including Salesforce, Databricks, Goldman Sachs, and Novo Nordisk.

There's an irony worth naming: CEO Dario Amodei earlier this month called on the global AI community to slow down the pace of releasing new capabilities to address safety concerns — and then Anthropic shipped its second frontier model in two weeks. The company's answer, presumably, is that an efficiency release isn't a capability release. That's partly true. But when a "medium" model starts matching the flagship's build quality in half the iterations, the frontier has moved anyway — just sideways instead of upward.

The bigger trend to watch is what the agent economy does with the savings. The same day agents got a 30% tax cut on each step, they also gained a new place to spend steps: MEXC's new CLI lets users' own AI agents turn trading intent directly into orders. Agents are becoming an execution layer, not just an advice layer — and in an execution layer, the cost of a failed action isn't a wasted token. It's a wasted trade, a botched deployment, a support ticket gone wrong. What to watch next: whether Haiku 5.5 follows the efficiency pattern, whether rivals answer on per-task cost rather than raw capability, and whether fewer failed tool calls start showing up as fewer agent incidents in the field.