Cloudflare has unveiled a new kind of AI model that does not write essays or answer questions but instead makes decisions. The Cloudflare decision models, called Clef and Clef-flash, were released on October 1, 2026, and are built to give autonomous AI agents a fast, reliable way to choose what to do next without calling a giant language model every time.

Unlike conventional large language models that generate text token by token, Clef outputs probabilities and structured choices. An agent can feed it a situation, for example whether an email needs escalation or which tool should handle a request, and receive a decision in as little as 39 milliseconds. The model weights have been published on Hugging Face under the Apache 2.0 license, making them available for researchers and enterprises to run locally.

A new layer between language models and software

The Cloudflare decision models fill a gap that has opened up as AI agents take on more operational work. Agents that browse the web, call APIs, and trigger workflows need to make hundreds of small judgments per task. Sending each of those judgments to a frontier model is slow and expensive, while a simple hand-written classifier is often too rigid. Clef is designed to sit between those extremes: more flexible than a classifier, far faster than a chat model.

According to The Register, which covered the release in its report on the new models, Clef is a 27-billion-parameter multimodal model based on a Qwen backbone, while the smaller Clef-flash uses a 9-billion-parameter variant tuned for latency-critical workloads. Both models accept text, JSON, images, and video, with the larger model carrying a vision encoder that lets it factor images into its decisions.

Built for speed on Cloudflare's edge network

Latency is the headline metric. Clef-flash reaches a decision in approximately 39 milliseconds, roughly 13 times faster than the rival Jev model from TypeSafe AI, which benchmarks at over 524 milliseconds. On the Decision Index benchmark, Clef scored 61.2 and Clef-flash 57.1, compared with 57.9 for Jev, suggesting the faster models sacrifice little accuracy for their speed.

The architecture achieves this by skipping token generation entirely. Clef uses a non-autoregressive, two-stage attention routing process that derives schema choices directly from the model's internal representations. In practical terms, the model never spells out an answer; it simply selects one, which is what makes the sub-40-millisecond response time possible.

Cloudflare is hosting the models on its Workers AI platform at $0.24 per million tokens and is rolling out a reinforcement learning service so companies can fine-tune Clef for their own decision patterns, initially with help from forward-deployed engineers. The company has confirmed that while the weights are open, the training datasets remain private, and it is seeking design partners for specific industry use cases.

Why agents matter in this release

The launch arrives at a moment when autonomous agents are moving from demos to deployment, and the surrounding debate is intensifying. OpenAI's newly launched always-on AI agents show how quickly agentic software is becoming a product category, while the White House's AI self-policing accord reflects growing concern about keeping those agents under control. Decision models like Clef add another control surface: a fast, auditable layer that can enforce rules, rank options, or defer to a human when confidence is low.

Cloudflare's edge network, spanning hundreds of cities, is a natural home for this kind of workload. Decisions made close to the user cut network round trips, and because the models run on shared infrastructure, small companies can access decision-grade AI without operating their own GPU clusters.

Why it matters

For readers, the takeaway is simple: AI is becoming less about one giant model doing everything and more about specialized parts working together. A voice agent, a coding assistant, or a customer service bot of the near future will likely combine a large reasoning model with fast components like these Cloudflare decision models handling the moment-to-moment choices. Cheaper, faster decisions mean agents can run longer, react quicker, and stay affordable at scale, which is exactly what the next wave of AI products needs to reach mainstream users.