On October 1, 2026, Cloudflare announced Clef and Clef-flash, its first open-source decision models, available immediately on Workers AI and downloadable from Hugging Face under the Apache 2.0 license. Alongside the models, the company debuted a reinforcement-learning fine-tuning service so customers can adapt them to their own workloads, according to the launch post on the Cloudflare blog.

The idea is straightforward. Most AI agents still ask a large language model a yes-or-no question and get back a full paragraph of text. Clef skips the paragraph. You hand it an input state and a set of typed questions, and it returns a probability for every allowed answer. Route the ticket, block the request, escalate to a human — a structured decision the agent can act on at once, with no free-form output to parse and no reasoning tokens to wait for.

The naming is deliberate. The company describes a decision model as analogous to a music clef, since it defines the domain of the context and the subsequent notes, the actions, that follow.

What decision models actually do

A decision model reads a situation and answers only fixed questions about it. Clef supports three question types: a yes-or-no check, a choice among named options with per-option probabilities, and a score against an ordered rubric. A single Workers AI request can carry up to 64 questions and as many as four images, MarkTechPost reported.

Speed is the point. Across Cloudflare's own benchmark runs, Clef answered in a median of 209.3 milliseconds. The smaller Clef-flash, built on Qwen 3.5-9B, hit a median of 38.8 milliseconds and 122.4 milliseconds at p95. Cloudflare sets those numbers beside Typesafe AI's Jev at 524.1 milliseconds median, per the Workers AI changelog.

Both variants carry a 64,000-token context window and a vision encoder. Unlike Jev, which reads only text, Clef can also read images and video. The models are built on frozen Qwen models with a small scoring head on top, according to the company's launch materials.

Crucially, the models are fully compatible with the Jev API. Switching from Jev means changing the endpoint and the model name, which makes experimentation cheap for teams already running decision models in production.

A category is forming in real time

Cloudflare is not alone in this race. The same day its announcement went live, Amazon's Strands team released Strands Decider 2B, a 2-billion-parameter open decision model designed to run locally and return confidence-scored choices in tens to low hundreds of milliseconds, as tracked by AI Agent Store's daily roundup.

Typesafe AI launched Jev, its first so-called System One model, on September 15, 2026. Open alternatives such as Kev-9B and Laya followed in the weeks after. One industry roundup called Clef the third such model in two days, after OpenAI's Decisions API and the Strands release, Startup Fortune noted.

The pattern is clear. Decision models are becoming the hot path inside AI agents. The expensive frontier model handles reasoning and action, while a small, calibrated decision model handles the cheap, high-volume choices: routing tickets, gating tool calls, verifying arguments before an action fires. The same trend is pushing self-hosted coding agents behind the firewall, and it pairs naturally with the new agent safety platforms meant to keep those agents in line.

The honest caveats

All of the performance claims are Cloudflare's own. Nobody has independently reproduced the benchmark results yet. And in Cloudflare's own tables, Jev still wins several of the harder reasoning tests, as Predictive Systems pointed out in its daily analysis.

There is also a pricing wrinkle. Hosted Clef costs about $0.24 per million tokens, compared with roughly $0.042 for Jev, according to reporting cited by The Register. The weights are free, so self-hosting is the budget play for teams that can run them.

Cloudflare has not made its training data public, the company told The Register. And the paired RL offering is rolling out first through a forward-deployed engineer team, with a self-serve stack stitching AI Gateway, Workers AI, and Containers together coming later, AI Weekly reported.

For agent builders, the practical takeaway is to prototype hybrid setups. A local decider gates risky tool calls, and the frontier model acts. The RL fine-tuning service, now open for design-partner signups, is the longer-term hook: it turns a generic decider into one tuned on your own data, captured and redeployed on Cloudflare's edge, OODA Loop summarized.

Decision models turn agent choices into numbers. Clef is the biggest infrastructure name yet to give the weights away, and the race to make every agent decision cheap, fast, and calibrated is officially on.