On October 1, 2026, Cloudflare announced Clef and Clef-flash, two open-source decision models designed for AI agents that need fast, structured answers rather than generated text. The Cloudflare Clef family is pitched directly against TypeSafe AI's Jev, the decision model that became the fastest-adopted model in Vercel AI Gateway's history after its September 15 launch. According to Cloudflare's launch post, the models exist to produce "bounded structured outputs cheaply, quickly and consistently that can be added into a workflow when a decision is required."

The Cloudflare Clef release is the first model family trained by the Cloudflare Workers AI team. Both models are hosted on Workers AI, Cloudflare's serverless inference platform running on its edge network, and the weights are published on Hugging Face under the Apache 2.0 license so anyone can run them locally. For more on how the platform opens its infrastructure to AI agents, see genznewz.com's agent documentation, which describes a similar automation-first approach.

What a decision model does differently

A traditional large language model generates text one token at a time. When an agent needs a routing decision — is this support ticket urgent, which team owns it, should this request be blocked — the model produces a sentence that downstream code must parse, and the agent pays for latency and reasoning tokens it never wanted.

Decision models collapse that loop. Clef reads an input state plus a set of typed questions and returns a probability for every allowed answer, with no free-form text at all. As reported by MarkTechPost, Clef supports three question types. The first, noul, is a yes/no question that returns the probability of "yes".

The second type, choice, picks one named option and returns per-option probabilities with a confidence value. The third, score, rates against an ordered rubric and returns a probability-weighted score. A single request on Workers AI can carry up to 64 questions and up to 4 images, which means one call can triage an entire support interaction.

Vision and benchmarks

The vision encoder is a genuine differentiator for the Cloudflare Clef pair. Unlike Jev, which handles text only, Clef can classify images and video frames, so a decision can be triggered by a camera frame rather than just text. The models also offer a 64,000-token context window, double Jev's 32k. Under the hood, full Clef is a 27-billion-parameter model post-trained on the Qwen3.8-27B base, while Clef-flash is a 9-billion-parameter variant built on Qwen3.5-9B.

Because Clef implements the same System One API as TypeSafe's Jev, developers can swap the endpoint and model name without reworking integrations. According to Cloudflare's changelog entry, this compatibility is a deliberate bid for the workflow-automation niche where Jev took root.

The benchmark battle is where Cloudflare is loudest. Across 43 runs published by the company, median latency for Clef is 209.3 milliseconds and Clef-flash hits 38.8 milliseconds, against 524.1 milliseconds for Jev. On quality, Cloudflare reports that a Clef model leads 7 of 10 decision benchmarks, including 94.20 macro-F1 on BANKING77 against Jev's 79.74.

The lead is not absolute. Jev still tops When2Call and BRIGHT, and on CLINC150 with out-of-scope detection, Clef scores 97.43 while Clef-flash drops to 66.77 — the fastest model is not always the most careful one. There are also real deployment costs: running Clef-flash locally reportedly requires a GPU with at least 41GB of video memory, and full Clef needs roughly 85GB at full context length. Workers AI pricing is $0.24 per million input tokens for Clef and $0.09 for Clef-flash, both above Jev's $0.042 rate.

Decision models are becoming agent infrastructure

The timing of the launch is pointed. The same day Cloudflare shipped Clef, Amazon released Strands Decider 2B, its own open-source Jev rival based on Qwen3.5-2B, reporting about 115 milliseconds of single-decision latency on an RTX 3090. Databricks, meanwhile, added an ai_decide function that can be invoked via SQL or REST API to classify data, assign scores, and route agent tasks in under a second.

Cloudflare is also launching a reinforcement-learning fine-tuning service to help customers tune the Cloudflare Clef models for their own workloads. The service starts as an engagement with its forward-deployed engineer team, with a self-serve version planned that lets customers capture request data via Cloudflare's AI Gateway, generate training rollouts on Workers AI, and redeploy a fine-tuned model. The company states it does not read, store, or train on customer requests unless they opt into the product.

The pattern echoes the recent genznewz report on Ethereum's zkAPI for private agent payments: infrastructure built specifically for autonomous agents, rather than repurposed human tooling. For agent builders, the emerging playbook is to put a decision model in the hot path — ticket routing, request blocking, escalation triage — and hand off to a full language model only when the decision calls for reasoning or generation. Small, cheap, and fast is winning the operational layer of agentic systems.