Cloudflare Launches Open-Weight Clef Decision Models, With Self-Reported Wins Over Jev

Released October 1, 2026 on Workers AI, Clef and Clef-flash return calibrated probabilities instead of text for AI agents, but Cloudflare's benchmark lead over TypeSafe's Jev is self-reported and the training data behind the models stays undisclosed.

EduFabTech Β· 4 October 2026 Β· 4 min read Β· 2 views
A schematic shows a support ticket flowing through Clef into calibrated yes/no probabilities, alongside stat tiles on Clef's benchmark wins, latency, and price versus Jev.
EduFabTech · Own work

Cloudflare released two open-weight "decision models," Clef and Clef-flash, on October 1, 2026, adding a second major entrant to a model category built to return structured judgments rather than generate text, according to the company's announcement on its blog. The models run on Cloudflare's Workers AI platform and are also downloadable from Hugging Face under the Apache 2.0 license.

A decision model takes a state β€” a support ticket, a web page, a trace of what an AI agent just did β€” together with a set of typed questions, and returns a calibrated probability for each allowed answer instead of writing a paragraph. Cloudflare describes the questions in three shapes: yes-or-no, a choice among named options, and a score against a rubric, according to its blog post. The company frames this as "smart if-statements" for software that routes support tickets, flags security incidents, or decides whether an autonomous agent should keep acting without asking a person first.

Clef, the larger of the two, is built on a frozen, specially post-trained Qwen3.8-27B backbone with an added vision encoder and a 64,000-token context window that accepts text and images; Clef-flash uses a frozen Qwen3.5-9B backbone, Cloudflare says. Both skip standard autoregressive text generation for a non-autoregressive, prefill-only pass that scores every candidate answer in parallel, which Cloudflare credits for latency far below typical chat models.

Grouped bar chart comparing Clef, Clef-flash, and Jev across the BFCL, API-Bank, and BANKING77 benchmarks, with a note that the scores are Cloudflare's self-reported numbers.
Grouped bar chart comparing Clef, Clef-flash, and Jev across the BFCL, API-Bank, and BANKING77 benchmarks, with a note that the scores are Cloudflare's self-reported numbers.EduFabTech · Own work

How Clef Compares With Jev

Cloudflare built Clef to be a drop-in, API-compatible alternative to Jev, a decision model from TypeSafe that established this category two weeks earlier, so that teams already using Jev's typed interface can swap models without rewriting application code. Across 43 of its own benchmark runs, Cloudflare reports a median latency of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash, against 524.1 milliseconds for Jev, and says a Clef model leads on seven of ten decision benchmarks in the Jev Decision Index, including a macro-F1 of 94.20 on the BANKING77 intent-classification set versus 79.74 for Jev, per the company's Workers AI changelog.

BenchmarkClefClef-flashJev
BFCL, case exact98.4798.7695.75
API-Bank, accuracy91.9393.1188.19
BANKING77, macro-F194.2090.9379.74

On TypeSafe's own workflow evaluations, Cloudflare reports that Clef beat Jev in three of four areas β€” invoice processing, customer service and security incidents β€” losing only on agent-trace observability, according to its blog post.

What Independent Reporting Adds

The Register's coverage attaches two qualifications absent from Cloudflare's own framing. First, the benchmark lead is Cloudflare's own measurement: "Cloudflare self-reported its own scores against the benchmark," and those scores have not yet been reproduced for ranking on the official Decision Index, the outlet reports. Second, despite the Apache 2.0 license on the published weights, the data used to train Clef is not public β€” Cloudflare AI Platform group product manager Michelle Chen confirmed to The Register that the training datasets are not released, meaning the models are open-weight rather than fully reproducible open-source releases.

Price is the other tradeoff. The Register reports that Clef costs $0.24 per million tokens on Cloudflare's hosted service, nearly six times Jev's $0.042 per million, so a team adopting Clef for its reported accuracy gains is also taking on materially higher inference costs unless it self-hosts the open weights instead.

Side-by-side panels contrast Cloudflare's reported benchmark and latency wins against The Register's caveats on unverified scores, undisclosed training data, and higher price.
Side-by-side panels contrast Cloudflare's reported benchmark and latency wins against The Register's caveats on unverified scores, undisclosed training data, and higher price.EduFabTech · Own work

Why This Matters for People Building AI Systems

Decision models target a specific cost problem in agentic pipelines: using a full general-purpose language model to answer a bounded question β€” route this ticket, flag this transaction, should this agent keep going β€” is slower and more expensive than the question requires. Cloudflare's own comparison found Clef completing a "fetch, render, and classify" task in 2.2 seconds, against 4.7 seconds for GPT-OSS-120B performing the same classification, according to its blog post. For engineers and researchers building or studying agentic systems, that is the practical argument for a separate decision layer: cheaper, faster, more consistent judgments at the many small branch points inside an agent's control loop, with a general-purpose model reserved for steps that need open-ended generation.

Cloudflare also announced a reinforcement learning fine-tuning service for Clef, starting with forward-deployed engineers working directly with individual customers and later a self-serve platform for capturing data and redeploying tuned models on Workers AI, per the company's blog post. The category itself β€” barely weeks old, with TypeSafe's Jev and Cloudflare's Clef as its two visible entrants β€” remains too new for independent, reproducible rankings, which is the gap The Register's reporting points to directly.

A quick question for readers

Source: Cloudflare

Sources (3)
  1. Cloudflare. Introducing Clef: our open-source decision models, and new RL fine-tuning platform. Cloudflare Blog, 2026. blog.cloudflare.com β†— Β· checked 4 Oct 2026
  2. Cloudflare. Introducing Clef: Cloudflare's first open-source decision models, now on Workers AI. Cloudflare Developers Changelog, 2026. developers.cloudflare.com β†— Β· checked 4 Oct 2026
  3. The Register. Cloudflare tries to outplay Jev with open-weight Clef models. The Register, 2026. theregister.com β†— Β· checked 4 Oct 2026