Cloudflare's Clef: When a Decision Model Beats an LLM Call
Most moderation, routing, and triage calls aren't generation problems. Clef proves it by answering in 38.8 milliseconds instead of seconds.
Key takeaways
- Most moderation, routing, and triage calls aren't generation problems.
- Clef proves it by answering in 38.8 milliseconds instead of seconds.
Cloudflare shipped Clef and Clef-flash on Workers AI on October 1, 2026, and the pitch is narrow on purpose: these are decision models that take your data and a typed question and return a probability, not a chat completion. Clef-flash answers in 38.8 milliseconds at the median. If your pipeline burns a full LLM call on something that's really a yes/no, a category pick, or a 1-10 score, you're paying generation prices for a classification problem, and Clef is the clearest evidence yet the two should be billed and built differently.
What a decision model actually returns#
Clef comes in two sizes: Clef at 27 billion parameters on a Qwen3.8-27B backbone, and Clef-flash at 9 billion parameters on Qwen3.5-9B. Both carry a 64,000-token context window and both accept images alongside text, so you can hand them a screenshot or a document scan as part of the state. The mechanism is what matters: instead of generating tokens one at a time, Clef does a single prefill pass over your input, then runs a schema head that scores a fixed set of options at once. You ask it noul questions (a 0-1 probability), choice questions (pick one of up to 255 options), or score questions (place the input on a scale). There's no sampling and nothing to parse out of a text response, because the answer is already shaped like your schema.
That's the whole difference from an LLM call. An LLM generates a sequence and you hope the JSON it produces matches your schema. A decision model skips generation and scores the schema directly. Cloudflare built Clef to be request-compatible with Jev, so teams already calling Jev for classification can point the same code at Clef with minimal changes.
The latency gap is the real headline#
On Cloudflare's own benchmark suite, Clef-flash posted a 38.8ms median latency with a 122.4ms p95. Clef, the bigger model, came in at 209.3ms median. Jev, the comparable general-purpose model, measured 524.1ms median on the same suite, roughly an order of magnitude slower than Clef-flash. On accuracy, Clef scored a macro-F1 of 94.20 on the BANKING77 intent-classification benchmark against 79.74 for Jev, so the speed isn't costing correctness on tasks the model is built for.
Put a moderation check or a fraud-flag in the hot path of a user request, and a few hundred milliseconds of LLM round trip is often the single biggest latency line item on the page. Swap it for a call that resolves in 40ms and that line item disappears from your waterfall.
Where this genuinely replaces an LLM call#
The honest list of candidates is anything that is already a bounded decision, meaning the set of possible answers is fixed before the request arrives:
- Moderation: is this content policy-violating, yes/no or a severity score.
- Routing: which of our five support queues does this ticket belong to.
- Fraud and risk flags: does this transaction match a known-bad pattern, as a probability.
- Triage and prioritization: score this input on a 1-10 urgency scale.
- Intent classification: pick the right tool or workflow branch from a known list.
All of these share a property: you could write the output options into an enum before the request arrives. That's the test. If your prompt's job is "pick from this list" or "score this on this scale," you're already describing a choice or score question, and paying LLM prices for an answer a schema head could score directly.
A pseudo-example of the shape (verify exact binding syntax against current Workers AI docs before shipping):
// Illustrative: Workers AI decision-model pattern
const result = await env.AI.run("@cf/cloudflare/clef-flash", {
state: { text: ticketBody },
questions: {
queue: {
type: "choice",
options: ["billing", "bug-report", "feature-request", "abuse"],
instructions: "Which support queue does this ticket belong to?",
},
urgent: { type: "noul", instructions: "Probability this needs a reply within 1 hour." },
},
});
// result.answers.queue -> { choice: "bug-report", confidence: 0.91 }
Typed questions in, probabilities out, no prompt engineering to coax valid JSON out of a chat model.
Where you still need the LLM#
Decision models fail the moment the task stops being bounded. If you need the system to write a reply, summarize a document, explain its reasoning in prose, or handle a request where the valid answers aren't knowable in advance, a schema head has nothing to score against. Clef doesn't generate text, so anything downstream that needs language output still routes to a generation model. The honest tradeoff is that Clef is less flexible by design: you trade open-ended reasoning for speed and determinism, and that only pays off when the task was never open-ended to begin with. Teams running a mixed pipeline, say a moderation gate in front of a chat model, win most, because the gate is a decision and the chat turn is generation, and conflating the two into one LLM call was always doing the expensive job twice.
If you're already running multiple providers behind a gateway, this is a routing decision, not a provider decision: see our AI gateway comparison for how LiteLLM, Portkey, and Cloudflare's own AI Gateway handle that split, and our guide to multi-provider LLM routing and failover for keeping both model types reachable through one path.
The decision, concretely#
- Is the output a fixed set of options, a probability, or a scale score? Use a decision model like Clef or Clef-flash. You get roughly 10x the latency improvement and lose nothing on accuracy for classification-shaped tasks.
- Does the task require generating new text or handling genuinely open-ended input? Keep the LLM call. A schema head can't write a sentence it wasn't asked to score.
- Is latency the complaint on a classification step inside a larger pipeline? Replace just that step. Clef-flash in front of a generation model is a gate, not a replacement for the model doing the writing.
- Are you already paying a general LLM to do intent classification or moderation? That's the easiest win here: same accuracy class, a fraction of the latency, lower price per token.
The call we'd make#
Default to a decision model for any step where you could write the valid outputs into an enum today, and keep the LLM for every step that has to produce language. Clef-flash's 38.8ms median makes the case alone: that's a different latency class for work that was never generation to begin with. The mistake isn't picking the wrong model, it's not noticing that half of what's running through your LLM gateway was a classification call wearing a chat completion's clothes.
Get the DevOps Troubleshooting Cheat Sheet
Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.
Claude Code Mods: Programmable, Unsandboxed, and Your Problem in CI
Mods make Claude Code a real runtime you can program. They also run with your full native permissions, no sandbox, so a bad one is a bad CI step.
GitLab's New Rate Limits: What to Fix Before Oct 19
GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.
More from Infrastructure
Explore more articles in this category
Perplexity Left DynamoDB for CobbleDB: When Should You?
Perplexity built its own key-value store because DynamoDB's read path and bill stopped fitting 50 KB search items. Here is the checklist for when leaving is justified.
The Terraform Lock File Is Code: Review It Before You Init
A DPRK-linked group is mailing DevOps candidates Terraform take-home repos whose lock file points at a fake registry. terraform init then runs the attacker's provider.
Redis vs Memcached: Choosing a Cache in 2026
Both are fast in-memory stores, and both get picked by habit more than by requirements. Here is what actually differs and when each one is the right call.
You might have missed
Evergreen posts worth revisiting.