Clef is not designed to write a polished paragraph. It is designed to return typed probabilities and scores from text, JSON, images, or video.
Cloudflare released Clef on October 1 as a family of decision models. Clef has 27 billion parameters, and Clef-flash has 9 billion. Both run on Workers AI, and Cloudflare has published the model weights under the Apache 2.0 license.
What makes a decision model different
The hosted Clef documentation accepts a state expressed as text or structured data and a set of typed questions. It supports null, choice, and score forms, then returns probabilities rather than only free-form prose. That shape fits routing, moderation, ranking, and policy checks where the application needs a defined answer space.
The model also accepts images and video. Cloudflare documents a 65,536-token context window, up to four hosted images and up to 64 questions in one request. Those ceilings don’t guarantee accuracy at the limit. Test the actual mix of text, image, and question types your product uses.
Price and licensing change the deployment choice
Cloudflare lists hosted Clef input at $0.24 per million tokens on Workers AI. The published Clef and Clef-flash weights allow teams to evaluate local or other hosted deployments under Apache 2.0. Self-hosting replaces a token price with hardware, engineering, monitoring, and availability costs.
Cloudflare identifies Qwen3.8-27B as the base for Clef and Qwen3.5-9B as the base for Clef-flash. The models use a typed-decision interface compatible with the Jev/System One API, which makes comparative testing easier than translating two unrelated schemas.
Do not turn the leaderboard into an independent verdict
Cloudflare reports favorable benchmark results against Jev. Those figures are company-reported. Our Jev guide explains that the smaller model is optimized for a narrow structured-decision job. A fair comparison must use the same prompts, labels, held-out cases, calibration metric, and hardware.
The same caution applies to Laya comparisons with Jev. Published numbers collected from different runs do not establish a winner. Choice-set size, multilingual inputs, fine-tuning, and abstention policy can reverse the result.
A reproducible CLEF evaluation
- Freeze the input schema, label set, and exact model versions.
- Create a held-out test set that resembles production decisions.
- Measure accuracy, calibration, abstention behavior, and latency.
- Report results by language, input type, and choice-set size.
- Count infrastructure and review cost per accepted decision.
- Publish failure examples, not only an aggregate score.
Clef expands the small but useful category of models that return structured decisions instead of essays. The open-weights and hosted options make it practical to test. The responsible claim is that it is testable, not that a vendor chart has already settled which model is best.