Cloudflare has released Clef and Clef-flash, its first open-weight decision models designed to handle structured classification tasks for AI agents without returning free-form text. Built on Qwen backbones and released under the Apache-2.0 license, the models are fully compatible with TypeSafe AI's Jev API. Cloudflare's self-reported benchmarks show that Clef-flash achieves a median latency of 38.8 milliseconds, significantly outperforming Jev's 524.1 milliseconds, though Jev retains its lead on knowledge-heavy reasoning benchmarks.
Model release and architecture
- ▪Clef is based on the Qwen3.8-27B model and Clef-flash is based on the Qwen3.5-9B model, with both keeping their respective base vision encoders
- ▪Cloudflare's Clef and Clef-flash models are available on Hugging Face under the Apache-2.0 license and run on Cloudflare's Workers AI platform
- ▪Cloudflare released Clef and Clef-flash, two open-weight decision models designed to classify inputs and return probabilities rather than free-form text
Training methodology for Clef
- ▪Cloudflare trained Clef and Clef-flash by freezing the base backbones and optimizing a joint schema head alongside rank-256 low-rank adapters
- ▪Cloudflare used its own variant of Reinforcement Learning for Calibrated Decisions to train Clef, a method that TypeSafe AI previously used for Jev
Compatibility and features compared to Jev
- ▪Clef and Clef-flash are fully compatible with TypeSafe AI's Jev API, allowing customers to switch models by changing the endpoint and model name from Jev to Clef
- ▪Cloudflare's Clef supports a 65,536-token context window and image inputs, whereas TypeSafe AI's Jev is limited to text inputs and a 32,000-token context window
Latency and processing speed
- ▪In a threat intelligence workflow, Clef classified a domain in 2.2 seconds, whereas a general-purpose language model, gpt-oss-120b, took 4.7 seconds
- ▪Cloudflare reported a median latency of 38.8 milliseconds for Clef-flash and 209.3 milliseconds for Clef, compared to 524.1 milliseconds for TypeSafe AI's Jev
Benchmark performance results
- ▪In Cloudflare's self-reported benchmarks, Clef-flash achieved a 97.73% exact-case score on the Home Appliances benchmark from Cloudflare's Decision Index suite compared to 52.27% for Jev
- ▪TypeSafe AI's Jev model maintains performance leads over Cloudflare's Clef on knowledge-heavy benchmarks, including GPQA Diamond, MMLU-Pro, and Big Bench Hard
- ▪Clef scored highest on 7 of 10 benchmarks in Cloudflare's 10-benchmark shortlist from the Decision Index 0.2.1 suite, including BANKING77 and CLINC150+OOS
Debatable claims
- ▪Specialized decision models are superior to general-purpose language models for automated workflows
- ▪Critical AI decision models should be released under open-source licenses
- ▪Low latency is more important than deep reasoning capability for operational AI agents
- ▪AI agents no longer require human oversight for operational decision-making
Story comments
Loading comments…