What changed
On October 1, 2026 Cloudflare released Clef and Clef-flash, its first in-house decision models from the Workers AI team. Both take state plus typed questions and score allowed answers instead of generating free-form text. The weights are available under Apache 2.0 and the models are hosted on Workers AI. Unlike Jev's current text-only interface, Clef includes a vision encoder and supports image inputs; Cloudflare also gives the family a 64k context window. The company is pairing the release with a hands-on reinforcement-learning fine-tuning service intended to evolve into a self-service platform.
Why it matters
The interesting shift is not another classifier benchmark. Multiple independent teams are now designing models specifically for software decisions rather than human conversation. Clef makes that layer deployable either through Workers AI or open weights, adds visual state to the decision loop, and gives builders an API-compatible alternative to Jev. Cloudflare's fine-tuning path also links production traffic, rollouts, sandboxed scoring and redeployment into one model-customisation workflow.
Clef returns probabilities, not prose
The models answer bounded yes/no, choice and score questions with typed probability outputs. That makes them suitable for routing, triage, policy and agent decisions where downstream software needs a finite answer rather than generated text.
Vision is the clearest capability difference
Cloudflare says Clef can classify images as well as text and supports a 64k context window. That broadens the state an agent can pass into a bounded decision without falling back to a general multimodal LLM.
The speed claims are promising but vendor-produced
Across Cloudflare's 43-evaluation set it reports median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash, versus 524.1 ms for Jev. Benchmark quality varies by task, and these measurements need independent replication.
Cloudflare is building a fine-tuning loop around the model
The initial RL service uses AI Gateway traffic capture, Workers AI rollouts, Containers for scoring and replay, a new Trainer component, and Workers AI/BYO Model for redeployment. It starts as a hands-on service rather than a self-service product.