# Cloudflare Clef adds vision and open weights to the decision-model race

Cloudflare has released Clef and Clef-flash as Apache-2.0 decision models that return typed probabilities rather than prose, adding image input, a 64k context window, Workers AI hosting and a path to workload-specific RL fine-tuning.

Bounded decision models are turning into a real model category. Cloudflare's entry is open-weight, multimodal and Jev-API compatible, while its fastest variant is aimed at latency-sensitive agent routing.

- Status: Active
- Published: 2026-10-04T09:22:05+13:00
- Updated: 2026-10-04T09:22:05+13:00
- Categories: Artificial Intelligence, AI Models, AI Agents, Open Models, Inference & APIs
- Tags: agent routing, Clef, Cloudflare, decision models, open models, Workers AI
- Canonical HTML: https://beyondthe.news/dossiers/cloudflare-clef-open-multimodal-decision-models-workers-ai

## What changed

On October 1, 2026 Cloudflare released Clef and Clef-flash, its first in-house decision models from the Workers AI team. Both take state plus typed questions and score allowed answers instead of generating free-form text. The weights are available under Apache 2.0 and the models are hosted on Workers AI. Unlike Jev's current text-only interface, Clef includes a vision encoder and supports image inputs; Cloudflare also gives the family a 64k context window. The company is pairing the release with a hands-on reinforcement-learning fine-tuning service intended to evolve into a self-service platform.

## Why it matters

The interesting shift is not another classifier benchmark. Multiple independent teams are now designing models specifically for software decisions rather than human conversation. Clef makes that layer deployable either through Workers AI or open weights, adds visual state to the decision loop, and gives builders an API-compatible alternative to Jev. Cloudflare's fine-tuning path also links production traffic, rollouts, sandboxed scoring and redeployment into one model-customisation workflow.

## Clef returns probabilities, not prose

The models answer bounded yes/no, choice and score questions with typed probability outputs. That makes them suitable for routing, triage, policy and agent decisions where downstream software needs a finite answer rather than generated text.

## Vision is the clearest capability difference

Cloudflare says Clef can classify images as well as text and supports a 64k context window. That broadens the state an agent can pass into a bounded decision without falling back to a general multimodal LLM.

## The speed claims are promising but vendor-produced

Across Cloudflare's 43-evaluation set it reports median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash, versus 524.1 ms for Jev. Benchmark quality varies by task, and these measurements need independent replication.

## Cloudflare is building a fine-tuning loop around the model

The initial RL service uses AI Gateway traffic capture, Workers AI rollouts, Containers for scoring and replay, a new Trainer component, and Workers AI/BYO Model for redeployment. It starts as a hands-on service rather than a self-service product.

## Key details

- Clef and Clef-flash launched October 1, 2026.
- Both are Apache-2.0 open-weight decision models and run on Workers AI.
- They return typed probabilities over predefined answers instead of free-form text.
- Clef supports image input and a 64k context window.
- Cloudflare says the API is compatible with Jev.
- Cloudflare reports 38.8 ms median latency for Clef-flash in its own 43-evaluation set; this is vendor-produced.
- A hands-on RL fine-tuning service is available, with self-service tooling planned later.

## Builder takeaways

- Use a bounded decider when the application needs a finite choice or calibrated score, not generated prose.
- The Jev-compatible API lowers the cost of benchmarking Clef against an existing decision-model integration.
- Image input makes Clef worth testing for moderation, visual routing and other workflows where state is not text-only.
- Benchmark on your own workload before accepting vendor latency or quality rankings.
- Treat the RL offering as early-stage: the current path involves Cloudflare's forward-deployed team rather than a finished self-service workflow.

## What to watch

- Independent benchmark replication across Jev, Strands Decider, CLM and GLiNER2.5-Decide.
- Workers AI pricing and production economics for sustained decision workloads.
- The self-service release of Cloudflare's Trainer/RL platform.
- Real deployments using image state rather than text-only classification.
- Whether Jev API compatibility becomes an informal interoperability layer across decision models.

## Uncertainties

- Cloudflare's benchmark suite and latency figures are vendor-produced.
- The self-service RL platform is not yet generally available.
- Production adoption and retention are not yet established.

## Sources

- [Introducing Clef: our open-source decision models, and new RL fine-tuning platform](https://blog.cloudflare.com/clef-decision-models/) — Cloudflare · primary · 2026-10-01T00:00:00+13:00. Primary release, architecture, benchmark and fine-tuning details.

