~8m14:09Clef: Open-source decision models, and new RL fine-tuning platform
Oct 2, 2026
Read: ~7m · You save: 7 min
Clef: Open-source decision models, and new RL fine-tuning platform
Cloudflare unveils Cleft & Cleft Flash: ultra-fast open-source decision models on Workers AI. Discover their speed, calibration, and RL fine-tuning platform!
Cloudflare released two open-source decision models, Clef and Clef Flash, on October 1, 2026, alongside a new reinforcement learning (RL) fine-tuning platform. These models are designed to provide typed, probability-scored answers against a defined schema, differentiating them from traditional language models that generate free-form text. The announcement, published on Cloudflare's blog, details the models' performance, architecture, and potential applications.
Performance Metrics
The released models demonstrate significant latency improvements compared to existing solutions. Clef achieves a median latency of 209.3 milliseconds, while Clef Flash records a median latency of 38.8 milliseconds. For comparison, Jev from Typesafe AI has a median latency of 524.1 milliseconds. At the P95 metric, Clef's latency is 238.6 milliseconds, and Clef Flash's is 122.4 milliseconds. Cloudflare highlights that Clef Flash's P95 latency is roughly triple its median, but still under an eighth of a second, indicating a predictable performance band. This speed advantage allows for decision models to be integrated into loops within agent workflows, whereas slower models necessitate user-facing spinners.
Architectural Innovations
Clef and Clef Flash are not trained from scratch. Cloudflare states they utilize QN as the base model, post-trained for decision use cases. QN3A-27B is frozen for Clef, and QN35-9B is frozen for Clef Flash. This freezing process means the backbone weights remain unchanged. The training jointly optimized a routing head alongside rank 256 low-rank adapters.
A key architectural feature contributing to Clef's speed is its inference process. It performs a prefill-only pass with QN, followed by scoring valid schema choices in parallel. Unlike autoregressive LLMs that generate text token by token, Clef skips the decode phase entirely. Cloudflare describes this as a non-autoregressive step, significantly accelerating the decision-making process. The architecture is characterized as a two-stage attention routing process, combining option-specific evidence routing, joint cross-field attention, and schema-bound scoring with a lexical prior. Currently, no detailed paper or reproducible details have been published alongside this description.
Calibration and Optimization
The probability scores provided by Clef are calibrated through post-training using label-smoothed cross-entropy and Brier loss. Cross-entropy trains the model to select the correct answer, while Brier loss ensures the model is honest about its confidence levels. Cloudflare also developed Reinforcement Learning for Calibrated Decisions (RLCD) as a secondary optimization target. RLCD provides partial credit to adjacent ordinal choices, rewards precise record outputs, and applies a penalty to prevent distribution shifts. This approach ensures that confidence curves are meaningful, treating near misses as such rather than complete failures.
Evaluation Benchmarks
Cloudflare conducted 43 evaluation benchmarks. On the Berkeley Function Calling Leaderboard (BFCL) exact match, Clef achieved 98.47% and Clef Flash 98.76%, compared to Jev's 95.75%. On tool NDCG@10, Clef scored 69.19% and Clef Flash 66.43%, against Jev's 65.28%. For API bank accuracy, Clef reached 91.93% and Clef Flash 93.11%, with Jev at 88.19%.
In the home appliances case exact match, Clef achieved 82.95% and Clef Flash 97.73%, significantly outperforming Jev (52.27%), Diffusion Gemma (42.05%), and Kev 9B (25.00%). On banking 77 macro F1, Clef scored 94.20% and Clef Flash 90.93%, compared to Jev's 79.74%. On CLNC50+OS macro F1, Clef achieved 97.43%, while Clef Flash scored 66.77% and Jev 89.27%. Cloudflare notes that on CLNC50+OS, the smaller model's performance drops significantly, while the larger model achieves the top score.
Cloudflare also acknowledges benchmarks where Jev leads: when-to-call accuracy (Jev 80.97% vs. Clef 72.37%) and Bright NDCG@10 (Jev 47.52% vs. Clef 45.91%). Diffusion Gemma leads fish and chips accuracy at 85.35% against Clef's 79.60%.
On Typesafe's own evaluation suite, Cloudflare states Clef won three out of four areas: invoice processing (64.7% vs. Jev 61.8%), customer service (76.3% vs. Jev 76.0%), and security incidents (62.9% vs. Jev 61.7%). Jev led agent trace observability at 71.6% against Clef's 68.5%. Cloudflare also claims Clef is currently the leader on the Jev decision index, with results available on a live demo site.
Feature Parity and API Compatibility
Clef offers a 64K context window, double that of Jev's 32K. Additionally, Clef includes a vision encoder capable of classifying visual content, a feature stated to be absent in Jev's current text-only classification capabilities. The models are fully Jev API compatible, allowing for integration as a configuration change rather than a complete rewrite. The API call involves an HTTP POST request to api.cloudflare.com/client/v4/accounts/{your_cloudflare_account_id}/run/cf/cloudflare/clef. The request body includes a state field and questions, which can be of types null, choice with named criteria, or score with an ordered criteria list.
Internal Testing and Use Cases
Cloudflare's threat intelligence team utilized Clef Flash for website domain classification within a workflow called "browser run." In this test, Clef Flash took 2.2 seconds to fetch, render, and classify a website. For comparison, GPT-O120B, described as Cloudflare's fastest general LLM, took 4.7 seconds and returned only two classifications. This internal test, while not a published benchmark, suggests a potential for significant latency and result savings. Cloudflare frames this as a roughly 2x saving in both latency and the number of classifications returned.
Fine-Tuning Platform
Cloudflare has introduced a reinforcement learning fine-tuning platform designed to enhance model performance for specific tasks. The pipeline includes five components: AI Gateway for data capture from traffic, Workers AI for generating rollouts against base Clef, Cloudflare containers as an RL sandbox, a new trainer component for weight updates, and Workers AI plus BYO (Bring Your Own) model redeploys. The BYO model work is labeled COG, stemming from Cloudflare's acquisition of Replicate.
Potential internal use cases for this fine-tuning include evaluating trust and safety submissions, triaging support requests, and classifying bots in their bot products. Cloudflare cites over 15 years of network data as fine-tuning material. The company acknowledges that fine-tuning may reduce general-purpose performance in favor of higher accuracy within a specific domain. The claim that a fine-tuned Clef is more accurate and faster than the generic model is asserted without supporting numbers.
Open Source and Availability
Both Clef and Clef Flash models are open-sourced on HuggingFace under the Apache 2.0 license, permitting modification and redistribution for commercial use. The weights are available for local execution. The hosted endpoint is documented on the Workers AI platform, which provides GPU inference across Cloudflare's global edge. Cloudflare has also implemented a policy stating that customer requests or responses are not read, stored, or trained on unless the fine-tuning product is used.
Unanswered Questions
Several details remain unpublished, including pricing for the models and the self-served platform, the general availability status of both models, specific parameter counts and hardware requirements for local deployment, the exact ship date for the trainer component, whether vision capabilities are live in the hosted API, image limits, rate limits, and the maximum number of questions per call. The specifics of the "browser run" workflow and customer ownership of weights trained by Cloudflare's FTE team are also not detailed. Cloudflare is inviting existing customers to participate as design partners for the fine-tuning and RL work, suggesting that some of these answers will emerge from that cohort.
Target Audience and Market Impact
Clef is positioned for use cases involving agents that currently rely on expensive general-purpose reasoning models for tasks such as routing, triage, moderation, domain classification, and severity scoring. The trend towards smaller, specialized classifiers replacing costly general calls is noted. Cloudflare's release of a calibrated, typed, API-compatible model with open-source weights is seen as a significant development in this area. Clef is described as Cloudflare's first Cloudflare-trained ML model from the Workers AI team, with the potential to disrupt agent utilization and align with Cloudflare's mission as an "agent cloud." The name "Clef" is derived from music theory, signifying the assignment of pitch names to a staff, analogous to the model's function of turning raw context into labeled, scored meaning.
The article concludes by emphasizing that Clef represents a decision model where speed, calibration methods, and licensing align. The median latency of 38.8 milliseconds, Brier calibration, and Apache 2.0 license with open weights are highlighted. Readers are encouraged to verify the published numbers through independent reruns on HuggingFace and the Workers AI platform.
Introduction and Latency Comparison
Cloudflare has released two open-source decision models, Cleft and Cleft Flash, on Workers AI, boasting significantly lower latency (38.8ms median for Cleft Flash) compared to competitors like Jev (524.1ms median). These models are Jev API compatible and available on HuggingFace under Apache 2.0. A new reinforcement learning fine-tuning platform is also introduced.
- Cloudflare released Cleft and Cleft Flash decision models on October 1st, 2026.
- Cleft Flash achieved a median latency of 38.8 milliseconds.
- Jev from Typesafe AI had a median latency of 524.1 milliseconds.
- Both models are open-sourced on HuggingFace under Apache 2.0.
- They are described as fully Jev API compatible.
- A new reinforcement learning fine-tuning platform is also launched.
What is a Decision Model?
Decision models differ from language models by returning typed, probability-scored answers against a defined schema, rather than free-form text. Cloudflare's example shows a domain classification with probabilities for 'fashion website', 'e-commerce', and 'fishing'. The latency table highlights Cleft at 209.3ms, Cleft Flash at 38.8ms, and Jev at 524.1ms for median latency, and provides P95 figures as well.
- Decision models return typed, probability-scored answers against a schema.
- Normal language models return free-form text.
- Example: Domain classification with probabilities for fashion (95%), e-commerce (85%), fishing (under 1%).
- Median latencies: Cleft (209.3ms), Cleft Flash (38.8ms), Jev (524.1ms).
- P95 latencies are also provided for comparison.
The Mechanism: How Cleft Achieves Speed
Cleft is built upon QN3A-27B (for Cleft) and QN35-9B (for Cleft Flash) base models, with their weights frozen. It uses a novel non-autoregressive approach, performing a prefill-only pass with QN and then scoring schema choices in parallel, avoiding the token-by-token decode step typical of LLMs. This architecture significantly reduces latency.
- Cleft uses QN3A-27B as the base model, with weights frozen.
- Cleft Flash uses QN35-9B as the base model, with weights frozen.
- Training jointly optimized a routing head alongside low-rank adapters.
- Inference uses a prefill-only pass with QN, skipping the autoregressive decode.
- The decision step is non-autoregressive, making it significantly faster.
- Architecture involves a two-stage attention routing process.
Calibration and RL Fine-Tuning
Calibration is crucial for the reliability of decision model probabilities. Cloudflare used label-smoothed cross-entropy and Brier loss during post-training on synthetic data to ensure outputs are valid against the schema and probabilities are honest. Reinforcement Learning for Calibrated Decisions (RLCD) was developed to provide partial credit for adjacent ordinal choices and reward precise outputs.
- Post-training used label-smoothed cross-entropy for schema validity.
- Brier loss was used for probability calibration on synthetic data.
- Aims for models to be honest about confidence (e.g., 95% should be correct ~95% of the time).
- Introduced Reinforcement Learning for Calibrated Decisions (RLCD).
- RLCD gives partial credit to adjacent ordinal choices.
- RLCD rewards fully precise record outputs and penalizes distribution shift.
Benchmark Performance
Cloudflare presents benchmark results across various tasks, including BFCL (function calling), NDCG, API bank accuracy, and intent classification (Banking 77, CLNC50). Cleft and Cleft Flash generally outperform Jev, though Jev leads in specific metrics like 'when to call' accuracy. Cloudflare also highlights benchmarks it does not lead, such as Diffusion Gemma on 'fish and chips' accuracy.
- 43 evaluation benchmarks were run.
- On BFCL case exact: Cleft (98.47%), Cleft Flash (98.76%), Jev (95.75%).
- On tool NDCG at 10: Cleft (69.19%), Cleft Flash (66.43%), Jev (65.28%).
- On API bank accuracy: Cleft (91.93%), Cleft Flash (93.11%), Jev (88.19%).
- Jev scores higher on 'when to call' accuracy (80.97% vs 72.37%).
- Cloudflare published benchmarks where it does not lead.
Integration and API Compatibility
Cleft offers a 64K context window and a vision encoder, making it a drop-in replacement for Jev (which has a 32K context window and text-only classification). Full Jev API compatibility simplifies integration. The API call involves an HTTP POST to a Cloudflare endpoint with state and questions, supporting various question types like null, choice, and score.
- Cleft has a 64K context window, double Jev's 32K.
- Cleft includes a vision encoder for image classification.
- Jev is described as text classification only.
- Full Jev API compatibility allows for easy swapping.
- API call is an HTTP POST to
api.cloudflare.com/client/v4/accounts/{account_id}/run/cf/cloudflare/cliff. - Request body includes 'state' and 'questions' (null, choice, score types).
Internal Test Case: Website Classification
An internal test showed Cleft taking 2.2 seconds to classify a website (fetch, render, classify), compared to GPT-O1 120B taking 4.7 seconds and returning fewer classifications. This internal test, while not a published benchmark, demonstrates the practical benefits of Cleft's speed, with most of the time attributed to the browser rendering rather than the model inference.
- Internal test: Cloudflare's threat intelligence team used Cleft for website domain classification.
- Cleft took 2.2 seconds (fetch, render, classify).
- GPT-O1 120B took 4.7 seconds and returned only two classifications.
- The 2.2s figure includes fetching and rendering a real page.
- This demonstrates a roughly 2x saving in latency and results.
- Most of the time is spent on browser rendering, not model inference.
The Fine-Tuning Platform
The fine-tuning loop involves AI Gateway for data capture, Workers AI for rollouts, Cloudflare Containers for RL sandbox, a new 'trainer' component for weight updates, and Workers AI with BYO model for redeployment. This pipeline, leveraging Cloudflare's acquisition of Replicate (for COG tooling), aims to improve accuracy in specific domains like trust and safety, support triaging, and bot detection.
- Five components in the fine-tuning pipeline: AI Gateway, Workers AI, Cloudflare Containers, Trainer, BYO model redeployment.
- Leverages COG tooling from Cloudflare's acquisition of Replicate.
- Targets: Evaluating trust and safety, triaging support requests, bot detection.
- Utilizes over 15 years of network data for fine-tuning.
- Fine-tuning may trade general-purpose performance for domain-specific accuracy.
- Builds on independent research, including contributions to VLM.
Availability and Licensing
The Cleft models are available on HuggingFace under Apache 2.0, allowing permissive commercial use and modification. The hosted endpoint is documented and uses the Workers AI platform for GPU inference globally. Cloudflare guarantees that customer data is not read, stored, or trained on unless the fine-tuning product is used.
- Weights are available on HuggingFace under Apache 2.0 license.
- Permissive commercial use, modification, and redistribution are allowed.
- Hosted endpoint uses the Workers AI platform for global GPU inference.
- Cloudflare guarantees data privacy unless fine-tuning is used.
Information Yet to Be Published
Key details yet to be published include pricing, general availability status (beta or GA), specific parameter counts and hardware requirements for local deployment, the ship date/price for the self-served platform, vision API live status, image limits, rate limits, and maximum questions per call. Cloudflare is also seeking design partners for the fine-tuning and RL work.
- Pricing is not yet published.
- General availability status (beta/GA) is unclear.
- Self-served platform ship date and price are unknown.
- Vision API live status and limits are not specified.
- Rate limits and max questions per call are not detailed.
- Cloudflare is inviting existing customers as design partners.
Target Audience and Positioning
Cleft is designed for agents that currently use general LLMs for tasks like routing, triage, and moderation, offering a faster, calibrated, and typed alternative. Cloudflare positions Cleft as a disruptive technology in the agent space, emphasizing its open-source nature and performance. The name 'Cleft' is inspired by music theory, signifying the model's role in interpreting and labeling data.
- Target users: Agents performing routing, triage, moderation, classification, scoring.
- Replaces expensive general LLM calls with smaller, specialized classifiers.
- Cloudflare calls Cleft its first Cloudflare-trained ML model from Workers AI team.
- Believes Cleft can disrupt agent usage.
- Name 'Cleft' relates to music theory, signifying pitch/meaning assignment.
- CF in Cleft refers to Cloudflare.
Verdict and Future Watch Points
In summary, Cleft represents a significant advancement in decision models, offering high speed (38.8ms median), Brier-calibrated probabilities, and an Apache 2.0 license. The key indicators to watch are pricing, detailed HuggingFace repository information, and independent benchmark reruns. The falsifiable nature of the published numbers makes this release particularly interesting.
- Cleft is the first decision model with speed, calibration, and license alignment.
- Key metrics: 38.8ms median latency, Brier calibrated probabilities, Apache 2.0 license.
- Future watch points: Pricing, HuggingFace repo details (parameter counts), independent benchmark reruns.
- Published numbers are falsifiable, making the release interesting.