Clef Demo

2026-10-03

How to Run Clef Locally (27B VRAM, Hugging Face) vs Online Playground

Run Clef locally with Hugging Face weights — 27B VRAM expectations, setup tradeoffs, and when to try Clef online instead of local GPUs.

  • Full 27B in higher precision often wants a data-center class GPU (or multi-GPU), not a single consumer card.
  • Quantization helps but can change calibration of decision probabilities — validate against a golden set.
  • Clef-flash is the practical local/prototype size when full 27B VRAM is out of reach.
  • How to use Cloudflare Clef Workers AI is usually faster for iteration.
  • Our Clef Playground lets you try Clef online with the same state / noul / choice / score schema — no local GPU provisioning.
  • Cache identical playground runs to save tokens while you design prompts.
  • Clef vs Jev decision model
  • Clef Workers AI / agent workflow tutorial
  • What is Cloudflare Clef model?
  • State & question schema

Ready to test? Open the Clef Decision Playground.

This is an independent demo tool for the Clef model. Not affiliated with Cloudflare. Clef model available under Apache 2.0 license.