clef-text-0.6b (Core ML)
A 0.70B-parameter text decision model for Apple silicon, distilled from Cloudflare's
clef-flash (9B) into Qwen3-0.6B
plus Clef's joint schema head. Same contract as Clef and the Jev / SystemOne API: a state and typed questions
(noul, choice, score) in, a probability for every option of every question out, one forward pass.
Runs on the GPU or 100% on the Neural Engine (every decoder op placed on the ANE at all four buckets). PyTorch weights:
FluidInference/clef-text-0.6b. Swift runtime:
ClefTextManager in FluidUse.
Packages
| File | Role | Size |
|---|---|---|
Decoder.mlpackage |
Qwen3-0.6B decoder + final norm, fp16; functions L256 L512 L1024 L2048 (shared weights) |
848 MB |
Head.mlpackage |
joint schema head, fp16, ≤ 16 questions / ≤ 96 options, same four functions | 197 MB |
embeddings.f16 |
token embeddings (tied: also the head's lexical option rows) | 297 MB |
tokenizer.json, config.json |
tokenizer; shapes and ids for the host |
Decoder inputs: hidden [1, L, 1024], cos / sin [L, 128], additive causal mask [1, 1, L, L] (all fp16) →
states [1, L, 1024].
Quality
Held-out test splits (official test / validation sets, 5,124 questions), gold accuracy:
| Task | this model | clef-flash 9B (teacher) |
|---|---|---|
| DBpedia-14 entity type | 99.7 | 100.0 |
| AG News topic | 90.7 | 92.0 |
| SST-2 sentiment | 90.0 | 94.0 |
| BANKING77 intent (77-way) | 86.4 | 94.5 |
| TweetEval offensive | 85.0 | 82.3 |
| Yelp stars (5-way) | 69.0 | 70.7 |
| Emotion (6-way) | 68.0 | 58.0 |
| BoolQ | 82.3 | 90.7 |
| MNLI | 72.3 | 84.3 |
| ARC-Easy | 80.0 | 100.0 |
| CommonsenseQA | 65.0 | 94.3 |
| ARC-Challenge | 60.7 | 97.3 |
| all | 79.1 | 88.2 |
Good at routing and classification (within 0–4 points of the teacher on topic, entity, sentiment, rating, toxicity, emotion); not a knowledge model — science / commonsense QA and NLI stay well below the 9B. Teacher agreement 86.4%, calibration ECE 0.046 (dev).
Core ML on the Neural Engine vs the PyTorch student on the same 5,118 test questions: 15 different answers (0.3%, all near-ties: PyTorch top-2 within 0.11), median probability drift 0.013, gold accuracy equal within ±0.7 per task.
Speed (M5 Pro)
Decoder at 512 tokens: 35 ms on the GPU (.cpuAndGPU), 64 ms on the Neural Engine (.cpuAndNeuralEngine),
200 ms CPU-only. On this Mac the GPU is ~2× faster; the Neural Engine leaves the GPU free for other work.
| Bucket | Neural Engine, decoder + head |
|---|---|
| 256 tokens | 29 ms |
| 512 tokens | 69 ms |
| 1,024 tokens | 196 ms |
| 2,048 tokens | 591 ms |
Swift runtime, 3-question support ticket (370 tokens): **48 ms on the GPU** (1,000 tickets in 57 s in the FluidUse
demo), ~77 ms on the Neural Engine — vs ~0.5 s for clef-flash Core ML
(9B, GPU).
Training
24,208 records from 12 public datasets (train splits; test splits held out) labelled by clef-flash (8-bit Core ML, 1 flip / 611 vs fp32), KL (T = 2) + 0.3 gold cross-entropy, LoRA r64 on the backbone + full head, 2 epochs. The head starts from clef-flash's head weights wherever the shapes match.
License and credits
Apache-2.0. Teacher: Cloudflare/clef-flash (Apache-2.0); base: Qwen/Qwen3-0.6B (Apache-2.0). By Fluid Inference.
- Downloads last month
- -