d3-lite — Core ML
vllm-sr/d3-lite (0.85B, part of Decision 3.0 by the vLLM Semantic Router team) converted to Core ML for Apple silicon: text, image and video requests on device. Same contract as upstream: one forward pass per question, the 255-way FP32 readout at the last prompt token, a probability for every option, no text generated.
Files
| File | What |
|---|---|
d3_E256 / d3_E640 / d3_E1536.mlpackage |
text graph (fp16 compute, int8 weights) per prompt-length bucket; input = token embeddings with image/video tokens spliced in |
vision_P320 / vision_P1536 / vision_P3072.mlpackage |
vision tower (fp32) per patch bucket; video frames go through it one frame pair at a time (block mask) |
embeddings.f16, pos_embed.f32, readout.f32, tokenizer.json, config.json, coreml_config.json |
host-side tables and config |
python/ |
Python host (coremltools): d3_coreml_text.py, d3_coreml_image.py, d3_coreml_video.py |
Swift: D3VisionManager in FluidUse loads this folder.
Checked against upstream
Same weights as upstream v3.1.0. Answers that differ from the upstream PyTorch runtime (bf16, MPS), same requests:
| flips | |
|---|---|
| Text (typed-decisions TEST, 20 requests) | 2 / 100 |
| Images (receipt checks) | 1 / 24 |
| Video (4-second street clips, 2 fps) | 0 / 28 |
Speed on an M5 Pro (24 GB), CPU_AND_GPU: ~46 ms per text question, 0.32 s per 4-second video look.
Python
from d3_coreml_text import D3CoreMLText # d3_coreml_image.D3CoreMLImage, d3_coreml_video.D3CoreMLVideo
m = D3CoreMLText("<vllm-sr/d3-lite snapshot>", "<this repo>")
m.system_one(state="...", questions={"route": {"type": "choice", "criteria": {...}}})
The Python host builds prompts with upstream's own runtime (d3_runtime.py in the upstream snapshot, run without weights), so prompts match upstream exactly.
License: Apache-2.0, as upstream.
- Downloads last month
- -