d3-lite — Core ML

vllm-sr/d3-lite (0.85B, part of Decision 3.0 by the vLLM Semantic Router team) converted to Core ML for Apple silicon: text, image and video requests on device. Same contract as upstream: one forward pass per question, the 255-way FP32 readout at the last prompt token, a probability for every option, no text generated.

Files

File What
d3_E256 / d3_E640 / d3_E1536.mlpackage text graph (fp16 compute, int8 weights) per prompt-length bucket; input = token embeddings with image/video tokens spliced in
vision_P320 / vision_P1536 / vision_P3072.mlpackage vision tower (fp32) per patch bucket; video frames go through it one frame pair at a time (block mask)
embeddings.f16, pos_embed.f32, readout.f32, tokenizer.json, config.json, coreml_config.json host-side tables and config
python/ Python host (coremltools): d3_coreml_text.py, d3_coreml_image.py, d3_coreml_video.py

Swift: D3VisionManager in FluidUse loads this folder.

Checked against upstream

Same weights as upstream v3.1.0. Answers that differ from the upstream PyTorch runtime (bf16, MPS), same requests:

flips
Text (typed-decisions TEST, 20 requests) 2 / 100
Images (receipt checks) 1 / 24
Video (4-second street clips, 2 fps) 0 / 28

Speed on an M5 Pro (24 GB), CPU_AND_GPU: ~46 ms per text question, 0.32 s per 4-second video look.

Python

from d3_coreml_text import D3CoreMLText      # d3_coreml_image.D3CoreMLImage, d3_coreml_video.D3CoreMLVideo
m = D3CoreMLText("<vllm-sr/d3-lite snapshot>", "<this repo>")
m.system_one(state="...", questions={"route": {"type": "choice", "criteria": {...}}})

The Python host builds prompts with upstream's own runtime (d3_runtime.py in the upstream snapshot, run without weights), so prompts match upstream exactly.

License: Apache-2.0, as upstream.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FluidInference/d3-lite-coreml

Finetuned
vllm-sr/d3-lite
Quantized
(5)
this model

Collection including FluidInference/d3-lite-coreml