Wald-4B GGUF

GGUF files of Wald-Q4B v2.1 (tag v2.1) for llama.cpp: an open 4B decision model that returns a calibrated probability for every option.

File Size Same answer as BF16 (JevBench public, one pass)
Wald-4B-v2.1-Q8_0.gguf 4.5 GB 231/231
Wald-4B-v2.1-Q6_K.gguf 3.5 GB 226/231
Wald-4B-v2.1-Q5_K_M.gguf 3.1 GB 226/231
Wald-4B-v2.1-Q4_K_M.gguf 2.7 GB 225/231
Wald-4B-v2-mmproj-F16.gguf (vision projector for llama.cpp's multimodal tools; the vision tower is unchanged since v2, so this file serves v2.1 too) 0.67 GB

Quick start

hf download org2ai/Wald-4B-GGUF --include "Wald-4B-v2.1-Q4_K_M.gguf" "v2.1/*" --local-dir ./wald-gguf
pip install "./wald-gguf/v2.1/server" transformers
wald-serve-native --gguf ./wald-gguf/Wald-4B-v2.1-Q4_K_M.gguf --tokenizer-dir ./wald-gguf/v2.1 --effort none --port 8000

How it works

  • One pass, text only: the bundled server reads a probability for every option from the model's option-letter logits (Qwen chat template), calibrated with the same frozen temperature table as the BF16 release.
  • --effort auto (Auto 0.7) also runs here but has not been parity-checked on GGUF; for images, use the BF16 release.
  • On a 1,000-request Decision Index sample (5,922 questions), Q8_0 gives the BF16 answer on 99.5 % of questions and Q4_K_M on 97.3 %. All numbers: evaluation/v2.1-gguf-parity.json.

Earlier versions

  • v2 (Wald-4B-v2-*.gguf, v2/): Wald-Q4B v2, unchanged in this repository.
  • v1.2 (Wald-4B-v1.2-*.gguf, server/): unchanged in this repository.

Apache-2.0.

Downloads last month
15,505
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for org2ai/Wald-4B-GGUF

Finetuned
Qwen/Qwen3.5-4B
Finetuned
org2ai/Wald-4B
Quantized
(2)
this model

Space using org2ai/Wald-4B-GGUF 1

Collection including org2ai/Wald-4B-GGUF