KV-cache quantization without any fork (recommended, 2026): upstream llama.cpp/Ollama now cover this natively — use -ctk q8_0 -ctv q8_0 (half KV memory, negligible quality loss: perplexity +0.002–0.05) or -ctk q4_0 -ctv q4_0 (quarter memory, ≈7.6% perplexity increase). In Ollama: OLLAMA_KV_CACHE_TYPE=q8_0 with OLLAMA_FLASH_ATTENTION=1. Keep K and V types symmetric to stay on the fast fused Flash-Attention path. Since April 2026, mainline llama.cpp also applies Hadamard rotation to KV activations (PR #21038), which greatly improves low-bit KV quality (opt-out: LLAMA_ATTN_ROT_DISABLE=1).

The RotorQuant/TurboQuant fork flow below is experimental/legacy: the TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork is unmaintained relative to mainline. It is NOT required to use this model.

Qwen3.5-27B-RotorQuant-MLX-2bit

MLX 2-bit weight quantization + RotorQuant 2-bit KV cache compression for Qwen/Qwen3.5-27B.

Dual compression for Apple Silicon: both the model weights and the KV cache are quantized to 2-bit, enabling long-context inference on memory-constrained Macs.

Overview

Qwen3.5-27B is a 27B-parameter hybrid transformer with 262K native context and built-in thinking mode (the model generates internal reasoning tokens before answering). Thinking mode makes KV cache compression especially valuable, since the reasoning chain can consume substantial cache memory.

This variant applies two layers of compression:

  1. MLX 2-bit weight quantization — reduces the 27B model from 54 GB (BF16) to approximately **8 GB**, making it loadable on Apple Silicon devices with limited unified memory.
  2. RotorQuant 2-bit KV cache — rotation-based isotropic quantization compresses the key-value cache with better quality and speed than standard approaches.

RotorQuant Advantages

Metric RotorQuant 2-bit Standard 2-bit

| Perplexity | 6.91 | 7.07 |

RotorQuant achieves lower perplexity (better quality) while also being faster — making it the preferred 2-bit KV cache method when quality matters.

Specifications

Property Value
Base model Qwen/Qwen3.5-27B
Parameters 27B
Architecture Hybrid Transformer
Native context 262,144 tokens
Thinking mode Yes
Weight quantization MLX 2-bit
KV cache method RotorQuant 2-bit (IsoQuant)
KV cache compression ~10x vs FP16
Runtime MLX (Apple Silicon)

Memory Estimates

Component Estimate
Model weights (MLX 2-bit) ~8 GB
KV cache at 128K context (2-bit RotorQuant) ~1.3 GB
Total at 128K context ~9.3 GB
Comparison: BF16 weights + FP16 KV at 128K ~66.8 GB

Quickstart

from mlx_lm import load, generate
from turboquant import IsoQuantCache

model_id = "majentik/Qwen3.5-27B-RotorQuant-MLX-2bit"

model, tokenizer = load(model_id)

# Apply 2-bit RotorQuant KV cache compression
cache = IsoQuantCache(bits=2)

prompt = "Explain the Riemann hypothesis in simple terms."
messages = [{"role": "user", "content": prompt}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)

response = generate(
    model,
    tokenizer,
    prompt=text,
    max_tokens=2048,
    kv_cache=cache,
)
print(response)

Quality Notes

  • 2-bit weights + 2-bit KV cache is the most aggressive quantization combination, but RotorQuant's rotation-based approach preserves more quality than standard methods (perplexity 6.91 vs 7.07).
  • For higher quality on Apple Silicon, consider 4-bit weight variants with 4-bit KV cache.
  • Thinking mode reasoning quality may be more sensitive to quantization since the model relies on both weight precision and cached reasoning tokens for its final answer.
  • Best suited for: prototyping, development, long-context exploration, and scenarios where running the model at all matters more than peak quality.

References

See Also

Quant trade-off (MLX lane)

Bits Approx size Use case Recommendation
2-bit ~7.0 GB Aggressive quantization Very low-RAM Macs
3-bit ~9.7 GB Lossy but small Low-RAM Macs
4-bit ~11 GB Balanced default Recommended for most Macs
5-bit ~14 GB Higher fidelity Quality-sensitive
6-bit ~16 GB Approaching FP16 quality High-fidelity
8-bit ~21 GB Near-lossless reference Fidelity-critical work

(Current variant — 2bit — is bolded.)

Variants in this family

(Showing 16 sibling variants under majentik/qwen3.5-27b-*. The current variant — RotorQuant-MLX-2bit — is bolded.)

Variant Runtime Approx size Use case
RotorQuant-GGUF-IQ4_XS llama.cpp ~23 GB Lossy 4-bit, low-RAM CPU/edge
RotorQuant-GGUF-Q2_K llama.cpp ~16 GB Lossy, low-RAM CPU/edge
RotorQuant-GGUF-Q3_K_M llama.cpp ~21 GB Smaller 3-bit, CPU-friendly
RotorQuant-GGUF-Q4_K_M llama.cpp ~30 GB Balanced default
RotorQuant-GGUF-Q5_K_M llama.cpp ~36 GB Higher fidelity, more RAM
RotorQuant-GGUF-Q8_0 llama.cpp ~57 GB Near-lossless reference
RotorQuant-MLX-2bit mlx-lm ~8.6 GB Apple Silicon, smallest
RotorQuant-MLX-4bit mlx-lm ~17 GB Apple Silicon balanced
RotorQuant-MLX-8bit mlx-lm ~32 GB Apple Silicon reference
TurboQuant-MLX-2bit mlx-lm ~8.6 GB Apple Silicon, smallest
TurboQuant-MLX-4bit mlx-lm ~17 GB Apple Silicon balanced
TurboQuant-MLX-8bit mlx-lm ~32 GB Apple Silicon reference

About the RotorQuant / TurboQuant labels

RotorQuant and TurboQuant are this project's release labels, not distinct quantization algorithms — for any given tier, both brand repos carry byte-identical weights produced with the standard MLX / llama.cpp quantizers. No brand-specific speedup is claimed or measured.

Downloads last month
226
Safetensors
Model size
27B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for majentik/Qwen3.5-27B-RotorQuant-MLX-2bit

Base model

Qwen/Qwen3.5-27B
Quantized
(225)
this model

Collection including majentik/Qwen3.5-27B-RotorQuant-MLX-2bit