Vishva007/Qwen3.5-9B-W4A16-AutoRound-LLM-Compressor

This is a W4A16 (4-bit weight, 16-bit activation) AutoRound-format quantized version of Qwen/Qwen3.5-9B, produced using AutoRound — Intel's sign gradient descent based quantization method designed for production-grade accuracy retention.

  • Vision Tower (quant_nontext_module): False (Kept in BF16 to preserve visual reasoning and OCR precision)
  • Special Modules (layer_config): Multi-Token Prediction (mtp, mtp.fc) kept in native bfloat16

Quantization Details

Parameter Value
Method AutoRound (W4A16, AutoRound format)
Group Size 32
Symmetric Yes
Iterations 800
Calibration Samples 512
Sequence Length 4096
Torch Compile Enabled

Key Notes

  • AutoRound format — Exported in the standard AutoRound format for broad ecosystem compatibility.
  • High accuracy configuration — 800 iterations with 512 calibration samples targets production-grade quality with minimal degradation from the base model.
  • W4A16 — Weights are quantized to 4-bit integers; activations remain in FP16 for inference stability.
  • ~50% memory reduction compared to the FP16 base model, enabling deployment on consumer and mid-range GPUs.

Usage

This model is compatible with transformers, AutoRound, vLLM, and SGLang — any backend supporting AutoRound-format weights works out of the box. For full model details, architecture, and capabilities, refer to the base model page.

🚀 Deploy on RunPod

One-click launch environments pre-configured with PyTorch, CUDA, and dependencies for fine-tuning or quantization.

🎁 Need GPU compute? Sign up via RunPod and get $5–$500 in free credits when you add your first $10.

PyTorch 2.14

Template CUDA Version Docker Image Template ID Deploy
PyTorch 2.14 (CUDA 12.6) 12.6 vishva123/cuda-12.6-pytorch-2.14-runpod d7lxsa4w9m Deploy to RunPod
PyTorch 2.14 (CUDA 13.0) 13.0 vishva123/cuda-13.0-pytorch-2.14-runpod yk0y6j6rpg Deploy to RunPod
PyTorch 2.14 (CUDA 13.2) 13.2 vishva123/cuda-13.2-pytorch-2.14-runpod gsp4gwx0nw Deploy to RunPod

PyTorch 2.13

Template CUDA Version Docker Image Template ID Deploy
PyTorch 2.13 (CUDA 12.6) 12.6 vishva123/cuda-12.6-pytorch-2.13-runpod gmlupxnxfk Deploy to RunPod
PyTorch 2.13 (CUDA 13.0) 13.0 vishva123/cuda-13.0-pytorch-2.13-runpod y3j8xvk4f4 Deploy to RunPod
PyTorch 2.13 (CUDA 13.2) 13.2 vishva123/cuda-13.2-pytorch-2.13-runpod vigpissn5w Deploy to RunPod

PyTorch 2.12

Template CUDA Version Docker Image Template ID Deploy
PyTorch 2.12 (CUDA 12.6) 12.6 vishva123/cuda-12.6-pytorch-2.12-runpod ctmz86zmf0 Deploy to RunPod
PyTorch 2.12 (CUDA 13.0) 13.0 vishva123/cuda-13.0-pytorch-2.12-runpod qjko5yiwzi Deploy to RunPod
PyTorch 2.12 (CUDA 13.2) 13.2 vishva123/cuda-13.2-pytorch-2.12-runpod ifg6xmye0f Deploy to RunPod

Downloads last month
266
Safetensors
Model size
2B params
Tensor type
I32
·
BF16
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Vishva007/Qwen3.5-9B-W4A16-AutoRound

Finetuned
Qwen/Qwen3.5-9B
Quantized
(515)
this model

Collection including Vishva007/Qwen3.5-9B-W4A16-AutoRound