Qwen 3 4B - Heretic (Abliterated)

💬 Community: Join the Abliterlitics Discord for discussion, model releases and support.

An abliterated version of Qwen 3 4B created using Heretic v1.2.0. This model has reduced refusals while maintaining model quality, making it suitable as an uncensored text encoder for image generation models like Z-Image and FLUX.2 Klein 4B. Available in five ComfyUI-native quantized formats (FP8, INT8, INT4, NVFP4, MXFP8), all produced with SVD-guided learned rounding for maximum fidelity.

Model Details

  • Base Model: Qwen/Qwen3-4B
  • Abliteration Method: Heretic v1.2.0
  • Trials: 200
  • Trial Selected: Trial 96
  • Refusals: 3/100 (vs 100/100 original)
  • KL Divergence: 0.0000 (zero measurable model damage)

Files

HuggingFace Format (for transformers, llama.cpp conversion)

model-00001-of-00002.safetensors
model-00002-of-00002.safetensors
config.json
tokenizer.json
tokenizer_config.json

ComfyUI Format (for Z-Image / FLUX.2 Klein 4B text encoder)

comfyui/qwen3-4b-heretic.safetensors              # bf16, 7.5GB
comfyui/qwen3-4b-heretic_fp8_e4m3fn.safetensors   # fp8 row-wise, 4.2GB
comfyui/qwen3-4b-heretic_int8.safetensors          # int8 ConvRot row-wise, 4.2GB
comfyui/qwen3-4b-heretic_int4.safetensors          # int4 W4A4 ConvRot, 2.5GB
comfyui/qwen3-4b-heretic_nvfp4.safetensors         # nvfp4, 2.7GB
comfyui/qwen3-4b-heretic_mxfp8.safetensors         # mxfp8, 4.3GB

Quality: All quantized variants use SVD-guided learned rounding (AdaRound via convert_to_quant), which optimizes each weight's rounding direction to minimize output reconstruction error — noticeably higher fidelity than naive round-to-nearest quantization.

GGUF Format (for llama.cpp and ComfyUI-GGUF)

Quant Size Notes
F16 ~7.5GB Lossless reference
Q8_0 ~4GB Excellent quality
Q6_K ~3GB Very good quality
Q5_K_M ~2.7GB Good quality
Q4_K_M ~2.3GB Recommended balance
Q3_K_M ~1.9GB For low VRAM only

Quantization Format Notes

All variants load natively in ComfyUI 0.30.0+ (no plugins) via the comfy_quant metadata embedded in each file.

Format Size Bits Notes
FP8 (E4M3, row-wise) 4.2GB 8 Best speed/quality balance; works on Ada/Hopper+
INT8 (ConvRot row-wise) 4.2GB 8 Hadamard-rotated; broad GPU support
MXFP8 4.3GB 8 Microscaling FP8 (E8M0 block scales); Blackwell-accelerated
INT4 (W4A4 ConvRot) 2.5GB 4 Smallest; Hadamard-rotated signed INT4
NVFP4 (E2M1) 2.7GB 4 NVIDIA FP4; Blackwell FP4 tensor cores for best perf

NVFP4/MXFP8 inference is fastest on Blackwell (RTX 5090/5080, SM100+), but ComfyUI also supports software dequantization on older GPUs (tested working on RTX 4090). INT8 and INT4 both use Hadamard rotation (ConvRot); INT4 W4A4 uses ComfyUI's convrot_w4a4 path.

Usage

With ComfyUI (Z-Image / FLUX.2 Klein 4B)

  1. Download a ComfyUI format file:

    • FP8 (recommended): comfyui/qwen3-4b-heretic_fp8_e4m3fn.safetensors (4.2GB)
    • INT4 (smallest): comfyui/qwen3-4b-heretic_int4.safetensors (2.5GB)
    • NVFP4: comfyui/qwen3-4b-heretic_nvfp4.safetensors (2.7GB)
    • INT8: comfyui/qwen3-4b-heretic_int8.safetensors (4.2GB)
    • MXFP8: comfyui/qwen3-4b-heretic_mxfp8.safetensors (4.3GB)
    • bf16 (full precision): comfyui/qwen3-4b-heretic.safetensors (7.5GB)
  2. Place in ComfyUI/models/text_encoders/

  3. In your Z-Image workflow, use the ClipLoader node and select the heretic file

With Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "DreamFast/qwen3-4b-heretic",
    device_map="auto",
    torch_dtype=torch.bfloat16
)
tokenizer = AutoTokenizer.from_pretrained("DreamFast/qwen3-4b-heretic")

prompt = "Describe a dramatic sunset over a cyberpunk city"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

With llama.cpp

llama-server -m qwen3-4b-heretic-Q4_K_M.gguf

Abliteration Process

Created using Heretic v1.2.0 with 200 optimization trials:

? Which trial do you want to use?
> [Trial  96] Refusals:  3/100, KL divergence: 0.0000  <-- selected
  [Trial  90] Refusals:  5/100, KL divergence: 0.0000
  [Trial  95] Refusals:  9/100, KL divergence: 0.0000
  [Trial 122] Refusals: 90/100, KL divergence: 0.0000
  ...

Trial 96 was selected for having the fewest refusals (3/100) with zero measurable KL divergence, indicating the abliteration surgically removed the refusal mechanism with no damage to model capabilities.

Limitations

  • This model inherits all limitations of the base Qwen 3 4B model
  • Abliteration reduces but does not completely eliminate refusals (3/100 remain)

License

This model is released under the Apache 2.0 License, following the base Qwen 3 4B model license.

Acknowledgments

Downloads last month
18,828
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
Input a message to start chatting with DreamFast/qwen3-4b-heretic.

Model tree for DreamFast/qwen3-4b-heretic

Finetuned
Qwen/Qwen3-4B
Quantized
(286)
this model
Finetunes
4 models

Space using DreamFast/qwen3-4b-heretic 1

Collection including DreamFast/qwen3-4b-heretic