Index-Translate-2B-GGUF

Official GGUF conversion of IndexTeam/Index-Translate-2B, part of the Index-Translate multilingual translation model family (150 languages, terminology/format-constrained translation, controlled dubbing translation, long-document translation).

Provided files

(sorted by size)

File Size Notes
mmproj-Q8_0 0.36 GB multi-modal projector supplement
mmproj-f16 0.67 GB multi-modal projector supplement
Q2_K 0.99 GB 2-bit, significant quality loss
Q3_K_S 1.05 GB 3-bit, noticeable quality loss
Q3_K_M 1.13 GB 3-bit, noticeable quality loss
Q3_K_L 1.20 GB 3-bit, noticeable quality loss
IQ4_XS 1.23 GB 4-bit I-quant, small
Q4_K_S 1.25 GB 4-bit, slightly smaller/faster
Q4_K_M 1.31 GB 4-bit, good balance, recommended
Q5_K_S 1.42 GB 5-bit, low quality loss
Q5_K_M 1.45 GB 5-bit, low quality loss
Q6_K 1.61 GB 6-bit, very low quality loss
Q8_0 2.08 GB 8-bit, near-lossless
f16 3.90 GB 16-bit, lossless conversion baseline

Multimodal Image Translation (mmproj)

This repository includes official native multi-modal projector weights (*.mmproj-*.gguf), enabling end-to-end image-to-text translation across 150 languages without any external OCR tools:

  • Index-Translate-2B.mmproj-Q8_0.gguf: 8-bit quantized projector, recommended for production and local use.
  • Index-Translate-2B.mmproj-f16.gguf: 16-bit lossless conversion baseline.

1. Launch llama-server with Vision Projector:

llama-server \
  -m Index-Translate-2B.Q4_K_M.gguf \
  --mmproj Index-Translate-2B.mmproj-Q8_0.gguf \
  -ngl 99 \
  -c 32768 \
  --port 8000 \
  --alias Index-Translate-2B

2. Run Image Translation:

Once the server is running, perform image translation with the official script:

python inference/llm/translate.py --image screenshot.png --target zh -m Index-Translate-2B

Or test it directly via the online web demo at https://index-translate.bilibili.com/?p=/site/translate.html. Full guide: Multimodal Image Translation Guide.

Usage

# Run the recommended Q4_K_M directly:
llama serve -hf IndexTeam/Index-Translate-2B-GGUF:Q4_K_M
llama cli -hf IndexTeam/Index-Translate-2B-GGUF:Q4_K_M

Translation prompt format (greedy decoding, temperature=0 recommended; use the chat template with enable_thinking: false):

请将以下文本翻译为{target-language},直接输出翻译结果,不要进行任何解释。

{source-text}
# llama.cpp: run the recommended Q4_K_M with the chat template and thinking disabled
llama-cli -hf IndexTeam/Index-Translate-2B-GGUF:Q4_K_M --jinja --chat-template-kwargs '{"enable_thinking":false}' --temp 0

Prompting & constrained translation (instTrans)

Beyond plain translation, the models follow the instTrans constrained-translation format. The official client wraps requests into the canonical structure 【源文】<text> + numbered 1. 【硬性要求】<hard constraints> + 2. 【注意】<soft constraints> + suffix instructions:

  • Hard constraints (binary, must hold): strict terminology glossary enforcement (e.g. 碳纤维:carbon fiber, 抗裂缝:crack resistance), and format/structure preservation for JSON/CSV/code/placeholders.
  • Soft constraints (graded): tone & style adaptation (e.g. formal business-email register), domain/word-sense disambiguation (e.g. plant → 工厂 in an industrial context), cross-sentence consistency, LaTeX preservation.
  • Syllable-controlled translation (dubbing): use the Index-Homura checkpoints (IndexTeam/Index-Homura-2B/9B), which strictly respect a target syllable budget and can be combined with glossaries.

Full prompt reference: github.com/bilibili/Index-Translate (Instruction Following section, docs/prompts.md, inference/llm/cases/).

See the base model card for the full instTrans constrained-translation format and serving presets.

Consistency validation

Before release, each quantization tier was validated on GPU (NVIDIA A100) against the F16 conversion: per-token KL divergence / RMS Δp via llama-perplexity, plus greedy-generation spot checks against the original BF16 weights (transformers reference). Generation outputs of Q4_K_M matched the reference almost verbatim.

Converted and published by the Index team, 2026-10-03.

Downloads last month
7,223
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IndexTeam/Index-Translate-2B-GGUF

Quantized
(16)
this model

Space using IndexTeam/Index-Translate-2B-GGUF 1

Collection including IndexTeam/Index-Translate-2B-GGUF

Paper for IndexTeam/Index-Translate-2B-GGUF