Index-Echo-S2TT-2B-FP4

Official NVFP4 (W4A4) quantization of IndexTeam/Index-Echo-S2TT-2B, part of the Index-Echo speech-to-text translation (S2TT) model family by bilibili.

This repository mirrors the original checkpoint layout; only the text LLM backbone (llm/) is quantized to NVFP4 - the audio tower, connector, and all speech-synthesis components remain in BF16. Use it exactly like the original repo (same infer.py / configs).

Quantization

  • Scheme: NVFP4 (W4A4: 4-bit floating-point weights with per-group-16 scales, 4-bit floating-point activations with calibrated per-tensor global scales), produced with llm-compressor (quantization_scheme recorded in recipe.yaml). Calibrated on a small bilingual translation corpus.
  • All Linear layers of the language model backbone (llm/) are quantized; the audio tower, connector, lm_head, embeddings and all other pipeline components are kept in BF16.
  • Format: compressed-tensors nvfp4-pack-quantized safetensors - load directly with vLLM (quantization="compressed-tensors") or transformers.

Consistency validation

Measured on an NVIDIA A100 (weight-dequantized execution) against the original BF16 checkpoint (greedy decoding, official translation prompt):

Metric BF16 FP4 Delta
Perplexity (fixed corpus) 4.8772 5.1599 +5.80%
zh->en generation identical - - no (semantically equivalent)
en->zh generation identical - - no (semantically equivalent)

Usage

Identical to the original checkpoint - clone this repo and follow the README / infer.py of the base model (IndexTeam/Index-Echo-S2TT-2B). The quantized LLM backbone loads via compressed-tensors; make sure compressed-tensors (or vLLM / a recent transformers) is installed.

Hardware note: full NVFP4 (W4A4) acceleration requires an NVIDIA Blackwell GPU (SM100+, e.g. B200 / RTX 50 series). On older GPUs vLLM falls back to weight-only dequantization, which still reduces memory but gives no FP4 speedup.

Quantized and published by the Index team, 2026-10-04.

Downloads last month
14
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IndexTeam/Index-Echo-S2TT-2B-FP4

Finetuned
(2)
this model

Collection including IndexTeam/Index-Echo-S2TT-2B-FP4