Historical Qwen2.5-3B X-Ray comparison — correction

#19
by tetracta - opened

Update · 12 September 2026. The X-Ray interpretations in the original post below were withdrawn on 6 September and remain withdrawn. The linked legacy deliveries are not current product evidence.

The current VG1 report gallery now distinguishes retained measurement examples from the historical archive. A refreshed report presentation is not a new scan of this pair. Knowledge scans remain unavailable; customer scan-to-report acceptance is still in progress. Current scope · Dated correction.

Original post — superseded; retained as historical context

We ran a layer-by-layer functional comparison between Qwen/Qwen2.5-3B and Qwen/Qwen2.5-3B-Instruct, plus a knowledge-delta probe. Sharing the findings here since they might be useful to people working with these checkpoints.

Structure — 80% of the functional difference sits in 14 of 37 stations, concentrated in the last five (stations 32–36, 2.72× a uniform spread; 1.41× what the probe grid alone would put there). Change is detectable at or before station 4 — that is the instrument's detection floor at this scale, not a finding. In this pair, instruction tuning is a late-concentrated, output-shaping change; early representation is largely preserved.

Knowledge & hallucination probe (corrected scoring, 2026-09-02) — every fact the base model answered correctly survived the fine-tune (19/20 → 20/20, zero broken, one repaired). Known-vs-fabricated separation went 0.9325 → 0.9475 (trajectory AUROC; output-only 0.8950 → 0.9475) — inside the interval at 20-vs-20, so: unchanged to slightly stronger. The model echoes fabricated names less often (13/20 → 18/20, paired 6 improved / 1 regressed, exact p = 0.125), but genuine refusals only went 0 → 2 — most of that shift is echo → confident fabrication, so we do not call it a hallucination win.

Full interactive reports (no login needed):

(The scans come from Model X-Ray, our checkpoint-inspection tool — currently in free open beta: https://www.tetracta.ai/xray.html)

Curious what others see in their own fine-tunes of these models — if you scan a checkpoint and find something interesting, we'd love to hear about it.

Correction (2026-09-02): an external review led us to re-audit this scan. The structure numbers above replace an earlier 'starts at layer 4, spreads across 33 of 37 (~89%)' reading (a function of the detection floor); the knowledge numbers replace 0.915 → 0.878 and '18/20 → 18/20', which came from two scoring bugs in the probe (space sub-token as answer target; double final-norm). Raw per-item data and both reports: https://huggingface.co/spaces/tetracta/model-xray-sample-reports

deleted
•
This comment has been hidden

Correction — 6 September 2026

We withdraw both the structural and knowledge conclusions in the opening post. The legacy evidence does not support the published onset, spread, concentration, location, AUROC, internal-readout, factual-change or fabrication-separation interpretations. Any portrait-visualization, lesion-response or legacy simulated-quantization finding linked from the opening post is also withdrawn.

The linked reports are historical and must not be used as current evidence. We are publishing no replacement figures. Validation remains pending. Correction record: https://www.tetracta.ai/model-xray/correction/

— Tetracta

tetracta changed discussion title from X-ray: what actually changed between Qwen2.5-3B base and Instruct (layer map + knowledge probes) to Historical Qwen2.5-3B X-Ray comparison — correction

Sign up or log in to comment