MDLM (LM1B)
MDLM (masked diffusion language model) with a factorized reverse process and a dense DiT backbone. This is the factorized baseline for E-MoE, trained with the same data, tokenizer and budget.
This checkpoint accompanies the paper E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models. Code: github.com/Arseny5/e-moe.
Model
| Architecture | DiT, 12 blocks, hidden size 768, 12 heads |
| Trainable parameters | 139.3M |
| Active parameters per token | 139.3M |
| Sequence length | 128 tokens |
| Tokenizer | bert-base-uncased |
| Diffusion | absorbing-state (masked) diffusion, log-linear noise schedule, no time conditioning |
| Weights | EMA (decay 0.9999), float32, model.safetensors |
Training data
LM1B (One Billion Word Benchmark), tokenized with bert-base-uncased and packed (wrapped) into blocks of 128 tokens.
Training
- Steps: 1M, global batch 512 sequences (128 per GPU × 4 GPUs), ≈ 65.5B tokens (~72 epochs)
- Optimizer: AdamW, lr 3e-4, betas (0.9, 0.999), weight decay 0, 2.5k warmup steps, gradient clipping 1.0
- Precision: bf16 mixed precision; attention via PyTorch SDPA (FlashAttention kernels), rotary embeddings from
flash-attn - Hardware: 4 × NVIDIA H200, ~80 h for 1M steps
- Training objective: the standard MDLM continuous-time ELBO (SUBS parameterization).
Results
Generative perplexity (↓, judged by GPT-2 Large) / unigram sample entropy, unconditional generation of 128 tokens on LM1B. Mean ± std over 5 disjoint groups of 1,000 samples. Best gen-PPL per row in bold.
| NFE | MDLM (this model) | SEDD | VADD | MDLM-MoE | E-MoE |
|---|---|---|---|---|---|
| 1 | 1440 ± 13 / 4.368 | 1604 ± 14 / 4.383 | 1279 ± 20 / 4.349 | 1293 ± 13 / 4.327 | 644 ± 6 / 4.352 |
| 2 | 993 ± 19 / 4.370 | 1065 ± 16 / 4.379 | 769 ± 11 / 4.334 | 816 ± 14 / 4.328 | 379 ± 5 / 4.349 |
| 4 | 469 ± 4 / 4.355 | 461 ± 2 / 4.347 | 372 ± 2 / 4.329 | 383 ± 2 / 4.323 | 235 ± 3 / 4.342 |
| 8 | 260 ± 1 / 4.351 | 242 ± 5 / 4.333 | 223.6 ± 5.5 / 4.326 | 214.7 ± 1.2 / 4.324 | 174.8 ± 2.1 / 4.340 |
| 16 | 179.4 ± 3.1 / 4.350 | 167.4 ± 3.5 / 4.329 | 165.7 ± 1.5 / 4.325 | 149.0 ± 3.1 / 4.322 | 148.4 ± 2.0 / 4.339 |
| 32 | 149.9 ± 1.9 / 4.349 | 138.0 ± 1.2 / 4.328 | 141.1 ± 2.5 / 4.324 | 125.1 ± 1.3 / 4.326 | 135.3 ± 1.8 / 4.338 |
| 64 | 133.7 ± 2.0 / 4.349 | 124.2 ± 1.5 / 4.326 | 126.6 ± 0.6 / 4.321 | 112.4 ± 1.1 / 4.325 | 130.1 ± 0.8 / 4.339 |
| 128 | 127.6 ± 1.3 / 4.349 | 118.9 ± 1.2 / 4.325 | 117.8 ± 0.6 / 4.317 | 107.0 ± 1.4 / 4.323 | 124.9 ± 1.6 / 4.337 |
Autoregressive baseline (same backbone and budget, 128 steps): 68.3 ± 0.9 / 4.320.
Usage
The checkpoint contains the backbone weights (model.safetensors) and the architecture/training configuration (config.json):
import json
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
repo = "ArsenyIvanov/mdlm-lm1b"
state_dict = load_file(hf_hub_download(repo, "model.safetensors"))
config = json.load(open(hf_hub_download(repo, "config.json")))
The model code and sampling scripts that load these weights will be released in github.com/Arseny5/e-moe.
Training checkpoint
training_checkpoint/step_1000000.ckpt is the full PyTorch Lightning checkpoint at step 1M: raw and EMA weights, AdamW optimizer state, learning-rate scheduler and loop state, for resuming training. It is a pickle file: load it only from this trusted repository, e.g. torch.load(path, map_location="cpu", weights_only=False). For inference, use model.safetensors (EMA weights).
License
Apache License 2.0.
Citation
@article{ivanov2026emoe,
title = {E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models},
author = {Ivanov, Arseny and Kolesov, Alexander and Korotin, Alexander and
Oseledets, Ivan and Goncharov, Mikhail},
journal = {arXiv preprint arXiv:2609.37533},
year = {2026}
}
- Downloads last month
- 5