UTR-LM-MLMSI

Minimal HuggingFace port of the MLM + MFE variant of UTR-LM -- an ESM2-style 5' UTR RNA language model pretrained on endogenous sequences from five species and a large synthetic library.

Architecture

Parameter Value
Layers 6
Attention heads 16
Embedding dimension 128
FFN hidden dimension 512 (GELU)
Vocabulary size 10
Positional encoding RoPE (base=10000)
Normalization LayerNorm
Architecture ESM2-style pre-LN Transformer with GELU FFN
Max sequence length 1024 tokens (1022 nucleotides + <cls> / <eos>)

Vocabulary: <pad> (0), <eos> (1), <unk> (2), A (3), G (4), C (5), T (6), <cls> (7), <mask> (8), <sep> (9)

Pretraining

  • Objective: Masked language modeling + MFE (minimum free energy) regression
  • Data: Endogenous 5' UTRs from five species (human, mouse, zebrafish, Drosophila, yeast) combined with the Cao et al. random 5' UTR synthetic library
  • Source checkpoint: ESM2SI_3.1_fiveSpeciesCao_6layers_16heads_128embedsize_4096batchToks_MLMLossMin.pkl

Checkpoint selection

Multiple ESM2SI checkpoints were available (versions 3.1, FS4.1, FS4.4, FS4.7). The 3.1 checkpoint was selected because it is the version specified in the original UTR-LM paper for translation efficiency (TE) and expression level (EL) downstream tasks (used in the MJ4_Finetune evaluation scripts). The FS4.x variants are later training runs but were not the ones reported in the original publication.

Parity Verification

All 7 representation levels (embedding + 6 transformer blocks) were verified to be bit-exact (max absolute difference = 0.00) against the original ESM2SI_3.1_fiveSpeciesCao_6layers_16heads_128embedsize_4096batchToks_MLMLossMin.pkl weights. Verified on GPU with PyTorch 2.7.1 / CUDA 12.9.

Related Models

See the full UTR-LM collection.

Model Pretraining Objective Notes
UTR-LM-MLM MLM Base model
UTR-LM-MLMSI MLM + MFE regression This model — recommended for TE / EL tasks
UTR-LM-MLMSS MLM + secondary structure —
UTR-LM-MLMSISS MLM + MFE + secondary structure Recommended for MRL tasks

Usage

Embedding generation

import torch
from transformers import AutoTokenizer, AutoModel

tokenizer = AutoTokenizer.from_pretrained("Taykhoom/UTR-LM-MLMSI", trust_remote_code=True)
model = AutoModel.from_pretrained("Taykhoom/UTR-LM-MLMSI", trust_remote_code=True)
model.eval()

sequences = ["ATGCATGCATGC", "GCTAGCTAGCTAGCTA"]
enc = tokenizer(sequences, return_tensors="pt", padding=True)

with torch.no_grad():
    out = model(**enc)

# CLS token embedding (position 0) - recommended for sequence-level tasks
cls_emb = out.last_hidden_state[:, 0, :]   # (batch, 128)

# All-token embeddings
token_emb = out.last_hidden_state           # (batch, seq_len, 128)

# Intermediate layer representations
out_all = model(**enc, output_hidden_states=True)
layer3_emb = out_all.hidden_states[3]       # after layer 3, shape (batch, seq_len, 128)

MLM logits

import torch
from transformers import AutoTokenizer, AutoModelForMaskedLM

tokenizer = AutoTokenizer.from_pretrained("Taykhoom/UTR-LM-MLMSI", trust_remote_code=True)
model = AutoModelForMaskedLM.from_pretrained("Taykhoom/UTR-LM-MLMSI", trust_remote_code=True)
model.eval()

enc = tokenizer(["ATGC<mask>ATGC"], return_tensors="pt")
with torch.no_grad():
    logits = model(**enc).logits   # (1, seq_len, 10)

Faster attention backends

# SDPA (PyTorch 2.0+)
model = AutoModel.from_pretrained(
    "Taykhoom/UTR-LM-MLMSI",
    trust_remote_code=True,
    attn_implementation="sdpa",
)

# Flash Attention 2 (requires flash-attn)
model = AutoModel.from_pretrained(
    "Taykhoom/UTR-LM-MLMSI",
    trust_remote_code=True,
    attn_implementation="flash_attention_2",
    dtype=torch.bfloat16,
)

Fine-tuning

The model follows standard HF conventions and can be fine-tuned with any Trainer-compatible setup. For sequence regression tasks, use the CLS token embedding as input to a prediction head (as done in the original UTR-LM paper).

Implementation Notes

The source tokenizer uses the DNA-style A/G/C/T alphabet. Convert U to T before tokenization when supplying RNA-spelled sequences; a literal U otherwise maps to <unk>. MFE was an auxiliary prediction target, not an input channel; this minimal port preserves the backbone and MLM head but omits the auxiliary regression head.

The original UTR-LM implementation uses eager scaled dot-product attention. This port additionally supports attn_implementation="sdpa" and attn_implementation="flash_attention_2".

Citation

@article{chu2024_utrlm,
  title   = {A 5' {UTR} Language Model for Decoding Untranslated Regions of {mRNA} and Function Predictions},
  author  = {Chu, Yanyi and Yu, Dan and Li, Yupeng and Huang, Kaixuan and Shen, Yue and Cong, Le and Zhang, Jason and Wang, Mengdi},
  journal = {Nature Machine Intelligence},
  volume  = {6},
  number  = {4},
  pages   = {449--460},
  year    = {2024},
  doi     = {10.1038/s42256-024-00823-9}
}

Credits

Original model and code by Yanyi Chu et al. Source: UTR-LM GitHub repository. The HF conversion code was authored primarily by Claude Code and reviewed manually by Taykhoom Dalal.

License

GPL-3.0, following the original UTR-LM repository.

Downloads last month
110
Safetensors
Model size
1.21M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including Taykhoom/UTR-LM-MLMSI