Instructions to use Taykhoom/UTR-LM-MLMSI with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Taykhoom/UTR-LM-MLMSI with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Taykhoom/UTR-LM-MLMSI", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
UTR-LM-MLMSI
Minimal HuggingFace port of the MLM + MFE variant of UTR-LM -- an ESM2-style 5' UTR RNA language model pretrained on endogenous sequences from five species and a large synthetic library.
Architecture
| Parameter | Value |
|---|---|
| Layers | 6 |
| Attention heads | 16 |
| Embedding dimension | 128 |
| FFN hidden dimension | 512 (GELU) |
| Vocabulary size | 10 |
| Positional encoding | RoPE (base=10000) |
| Normalization | LayerNorm |
| Architecture | ESM2-style pre-LN Transformer with GELU FFN |
| Max sequence length | 1024 tokens (1022 nucleotides + <cls> / <eos>) |
Vocabulary: <pad> (0), <eos> (1), <unk> (2), A (3), G (4), C (5), T (6), <cls> (7), <mask> (8), <sep> (9)
Pretraining
- Objective: Masked language modeling + MFE (minimum free energy) regression
- Data: Endogenous 5' UTRs from five species (human, mouse, zebrafish, Drosophila, yeast) combined with the Cao et al. random 5' UTR synthetic library
- Source checkpoint:
ESM2SI_3.1_fiveSpeciesCao_6layers_16heads_128embedsize_4096batchToks_MLMLossMin.pkl
Checkpoint selection
Multiple ESM2SI checkpoints were available (versions 3.1, FS4.1, FS4.4, FS4.7). The 3.1 checkpoint was selected because it is the version specified in the original UTR-LM paper for translation efficiency (TE) and expression level (EL) downstream tasks (used in the MJ4_Finetune evaluation scripts). The FS4.x variants are later training runs but were not the ones reported in the original publication.
Parity Verification
All 7 representation levels (embedding + 6 transformer blocks) were verified
to be bit-exact (max absolute difference = 0.00) against the original
ESM2SI_3.1_fiveSpeciesCao_6layers_16heads_128embedsize_4096batchToks_MLMLossMin.pkl
weights. Verified on GPU with PyTorch 2.7.1 / CUDA 12.9.
Related Models
See the full UTR-LM collection.
| Model | Pretraining Objective | Notes |
|---|---|---|
| UTR-LM-MLM | MLM | Base model |
| UTR-LM-MLMSI | MLM + MFE regression | This model — recommended for TE / EL tasks |
| UTR-LM-MLMSS | MLM + secondary structure | — |
| UTR-LM-MLMSISS | MLM + MFE + secondary structure | Recommended for MRL tasks |
Usage
Embedding generation
import torch
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("Taykhoom/UTR-LM-MLMSI", trust_remote_code=True)
model = AutoModel.from_pretrained("Taykhoom/UTR-LM-MLMSI", trust_remote_code=True)
model.eval()
sequences = ["ATGCATGCATGC", "GCTAGCTAGCTAGCTA"]
enc = tokenizer(sequences, return_tensors="pt", padding=True)
with torch.no_grad():
out = model(**enc)
# CLS token embedding (position 0) - recommended for sequence-level tasks
cls_emb = out.last_hidden_state[:, 0, :] # (batch, 128)
# All-token embeddings
token_emb = out.last_hidden_state # (batch, seq_len, 128)
# Intermediate layer representations
out_all = model(**enc, output_hidden_states=True)
layer3_emb = out_all.hidden_states[3] # after layer 3, shape (batch, seq_len, 128)
MLM logits
import torch
from transformers import AutoTokenizer, AutoModelForMaskedLM
tokenizer = AutoTokenizer.from_pretrained("Taykhoom/UTR-LM-MLMSI", trust_remote_code=True)
model = AutoModelForMaskedLM.from_pretrained("Taykhoom/UTR-LM-MLMSI", trust_remote_code=True)
model.eval()
enc = tokenizer(["ATGC<mask>ATGC"], return_tensors="pt")
with torch.no_grad():
logits = model(**enc).logits # (1, seq_len, 10)
Faster attention backends
# SDPA (PyTorch 2.0+)
model = AutoModel.from_pretrained(
"Taykhoom/UTR-LM-MLMSI",
trust_remote_code=True,
attn_implementation="sdpa",
)
# Flash Attention 2 (requires flash-attn)
model = AutoModel.from_pretrained(
"Taykhoom/UTR-LM-MLMSI",
trust_remote_code=True,
attn_implementation="flash_attention_2",
dtype=torch.bfloat16,
)
Fine-tuning
The model follows standard HF conventions and can be fine-tuned with any Trainer-compatible setup. For sequence regression tasks, use the CLS token embedding as input to a prediction head (as done in the original UTR-LM paper).
Implementation Notes
The source tokenizer uses the DNA-style A/G/C/T alphabet. Convert U to
T before tokenization when supplying RNA-spelled sequences; a literal U
otherwise maps to <unk>. MFE was an auxiliary prediction target, not an
input channel; this minimal port preserves the backbone and MLM head but
omits the auxiliary regression head.
The original UTR-LM implementation uses eager scaled dot-product attention.
This port additionally supports attn_implementation="sdpa" and
attn_implementation="flash_attention_2".
Citation
@article{chu2024_utrlm,
title = {A 5' {UTR} Language Model for Decoding Untranslated Regions of {mRNA} and Function Predictions},
author = {Chu, Yanyi and Yu, Dan and Li, Yupeng and Huang, Kaixuan and Shen, Yue and Cong, Le and Zhang, Jason and Wang, Mengdi},
journal = {Nature Machine Intelligence},
volume = {6},
number = {4},
pages = {449--460},
year = {2024},
doi = {10.1038/s42256-024-00823-9}
}
Credits
Original model and code by Yanyi Chu et al. Source: UTR-LM GitHub repository. The HF conversion code was authored primarily by Claude Code and reviewed manually by Taykhoom Dalal.
License
GPL-3.0, following the original UTR-LM repository.
- Downloads last month
- 110