S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens
S2Tok reconstructs a scene as 3D Gaussians from an uncalibrated image stream, one frame at a time. It keeps a persistent, size-adaptive set of spatial tokens as its scene state: each new frame updates the existing tokens, and a learned admission module adds new tokens only for content that is not yet represented. A hierarchical decoder turns the tokens into non-pixel-aligned Gaussians, so the scene can be rendered from new views at any point of the stream without keeping past frames.
⚠️ Research release. This model is released for research purposes under the terms below. It is not a production system and does not include any Applied Intuition proprietary data, product code, or production checkpoints.
Model Details
| Developed by | Applied Intuition — AI Research |
| Model type | Streaming feed-forward transformer (DA3-Giant ViT-G/14 backbone) with a persistent token memory, a learned admission module and a 3D Gaussian decoder |
| Paper | S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens |
| Code | https://github.com/Applied-Intuition-Open-Source/S2Tok |
| Project page | https://s2tok.github.io/ |
| Initialized from | ZipSplat (CC BY-NC 4.0), itself initialized from DA3-Giant (CC BY-NC 4.0) |
| Training data | DL3DV-10K and RealEstate10K |
| License (weights) | CC BY-NC 4.0 — see WEIGHTS_LICENSE.md |
| Contact | GitHub Issues on the code repo |
Checkpoints
| Checkpoint | Description | Size | Link |
|---|---|---|---|
s2tok_252.pt |
1.72B params. 252x252 input, trained on 12–24 frame sequences from RealEstate10K and DL3DV-10K. | 6.87 GB | s2tok_252.pt |
The file holds model (inference weights) and config (architecture settings), and loads with
torch.load(..., weights_only=True). SHA-256 checksums are in SHA256SUMS.
Intended Use & Limitations
Intended use: research on streaming and feed-forward 3D reconstruction and novel-view synthesis.
Out of scope / limitations:
- Frames are center-cropped and resized to 252x252. Gaussians are expressed in the camera frame of the first image, up to a global scale.
- The memory holds at most 4,000 tokens (128k Gaussians); once it is full, no new content is admitted.
- Trained on 12–24 frame sequences of static scenes and evaluated at 12 and 50 views; retention over much longer streams is not established.
Commercial use is not permitted under CC BY-NC 4.0. The weights are also subject to the terms of their upstream model and datasets (see License).
How to Use
from huggingface_hub import hf_hub_download
from s2tok import load_model
path = hf_hub_download("AppliedIntuitionResearch/S2Tok", "s2tok_252.pt")
model = load_model(path, device="cuda")
memory = model.init_memory()
for image in frames: # [3, 252, 252] RGB tensors in [0, 1]
out, memory = model.step(image.cuda(), memory)
out.gaussians.save_ply("scene.ply") # Gaussians after the last frame (3DGS PLY)
Streaming demo with an interactive viewer:
python demo.py --model_path s2tok_252.pt --seq_path path/to/video.mp4 --output_dir outputs/my_scene
Full installation and usage instructions: see the GitHub repository.
License
The weights are released under CC BY-NC 4.0, non-commercial use only (WEIGHTS_LICENSE.md, full text in LICENSE). The non-commercial restriction is also required by their upstream sources:
- fine-tuned from the released ZipSplat checkpoint (CC BY-NC 4.0), which is initialized from DA3-Giant (CC BY-NC 4.0); depth labels for training were produced with DA3-Nested-Giant-Large (CC BY-NC 4.0);
- trained on DL3DV-10K (DL3DV-10K Terms of Use and CC BY-NC 4.0: non-commercial research and education only) and RealEstate10K (Google; frames from YouTube videos).
Use of the weights must also comply with the terms of these models and datasets. No training data is included in this repository.
Citation
@article{li2026s2tok,
title = {S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens},
author = {Li, Fang and Yenphraphai, Jiraphon and Herau, Quentin and Meng, Depu and Hu, Yihan and Xu, Tianshuo and Ahuja, Narendra and Zhan, Wei},
journal = {arXiv preprint arXiv:2610.08978},
year = {2026}
}
Acknowledgments
This model builds on ZipSplat (code Apache-2.0, weights CC BY-NC 4.0), Depth Anything 3 and DINOv2 (Apache-2.0). It is trained on DL3DV-10K and RealEstate10K. We thank the authors for making their work available.
Model tree for AppliedIntuitionResearch/S2Tok
Base model
veichta/zipsplat