OpenWAM Pretraining
OpenWAM video pretraining model weights, trained using single-view and multi view video data with text conditioning.
Training objective of this revision: chunk-causal (autoregressive) video prediction. Earlier revisions of this repository were trained with a bidirectional prefix/suffix objective; this revision continues from them with a teacher-forced chunk-causal objective:
- The first latent frame is a clean context frame; later latent frames form chunks of 4.
- The model sees a noisy copy and a clean copy of the clip. A chunk being denoised attends to its own noisy frames and to the clean frames of earlier chunks only; clean frames attend to clean frames of their own and earlier chunks. No chunk attends to a later one.
- Timesteps are sampled per frame. For half of the samples the clean copy is noised as an augmentation.
Use these weights chunk by chunk: denoise one chunk at a time while attending to the clean context and the previously generated chunks. Generating all future frames in one bidirectional pass is not the mode this revision was trained for and gives worse results than the earlier revisions; those remain available through the repository history.
This repository contains model weights only, with a portable reference configuration and checksums. Optimizer and scheduler state, datasets, and frontend assets are not included. The PyTorch payload contains model_state_dict. In the reference configuration the objective is selected by training.use_teacher_forcing, training.chunk_size, training.window_size and training.noisy_video_condition_prob.
See the OpenWAM pretraining source and documentation for dataset processing and training instructions. Configure the external video frontend, text encoder, tokenizer, and local asset paths before use.