Generative Models
This page is a reference manual. A world model has to answer "what happens next", and there is usually more than one possible future. That calls for a model that can represent a distribution and sample from it, not just give a point estimate. From the RSSM in Lab 3 to Genie, Sora and Cosmos, generative models are the other half of world models. This page only covers what the main tracks actually use.
Why world models need generative modelsâ
For a deterministic predictor trained with MSE, the optimum is the conditional mean . When the future is multimodal (the ball might bounce left or right), the mean is a blurry "average future"; the deterministic baseline in Lab 3 and the long-range predictions in Lab 7 both reproduce this first-hand. A generative model instead learns the whole conditional distribution and then samples from it, so every sample is one sharp possible future.
Four model families answer the same question, "how do you represent a high-dimensional distribution and sample from it", with different trade-offs:
| Family | How the distribution is represented | Sampling | Where it shows up in world models |
|---|---|---|---|
| VAE | latent + decoder , trained with the ELBO | one step: sample , then decode | the V in World Models 2018, the stochastic state of the RSSM |
| Autoregressive (AR) | chain rule , often over VQ discrete tokens | token by token, sequentially | token world models such as IRIS, Genie, GAIA-1 |
| Diffusion | learns to denoise step by step, turning noise back into data | many iterative denoising steps | DIAMOND, GameNGen, Sora, Cosmos |
| Flow matching | learns a velocity field that "transports" noise to data | solve an ODE, possibly in very few steps | newer video models and the action heads of VLAs |
Minimal knowledge setâ
- VAE and ELBO (introduced in Deep Learning): . The reconstruction term asks to carry information; the KL term asks the posterior to stay close to the prior. The RSSM's training objective is the sequential version of this (Lab 3). Known weakness: pixel-wise Gaussian/Bernoulli likelihoods make samples blurry.
- Discrete tokens and VQ-VAE: a codebook cuts a continuous image into discrete tokens , after which a Transformer can model them autoregressively like a language model: . The upside is reusing LLM training and sampling tools directly; the cost is that the tokenizer's reconstruction quality is a ceiling, and token-by-token sampling is slow.
- Diffusion models (DDPM): the forward process adds Gaussian noise step by step, ; the network learns to predict the noise, Sampling starts from pure noise and denoises repeatedly. is the conditioning: text, past frames, actions. Set and the diffusion model becomes a world model.
- The score view: predicting the noise is equivalent (up to a scale) to estimating the score . This is the language that unifies DDPM, score-based SDEs and ODE sampling, and you will meet it again and again in papers.
- Flow matching / Rectified Flow: take the straight-line interpolation between noise and data and regress its velocity: Sampling integrates the ODE numerically from to . The straighter the paths, the fewer steps you need, which is crucial for real-time interactive world models (Module 10b).
- Classifier-free guidance (CFG): drop the condition at random during training and extrapolate at sampling time, . makes samples "obey the condition" more closely (e.g. follow the action more strictly), at the cost of diversity.
- Latent diffusion: first compress images/video into a low-dimensional latent with an autoencoder, then run diffusion in the latent. Stable Diffusion, Sora and Cosmos all use this structure, which shares its lineage with the world-model idea of "predicting the future in latent space".
- DiT: replace the diffusion U-Net with a Transformer over (spatiotemporal) patch tokens. Almost every contemporary video world model uses it as its backbone.
Minimal runnable exampleâ
The core of diffusion training is only a few lines. Below is a DDPM training step and sampling loop that works for any data tensor x0 (engineering details stripped so it maps onto the formulas):
import torch
T = 100
betas = torch.linspace(1e-4, 0.02, T)
alphas = 1 - betas
abar = torch.cumprod(alphas, 0) # \bar\alpha_t
def train_step(model, x0, cond, opt):
t = torch.randint(0, T, (x0.shape[0],))
eps = torch.randn_like(x0)
a = abar[t].view(-1, *[1] * (x0.dim() - 1))
xt = a.sqrt() * x0 + (1 - a).sqrt() * eps # forward noising
loss = ((model(xt, t, cond) - eps) ** 2).mean() # predict the noise
opt.zero_grad(); loss.backward(); opt.step()
return loss.item()
@torch.no_grad()
def sample(model, shape, cond):
x = torch.randn(shape) # start from pure noise
for t in reversed(range(T)):
tt = torch.full((shape[0],), t)
eps = model(x, tt, cond)
x = (x - betas[t] / (1 - abar[t]).sqrt() * eps) / alphas[t].sqrt()
if t > 0:
x = x + betas[t].sqrt() * torch.randn_like(x) # stochastic term of the reverse process
return x
Set cond to "the last few frames + the current action" and x0 to "the next frame", and this is the skeleton of the training loop of diffusion world models such as DIAMOND / GameNGen. Swap Lab 7's ConvGRU decoder for a conditional denoiser like this and long-range predictions no longer blur into the mean (but new problems appear: slow sampling and closed-loop drift).
When to consultâ
| Main-track module | Generative-model knowledge used |
|---|---|
| Track A 04 Representation Learning | autoencoders, VQ discretization, generative vs discriminative representations |
| Track A 06 RSSM, 07 World Models 2018 | VAE, ELBO, KL balancing, MDNs (mixture density networks) as multimodal prediction |
| Track A 10 Video World Models | diffusion, flow matching, latent diffusion, DiT |
| Track A 10b Interactive World Models | conditional diffusion, few-step sampling, CFG, autoregressive token models |
| Track B 12 VLA & World-Action Models | diffusion / flow matching as action-generation heads (diffusion policy) |
| Lab 3, Lab 7 | worked counterexamples of "deterministic prediction â blurry mean" |
Best external resourcesâ
- Stanford CS236 Deep Generative Models (deepgenerativemodels.github.io): a systematic course on VAEs, autoregressive models, GANs, flows and diffusion.
- MIT 6.S184 Generative AI with Stochastic Differential Equations (diffusion.csail.mit.edu): flow matching and diffusion taught from ODEs/SDEs, with notes and exercises; the most direct material for understanding why flow matching allows few-step sampling.
- Lilian Weng, What are Diffusion Models? (blog): the survey blog post with the most complete derivations.
- Calvin Luo, Understanding Diffusion Models: A Unified Perspective (arXiv:2208.11970): derives everything from the ELBO through to score matching, placing VAEs and diffusion in one framework.
- CIS6280 L07 / L11 (course homepage): latent-variable generative models and diffusion/flow matching in the world-model context.
Key papersâ
- van den Oord et al. (2017), Neural Discrete Representation Learning (VQ-VAE, arXiv:1711.00937).
- Ho et al. (2020), Denoising Diffusion Probabilistic Models (DDPM, arXiv:2006.11239).
- Song et al. (2021), Score-Based Generative Modeling through Stochastic Differential Equations (arXiv:2011.13456).
- Rombach et al. (2022), High-Resolution Image Synthesis with Latent Diffusion Models (arXiv:2112.10752).
- Ho & Salimans (2022), Classifier-Free Diffusion Guidance (arXiv:2207.12598).
- Lipman et al. (2023), Flow Matching for Generative Modeling (arXiv:2210.02747); Liu et al. (2023), Flow Straight and Fast (Rectified Flow, arXiv:2209.03003).
- Peebles & Xie (2023), Scalable Diffusion Models with Transformers (DiT, arXiv:2212.09748).
Self-checkâ
You are ready when you can:
- Explain why a deterministic predictor trained with MSE outputs blur on multimodal futures, while diffusion samples are sharp.
- Write down DDPM's forward noising formula and the noise-prediction loss, and say what the condition corresponds to in a world model.
- State the flow-matching training objective, and why "straighter paths â fewer sampling steps" matters for real-time world models.
- Compare the strengths and weaknesses of the VQ + autoregressive route and the diffusion route (sampling speed, quality ceiling, reuse of the LLM toolchain).
Nextâ
Back to the main track: Track A Module 10 ¡ Video World Models â Module 10b ¡ Interactive World Models