Skip to main content

Generative Models

This page is a reference manual. A world model has to answer "what happens next", and there is usually more than one possible future. That calls for a model that can represent a distribution and sample from it, not just give a point estimate. From the RSSM in Lab 3 to Genie, Sora and Cosmos, generative models are the other half of world models. This page only covers what the main tracks actually use.

Why world models need generative models​

For a deterministic predictor trained with MSE, the optimum is the conditional mean E[xt+1âˆŖx≤t,at]\mathbb{E}[x_{t+1} \mid x_{\le t}, a_t]. When the future is multimodal (the ball might bounce left or right), the mean is a blurry "average future"; the deterministic baseline in Lab 3 and the long-range predictions in Lab 7 both reproduce this first-hand. A generative model instead learns the whole conditional distribution p(xt+1âˆŖx≤t,at)p(x_{t+1} \mid x_{\le t}, a_t) and then samples from it, so every sample is one sharp possible future.

Four model families answer the same question, "how do you represent a high-dimensional distribution and sample from it", with different trade-offs:

FamilyHow the distribution is representedSamplingWhere it shows up in world models
VAElatent zz + decoder p(xâˆŖz)p(x \mid z), trained with the ELBOone step: sample zz, then decodethe V in World Models 2018, the stochastic state of the RSSM
Autoregressive (AR)chain rule ∏ip(xiâˆŖx<i)\prod_i p(x_i \mid x_{<i}), often over VQ discrete tokenstoken by token, sequentiallytoken world models such as IRIS, Genie, GAIA-1
Diffusionlearns to denoise step by step, turning noise back into datamany iterative denoising stepsDIAMOND, GameNGen, Sora, Cosmos
Flow matchinglearns a velocity field that "transports" noise to datasolve an ODE, possibly in very few stepsnewer video models and the action heads of VLAs

Minimal knowledge set​

  • VAE and ELBO (introduced in Deep Learning): log⁥p(x)â‰ĨEq(zâˆŖx)[log⁥p(xâˆŖz)]−DKL(q(zâˆŖx) âˆĨ p(z))\log p(x) \geq \mathbb{E}_{q(z|x)}[\log p(x \mid z)] - D_{KL}(q(z \mid x) \,\|\, p(z)). The reconstruction term asks zz to carry information; the KL term asks the posterior to stay close to the prior. The RSSM's training objective is the sequential version of this (Lab 3). Known weakness: pixel-wise Gaussian/Bernoulli likelihoods make samples blurry.
  • Discrete tokens and VQ-VAE: a codebook cuts a continuous image into discrete tokens xâ†Ļ(k1,â€Ļ,kn)x \mapsto (k_1, \dots, k_n), after which a Transformer can model them autoregressively like a language model: p(k1,â€Ļ,kn)=∏ip(kiâˆŖk<i)p(k_1, \dots, k_n) = \prod_i p(k_i \mid k_{<i}). The upside is reusing LLM training and sampling tools directly; the cost is that the tokenizer's reconstruction quality is a ceiling, and token-by-token sampling is slow.
  • Diffusion models (DDPM): the forward process adds Gaussian noise step by step, xt=αˉt x0+1−αˉt Īĩx_t = \sqrt{\bar\alpha_t}\, x_0 + \sqrt{1 - \bar\alpha_t}\, \epsilon; the network learns to predict the noise, L=Ex0,Īĩ,tâˆĨĪĩ−Īĩθ(xt,t,c)âˆĨ2\mathcal{L} = \mathbb{E}_{x_0, \epsilon, t} \big\| \epsilon - \epsilon_\theta(x_t, t, c) \big\|^2 Sampling starts from pure noise and denoises repeatedly. cc is the conditioning: text, past frames, actions. Set c=(x≤t,at)c = (x_{\le t}, a_t) and the diffusion model becomes a world model.
  • The score view: predicting the noise is equivalent (up to a scale) to estimating the score ∇xlog⁥pt(x)\nabla_x \log p_t(x). This is the language that unifies DDPM, score-based SDEs and ODE sampling, and you will meet it again and again in papers.
  • Flow matching / Rectified Flow: take the straight-line interpolation xt=(1−t) x0+t x1x_t = (1 - t)\, x_0 + t\, x_1 between noise x0x_0 and data x1x_1 and regress its velocity: L=Ex0,x1,tâˆĨvθ(xt,t,c)−(x1−x0)âˆĨ2\mathcal{L} = \mathbb{E}_{x_0, x_1, t} \big\| v_\theta(x_t, t, c) - (x_1 - x_0) \big\|^2 Sampling integrates the ODE x˙=vθ(x,t,c)\dot x = v_\theta(x, t, c) numerically from t=0t=0 to t=1t=1. The straighter the paths, the fewer steps you need, which is crucial for real-time interactive world models (Module 10b).
  • Classifier-free guidance (CFG): drop the condition at random during training and extrapolate at sampling time, Īĩ^=Īĩθ(xt,∅)+w (Īĩθ(xt,c)−Īĩθ(xt,∅))\hat\epsilon = \epsilon_\theta(x_t, \varnothing) + w\,\big(\epsilon_\theta(x_t, c) - \epsilon_\theta(x_t, \varnothing)\big). w>1w > 1 makes samples "obey the condition" more closely (e.g. follow the action more strictly), at the cost of diversity.
  • Latent diffusion: first compress images/video into a low-dimensional latent with an autoencoder, then run diffusion in the latent. Stable Diffusion, Sora and Cosmos all use this structure, which shares its lineage with the world-model idea of "predicting the future in latent space".
  • DiT: replace the diffusion U-Net with a Transformer over (spatiotemporal) patch tokens. Almost every contemporary video world model uses it as its backbone.

Minimal runnable example​

The core of diffusion training is only a few lines. Below is a DDPM training step and sampling loop that works for any data tensor x0 (engineering details stripped so it maps onto the formulas):

import torch

T = 100
betas = torch.linspace(1e-4, 0.02, T)
alphas = 1 - betas
abar = torch.cumprod(alphas, 0) # \bar\alpha_t

def train_step(model, x0, cond, opt):
t = torch.randint(0, T, (x0.shape[0],))
eps = torch.randn_like(x0)
a = abar[t].view(-1, *[1] * (x0.dim() - 1))
xt = a.sqrt() * x0 + (1 - a).sqrt() * eps # forward noising
loss = ((model(xt, t, cond) - eps) ** 2).mean() # predict the noise
opt.zero_grad(); loss.backward(); opt.step()
return loss.item()

@torch.no_grad()
def sample(model, shape, cond):
x = torch.randn(shape) # start from pure noise
for t in reversed(range(T)):
tt = torch.full((shape[0],), t)
eps = model(x, tt, cond)
x = (x - betas[t] / (1 - abar[t]).sqrt() * eps) / alphas[t].sqrt()
if t > 0:
x = x + betas[t].sqrt() * torch.randn_like(x) # stochastic term of the reverse process
return x

Set cond to "the last few frames + the current action" and x0 to "the next frame", and this is the skeleton of the training loop of diffusion world models such as DIAMOND / GameNGen. Swap Lab 7's ConvGRU decoder for a conditional denoiser like this and long-range predictions no longer blur into the mean (but new problems appear: slow sampling and closed-loop drift).

When to consult​

Main-track moduleGenerative-model knowledge used
Track A 04 Representation Learningautoencoders, VQ discretization, generative vs discriminative representations
Track A 06 RSSM, 07 World Models 2018VAE, ELBO, KL balancing, MDNs (mixture density networks) as multimodal prediction
Track A 10 Video World Modelsdiffusion, flow matching, latent diffusion, DiT
Track A 10b Interactive World Modelsconditional diffusion, few-step sampling, CFG, autoregressive token models
Track B 12 VLA & World-Action Modelsdiffusion / flow matching as action-generation heads (diffusion policy)
Lab 3, Lab 7worked counterexamples of "deterministic prediction → blurry mean"

Best external resources​

  • Stanford CS236 Deep Generative Models (deepgenerativemodels.github.io): a systematic course on VAEs, autoregressive models, GANs, flows and diffusion.
  • MIT 6.S184 Generative AI with Stochastic Differential Equations (diffusion.csail.mit.edu): flow matching and diffusion taught from ODEs/SDEs, with notes and exercises; the most direct material for understanding why flow matching allows few-step sampling.
  • Lilian Weng, What are Diffusion Models? (blog): the survey blog post with the most complete derivations.
  • Calvin Luo, Understanding Diffusion Models: A Unified Perspective (arXiv:2208.11970): derives everything from the ELBO through to score matching, placing VAEs and diffusion in one framework.
  • CIS6280 L07 / L11 (course homepage): latent-variable generative models and diffusion/flow matching in the world-model context.

Key papers​

  • van den Oord et al. (2017), Neural Discrete Representation Learning (VQ-VAE, arXiv:1711.00937).
  • Ho et al. (2020), Denoising Diffusion Probabilistic Models (DDPM, arXiv:2006.11239).
  • Song et al. (2021), Score-Based Generative Modeling through Stochastic Differential Equations (arXiv:2011.13456).
  • Rombach et al. (2022), High-Resolution Image Synthesis with Latent Diffusion Models (arXiv:2112.10752).
  • Ho & Salimans (2022), Classifier-Free Diffusion Guidance (arXiv:2207.12598).
  • Lipman et al. (2023), Flow Matching for Generative Modeling (arXiv:2210.02747); Liu et al. (2023), Flow Straight and Fast (Rectified Flow, arXiv:2209.03003).
  • Peebles & Xie (2023), Scalable Diffusion Models with Transformers (DiT, arXiv:2212.09748).

Self-check​

You are ready when you can:

  1. Explain why a deterministic predictor trained with MSE outputs blur on multimodal futures, while diffusion samples are sharp.
  2. Write down DDPM's forward noising formula and the noise-prediction loss, and say what the condition cc corresponds to in a world model.
  3. State the flow-matching training objective, and why "straighter paths → fewer sampling steps" matters for real-time world models.
  4. Compare the strengths and weaknesses of the VQ + autoregressive route and the diffusion route (sampling speed, quality ceiling, reuse of the LLM toolchain).

Next​

Back to the main track: Track A Module 10 · Video World Models → Module 10b · Interactive World Models