Latent Dynamics
Why this matters
Dynamics in pixel space are both high-dimensional and full of irrelevant detail. Latent dynamics splits the world model into two steps — "learn a representation + learn dynamics inside it" — the shared skeleton from PlaNet/Dreamer to modern video models, and the first world model you will train with your own hands.
Visual Intuition
The encoder compresses the observation into a latent state ; the dynamics model predicts in latent space; the decoder translates latent states back into observations for supervision. Prediction happens entirely in the low-dimensional latent space — the key to computational efficiency and predictability.
Core Idea
A latent-dynamics model has three learnable components: an encoder (inferring the current latent state), a transition model (predicting in latent space), and a decoder (providing the training signal). The whole system is trained jointly under the VAE/ELBO framework: the reconstruction term keeps the representation faithful, and the KL term aligns the posterior with the prior (i.e., the transition model's prediction).
This structure answers Module 02's history-compression question: is the learned belief state — it compresses the observation history (via an RNN or a filtered sequence), retains what prediction needs, at a dimensionality far below pixels.
The choice of training objective shapes the model's character: a one-step objective (predict only the next frame) is simple and stable but errors compound at rollout time; a multi-step objective (roll out open-loop for several steps during training) optimizes the deployment scenario directly but is harder to train. This one-step-vs-rollout fault line erupts fully in Module 12.
Key Concepts
- Encoder / Decoder: bidirectional translators between observations and latent states.
- Latent-state design criteria: inferable, predictable, information-sufficient — the three pull against each other.
- ELBO: the variational objective with a reconstruction term and a KL regularizer.
- One-step vs rollout objectives: the mismatch between the training objective and the deployment scenario.
- Error compounding: tiny per-step errors are compounded and amplified through recursive rollout.
Core Equations
The ELBO for latent dynamics:
University Lecture
| Course | Lecture | Link |
|---|---|---|
| UPenn CIS 6280 World Models | L07 Latent-Variable and Adversarial Models (VAE/ELBO) | slides |
| UPenn CIS 6280 World Models | L08 Latent World Models | slides |
Papers
- Must Read: Hafner et al. (2019), Learning Latent Dynamics for Planning from Pixels (PlaNet, arXiv:1811.04551) — the paper that established the latent-dynamics paradigm.
- Recommended: Kingma & Welling (2014), Auto-Encoding Variational Bayes (arXiv:1312.6114) — the original source of the ELBO derivation.
- Optional: Deisenroth & Rasmussen (2011), PILCO (ICML; search the title) — the pioneer of non-image latent dynamics with uncertainty propagation.
Hands-on
Lab 2 is this module's full hands-on: train an encoder + MLP/GRU dynamics on image rollouts collected in the Lab 0 environment, and plot one-step and 10-step open-loop prediction error curves — you will see error compounding with your own eyes.
Check Your Understanding
- Why not learn directly in pixel space?
- What does the KL term of the ELBO do in latent dynamics?
- If the one-step error is small, why can a 10-step rollout still diverge completely?
Show answer
- Pixel space is high-dimensional and full of prediction-irrelevant stochastic detail (textures, noise); modeling there directly wastes capacity and confronts the dynamics learner with irrelevant randomness. The latent space first compresses out a "predictable core," and dynamics are easier to learn in a low-dimensional space.
- The KL term pulls the posterior toward the prior — forcing "states that can be inferred" to coincide with "states that the dynamics can predict." Without it, the encoder could encode unpredictable information and rollout would fail immediately.
- Rollout is closed-loop: the model's prediction error at step becomes the input at step , so errors are compounded and amplified by the dynamics; moreover, the input distribution during rollout drifts away from the training distribution (OOD), so one-step accuracy cannot guarantee multi-step stability.
Takeaway
- Latent dynamics = encoder + latent-space transition + decoder; is the learned belief state.
- The ELBO's KL term aligns "inferable" with "predictable" — the invisible pillar of latent dynamics.
- One-step accuracy does not imply usable rollouts — error compounding is a theme running through the whole course.
Next Module
Module 06: RSSM — add the hybrid "deterministic memory + stochasticity" structure to latent dynamics.