World Models 2018
Why this mattersâ
This course takes its name from this 2018 paper. With a minimalist three-stage architecture it proved that "an agent can train inside its own learned dream," and it exposed the typical failure modes of simple world models â understanding its limits is understanding the motivation for every later improvement.
Visual Intuitionâ
The three components are trained independently: the VAE does the seeing (compressing observations), the MDN-RNN does the dreaming (predicting latent evolution), and the Controller â only a few hundred parameters â decides on . Dream training means using M as the environment.
Core Ideaâ
Ha & Schmidhuber's World Models (2018) splits the agent into three independent modules: Vision (VAE) â Memory (MDN-RNN) â Controller (linear / CMA-ES). The key innovation is the MDN (Mixture Density Network): the RNN outputs the parameters of a mixture of Gaussians rather than a single-point prediction, letting the model express stochastic futures â the embryo of RSSM's stochastic state.
The most famous experiment is dream training on CarRacing: a policy trained entirely on imagined trajectories generated by the MDN-RNN, deployed zero-shot back into the real environment, still completes the track. This directly demonstrates Module 01's central claim: a world model can trade real trial-and-error for imagined trial-and-error.
The paper is equally honest about failure modes: in the VizDoom experiment, the policy learned to exploit loopholes in the dream (in the dream, monsters never shoot fireballs) â the first famous demonstration of model exploitation, and the subject of Module 12. The paper's intellectual lineage traces back to Schmidhuber's 1990s work on "curious model-building control systems."
Key Conceptsâ
- VAE (Vision): compresses pixels into a 32-dimensional .
- MDN-RNN (Memory): outputs mixture-of-Gaussians parameters, predicting the stochastic next latent state and termination probability.
- Controller: a small-parameter policy reading directly, evolved with CMA-ES.
- Dream training: training the policy on in-model rollouts with zero real interaction.
- Model exploitation: the policy games the model's flaws â invincible in the dream, crashing in reality.
Core Equationsâ
The MDN-RNN outputs a mixture of Gaussians:
University Lectureâ
| Course | Lecture | Link |
|---|---|---|
| UPenn CIS 6280 World Models | L01âL02 (historical context) and L08 Latent World Models | L08 slides |
Papersâ
- Must Read: Ha & Schmidhuber (2018), World Models (arXiv:1803.10122) â the full paper is this module, with a companion interactive blog.
- Recommended: Schmidhuber (1991), Curious Model-Building Control Systems (search the title) â the intellectual source; see the coupling of models and curiosity.
- Optional: Oh et al. (2015), Action-Conditional Video Prediction using Deep Networks in Atari Games (NeurIPS; search the title) â a parallel work on action-conditioned prediction from the same era.
Hands-onâ
Lab 3: the tiny RSSM implementation keeps World Models' "stage-wise training" option â you can train each component separately first, then jointly, and compare against end-to-end ELBO training.
Open Lab: Open in Colab: lab03_tiny_rssm
Check Your Understandingâ
- Why train the three components separately rather than end-to-end?
- How is the MDN's mixture of Gaussians better than single-point MSE prediction?
- What fundamental risk of model-based decision-making does "the policy exploiting dream loopholes" expose?
Show answer
- The engineering reality of 2018: end-to-end training of all three components was computationally expensive and unstable; divide-and-conquer made each step simple on its own (unsupervised VAE, supervised RNN, gradient-free CMA-ES) and fits the "replaceable modules" interface philosophy. The cost is inconsistent component objectives â the Dreamer family's end-to-end ELBO is exactly the response to this.
- Single-point MSE prediction outputs the "average" of stochastic futures (blurry, unrealistic); a mixture of Gaussians can express multimodal distributions (the car might turn left or right), so sampled dreams resemble real trajectories â otherwise the policy learns wrong dynamics from a blurry dream.
- Decision-time optimization actively pushes rollouts toward the regions of largest model error (an optimizer naturally seeks the model's loopholes) â model exploitation / model bias. Remedies include short horizons, uncertainty penalties, and closed-loop correction with real data â expanded in Modules 09 and 12.
Takeawayâ
- World Models 2018 first demonstrated "dream training" end-to-end with VAE + MDN-RNN + Controller.
- Stage-wise independent training is simple but misaligned in objectives; end-to-end joint training (PlaNet/Dreamer) is the next step.
- The dream-loophole experiment foreshadows model bias â the eternal theme of using models for decisions.
Next Moduleâ
Module 08: Dreamer â the benchmark system uniting RSSM, imagination, and actor-critic.