Dreamer
Why this matters
The Dreamer family proved that "imagining in a learned latent space and training the policy there" can sweep Atari, DM Control, robotics, and even Minecraft. It is the engineering pinnacle of MBRL and an unavoidable system for understanding how world models serve decision-making.
Visual Intuition
Three loops run simultaneously: the world model learns an RSSM from replay data; the actor-critic trains entirely on the RSSM's imagination rollouts; the real environment is touched only to collect data. Real interaction and policy learning are decoupled by imagination.
Core Idea
DreamerV1 (Dream to Control, 2020) took the key step: instead of a planner (PlaNet's CEM), train an actor-critic in imagination. Imagination rollouts are fully differentiable (the RSSM is a neural network), so value gradients can backpropagate through the dynamics directly into the actor — policy learning becomes far more efficient than gradient-free evolution or sampling-based planning.
DreamerV2 replaced the stochastic state with discrete categoricals, proving the power of discrete latent variables by reaching human-level Atari within a 200M-frame budget. DreamerV3 (2023) solved "generality": symlog prediction (unifying reward scales), percentile return normalization, free bits, unimix and other tricks let one set of hyperparameters work tuning-free across 150+ tasks — the first time a world model showed "foundation model"-style transfer.
Reading the three generations together, the evolution is: policy learning in imagination (V1) → latent-variable form and robustness (V2) → cross-domain universal hyperparameters and scale (V3). This thread is also a microcosm of the field's journey "from demo to system."