Planning with World Models
Why this mattersâ
With a world model, decisions do not require a policy network: at each step, search action sequences inside the model, execute the first action, and plan again. The MPC paradigm is simple, flexible, and interpretable â the standard toolkit of latent planning, and the most direct touchstone of world-model quality.
Visual Intuitionâ
Each MPC step does three things: from the current state, roll out multiple candidate action sequences inside the model (a branching tree), pick the best sequence by cumulative reward, but execute only the first action â once the next real observation arrives, plan again from scratch (receding horizon).
Core Ideaâ
The core of model predictive control (MPC) is the receding horizon: plan steps but execute only one, using the latest observation to continually correct model error. This "open-loop planning + closed-loop execution" structure makes MPC inherently more robust to model error than a pure open-loop rollout â errors are reset by real observations at every step.
The dynamics learned by a world model are usually non-differentiable, high-dimensional, and stochastic, so gradient planning does not apply and zero-order sampling planners became the standard. CEM (cross-entropy method) maintains a Gaussian distribution over action sequences: sample â evaluate inside the model â refit the distribution on the elites, iterating a few rounds until convergence. MPPI derives a soft weighted update from information theory (exponential weights instead of a hard elite cutoff), is GPU-parallel-friendly, and is common in real-time robot control.
TD-MPC represents a third route: MPC + a learned value function. Within a short horizon, roll out in latent space with CEM/MPPI, and attach a value function at the horizon's end as a backstop â the trade-off between planning depth and model error is automatically resolved by value learning. This is also the bridge from "planning" smoothly over to "policy learning" (Module 11).
Key Conceptsâ
- Receding horizon: plan H steps, execute 1 step, replan.
- CEM: a distribution-iterating zero-order planner â elite selection + resampling.
- MPPI: exponentially importance-weighted sampling MPC with good real-time properties.
- Latent planning: rolling out in the latent space of an RSSM/encoder, not in pixels or the true state.
- Value expansion (TD-MPC): a value function at the end of the planning horizon â short planning, long sight.
Core Equationsâ
The per-step MPC optimization:
The MPPI weight update:
University Lectureâ
| Course | Lecture | Link |
|---|---|---|
| UPenn CIS 6280 World Models | L09 Planning and Control with World Models (MPC/CEM/MPPI) | slides |
| MIT 16.485 VNAV | L08âL10 Trajectory Optimization (the classical counterpart in continuous spaces) | course homepage |
Papersâ
- Must Read: Hansen, Wang & Su (2022), Temporal Difference Learning for Model Predictive Control (TD-MPC, arXiv:2203.04955).
- Recommended: Williams et al. (2017), Information Theoretic MPC for Model-Based Reinforcement Learning (MPPI, arXiv:1707.02342); Chua et al. (2018), Deep RL in a Handful of Trials (PETS, arXiv:1805.12114) â the classic combination of probabilistic ensembles + CEM.
- Optional: Schrittwieser et al. (2020), Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (MuZero, arXiv:1911.08265) â the pinnacle of the tree-search route.
Hands-onâ
Lab 4: implement CEM and MPPI on the latent dynamics from Labs 2/3, and systematically compare how horizon and sample count affect return.
Open Lab: Open in Colab: lab04_mpc_cem_planning
Check Your Understandingâ
- Why is MPC more robust to model error than open-loop rollout?
- How do the update rules of CEM and MPPI differ?
- Why can TD-MPC use a shorter planning horizon?
Show answer
- Because after every executed action it re-anchors the state with a real observation before replanning â model error never gets a chance to compound over time. Open-loop rollout errors accumulate all the way, and the longer the horizon, the more absurd the result.
- CEM refits the Gaussian on the top-return elite subset (a hard cutoff); MPPI softly weights all samples by , with temperature controlling exploration. MPPI uses information more fully, needs no elite fraction, and is empirically smoother.
- Because the value function at the horizon's end compresses "the long-term return beyond" into one estimate: planning only needs to cover short-term detail while the learned value covers the long term. The tension between planning depth and model error is thereby relieved.
Takeawayâ
- MPC = plan inside the model + execute only the first step + keep resetting with real observations â the natural antidote to model error.
- CEM/MPPI are the standard zero-order planners on learned dynamics; vectorized, they are extremely fast on GPUs.
- TD-MPC bridges planning and policy learning with a value function â the bridge to Module 11.
Next Moduleâ
Module 10: Video World Models â upgrading "predict the future" to "generate the future."