World Models & Spatial Intelligence
The course in 10 secondsโ
A world model is an agent's predictable representation of the world: given the current state and an action, it answers "what happens next." Spatial intelligence grounds that ability in the 3D/4D physical world: perceiving geometry, estimating one's own state, remembering the environment, and navigating and acting within it. This open course places both in a single framework: represent the world โ predict the world โ act in the world.
The figure above is the course's global map: the left half is the "world model" thread (representation, prediction, planning), the right half is the "spatial intelligence" thread (geometry, 3D representations, SLAM, robotics), and the two converge at "action and evaluation."
Two tracks, how to chooseโ
| ๐งโ๐ฌ Track A ยท World Model Scientist | ๐ค Track B ยท Spatial & Embodied Intelligence | |
|---|---|---|
| Goal | Build models: learn to construct, train, and evaluate world models | Build systems: assemble spatial-intelligence systems from world models and 3D representations |
| Audience | Learners aiming at world models / generative models / MBRL research | Robotics, autonomous driving, AR/VR, spatial computing |
| Core content | Representation learning โ latent dynamics โ RSSM/Dreamer โ planning โ video/interactive world models โ evaluation | Camera geometry โ depth/point clouds โ NeRF/3DGS โ state estimation/SLAM โ navigation/VLA |
| Labs | Labs 0โ4, 7, 9, 10 | Labs 0, 1, 5, 6, 8, 9, 10 |
| Exit | Reproduce a Dreamer-class system; complete the Capstone world model | Ship a perception โ state-estimation โ planning closed loop; complete the spatial Capstone |
| Suggested time | 14 modules, about 8โ10 weeks | 13 modules, about 8โ10 weeks |
The two tracks are parallel, not sequential: building video world models does not require SLAM, and doing SLAM does not require diffusion. If you are unsure, read the first two modules of both tracks before deciding; you can also switch tracks at any time.
How to use each moduleโ
Every module follows the same template. Reading the sections in their fixed order builds stable expectations:
- Why this matters: one paragraph on why the topic is worth learning;
- Visual Intuition: build intuition from the figures first, then read the text;
- Core Idea / Key Concepts / Core Equations: the core principles, terminology, and formulas;
- University Lecture: official links to the corresponding lectures at top universities worldwide, for deeper study;
- Papers: tiered as Must Read / Recommended / Optional โ you do not need to read them all;
- Hands-on: the corresponding Colab lab โ hands-on work is mandatory, not optional;
- Check Your Understanding: self-test questions; if you cannot answer, go back and reread;
- Takeaway / Next Module: a three-sentence summary plus an explicit next step.
The progress bar at the top of each module (ModuleProgress) tracks your learning progress.
Prerequisites self-checkโ
This course assumes graduate-level machine-learning background. If you can answer "yes" to every item below, start right away; otherwise, fill gaps on demand from Foundations โ Foundations is a reference manual, not a mandatory starting point:
- Can do matrix factorization and least squares, and understand the geometric meaning of eigenvalues โ Linear Algebra
- Comfortable with conditional distributions, Bayes' rule, Gaussian distributions, and covariance โ Probability
- Can write training loops in PyTorch and manage DataLoaders and GPU memory โ PyTorch
- Understand backpropagation, VAE/ELBO, and common regularization techniques โ Deep Learning
- Can explain attention, causal masking, and how ViT tokenizes images โ Transformers
- Can write the pinhole camera projection and understand epipolar geometry and feature matching โ Computer Vision
- Know the basic forms of MDPs, the Bellman equation, and policy gradients โ Reinforcement Learning
Provenanceโ
This course is built on the skeleton of UPenn CIS 6280: World Models (first offered Fall 2026), and integrates lectures and assignment designs from 11 university courses including Stanford CS231A, CMU 16-825, MIT VNAV, and ETH/UZH VAMR. The University Lecture table in each module is the provenance annotation.
Next Moduleโ
- Track A โ Module 01: What is a World Model?
- Track B โ Module 01: What is Spatial Intelligence?