Skip to main content

World Models & Spatial Intelligence

The course in 10 secondsโ€‹

A world model is an agent's predictable representation of the world: given the current state and an action, it answers "what happens next." Spatial intelligence grounds that ability in the 3D/4D physical world: perceiving geometry, estimating one's own state, remembering the environment, and navigating and acting within it. This open course places both in a single framework: represent the world โ†’ predict the world โ†’ act in the world.

Course overview

The figure above is the course's global map: the left half is the "world model" thread (representation, prediction, planning), the right half is the "spatial intelligence" thread (geometry, 3D representations, SLAM, robotics), and the two converge at "action and evaluation."

Two tracks, how to chooseโ€‹

๐Ÿง‘โ€๐Ÿ”ฌ Track A ยท World Model Scientist๐Ÿค– Track B ยท Spatial & Embodied Intelligence
GoalBuild models: learn to construct, train, and evaluate world modelsBuild systems: assemble spatial-intelligence systems from world models and 3D representations
AudienceLearners aiming at world models / generative models / MBRL researchRobotics, autonomous driving, AR/VR, spatial computing
Core contentRepresentation learning โ†’ latent dynamics โ†’ RSSM/Dreamer โ†’ planning โ†’ video/interactive world models โ†’ evaluationCamera geometry โ†’ depth/point clouds โ†’ NeRF/3DGS โ†’ state estimation/SLAM โ†’ navigation/VLA
LabsLabs 0โ€“4, 7, 9, 10Labs 0, 1, 5, 6, 8, 9, 10
ExitReproduce a Dreamer-class system; complete the Capstone world modelShip a perception โ†’ state-estimation โ†’ planning closed loop; complete the spatial Capstone
Suggested time14 modules, about 8โ€“10 weeks13 modules, about 8โ€“10 weeks

The two tracks are parallel, not sequential: building video world models does not require SLAM, and doing SLAM does not require diffusion. If you are unsure, read the first two modules of both tracks before deciding; you can also switch tracks at any time.

How to use each moduleโ€‹

Every module follows the same template. Reading the sections in their fixed order builds stable expectations:

  1. Why this matters: one paragraph on why the topic is worth learning;
  2. Visual Intuition: build intuition from the figures first, then read the text;
  3. Core Idea / Key Concepts / Core Equations: the core principles, terminology, and formulas;
  4. University Lecture: official links to the corresponding lectures at top universities worldwide, for deeper study;
  5. Papers: tiered as Must Read / Recommended / Optional โ€” you do not need to read them all;
  6. Hands-on: the corresponding Colab lab โ€” hands-on work is mandatory, not optional;
  7. Check Your Understanding: self-test questions; if you cannot answer, go back and reread;
  8. Takeaway / Next Module: a three-sentence summary plus an explicit next step.

The progress bar at the top of each module (ModuleProgress) tracks your learning progress.

Prerequisites self-checkโ€‹

This course assumes graduate-level machine-learning background. If you can answer "yes" to every item below, start right away; otherwise, fill gaps on demand from Foundations โ€” Foundations is a reference manual, not a mandatory starting point:

  • Can do matrix factorization and least squares, and understand the geometric meaning of eigenvalues โ†’ Linear Algebra
  • Comfortable with conditional distributions, Bayes' rule, Gaussian distributions, and covariance โ†’ Probability
  • Can write training loops in PyTorch and manage DataLoaders and GPU memory โ†’ PyTorch
  • Understand backpropagation, VAE/ELBO, and common regularization techniques โ†’ Deep Learning
  • Can explain attention, causal masking, and how ViT tokenizes images โ†’ Transformers
  • Can write the pinhole camera projection and understand epipolar geometry and feature matching โ†’ Computer Vision
  • Know the basic forms of MDPs, the Bellman equation, and policy gradients โ†’ Reinforcement Learning

Provenanceโ€‹

This course is built on the skeleton of UPenn CIS 6280: World Models (first offered Fall 2026), and integrates lectures and assignment designs from 11 university courses including Stanford CS231A, CMU 16-825, MIT VNAV, and ETH/UZH VAMR. The University Lecture table in each module is the provenance annotation.

Next Moduleโ€‹