Reinforcement Learning
This page is a reference manual. To borrow CIS6280's official disclaimer: this is not an RL course. You only need enough vocabulary to follow "how world models serve decision-making."
Minimal knowledge set
- MDP: the five-tuple ; a POMDP is its partially observable version (expanded in Track A Module 02).
- Bellman equation: — the recursive definition of the value function.
- Value / Policy: the relationship among , , and ; TD-MPC's "model + value function" combination rests on this distinction.
- Model-free vs model-based: everything in this course sits on the model-based side; understanding the contrast is enough.
- Policy gradients and actor-critic: Dreamer trains an actor-critic in imagination; you only need to know the structure, not to derive it.
When to consult
| Main-track module | RL used |
|---|---|
| Track A 02 | MDP/POMDP formalism |
| Track A 08 Dreamer | Actor-critic, return estimation |
| Track A 09 Planning | The return-maximization view of trajectory optimization |
| Track A 11 World Model + Policy | Dyna architecture, model bias |
Best external resources
- Sutton & Barto, Reinforcement Learning: An Introduction (free PDF): Chapters 1–6 plus Chapter 8 (Dyna) cover everything needed.
- OpenAI Spinning Up (spinningup.openai.com): the friendliest introduction to policy gradients and actor-critic.
- CIS6280 L09–L10 (course homepage): RL machinery explained in the context of "using world models."
Self-check
You are ready when you can write the Bellman equation and explain each term, articulate where the sample-efficiency gap between model-free and model-based comes from, and sketch what each of the two actor-critic networks outputs.
Next
Back to the main tracks: Track A Module 02