Observation, State and POMDP
Why this matters
Are you modeling what is "seen" or what is "real"? Confusing observation with state is the most common beginner mistake in world models — and the root of every difficulty in partially observable environments (robotics, first-person video).
Visual Intuition
The true state hides behind the scenes (the ball's real velocity and position); the observation is only its noisy projection (a frame of pixels). The agent can only infer the state from its observation history — and that inference is the belief state .
Core Idea
The state is a complete description of the world right now: given and the action , the distribution over the future is uniquely determined (the Markov property). An observation is a lossy sample of the state passed through the observation model — a camera cannot see occluded objects; radar cannot measure color.
In the fully observable case (MDP) , and deciding requires only the current frame; in the partially observable case (POMDP), a single-frame observation does not determine the state, so the agent must maintain a belief state: , the posterior over the state conditioned on the entire history.
The belief state is the classical prototype of the "learned latent state": it is itself a sufficient statistic and can be updated recursively (the Bayes filter). All the later latent-dynamics modules (05–08) are, in essence, learning a differentiable, compact belief state.
Key Concepts
- Markov property: the state is a sufficient statistic for the future; history can be discarded.
- Observation model: , the bridge between "real" and "seen."
- Belief state : the posterior over the state under partial observability, updated recursively.
- History compression: compressing into a fixed-dimensional — the motivation for representation learning.
- Sufficient statistic: whether a belief or latent state retains all information needed for prediction.
Core Equations
The POMDP seven-tuple:
where is the state space, the action space, the observation space, the transition, the observation model, the reward, and the discount.
Recursive belief update (Bayes filter):
University Lecture
| Course | Lecture | Link |
|---|---|---|
| UPenn CIS 6280 World Models | L02 History, Foundations, Probabilistic Formulation | slides |
Papers
- Must Read: Kaelbling, Littman & Cassandra (1998), Planning and Acting in Partially Observable Stochastic Domains (classic journal paper; search the title for the PDF) — the bible of POMDPs; §2 suffices.
- Recommended: Ha & Schmidhuber (2018), World Models (arXiv:1803.10122) §2 — see how an RNN compresses history into a hidden state.
Hands-on
Try one change in the Lab 0 environment: make the observation contain position but not velocity — you have just manufactured a POMDP by hand, and a random policy's performance will degrade immediately.
Check Your Understanding
- Is a single Atari game frame a state or an observation?
- What is the relationship between a belief state and an RNN's hidden state?
- Why is "remembering history" necessary in a POMDP?
Show answer
- An observation. A single frame does not reveal state components such as the ball's velocity direction or the opponent's intent; this is exactly why Atari agents usually stack 4 frames — using short-term history to approximate a sufficient statistic.
- An RNN's hidden state is a learnable, deterministic approximation of a belief state: it likewise compresses history into a fixed-dimensional vector and updates recursively, but discards the explicit representation of uncertainty (RSSM will bring that back).
- Because a single-frame observation is not a sufficient statistic: two different true states can produce the same observation (perceptual aliasing). Only history can disambiguate — for instance, estimating velocity from consecutive frames.
Takeaway
- Observation ≠ state: an observation is a noisy, lossy projection of the state; confusing the two is a common error.
- The belief state is the correct "internal state" under partial observability and is updated recursively.
- Learned latent states are differentiable approximations of belief states — this foreshadows Modules 05–08.
Next Module
Module 03: State Space Models — the closed-form solution of recursive belief updating in the Gaussian case.