Skip to main content

Observation, State and POMDP

Why this matters​

Are you modeling what is "seen" or what is "real"? Confusing observation with state is the most common beginner mistake in world models — and the root of every difficulty in partially observable environments (robotics, first-person video).

Visual Intuition​

Observation vs state

The true state sts_t hides behind the scenes (the ball's real velocity and position); the observation oto_t is only its noisy projection (a frame of pixels). The agent can only infer the state from its observation history — and that inference is the belief state btb_t.

Core Idea​

The state is a complete description of the world right now: given sts_t and the action ata_t, the distribution over the future is uniquely determined (the Markov property). An observation is a lossy sample of the state passed through the observation model p(ot∣st)p(o_t \mid s_t) — a camera cannot see occluded objects; radar cannot measure color.

In the fully observable case (MDP) ot=sto_t = s_t, and deciding requires only the current frame; in the partially observable case (POMDP), a single-frame observation does not determine the state, so the agent must maintain a belief state: bt=p(st∣o1:t,a1:t−1)b_t = p(s_t \mid o_{1:t}, a_{1:t-1}), the posterior over the state conditioned on the entire history.

The belief state is the classical prototype of the "learned latent state": it is itself a sufficient statistic and can be updated recursively (the Bayes filter). All the later latent-dynamics modules (05–08) are, in essence, learning a differentiable, compact belief state.

Key Concepts​

  • Markov property: the state is a sufficient statistic for the future; history can be discarded.
  • Observation model: p(ot∣st)p(o_t \mid s_t), the bridge between "real" and "seen."
  • Belief state btb_t: the posterior over the state under partial observability, updated recursively.
  • History compression: compressing o1:t,a1:t−1o_{1:t}, a_{1:t-1} into a fixed-dimensional btb_t — the motivation for representation learning.
  • Sufficient statistic: whether a belief or latent state retains all information needed for prediction.

Core Equations​

The POMDP seven-tuple:

(S,A,O,P,Ω,R,γ)(\mathcal{S}, \mathcal{A}, \mathcal{O}, P, \Omega, R, \gamma)

where S\mathcal{S} is the state space, A\mathcal{A} the action space, O\mathcal{O} the observation space, P(s′∣s,a)P(s' \mid s,a) the transition, Ω(o∣s′,a)\Omega(o \mid s', a) the observation model, RR the reward, and γ\gamma the discount.

Recursive belief update (Bayes filter):

bt+1(s′)∝Ω(ot+1∣s′,at)∑sP(s′∣s,at) bt(s)b_{t+1}(s') \propto \Omega(o_{t+1} \mid s', a_t) \sum_{s} P(s' \mid s, a_t)\, b_t(s)

University Lecture​

CourseLectureLink
UPenn CIS 6280 World ModelsL02 History, Foundations, Probabilistic Formulationslides

Papers​

  • Must Read: Kaelbling, Littman & Cassandra (1998), Planning and Acting in Partially Observable Stochastic Domains (classic journal paper; search the title for the PDF) — the bible of POMDPs; §2 suffices.
  • Recommended: Ha & Schmidhuber (2018), World Models (arXiv:1803.10122) §2 — see how an RNN compresses history into a hidden state.

Hands-on​

Open in ColabOpen in Colab: lab00_tiny_world

Try one change in the Lab 0 environment: make the observation contain position but not velocity — you have just manufactured a POMDP by hand, and a random policy's performance will degrade immediately.

Check Your Understanding​

  1. Is a single Atari game frame a state or an observation?
  2. What is the relationship between a belief state and an RNN's hidden state?
  3. Why is "remembering history" necessary in a POMDP?
Show answer
  1. An observation. A single frame does not reveal state components such as the ball's velocity direction or the opponent's intent; this is exactly why Atari agents usually stack 4 frames — using short-term history to approximate a sufficient statistic.
  2. An RNN's hidden state is a learnable, deterministic approximation of a belief state: it likewise compresses history into a fixed-dimensional vector and updates recursively, but discards the explicit representation of uncertainty (RSSM will bring that back).
  3. Because a single-frame observation is not a sufficient statistic: two different true states can produce the same observation (perceptual aliasing). Only history can disambiguate — for instance, estimating velocity from consecutive frames.

Takeaway​

  • Observation ≠ state: an observation is a noisy, lossy projection of the state; confusing the two is a common error.
  • The belief state btb_t is the correct "internal state" under partial observability and is updated recursively.
  • Learned latent states are differentiable approximations of belief states — this foreshadows Modules 05–08.

Next Module​

Module 03: State Space Models — the closed-form solution of recursive belief updating in the Gaussian case.