NeRF and Gaussian Splatting
Why this matters
NeRF and 3D Gaussian Splatting are the twin stars of modern 3D scene representation: one implicit and continuous, the other explicit and real-time. They are the storage format of the "spatial world model" — what can be rendered can predict novel views, and predicting novel views is the spatial version of a world model.
Visual Intuition
NeRF represents a scene with an MLP: given a 3D position and viewing direction, it outputs density and color. To render, it samples points along each pixel's ray and accumulates density-weighted colors via the volume-rendering equation — the entire pipeline is differentiable and trained directly on photometric error.
3DGS represents a scene explicitly with hundreds of thousands of anisotropic Gaussian primitives, each with a position, covariance, color, and opacity. Rendering projects them onto the image plane, sorts them, and alpha-blends — pure rasterization, in real time.
Core Idea
NeRF (2020) represents a scene as a continuous function : given a spatial position and viewing direction, it outputs color and volume density. Novel views are synthesized by volume rendering: integrating density-weighted color along each ray. Training requires no 3D annotation whatsoever — the photometric reconstruction error on multi-view photos is the entire supervision; multi-view consistency (the core constraint from Module 01) becomes the training signal. Positional encoding fixes the spectral bias of MLPs, letting the network express high-frequency detail.
3D Gaussian Splatting (2023) takes the opposite, explicit route: the scene is a set of hundreds of thousands of 3D Gaussians, rendering is classical splatting rasterization, and training allocates capacity automatically through adaptive density control (splitting large Gaussians, cloning small ones, pruning transparent ones). Compared to NeRF it trains an order of magnitude faster and renders in real time, at the cost of memory footprint and discretization artifacts.
From the world-model perspective, both are queryable world states: given any camera pose they produce an observation — exactly the spatial version of the state → observation interface. Their limitations (per-scene optimization, static-scene assumption) are addressed respectively by feed-forward reconstruction (DUSt3R/VGGT) and by the 4D representations of Module 05.
Key Concepts
- Radiance field: a continuous field mapping position + direction → color + density.
- Volume-rendering equation: a differentiable integral of density-weighted color along a ray.
- Positional encoding: the encoding trick that lets an MLP learn high-frequency detail.
- 3D Gaussian parameterization: position / covariance / spherical-harmonic color / opacity.
- Adaptive density control: the primitive creation/removal mechanism in 3DGS training.
Core Equations
Volume-rendering equation: