Camera and Geometry
Why this matters
The camera is the most basic "world → observation" model (recall the observation equation from Track A Module 02). Without understanding projection and multi-view constraints, everything later — depth, NeRF, SLAM — is just a pile of APIs. Geometry is the native language of spatial intelligence.
Visual Intuition
A world point is first transformed by the extrinsics into the camera frame, then projected to a pixel by the intrinsics . Two views of the same 3D point induce the epipolar constraint — the source of multi-view supervision without annotations.
Core Idea
The pinhole camera model projects a 3D point to a 2D pixel: the intrinsic matrix encodes focal length and principal point, and the extrinsics encode the camera's pose in the world. Real lenses add distortion parameters; calibration is the process of solving for these parameters from known patterns such as checkerboards — Zhang's method is the industrial standard.
Epipolar geometry captures the intrinsic constraint between two views: the images of the same 3D point satisfy . The fundamental matrix (intrinsics unknown) / essential matrix (intrinsics known) depends only on the relative pose of the two cameras, not on the scene — which means camera motion can be recovered from point correspondences (the eight-point algorithm, Nistér's five-point algorithm + RANSAC), and 3D structure can then be triangulated. This is the skeleton of SfM (Structure from Motion), and the shared classical core from COLMAP to SLAM front ends.
Another layer of geometric discipline is coordinate-frame management: once the transform chain among world, camera, and body frames gets tangled, every downstream result is wrong. MIT VNAV treats rotation with the rigorous language of the Lie groups — rotation matrices are singularity-free but constrained, quaternions are singularity-free and well suited to optimization; this is the common-sense starting point for engineering implementation.
Key Concepts
- Intrinsics / extrinsics : the camera's own optical parameters and its pose in the world.
- Calibration: solving for intrinsics from known patterns (Zhang's method).
- Epipolar constraint: , the geometric bond between two views.
- Triangulation: recovering 3D points from intersecting multi-view rays.
- Rotation representations: rotation matrices / quaternions / Lie groups, each with its own trade-off between singularities and constraints.
Core Equations
Pinhole projection:
University Lecture
| Course | Lecture | Link |
|---|---|---|
| Stanford CS231A | L02–L05 (Camera Models, Calibration, multi-view geometry) + PS1/PS2 | Course page |
| ETH/UZH VAMR | L02–L03 (perspective projection, DLT, PnP) and Ex01–Ex02 | Course page |
Papers
- Must Read: Hartley & Zisserman, Multiple View Geometry in Computer Vision (2004, textbook; search the title) — Chapters 6 and 9 are the full version of this module.
- Recommended: Zhang (2000), A Flexible New Technique for Camera Calibration (IEEE TPAMI; search the title); Nistér (2004), An Efficient Solution to the Five-Point Relative Pose Problem (search the title).
Hands-on
Lab 0's custom environment is a warm-up for understanding the "observation interface." To build geometric intuition, work through VAMR's public exercises (hand-written implementations of perspective projection and DLT).
Check Your Understanding
- With unknown intrinsics, is the two-view constraint expressed with or , and why?
- Why can't absolute scale be recovered from a single image?
- What problem does RANSAC solve in the eight-point pipeline?
Show answer
- Use the fundamental matrix . The essential matrix depends on the intrinsics; when intrinsics are unknown, only can be estimated (at the pixel-coordinate level). Once intrinsics are known, coordinates can be normalized to estimate and decompose it into .
- Pinhole projection folds depth into the scale factor : scaling a 3D point by λ and its distance by λ leaves the projected pixel unchanged. A single view carries only relative geometry; absolute scale requires extra information (a known object size, a stereo baseline, an IMU).
- Point correspondences contain many outliers, and the eight-point algorithm is noise-sensitive. RANSAC repeatedly samples minimal subsets to propose hypotheses, counts inliers, and keeps the largest consensus set — rescuing the estimate from dirty data.
Takeaway
- The camera model is the "world → image" observation equation; calibration is system identification.
- Epipolar geometry provides annotation-free multi-view constraints — the geometric foundation of SfM/SLAM.
- Coordinate-frame discipline and the choice of rotation representation are the first lessons of engineering practice.
Next Module
Module 03: Depth and Point Clouds — from 2D pixels to the first explicit 3D representations.