Skip to main content

Computer Vision

This page is a reference manual. Most world models take images and video as input, and spatial intelligence is built on recovering geometry from pixels. This page covers only what the main tracks actually use.

Minimal knowledge set​

  • Image formation and the pinhole camera: a 3D point is moved into the camera frame by the extrinsics [R∣t][R \mid t], then projected to pixels by the intrinsics KK: λ x~=K[R∣t] X~\lambda\, \tilde{x} = K [R \mid t]\, \tilde{X}. Depth is lost in the projection, which is the root of every difficulty in monocular vision.
  • Multi-view geometry: corresponding points in two views satisfy the epipolar constraint, and with two known camera poses you can triangulate a 3D point. SLAM, SfM, and NeRF training data all rely on this.
  • Features and matching: corners, SIFT / ORB descriptors, and RANSAC to reject outliers. This is exactly the front end of ORB-SLAM.
  • Convolutional features and visual encoders: translation equivariance and hierarchical features in CNNs. Pretrained encoders (ResNet, DINO, ViT) compress an image into a vector you can predict on.
  • Dense prediction: depth estimation, optical flow, and segmentation produce one output per pixel. They are common intermediate representations for 4D scenes and navigation.

When to consult​

Main-track moduleComputer vision used
Track B 02 Camera and GeometryPinhole model, intrinsics and extrinsics, epipolar geometry
Track B 03 Depth and Point CloudsTriangulation, depth estimation, back-projection to point clouds
Track B 07 SLAM and VIOFeature matching, RANSAC, bundle adjustment
Track A 04 Representation LearningVisual encoders, pretrained features

Best external resources​

  • Stanford CS231A (course page · course notes): camera models, epipolar geometry, and SfM, the closest match to Track B. Start with Lecture 2, Camera Models.
  • Stanford CS231n course notes (cs231n.github.io): the classic introduction to convolutional networks and visual representations.
  • Richard Szeliski, Computer Vision: Algorithms and Applications, 2nd ed. (szeliski.org/Book): free PDF. Chapter 2 (image formation), Chapter 7 (feature detection and matching), and Chapter 11 (SfM and SLAM) are the most relevant.
  • Dive into Deep Learning, Chapter 7 (d2l.ai): convolutional networks, explained with code.

Self-check​

You are ready when you can write the pinhole projection and say what the intrinsics and extrinsics each control, explain why two views can triangulate a point but one cannot, and say what problem RANSAC solves.

Next​

Back to the main tracks: Track B Module 02 or Track A Module 04