Skip to main content

PyTorch

This page is a reference manual. Every lab (except Lab 0) uses PyTorch. The goal is not framework mastery, but never getting stuck on engineering details.

Minimal knowledge setโ€‹

  • Tensors and autograd: the standard loop of tensor.requires_grad, loss.backward(), optimizer.step(); understand when the computation graph is built and when it is freed.
  • nn.Module: modular implementations of encoder/decoder/dynamics; forward and parameter registration.
  • DataLoader: Dataset/DataLoader, shuffling and batching; this is what feeds the rollout data in Lab 2.
  • GPU memory basics: batch size, mixed precision (torch.autocast), and checkpoint recomputation; the free Colab T4 has only 15GB.
  • Device management: consistency of .to(device) โ€” 90% of lab errors come from mixing CPU and GPU tensors.

When to consultโ€‹

ScenarioWhat to look up
Training the encoder in Lab 2nn.Module, the reparameterization trick for VAEs
Rollout prediction in Labs 2/3Batch-dimension management for sequential data, pack_padded_sequence (optional)
Any lab hits OOMMixed precision, gradient checkpointing, smaller batches
Getting things running on ColabSelecting a GPU runtime, sanity check with torch.cuda.is_available()

Best external resourcesโ€‹

  • The official PyTorch 60-minute tutorial (pytorch.org/tutorials): one pass through and you can start building.
  • Official PyTorch Docs: the torch.distributions section is a must-read โ€” RSSM's prior/posterior are written directly with it.
  • Karpathy's Zero to Hero (YouTube series): hand-writing backpropagation and GPT from scratch โ€” engineering and principles together.

Self-checkโ€‹

You are ready when you can write the full "define model โ†’ forward โ†’ compute loss โ†’ backward โ†’ update" loop without a tutorial, and port a training script from CPU to GPU without blowing up memory.

Nextโ€‹

Back to the main tracks: Track A Module 05 or start Lab 0