PyTorch
This page is a reference manual. Every lab (except Lab 0) uses PyTorch. The goal is not framework mastery, but never getting stuck on engineering details.
Minimal knowledge setโ
- Tensors and autograd: the standard loop of
tensor.requires_grad,loss.backward(),optimizer.step(); understand when the computation graph is built and when it is freed. nn.Module: modular implementations of encoder/decoder/dynamics;forwardand parameter registration.- DataLoader:
Dataset/DataLoader, shuffling and batching; this is what feeds the rollout data in Lab 2. - GPU memory basics: batch size, mixed precision (
torch.autocast), andcheckpointrecomputation; the free Colab T4 has only 15GB. - Device management: consistency of
.to(device)โ 90% of lab errors come from mixing CPU and GPU tensors.
When to consultโ
| Scenario | What to look up |
|---|---|
| Training the encoder in Lab 2 | nn.Module, the reparameterization trick for VAEs |
| Rollout prediction in Labs 2/3 | Batch-dimension management for sequential data, pack_padded_sequence (optional) |
| Any lab hits OOM | Mixed precision, gradient checkpointing, smaller batches |
| Getting things running on Colab | Selecting a GPU runtime, sanity check with torch.cuda.is_available() |
Best external resourcesโ
- The official PyTorch 60-minute tutorial (pytorch.org/tutorials): one pass through and you can start building.
- Official PyTorch Docs: the
torch.distributionssection is a must-read โ RSSM's prior/posterior are written directly with it. - Karpathy's Zero to Hero (YouTube series): hand-writing backpropagation and GPT from scratch โ engineering and principles together.
Self-checkโ
You are ready when you can write the full "define model โ forward โ compute loss โ backward โ update" loop without a tutorial, and port a training script from CPU to GPU without blowing up memory.
Nextโ
Back to the main tracks: Track A Module 05 or start Lab 0