Lecture 06: Automatic Differentiation with PyTorch and Project Discussion
From Bernoulli likelihoods to loss functions, PyTorch autograd, model construction, training and evaluation loops, debugging with hooks, and final-project planning.
1. Today’s Roadmap
- Automatic differentiation in PyTorch
- A closer look at the PyTorch API
- Final-project discussion
The main technical idea is that PyTorch records the tensor operations used in a forward pass and uses that computation graph to evaluate derivatives during the backward pass. This lets us express the model and loss in ordinary tensor code instead of deriving and implementing every parameter gradient by hand.
2. PyTorch at a Glance
PyTorch is a Python-first framework for tensor computation and deep learning. Its core features include:
- Automatic differentiation through
torch.autograd - Dynamic computation graphs constructed from the operations actually run
- GPU acceleration through CUDA on supported hardware
- NumPy-like tensor operations and straightforward NumPy interoperability
- Neural-network layers, losses, optimizers, data loaders, and other training utilities
PyTorch is implemented with a high-performance C++ core while exposing a Python interface for day-to-day model development.
2.1 Installation and learning resources
Use the command generated by the official PyTorch installation selector because the correct build depends on the operating system and compute platform. After installation, remember that the package is imported with:
Useful resources mentioned in lecture include the PyTorch tutorials, the PyTorch discussion forum, and the video/tutorial The Fundamentals of Autograd.
3. From a Distribution to a Loss Function
The lecture revisited binary classification to show how a statistical model becomes an objective that autograd can differentiate.
For an input vector
, and bias
Here
. The Bernoulli likelihood for one example is
Taking the logarithm turns the product into a sum:
Training code typically minimizes a loss, so we negate the log-likelihood:
This is binary cross-entropy. In the lecture notebook, the direct formula,
torch.distributions.Bernoulli(probs=y_hat).log_prob(y), and nn.BCELoss()
produce the same per-example quantity after accounting for the minus sign and
the chosen reduction.
3.1 A numerically safer implementation
When the model returns a raw logit nn.BCEWithLogitsLoss() over applying a sigmoid and then calling nn.BCELoss().
The combined function computes the same objective but uses a numerically stable
implementation.
For mutually exclusive classes, use F.cross_entropy(logits, targets). It expects
raw, unnormalized logits, not probabilities and not a separately computed
softmax.
4. How Autograd Works
Autograd is PyTorch’s automatic differentiation engine. During the forward pass,
operations involving tensors that require gradients are recorded in a directed
acyclic computation graph. Calling .backward() traverses that graph in reverse,
applies the chain rule, and accumulates derivatives in the .grad fields of leaf
tensors such as model parameters.
Consider the scalar computation
The chain rule gives
PyTorch performs this bookkeeping automatically:
4.1 Dynamic graphs
PyTorch rebuilds the computation graph on each forward pass. Consequently,
ordinary Python control flow can change the operations executed from one pass to
the next. After a normal backward pass, the graph is released; requesting another
backward pass through the same graph requires retain_graph=True, which should be
used only when the algorithm genuinely needs it.
4.2 Gradients accumulate
Calling .backward() adds newly computed gradients to existing .grad
values. It does not overwrite them. This behavior is useful when deliberately
accumulating gradients over several minibatches, but a standard training loop
must clear the old gradients before each backward pass:
In current PyTorch versions, optimizer.zero_grad() may set gradients to None
rather than filling them with numerical zeros. The training-loop meaning is the
same: stale gradients will not be added to the next minibatch’s gradients.
5. The PyTorch Workflow
The lecture organized ordinary PyTorch usage into three steps: definition, creation, and training.
5.1 Step 1: define the model
A model usually subclasses torch.nn.Module. Layers assigned as attributes in
__init__ are registered as submodules, and their trainable tensors become model
parameters. The forward method specifies how those layers are used.
There is no need to manually implement the backward pass for this network.
Autograd knows the derivatives of the tensor operations used in forward and
constructs the appropriate backward graph at runtime.
5.2 Step 2: create the model and optimizer
Instantiating the class creates the model parameters. Move the model to the selected device, then give those parameters to an optimizer.
The model and every tensor used together must be on the same device. If the model is on a GPU, each feature and target minibatch must be moved there as well.
5.3 Step 3: train the model
Each minibatch follows the same cycle:
- Forward: compute logits from the inputs.
- Loss: compare logits with targets.
- Clear gradients: remove gradients left by the previous iteration.
- Backward: compute new parameter gradients with autograd.
- Step: let the optimizer update the parameters.
Calling model(features) is preferable to calling model.forward(features)
directly. The module call machinery invokes forward while also preserving the
behavior of registered hooks and other nn.Module features.
5.4 Evaluate without building a backward graph
model.eval() and torch.no_grad() solve different problems:
| Tool | What it changes |
|---|---|
model.eval() |
Switches modules such as Dropout and BatchNorm to evaluation behavior |
torch.no_grad() |
Prevents operations from being recorded for backward, reducing autograd overhead |
Evaluation code commonly needs both. Calling model.eval() alone does not
disable gradient tracking.
6. Inspecting Intermediate Activations with Hooks
Printing a model shows its module structure, but it does not expose the values produced inside a particular layer during a forward pass. A forward hook can capture those intermediate activations without rewriting the model.
The hook receives the module, its inputs, and its output whenever that module is called. Detaching the saved tensor prevents the debugging list from retaining an unneeded autograd history. Keep the returned handle and remove it after use; otherwise repeated registrations can make the hook run multiple times and can retain extra memory.
7. Common PyTorch Pitfalls
| Symptom | Likely cause | Fix |
|---|---|---|
| Gradients grow unexpectedly across iterations | .backward() accumulates into .grad |
Call optimizer.zero_grad() before loss.backward() unless accumulation is intentional |
| Device mismatch error | Model and minibatch are on different devices | Move the model, features, and targets to the same device |
| Cross-entropy behaves incorrectly | Probabilities or log-probabilities were passed where logits are expected | Pass raw logits to F.cross_entropy |
| Evaluation remains stochastic | Model is still in training mode | Call model.eval() before validation or testing |
| Evaluation uses unnecessary memory | Grad tracking is still enabled | Use with torch.no_grad(): or an appropriate inference-mode context |
| Hooks fire repeatedly | Old hook registrations remain active | Store the handle and call handle.remove() |
| No gradient appears on an intermediate tensor | By default, .grad is retained for leaf tensors |
Inspect leaf parameters, call .retain_grad() when appropriate, or register a hook |
8. Final-Project Discussion
The project is an opportunity to apply or investigate deep generative modeling in a focused, reproducible way. Teams should consist of 2–4 students, with teams of 3–4 strongly encouraged. Once a team is formed, email the instructors.
8.1 Possible project directions
- Apply a deep generative model to a new domain or dataset, including research data.
- Reproduce and extend a result from a paper discussed in class or found in a venue such as ICML, NeurIPS, ICLR, or AISTATS.
- Build and benchmark a tool, such as an evaluation metric, interpretability method, or efficient-training technique.
- Compare architectures or training methods on a shared task.
Examples from prior years included breast-cancer detection from ultrasound with an LSTM, DCGAN-generated surrealist art, chest X-ray diagnosis, sentiment analysis with BERT, and audio-to-video animation.
8.2 Dataset sources
| Modality | Possible sources |
|---|---|
| Vision | MNIST, CIFAR-10/100 for demonstrations, ImageNet, or a new image dataset |
| Text | Hugging Face Datasets, OpenWebText, or a new corpus |
| Tabular | UCI Machine Learning Repository or Kaggle |
| Research data | A team member’s research-group data, when permitted |
Before committing to a dataset, check its license and any human-subjects, privacy, or IRB constraints. The lecture also mentioned Summand as an option for exploratory data analysis.
8.3 Deliverables and Fall 2026 timeline
| Deliverable | Weight | Due | Expected content |
|---|---|---|---|
| Proposal | 5% | Friday, October 9 | Problem statement, literature review of 4+ papers, dataset, and activity plan |
| Midway report | 5% | Friday, November 6 | Working implementation and preliminary results, not only a plan |
| In-class presentation | 5% | Tuesday, December 1 or Thursday, December 3 | Team presentation; graded individually |
| Final report | 15% | Tuesday, December 8 | Eight-page ICML-format paper with a baseline, ablation study, and error analysis |
8.4 A useful proposal checklist
A strong proposal should make the following items concrete:
- Question: What exact problem or hypothesis will the project address?
- Prior work: Which four or more papers define the baseline and opportunity?
- Data: Is the dataset accessible, appropriately licensed, and large enough?
- Method: What model, comparison, or tool will the team implement?
- Evaluation: Which metrics, baselines, ablations, and error analyses will determine whether the idea worked?
- Execution: Who owns each task, and what must be complete by the midway report?
9. Key Takeaways
- A probabilistic model yields a likelihood; negating its log gives a loss that can be minimized.
- Autograd records forward computations and applies the chain rule in reverse when
.backward()is called. - Parameter gradients accumulate, so a normal training iteration clears them before the backward pass.
- A complete PyTorch iteration is: forward, loss, zero gradients, backward, step.
model.eval()changes layer behavior, whiletorch.no_grad()disables gradient recording; evaluation commonly uses both.- Forward hooks provide a controlled way to inspect intermediate activations.
- The course project should define a clear question, use a feasible dataset, make measurable progress by the midpoint, and include baselines, ablations, and error analysis in the final report.