Lecture 06: Automatic Differentiation with PyTorch and Project Discussion

From Bernoulli likelihoods to loss functions, PyTorch autograd, model construction, training and evaluation loops, debugging with hooks, and final-project planning.

1. Today’s Roadmap

The main technical idea is that PyTorch records the tensor operations used in a forward pass and uses that computation graph to evaluate derivatives during the backward pass. This lets us express the model and loss in ordinary tensor code instead of deriving and implementing every parameter gradient by hand.


2. PyTorch at a Glance

PyTorch is a Python-first framework for tensor computation and deep learning. Its core features include:

PyTorch is implemented with a high-performance C++ core while exposing a Python interface for day-to-day model development.

2.1 Installation and learning resources

Use the command generated by the official PyTorch installation selector because the correct build depends on the operating system and compute platform. After installation, remember that the package is imported with:

import torch print(torch.__version__) print(torch.cuda.is_available())

Useful resources mentioned in lecture include the PyTorch tutorials, the PyTorch discussion forum, and the video/tutorial The Fundamentals of Autograd.


3. From a Distribution to a Loss Function

The lecture revisited binary classification to show how a statistical model becomes an objective that autograd can differentiate.

For an input vector \mathbf{x}, weights

\mathbf{w}

, and bias b, define

z = \mathbf{w}^{\top}\mathbf{x} + b, \qquad \hat{y} = \sigma(z), \qquad y \mid \mathbf{x} \sim \operatorname{Bernoulli}(\hat{y}).

Here \hat{y} is interpreted as

P(y=1\mid\mathbf{x})

. The Bernoulli likelihood for one example is

P(y \mid \mathbf{x}) = \hat{y}^{,y}(1-\hat{y})^{1-y}.

Taking the logarithm turns the product into a sum:

\log P(y \mid \mathbf{x}) = y\log\hat{y} + (1-y)\log(1-\hat{y}).

Training code typically minimizes a loss, so we negate the log-likelihood:

\mathcal{L}_{\mathrm{BCE}} = -\left[y\log\hat{y} + (1-y)\log(1-\hat{y})\right].

This is binary cross-entropy. In the lecture notebook, the direct formula, torch.distributions.Bernoulli(probs=y_hat).log_prob(y), and nn.BCELoss() produce the same per-example quantity after accounting for the minus sign and the chosen reduction.

3.1 A numerically safer implementation

When the model returns a raw logit z, prefer nn.BCEWithLogitsLoss() over applying a sigmoid and then calling nn.BCELoss(). The combined function computes the same objective but uses a numerically stable implementation.

logits = model(features) loss = torch.nn.functional.binary_cross_entropy_with_logits( logits, targets.float(), )

For mutually exclusive classes, use F.cross_entropy(logits, targets). It expects raw, unnormalized logits, not probabilities and not a separately computed softmax.


4. How Autograd Works

Autograd is PyTorch’s automatic differentiation engine. During the forward pass, operations involving tensors that require gradients are recorded in a directed acyclic computation graph. Calling .backward() traverses that graph in reverse, applies the chain rule, and accumulates derivatives in the .grad fields of leaf tensors such as model parameters.

Consider the scalar computation

x \xrightarrow{\;z=wx+b\;} z \xrightarrow{\;a=\sigma(z)\;} a \xrightarrow{\;\mathcal{L}=\operatorname{BCE}(a,y)\;} \mathcal{L}.

The chain rule gives

\frac{\partial \mathcal{L}}{\partial w} = \frac{\partial \mathcal{L}}{\partial a} \frac{\partial a}{\partial z} \frac{\partial z}{\partial w}.

PyTorch performs this bookkeeping automatically:

import torch x = torch.tensor(2.0) y = torch.tensor(1.0) w = torch.tensor(0.5, requires_grad=True) b = torch.tensor(0.0, requires_grad=True) z = w * x + b loss = torch.nn.functional.binary_cross_entropy_with_logits(z, y) loss.backward() print(w.grad) # d(loss) / d(w) print(b.grad) # d(loss) / d(b)

4.1 Dynamic graphs

PyTorch rebuilds the computation graph on each forward pass. Consequently, ordinary Python control flow can change the operations executed from one pass to the next. After a normal backward pass, the graph is released; requesting another backward pass through the same graph requires retain_graph=True, which should be used only when the algorithm genuinely needs it.

4.2 Gradients accumulate

Calling .backward() adds newly computed gradients to existing .grad values. It does not overwrite them. This behavior is useful when deliberately accumulating gradients over several minibatches, but a standard training loop must clear the old gradients before each backward pass:

optimizer.zero_grad() loss.backward() optimizer.step()

In current PyTorch versions, optimizer.zero_grad() may set gradients to None rather than filling them with numerical zeros. The training-loop meaning is the same: stale gradients will not be added to the next minibatch’s gradients.


5. The PyTorch Workflow

The lecture organized ordinary PyTorch usage into three steps: definition, creation, and training.

5.1 Step 1: define the model

A model usually subclasses torch.nn.Module. Layers assigned as attributes in __init__ are registered as submodules, and their trainable tensors become model parameters. The forward method specifies how those layers are used.

import torch import torch.nn as nn import torch.nn.functional as F class MultilayerPerceptron(nn.Module): def __init__(self, num_features, num_classes, num_h1=128, num_h2=256): super().__init__() self.linear_1 = nn.Linear(num_features, num_h1) self.linear_2 = nn.Linear(num_h1, num_h2) self.linear_out = nn.Linear(num_h2, num_classes) def forward(self, x): x = F.relu(self.linear_1(x)) x = F.relu(self.linear_2(x)) logits = self.linear_out(x) return logits

There is no need to manually implement the backward pass for this network. Autograd knows the derivatives of the tensor operations used in forward and constructs the appropriate backward graph at runtime.

5.2 Step 2: create the model and optimizer

Instantiating the class creates the model parameters. Move the model to the selected device, then give those parameters to an optimizer.

random_seed = 1 learning_rate = 0.1 num_features = 28 * 28 num_classes = 10 torch.manual_seed(random_seed) device = torch.device("cuda" if torch.cuda.is_available() else "cpu") model = MultilayerPerceptron(num_features, num_classes).to(device) optimizer = torch.optim.SGD(model.parameters(), lr=learning_rate)

The model and every tensor used together must be on the same device. If the model is on a GPU, each feature and target minibatch must be moved there as well.

5.3 Step 3: train the model

for epoch in range(num_epochs): model.train() for features, targets in train_loader: features = features.view(-1, 28 * 28).to(device) targets = targets.to(device) # Forward pass and loss logits = model(features) loss = F.cross_entropy(logits, targets) # Backward pass and parameter update optimizer.zero_grad() loss.backward() optimizer.step()

Each minibatch follows the same cycle:

  1. Forward: compute logits from the inputs.
  2. Loss: compare logits with targets.
  3. Clear gradients: remove gradients left by the previous iteration.
  4. Backward: compute new parameter gradients with autograd.
  5. Step: let the optimizer update the parameters.

Calling model(features) is preferable to calling model.forward(features) directly. The module call machinery invokes forward while also preserving the behavior of registered hooks and other nn.Module features.

5.4 Evaluate without building a backward graph

model.eval() num_correct = 0 num_examples = 0 with torch.no_grad(): for features, targets in test_loader: features = features.view(-1, 28 * 28).to(device) targets = targets.to(device) logits = model(features) predictions = logits.argmax(dim=1) num_correct += (predictions == targets).sum().item() num_examples += targets.size(0) accuracy = num_correct / num_examples

model.eval() and torch.no_grad() solve different problems:

Tool What it changes
model.eval() Switches modules such as Dropout and BatchNorm to evaluation behavior
torch.no_grad() Prevents operations from being recorded for backward, reducing autograd overhead

Evaluation code commonly needs both. Calling model.eval() alone does not disable gradient tracking.


6. Inspecting Intermediate Activations with Hooks

Printing a model shows its module structure, but it does not expose the values produced inside a particular layer during a forward pass. A forward hook can capture those intermediate activations without rewriting the model.

outputs = [] def save_output(module, inputs, output): outputs.append(output.detach()) handle = model.linear_2.register_forward_hook(save_output) _ = model(features) second_hidden_layer_output = outputs[-1] handle.remove()

The hook receives the module, its inputs, and its output whenever that module is called. Detaching the saved tensor prevents the debugging list from retaining an unneeded autograd history. Keep the returned handle and remove it after use; otherwise repeated registrations can make the hook run multiple times and can retain extra memory.


7. Common PyTorch Pitfalls

Symptom Likely cause Fix
Gradients grow unexpectedly across iterations .backward() accumulates into .grad Call optimizer.zero_grad() before loss.backward() unless accumulation is intentional
Device mismatch error Model and minibatch are on different devices Move the model, features, and targets to the same device
Cross-entropy behaves incorrectly Probabilities or log-probabilities were passed where logits are expected Pass raw logits to F.cross_entropy
Evaluation remains stochastic Model is still in training mode Call model.eval() before validation or testing
Evaluation uses unnecessary memory Grad tracking is still enabled Use with torch.no_grad(): or an appropriate inference-mode context
Hooks fire repeatedly Old hook registrations remain active Store the handle and call handle.remove()
No gradient appears on an intermediate tensor By default, .grad is retained for leaf tensors Inspect leaf parameters, call .retain_grad() when appropriate, or register a hook

8. Final-Project Discussion

The project is an opportunity to apply or investigate deep generative modeling in a focused, reproducible way. Teams should consist of 2–4 students, with teams of 3–4 strongly encouraged. Once a team is formed, email the instructors.

8.1 Possible project directions

Examples from prior years included breast-cancer detection from ultrasound with an LSTM, DCGAN-generated surrealist art, chest X-ray diagnosis, sentiment analysis with BERT, and audio-to-video animation.

8.2 Dataset sources

Modality Possible sources
Vision MNIST, CIFAR-10/100 for demonstrations, ImageNet, or a new image dataset
Text Hugging Face Datasets, OpenWebText, or a new corpus
Tabular UCI Machine Learning Repository or Kaggle
Research data A team member’s research-group data, when permitted

Before committing to a dataset, check its license and any human-subjects, privacy, or IRB constraints. The lecture also mentioned Summand as an option for exploratory data analysis.

8.3 Deliverables and Fall 2026 timeline

Deliverable Weight Due Expected content
Proposal 5% Friday, October 9 Problem statement, literature review of 4+ papers, dataset, and activity plan
Midway report 5% Friday, November 6 Working implementation and preliminary results, not only a plan
In-class presentation 5% Tuesday, December 1 or Thursday, December 3 Team presentation; graded individually
Final report 15% Tuesday, December 8 Eight-page ICML-format paper with a baseline, ablation study, and error analysis

8.4 A useful proposal checklist

A strong proposal should make the following items concrete:

  1. Question: What exact problem or hypothesis will the project address?
  2. Prior work: Which four or more papers define the baseline and opportunity?
  3. Data: Is the dataset accessible, appropriately licensed, and large enough?
  4. Method: What model, comparison, or tool will the team implement?
  5. Evaluation: Which metrics, baselines, ablations, and error analyses will determine whether the idea worked?
  6. Execution: Who owns each task, and what must be complete by the midway report?

9. Key Takeaways