Lecture 07: Multi-layer Perceptrons and Backpropagation

From binary to multi-class classification with softmax, stacking fully-connected layers into a multi-layer perceptron, the nonlinear activations that finally solve XOR, and backpropagation in matrix form.

Section 1: From Binary to Multi-Class Classification

Moving from Two Classes to K Classes

Binary classification (2 categories). When deciding between just two things — “spam” or “not spam” — we use the sigmoid, which squashes the network’s output into a probability between 0 and 1. This is the setup from Lecture 6:

y \mid x \sim \mathrm{Bernoulli}(\hat{y}), \qquad \hat{y} = \sigma(z).

Multi-class classification (K categories). With multiple categories, the model outputs a raw, unnormalized score — a logit — for every category, and the label follows a categorical distribution:

y \mid x \sim \mathrm{Categorical}(\hat{y}_1, \dots, \hat{y}_K), \qquad \text{one logit } z_k \text{ per class}.

The Softmax Function

We need a way to turn those raw scores into probabilities. Softmax is a fixed mathematical rule — it has no trainable parameters, so it does not “learn” — that makes every score positive with an exponential and divides by the total, so the outputs lie between 0 and 1 and sum to exactly 1:

\hat{y}_k = \mathrm{softmax}(z)_k = \frac{\exp(z_k)}{\sum_j \exp(z_j)}.

The sigmoid is exactly the K = 2 special case of softmax, so nothing from Lecture 6 is being discarded — it is being generalized.

From Likelihood to Cross-Entropy Loss

Likelihood. This measures the probability the model assigned to the correct answer. When the label is a one-hot vector such as [0, 1, 0], the likelihood is just the predicted score in the correct slot:

P(y \mid x) = \prod_k \hat{y}_k^{\,y_k} = \hat{y}_{\text{true class}}.

Taking logs turns that product into a sum:

\log P(y \mid x) = \sum_k y_k \log \hat{y}_k.

Cross-entropy loss. To train the model we want a loss that goes down as the model gets better, so we take the negative log-likelihood:

L = -\sum_k y_k \log \hat{y}_k.

This is the categorical cross-entropy, and it is the same recipe as Lecture 6 with more classes. In PyTorch, nn.CrossEntropyLoss takes the raw logits directly, just as BCEWithLogitsLoss did — the softmax is applied internally for numerical stability, so you should not apply it yourself first.

The Gradient: Same Pattern, More Classes

To fix its errors, the network needs to know how to adjust its raw scores. Softmax combined with cross-entropy has a simple closed-form gradient, elementwise in k with y one-hot:

\frac{\partial L}{\partial z_k} = \hat{y}_k - y_k, \qquad \text{that is} \qquad \frac{\partial L}{\partial z} = \mathrm{softmax}(z) - \mathrm{onehot}(y).

This is the same \hat{y} - y pattern as the binary case, reached by the same derivation strategy: the chain rule through softmax’s own derivative. Section 4 shows that this expression is exactly the starting point \delta^{(L)} of the backward pass.

Section 2: Multi-layer Perceptron Architecture

Stacking Fully-Connected Layers

An MLP is a computation graph built from multiple fully-connected layers, and as the lecture puts it, there is “nothing new, really” in the individual pieces: each layer is the same affine map followed by an activation that we have seen since Lecture 4. What is new is that the output of one layer becomes the input of the next.

If a network only uses straight-line math, it cannot solve complex, overlapping patterns such as the famous XOR problem. We must introduce bends and curves into the math using nonlinear activation functions such as ReLU — which is what Section 3 takes up in detail.

Counting the Trainable Parameters

A question worth answering carefully, because it recurs every time you design a network: how many trainable parameters does an MLP actually have? For a fully-connected layer mapping n_{\text{in}} units to n_{\text{out}} units, the count is n_{\text{in}} \times n_{\text{out}} weights plus n_{\text{out}} biases.

Counting parameters in the lecture's network. Two inputs feed a hidden layer of five units, then a hidden layer of four, then three outputs. Each layer contributes its weight matrix plus one bias per receiving unit.

Applying that rule layer by layer to the network above:

Layer Weights Biases Total
Layer 1 (2 \to 5) 2 \times 5 = 10 5 15
Layer 2 (5 \to 4) 5 \times 4 = 20 4 24
Layer 3 (4 \to 3) 4 \times 3 = 12 3 15
Total 42 12 54

So the network has 54 trainable parameters. Note also that because the three outputs are mutually exclusive classes, this is a multi-class problem and the output layer uses the softmax of Section 1.

Updating the Early Layers

Counting parameters raises the obvious follow-up: how does a weight in the first layer get updated, when it only affects the loss indirectly? The network has to account for every path along which that weight influences the final prediction, multiplying the local derivatives backward through the network with the chain rule and summing over paths. Section 4 makes that statement precise and shows that “sum over paths” is exactly a matrix multiplication.

Section 3: Nonlinear Activation Functions

Why Stacking Layers Is Not Enough

Section 2 stacked fully-connected layers into a computation graph, but stacking alone buys us nothing. Consider two linear layers with no activation between them:

\begin{aligned} z_2 &= W_2 (W_1 x + b_1) + b_2 \\[4pt] &= (W_2 W_1) x + (W_2 b_1 + b_2) \\[4pt] &= \tilde{W} x + \tilde{b}. \end{aligned}

The composition collapses back into a single linear map. A hundred such layers is still one linear layer with a more expensive parameterization, so the decision boundary is still a single hyperplane — exactly the limitation that the perceptron ran into.

XOR makes the failure concrete. The four clusters alternate labels along both axes, so no single line separates them. Training a one-hidden-layer MLP with the identity activation produces exactly one straight boundary and misclassifies half the data. Swapping the identity for ReLU on the same architecture produces a piecewise-linear boundary that bends around the clusters and separates them. The nonlinearity, not the depth, is what buys the expressiveness.

A Selection of Common Activations

The classical choices are the identity, the logistic sigmoid, tanh, and hard tanh:

\begin{aligned} \sigma(z) &= \frac{1}{1 + e^{-z}}, \\[6pt] \tanh(z) &= \frac{\exp(z) - \exp(-z)}{\exp(z) + \exp(-z)}, \\[6pt] \mathrm{HardTanh}(z) &= \begin{cases} 1 & \text{if } z > 1\\ -1 & \text{if } z < -1\\ z & \text{otherwise.} \end{cases} \end{aligned}

The modern rectifier family replaces saturation on the positive side with a straight line:

\begin{aligned} \mathrm{ReLU}(z) &= \max(0, z), \\[4pt] \mathrm{LeakyReLU}(z) &= \max(0, z) + \alpha \min(0, z), \\[4pt] \mathrm{ELU}(z) &= \max(0, z) + \min\big(0, \alpha (e^{z} - 1)\big). \end{aligned}

Leaky ReLU fixes \alpha (the slides use \alpha = 0.025); PReLU is the same formula with \alpha promoted to a trainable parameter, learned by gradient descent alongside the weights.

Activations and their derivatives. The top row plots five common activation functions; the bottom row plots their derivatives, which are the factors that backpropagation multiplies together. The sigmoid's derivative never exceeds 0.25 and decays to zero in both tails, while the rectifier family passes a derivative of exactly 1 on the positive side.

Why Tanh Is Usually Preferred Over Sigmoid

Tanh has three advantages over the logistic sigmoid:

  1. Mean centering. Tanh is centered at zero, while the sigmoid outputs live in (0, 1)

    and so carry a positive bias into the next layer.

  2. Positive and negative values. A unit can push the next layer in either direction.
  3. Larger gradients. \tanh'(0) = 1 against \sigma'(0) = 0.25, so the signal that survives a backward pass through a tanh layer is up to four times stronger.

Tanh also has a convenient derivative, expressible in terms of its own output:

\frac{d}{dz}\tanh(z) = 1 - \tanh(z)^2.

These advantages only pay off if the unit actually operates near zero, where the derivative is largest. That is why it matters to normalize inputs to mean zero and to use random weight initialization centered at zero: both keep pre-activations in the high-gradient region instead of the flat tails.

The bottom row of the figure above also previews the gradient-flow problem that Section 4 makes precise. Backpropagation multiplies one activation derivative per layer, and \sigma' \le 0.25 everywhere, so a deep stack of sigmoids shrinks the gradient exponentially with depth. ReLU’s derivative of exactly 1 on the positive side is the main reason rectifiers replaced saturating activations in deep networks.

The Price: The Loss Is No Longer Convex

Every model in the course so far — linear regression, Adaline, logistic regression, softmax regression — has a convex loss, so gradient descent converges to the global minimum regardless of where it starts. Once nonlinear activations are inserted between layers, that guarantee is gone: the deep loss is non-convex almost always.

The practical consequence is a loss of reproducibility that surprises people the first time they see it. Repeat the same training run, changing only the random seed for weight initialization or the shuffling of the dataset, and you land in a different local minimum with different final weights. Visualizations of deep loss surfaces show a landscape of ridges, valleys and basins rather than a single bowl.

This is less alarming than it sounds. As LeCun argued in “Who’s Afraid of Non-Convex Loss Functions?”, convexity is overrated: choosing a suitable architecture matters more than insisting on a convex objective, particularly when convexity restricts you to architectures that cannot express the problem in the first place. Even for shallow convex models such as SVMs, swapping in a non-convex loss can improve both accuracy and speed. In practice, different local minima of a deep network often achieve comparable loss, so we trade a uniqueness guarantee we cannot have for expressiveness we need.

Why We Initialize Randomly: Breaking Symmetry

A natural question at this point: what happens if we initialize an MLP to all-zero weights?

The answer is that the units stay symmetric forever. If every unit in a hidden layer starts with identical weights, every unit computes the same pre-activation, receives the same gradient in the backward pass, and therefore applies the same update. A layer of n identical units is only as expressive as a layer with one unit, and no amount of training breaks the tie, because the update rule itself is symmetric in the units.

This is exactly why we initialize randomly: random initialization breaks the symmetry so that different units can specialize on different features. Note the contrast with the perceptron of Lecture 4, where all-zero initialization was perfectly fine — a single unit has no siblings to be symmetric with. Combined with the requirement above that weights stay centered at zero, the constraints on initialization are becoming specific enough to deserve their own treatment, which is where principled schemes such as Xavier and He initialization come in later in the course.

Section 4: Backpropagation

Having introduced multilayer perceptrons and their activation functions, we now consider how to learn their parameters. Training requires the gradient of the loss with respect to every weight and bias. Backpropagation computes these gradients by applying the chain rule backward through the network — the algorithm Rumelhart, Hinton and Williams introduced for learning internal representations. This section covers the matrix formulation, the XOR example, and gradient flow.

Notation and the Forward Pass

Consider a network with L parameterized layers. For a single example, let a^{(0)} = x and write

z^{(l)} = W^{(l)} a^{(l-1)} + b^{(l)}, \qquad a^{(l)} = \phi_l\left(z^{(l)}\right).

We use column vectors throughout. If layer l has n_l units, then W^{(l)} \in \mathbb{R}^{n_l \times n_{l-1}}, while z^{(l)}, a^{(l)} and b^{(l)} have length n_l.

During the forward pass we compute the prediction \hat{y} = a^{(L)} and the scalar loss L(\hat{y}, y). We retain the intermediate activations and pre-activations for the later backward pass.

Propagating the Loss Gradient Backward

Let \delta^{(l)} denote the gradient of the loss with respect to the pre-activation z^{(l)}:

\delta^{(l)} = \frac{\partial L}{\partial z^{(l)}}.

Using the forward equation, each downstream pre-activation is

z_j^{(l+1)} = \sum_i W_{ji}^{(l+1)} a_i^{(l)} + b_j^{(l+1)}.

Applying the chain rule and summing over every downstream unit the activation reaches,

\frac{\partial L}{\partial a_i^{(l)}} = \sum_j \frac{\partial L}{\partial z_j^{(l+1)}} \frac{\partial z_j^{(l+1)}}{\partial a_i^{(l)}} = \sum_j W_{ji}^{(l+1)} \delta_j^{(l+1)}.

This sum over downstream units is precisely a matrix-vector product. Applying the local activation derivative yields

\delta^{(l)} = \left[\left(W^{(l+1)}\right)^{\top} \delta^{(l+1)}\right] \odot \phi_l'\left(z^{(l)}\right).

Here \odot denotes elementwise multiplication. The dimensions must agree: an n_l \times n_{l+1} matrix multiplies a vector of length n_{l+1} to produce a vector of length n_l.

Gradients of Weights and Biases

Within a layer, \partial z_i^{(l)} / \partial W_{ij}^{(l)} = a_j^{(l-1)}, so the chain rule gives

\frac{\partial L}{\partial W_{ij}^{(l)}} = \delta_i^{(l)} a_j^{(l-1)}.

Collecting these entries gives an outer product:

\begin{aligned} \frac{\partial L}{\partial W^{(l)}} &= \delta^{(l)} \left(a^{(l-1)}\right)^{\top}, \\[4pt] \frac{\partial L}{\partial b^{(l)}} &= \delta^{(l)}. \end{aligned}

Each weight gradient combines the signal arriving along that connection with the loss sensitivity of the receiving unit. The resulting matrix has the same shape as W^{(l)}.

Starting at the Output Layer

Assume a binary classification task with a sigmoid output and binary cross-entropy:

\hat{y} = \sigma\left(z^{(L)}\right), \qquad L = -y \log \hat{y} - (1 - y) \log (1 - \hat{y}).

Since \sigma'(z) = \sigma(z)\left(1 - \sigma(z)\right), the chain rule simplifies the output delta to

\delta^{(L)} = \left(-\frac{y}{\hat{y}} + \frac{1-y}{1-\hat{y}}\right) \hat{y}\left(1 - \hat{y}\right) = \hat{y} - y.

The same delta holds not only for sigmoid with BCE, but also for softmax with categorical cross-entropy, which is the closed form Section 1 arrived at from the categorical likelihood.

A Worked Example: Forward and Backward Passes on XOR

Consider the network

x \in \mathbb{R}^2 \;\to\; \left[\mathrm{Linear}(2,2),\ \mathrm{ReLU}\right] \;\to\; h \in \mathbb{R}^2 \;\to\; \left[\mathrm{Linear}(2,1),\ \mathrm{sigmoid}\right] \;\to\; \hat{y} \in (0,1),

evaluated with binary cross-entropy. Its parameters and the selected training example are

\begin{aligned} W^{(1)} &= \begin{bmatrix} 1 & 1 \\ 1 & 1 \end{bmatrix}, & b^{(1)} &= \begin{bmatrix} 0 \\ -1 \end{bmatrix}, \\[4pt] W^{(2)} &= \begin{bmatrix} 1 & -2 \end{bmatrix}, & b^{(2)} &= -0.5, \\[4pt] x &= \begin{bmatrix} 1 \\ 0 \end{bmatrix}, & y &= 1. \end{aligned}

The target is one because exactly one input bit is one. We carry extra decimal places to keep the intermediate values consistent.

Forward pass. The hidden pre-activations and activations are

z^{(1)} = W^{(1)} x + b^{(1)} = \begin{bmatrix} 1 \\ 0 \end{bmatrix}, \qquad h = \mathrm{ReLU}\left(z^{(1)}\right) = \begin{bmatrix} 1 \\ 0 \end{bmatrix}.

Consequently,

z^{(2)} = W^{(2)} h + b^{(2)} = 1(1) - 2(0) - 0.5 = 0.5, \hat{y} = \sigma(0.5) \approx 0.622459, \qquad L = -\log \hat{y} \approx 0.474077.

Backward pass. At the output,

\delta^{(2)} = \hat{y} - y \approx -0.377541,

so the output-layer gradients are

\begin{aligned} \frac{\partial L}{\partial W^{(2)}} &= \delta^{(2)} h^{\top} \approx \begin{bmatrix} -0.377541 & 0 \end{bmatrix}, \\[4pt] \frac{\partial L}{\partial b^{(2)}} &\approx -0.377541. \end{aligned}

Next, propagate through the output weights:

\frac{\partial L}{\partial h} = \frac{\partial L}{\partial z^{(2)}} \frac{\partial z^{(2)}}{\partial h} = \left(W^{(2)}\right)^{\top} \delta^{(2)} \approx \begin{bmatrix} -0.377541 \\ 0.755081 \end{bmatrix}.

For ReLU the derivative is one at positive inputs and zero at negative inputs, so the activation mask is (1, 0)^{\top}, giving

\delta^{(1)} = \frac{\partial L}{\partial z^{(1)}} = \frac{\partial L}{\partial h} \odot \begin{bmatrix} 1 \\ 0 \end{bmatrix} \approx \begin{bmatrix} -0.377541 \\ 0 \end{bmatrix}.

Finally,

\begin{aligned} \frac{\partial L}{\partial W^{(1)}} &= \delta^{(1)} x^{\top} \approx \begin{bmatrix} -0.377541 & 0 \\ 0 & 0 \end{bmatrix}, \\[4pt] \frac{\partial L}{\partial b^{(1)}} &\approx \begin{bmatrix} -0.377541 \\ 0 \end{bmatrix}. \end{aligned}

There are two distinct reasons for zero weight gradients here. The second input is zero, so all gradients in the second column of W^{(1)} vanish for this example. Separately, the second hidden unit’s ReLU mask is zero, so its incoming weights and bias also receive zero gradient, and its outgoing weight receives zero gradient because its activation is zero.

Why Depth Makes Gradients Fragile

Because backpropagation relies on the chain rule, the gradient reaching the first layer depends on the product of the local derivatives of every layer that follows:

\frac{\partial L}{\partial z^{(1)}} = \frac{\partial L}{\partial z^{(L)}} \frac{\partial z^{(L)}}{\partial z^{(L-1)}} \cdots \frac{\partial z^{(2)}}{\partial z^{(1)}}.

Each factor describes how one layer’s pre-activation changes with the preceding layer’s, including the effect of both the weights and the activation function. To reach an early layer, the gradient must pass through all of them.

This is where the problem arises. Repeated multiplication by small factors makes gradients in early layers extremely small — the vanishing-gradient problem — and repeated multiplication by large factors makes them very large — the exploding-gradient problem. For intuition, a scalar factor of 0.5 repeated twenty times gives about 9.5 \times 10^{-7}, whereas 2 repeated twenty times gives about 1.05 \times 10^{6}.

The activation derivatives plotted in Section 3 explain part of this sensitivity:

Activation Derivative Implication for gradient flow
Sigmoid \sigma(z)(1-\sigma(z)) \le 1/4 Can strongly attenuate gradients, especially in saturation.
Tanh 1 - \tanh^2(z) \le 1 Has derivative one at zero, but still saturates at large magnitudes.
ReLU One for z > 0, zero for z < 0 Preserves the local gradient on its active branch and blocks it on its inactive branch.

Activation derivatives are only part of the product, though: the weight matrices affect both its magnitude and its direction. ReLU therefore does not guarantee stable gradients, and tanh does not eliminate vanishing gradients. This motivates the later study of initialization and normalization.

Section 5: Practical Considerations for Training MLPs

Reading Training and Validation Curves

Checking the training and test loss or error is a simple and useful diagnosis. The curves carry a lot of information, especially about overfitting and underfitting. A validation plateau, or an error that starts increasing, is a sign of overfitting. Minibatch losses are noisy, so averaging or smoothing can reveal the underlying trend. The curves also say something about the learning rate: if it is too high, the error tends to oscillate.

One caveat worth knowing about is grokking: on some small algorithmic tasks, strong generalization emerges long after the training data have been fit. This is an example of delayed validation improvement, not a guarantee that every stalled model will improve if trained longer.

Parameters versus Hyperparameters

Parameters are learned from training data through optimization; in the MLP considered here they are the entries of the weight matrices and bias vectors. Hyperparameters specify the model or the training procedure and are chosen by the practitioner.

Learned parameters Model and training choices
Weights and biases Hidden-layer count and width; activation and loss functions
  Learning rate, schedule, optimizer, minibatch size, and epoch budget
  Initialization, input normalization, random seed, and regularization settings

Underfitting, Overfitting, and Model Capacity

A low training error is not always a positive signal. Limited capacity prevents a good fit to the training data; increasing capacity lowers training error, but excessive capacity may increase generalization error, which is usually estimated with test-set error. The result is a gap between fitting observed examples and predicting unseen ones: as training error falls, the model improves on data it has seen, but that improvement is not guaranteed to transfer.

A Practical Training Routine

Before a long run, check input values, label encoding, tensor shapes, and whether the output is compatible with the loss. Record the configuration and the seed, log loss and accuracy, and measure epoch times. If the loss behaves unexpectedly, inspect the setup before spending more computation. Change one hyperparameter at a time when diagnosing a problem, and repeat promising comparisons across runs.

The lecture summarizes each minibatch update as four steps:

  1. Clear previously accumulated parameter gradients (zero_grad()).
  2. Run the forward pass and compute the loss.
  3. Compute gradients with backpropagation (backward()).
  4. Apply the optimizer update (step()).

Keeping these roles separate makes debugging easier: the forward pass determines the prediction, the backward pass computes sensitivities, and the optimizer changes the parameters.

Why Deep Instead of Merely Wide?

A classical universal approximation result states that a network with one hidden layer and a suitable activation can approximate any continuous function on a compact domain arbitrarily well, given sufficient width. Cybenko established such a result for continuous sigmoidal activations, and related approximation bounds were developed by Barron and by Csáji. But these theorems only say that such a network exists; they do not say that the required width is practical.

Depth offers a different route: expressing a function as a composition of simpler transformations. Connected layers reuse intermediate features, so the same expressiveness can be reached with fewer parameters, and earlier features become the inputs from which later ones are built. That composition also acts as a kind of regularization, since later layers are constrained by the behavior of earlier ones. This is one reason deep learning is concerned with representation learning rather than simply adding layers.

Architecture also introduces an inductive bias: assumptions that favor certain representations or solutions. A layered composition constrains how later features are built from earlier ones. Adding layers does not automatically reduce parameter count, computation time, or overfitting — those outcomes depend on layer widths, the target task, and the training procedure. Greater depth also lengthens the chain through which gradients must pass, creating exactly the optimization difficulties discussed above.

From Fully Connected Layers to CNNs

The next lecture introduces convolutional neural networks. The useful connection is that, after arranging inputs and outputs as vectors, a convolutional layer is the same equation as a fully-connected one, z = W x + b.

The difference lies in the structure of W. Local receptive fields make it sparse: an output depends only on a local input region. Weight sharing ties many entries together: the same filter coefficients are reused at different spatial positions, and a per-channel bias is shared across positions.

The chain rule is unchanged. The gradient with respect to the input of this linear operation is still \left(W\right)^{\top} \delta, so the backward pass carries over directly. One thing does change: because a shared weight is used at every position, it receives a gradient contribution from each of them, and those contributions are summed — the weight-sharing rule from Lecture 5.

Backpropagation therefore connects these architectures through one principle: propagate loss sensitivities backward through local operations, and add all contributions when a variable or parameter influences the loss through multiple paths.

Key Takeaways