Lecture 04: Single-layer networks
From inputs and weights to the perceptron, decision boundaries, the Perceptron Learning Algorithm, and multi-neuron networks.
1. Today’s Roadmap
- Terminology & Notations
- Perceptrons
- Threshold Functions
- Perceptron Learning Algorithm
- Notational Conventions for Neural Networks
A single-layer network combines inputs using weights and bias, applies an activation function, and produces an output. The perceptron creates a linear decision boundary and the Perceptron Learning Algorithm updates the model for misclassifications.
2. Terminology & Notation
A single-layer network takes an input and computes a net input, also called the pre-activation or weighted input:
\[z = \sum_{i=0}^{n} w_i x_i.\]The net input is passed through an activation function to produce the activation:
\[a = a(z).\]The activation function determines how the neuron transforms its net input. For example, a perceptron uses a threshold function, while linear regression uses the identity function. A non-linear activation function such as the sigmoid can be written as $\sigma(z)$.
The predicted label output is then given by
\[\hat{y} = f(a).\]The bias is included as $w_0$, with the corresponding input $x_0 = 1$.
3. Perceptrons
3.1 Continuous Weighting
A perceptron combines many inputs using continuous weights before applying an activation function. The weighted inputs are summed to produce the net input, which is then passed through an activation function to determine the output of the perceptron.
\[net = \sum_{i=0}^{n} w_i x_i.\]3.2 Activation Function
The activation function is applied to the net input of the perceptron. Different activation functions result in different outputs.
In Deep Learning, the outputs are continuous and can take multiple values that signify the strength of a signal, which is why the sigmoid function (see below) applied on the activation makes the perceptron DL.
The sigmoid function:
\[\sigma(net) = \frac{1}{1+e^{-net}}.\]A perceptron with a sigmoid activation function turns it from linear regression to logistic regression. Instead of providing a class defined by a threshold, it gives the probability of something.
3.3 Separation Boundary
An optimal estimate for a good separation boundary (with parameters θ) can be found using gradient ascent (MLE).
- MLE: Maximize log likelihood of observing the data given some parameters
- Gradient ascent: Shift the parameters in the direction of increasing log likelihood.
4. Threshold Functions
Threshold functions are used for discrete classification models and lead to the classic perceptron. Instead of continuous, the threshold function determines class by whether the net input reaches a threshold.
4.1 Classic Perceptron Limitations
The classic perceptron is a binary linear classifier.
- It cannot have non-linear decision boundaries.
- It cannot solve multi-class problems.
- It contains multiple (often not optimal) solutions so there is no clear single solution.
- Does not converge if a linearly separable solution does not exist (but is guaranteed to converge otherwise)
5. Perceptron Learning Algorithm
5.1 Pseudocode
The Perceptron Learning Algorithm (PLA) updates weights based on if the perceptron is correct in classifying each training example.
Let the dataset be
\[D = (\langle x^{[1]}, y^{[1]}\rangle, \langle x^{[2]}, y^{[2]}\rangle, \ldots, \langle x^{[n]}, y^{[n]}\rangle) \in (\mathbb{R}^m \times \{0,1\})^n.\]Initialize the weight vector to zero:
\[\mathbf{w} \leftarrow \mathbf{0}^m.\]For every training epoch, the algorithm considers each training example $(\mathbf{x}^{[i]},y^{[i]})$.
The predicted label is calculated using the function:
\[\hat{y}^{[i]} = \sigma(\mathbf{x}^{[i]T}\mathbf{w}).\]The error is calculated as
\[\mathrm{err} = y^{[i]}-\hat{y}^{[i]}.\]Finally, the weight vector is set to:
\[\mathbf{w} \leftarrow \mathbf{w}+\mathrm{err}\times\mathbf{x}^{[i]}.\]5.2 Geometric Intuition
The weight vector is perpendicular to the decision boundary.
- Decision boundary line: $wx + b = 0$ for some $x$ on the decision boundary.
- If we take any two points on it and subtract them, we get a line on the decision boundary. So $(wx_1 + b) - (wx_2 + b) = w(x_1-x_2) = 0$
- The dot product between the weight and the line along the decision boundary can only be 0 if the weight is perpendicular to the decision boundary.
The intuition behind adding $y_i x_i$ to the weight is to rotate the weight vector toward or away from the misclassified point to get the point closer to the decision boundary.
The intuition behind adding $y_i$ to the bias term is to shift the boundary toward the misclassified point in order to send it to the other side of the line.
If this is done for all points, eventually we reach a boundary that satisfies all points if the data is linearly separable.
5.3 PLA, MLE & MAP
The perceptron learning algorithm involves neither MLE nor MAP. This is because:
- The perceptron learning algorithm is not a probabilistic model, rather it is deterministic.
- If a datapoint lies to one side of the decision boundary, it is definitively classified as either -1 or 1. There is no likelihood or priors involved here.
- It also does not try to maximize any of its parameters. Several suitable decision boundaries can be found by the PLA for some linearly separable data, but it does not try to find the one that’s the most distant from all neighboring points.
- You cannot optimize over the threshold function because the gradient is either undefined or 0
6. Notational Conventions for Neural Networks
6.1 Single Connected Layer
The inputs are weighted differently based on the node that they are feeding into.
- A scaled set of inputs may need to result in a certain value for the first component of the output and a different value for the second component of the output. Thus, even though the same inputs may feed into the same activations, the weight vector applied at each neuron may be different.
Example:
- Assume Input = $[3,4]$ and expected output = $[4,3]$
- Same weighting would cause Input $[3,4]$, Output $[4,4]$
A fully connected layer converts the outputted activation as inputs to the next layer of neurons having their own associated weights and activations.
7. Summary
- Terminology & Notation: Single-layer networks combine inputs with weights and bias, then apply an activation function.
- Perceptrons: Perceptrons use continuous weighting and activation functions to produce outputs and define a separation boundary.
- Threshold Functions: The classic perceptron uses a threshold function for binary classification and is limited to linear decision boundaries.
- Perceptron Learning Algorithm: The PLA updates the weights and bias based on misclassified points and converges when the data is linearly separable.
- Notational Conventions for Neural Networks: Multiple neurons can apply different weights to the same inputs, and fully connected layers pass activations to the next layer.