Lecture 04: Single-layer networks

From inputs and weights to the perceptron, decision boundaries, the Perceptron Learning Algorithm, and multi-neuron networks.

1. Today’s Roadmap

A single-layer network combines inputs using weights and bias, applies an activation function, and produces an output. The perceptron creates a linear decision boundary and the Perceptron Learning Algorithm updates the model for misclassifications.


2. Terminology & Notation

A single-layer network takes an input and computes a net input, also called the pre-activation or weighted input:

\[z = \sum_{i=0}^{n} w_i x_i.\]

The net input is passed through an activation function to produce the activation:

\[a = a(z).\]

The activation function determines how the neuron transforms its net input. For example, a perceptron uses a threshold function, while linear regression uses the identity function. A non-linear activation function such as the sigmoid can be written as $\sigma(z)$.

The predicted label output is then given by

\[\hat{y} = f(a).\]

The bias is included as $w_0$, with the corresponding input $x_0 = 1$.

3. Perceptrons

3.1 Continuous Weighting

A perceptron combines many inputs using continuous weights before applying an activation function. The weighted inputs are summed to produce the net input, which is then passed through an activation function to determine the output of the perceptron.

\[net = \sum_{i=0}^{n} w_i x_i.\]

3.2 Activation Function

The activation function is applied to the net input of the perceptron. Different activation functions result in different outputs.

In Deep Learning, the outputs are continuous and can take multiple values that signify the strength of a signal, which is why the sigmoid function (see below) applied on the activation makes the perceptron DL.

The sigmoid function:

\[\sigma(net) = \frac{1}{1+e^{-net}}.\]

A perceptron with a sigmoid activation function turns it from linear regression to logistic regression. Instead of providing a class defined by a threshold, it gives the probability of something.

3.3 Separation Boundary

An optimal estimate for a good separation boundary (with parameters θ) can be found using gradient ascent (MLE).

4. Threshold Functions

Threshold functions are used for discrete classification models and lead to the classic perceptron. Instead of continuous, the threshold function determines class by whether the net input reaches a threshold.

4.1 Classic Perceptron Limitations

The classic perceptron is a binary linear classifier.

5. Perceptron Learning Algorithm

5.1 Pseudocode

The Perceptron Learning Algorithm (PLA) updates weights based on if the perceptron is correct in classifying each training example.

Let the dataset be

\[D = (\langle x^{[1]}, y^{[1]}\rangle, \langle x^{[2]}, y^{[2]}\rangle, \ldots, \langle x^{[n]}, y^{[n]}\rangle) \in (\mathbb{R}^m \times \{0,1\})^n.\]

Initialize the weight vector to zero:

\[\mathbf{w} \leftarrow \mathbf{0}^m.\]

For every training epoch, the algorithm considers each training example $(\mathbf{x}^{[i]},y^{[i]})$.

The predicted label is calculated using the function:

\[\hat{y}^{[i]} = \sigma(\mathbf{x}^{[i]T}\mathbf{w}).\]

The error is calculated as

\[\mathrm{err} = y^{[i]}-\hat{y}^{[i]}.\]

Finally, the weight vector is set to:

\[\mathbf{w} \leftarrow \mathbf{w}+\mathrm{err}\times\mathbf{x}^{[i]}.\]

5.2 Geometric Intuition

The weight vector is perpendicular to the decision boundary.

The intuition behind adding $y_i x_i$ to the weight is to rotate the weight vector toward or away from the misclassified point to get the point closer to the decision boundary.

The intuition behind adding $y_i$ to the bias term is to shift the boundary toward the misclassified point in order to send it to the other side of the line.

If this is done for all points, eventually we reach a boundary that satisfies all points if the data is linearly separable.

5.3 PLA, MLE & MAP

The perceptron learning algorithm involves neither MLE nor MAP. This is because:

6. Notational Conventions for Neural Networks

6.1 Single Connected Layer

The inputs are weighted differently based on the node that they are feeding into.

Example:

A fully connected layer converts the outputted activation as inputs to the next layer of neurons having their own associated weights and activations.

7. Summary