
Basics to Machine Learning · v1.1.27
2026-09-22 13:30:36
The previous unit assembled a complete supervised pipeline.
All of it applies here unchanged. What changes is the model family.
The task
The output
What the output must be decides the architecture.
An image is not a bag of numbers. Three properties of the data are ignored by a fully connected layer:
A tree is a tree wherever it appears in the frame.
A function \(\mathbf{f}\) is invariant to \(\mathbf{t}\) when \(\mathbf{f}\big[\mathbf{t}[\mathbf{x}]\big] = \mathbf{f}[\mathbf{x}]\)

A photograph of a mountain.

The same scene, horizontally shifted.
The output does not move when the input does.
A function \(\mathbf{f}\) is equivariant to \(\mathbf{t}\) when \(\mathbf{f}\big[\mathbf{t}[\mathbf{x}]\big] = \mathbf{t}\big[\mathbf{f}[\mathbf{x}]\big]\)

One label per pixel.

The overlay for a translated input.
The output moves exactly as the input does.
Classification wants invariance
Segmentation wants equivariance
Neither implies the other: an equivariant output moves, an invariant one does not.
A \(224 \times 224\) RGB image is \(150{,}528\) input dimensions.
\[\approx 2.2 \times 10^{10} \text{ parameters, in a single layer}\]
Prohibitive before training has even been considered.
Structure is discarded
Symmetry must be learned
Possible in principle, and wasteful in every respect.
Each output is a weighted sum of nearby inputs, with the same weights at every position:
\[z_i = \omega_1 x_{i-1} + \omega_2 x_i + \omega_3 x_{i+1}\]

\(z_2 = \omega_1 x_1 + \omega_2 x_2 + \omega_3 x_3\)

\(z_3 = \omega_1 x_2 + \omega_2 x_3 + \omega_3 x_4\)
Translating the input translates the output identically — equivariance, by construction.

Zero-padding — out-of-range inputs taken as zero, output size kept.

Valid — overhanging positions discarded, output shrinks.
The first output has no previous input, the last no subsequent one.
Stride
Kernel size

Size five, dilation one.

Size three, dilation two.
Two readings: sample every second input, or zero-pad the kernel.
A kernel of size \(K\) at dilation rate \(r\) spans \(K' = r(K-1)+1\) positions.
Output size
Input length \(D\), padding \(p\) each side, stride \(s\):
\[D_{out} = \left\lfloor \frac{D + 2p - K'}{s} \right\rfloor + 1\]
Valid convolution: \(p = 0\). Size-preserving: \(p = (K'-1)/2\).
Model size
One kernel and one bias:
\[K + 1\]
The \(K\) kernel weights and the bias, whatever \(D\), \(p\), \(s\) or \(r\) may be.
Model size is independent of \(D\), and of the dilation rate.
A convolution, a bias, and an activation function:
\[h_i = \mathrm{a}\Big[\beta + \sum_{j=1}^{3} \omega_j\, x_{i+j-2}\Big]\]
The fully connected layer, for comparison:
\[h_i = \mathrm{a}\Big[\beta_i + \sum_{j=1}^{D} \omega_{ij}\, x_j\Big]\]
Four trainable parameters against \(D^2 + D\).

Six inputs, six hidden units: 36 independent weights.

The same weighted sum of three neighbours at every position.
Banded weights, tied across positions.
One convolution loses information: neighbours are averaged, and the ReLU clips what falls below zero.
Width is recovered in the channel dimension, not the spatial one.
Many in, one out
One kernel per input channel; the responses are summed into a single output channel.
Many out, one in
Repeat with different weights, once per output channel.
\(C_i\) in, \(C_o\) out
\[\boldsymbol{\Omega} \in \mathbb{R}^{C_i \times C_o \times K}, \quad \boldsymbol{\beta} \in \mathbb{R}^{C_o}\]
Each output unit is a weighted sum over all \(C_i\) channels at \(K\) positions.
No count here depends on the spatial size of the input.

One convolution combines both input channels at the three closest positions.

Output channels are parallel readings of the same input.
At each position \(i\), stack the \(K\) neighbouring columns of \(\mathbf{X} \in \mathbb{R}^{C_i \times D}\) into one patch vector:
\[\mathbf{x}_i = \operatorname{vec}\big[\mathbf{X}_{:,\, i-\lfloor K/2\rfloor \,:\, i+\lfloor K/2\rfloor}\big] \in \mathbb{R}^{C_i K}\]
One matrix, one bias, reused at every position:
\[\mathbf{f}_i = \mathbf{a}\big[\boldsymbol{\beta} + \boldsymbol{\Omega}\,\mathbf{x}_i\big], \qquad \boldsymbol{\Omega} \in \mathbb{R}^{C_o \times C_i K}, \quad \boldsymbol{\beta} \in \mathbb{R}^{C_o}\]
\[\mathbf{F} = \big[\,\mathbf{f}_1\;\mathbf{f}_2\;\cdots\;\mathbf{f}_D\,\big] \in \mathbb{R}^{C_o \times D}\]
A convolution is one matrix–vector product per position, with the matrix held fixed.
The receptive field of a unit \(h^{(k)}_i\) is the set of input positions it depends on:
\[\mathcal{R}\big[h^{(k)}_i\big] = \left\{\, j \;:\; \frac{\partial h^{(k)}_i}{\partial x_j} \neq 0 \,\right\}\]
A local operator becomes global by being stacked.

First-layer units cover three input positions.

Second-layer units combine three of those, covering five.
Depth buys context; the parameter count per layer does not change.
MNIST-1D: a 40D input, ten classes, three hidden layers of fifteen channels.

Convolutional: 2,050 parameters, about 17% test error.

Fully connected: 59,065 parameters, about 40% test error.
What remains is the inductive bias: every position is processed alike, and the data were built by translating a template.
The architecture forbids most of what the other model can say.
The kernel gains a second index and is translated along both axes:
\[h_{ij} = \mathrm{a}\Big[\beta + \sum_{m=1}^{3}\sum_{n=1}^{3} \omega_{mn}\, x_{i+m-2,\, j+n-2}\Big]\]

\(h_{23}\) is a weighted sum of the nine positions \(x_{12}\) to \(x_{34}\).

Other outputs come from translating the same kernel.

The hidden layer is a tensor: height, width, channels.
The output of a convolution is a feature map; pooling aggregates non-overlapping regions of it.

Sub-sampling: keep every other position.

Max pooling: the maximum of each \(2\times2\) block.

Mean pooling: the average of each block.
Applied per channel: half the height and width, the same channel count.

Duplication: each value copied four times.

Max unpooling: values returned to where the maxima came from.

Bilinear interpolation: fill between known samples.
Max unpooling requires the positions recorded by an earlier max pooling layer.

Stride two down: one output per three inputs.

Transposed convolution: one input, three of twice as many outputs.
Its weight matrix is the transpose of the downsampling one.

Channel depth and spatial resolution are adjusted independently.
A typical convolutional network alternates convolution with downsampling by a factor of two.
Geometry is progressively exchanged for channel depth.
ImageNet — 1,281,167 training images, 1000 categories, a \(224 \times 224\) RGB input.
The task that made the case for convolutional architectures.

16.4% top-five error, against 25% the year before.

6.8% top-five error. The decisive change from AlexNet was depth.

A \(7\times7\) grid, one class per cell.

A VGG-derived encoder down to \(14\times14\), a fully connected bottleneck reformed to \(7\times7\times512\), and a mirrored decoder to \(224\times224\times21\).
Encoder
Decoder
Also termed an hourglass network, for its shape.
Classification improved as networks got deeper — but only up to a point.
Every network so far passes each layer’s output to the next, and nothing else:
\[\mathbf{h}_1 = \mathbf{f}_1[\mathbf{x}], \quad \mathbf{h}_2 = \mathbf{f}_2[\mathbf{h}_1], \quad \mathbf{h}_3 = \mathbf{f}_3[\mathbf{h}_2], \quad \mathbf{y} = \mathbf{f}_4[\mathbf{h}_3]\]
Equivalently, one chain of nested functions:
\[\mathbf{y} = \mathbf{f}_4\big[\mathbf{f}_3\big[\mathbf{f}_2[\mathbf{f}_1[\mathbf{x}]]\big]\big]\]
Depth helped from AlexNet to VGG. It does not keep helping.

CIFAR-10 test error: 20 layers beat 56.

Training error: the same ordering.
Worse on the data it was fitted to — a failure of optimization, not of generalization.

One hidden layer: \(\partial y/\partial x\) changes slowly with the input.

Twenty-four layers: a tiny change gives a completely different gradient.
The gradient is correct. It is correct only infinitesimally.

Autocorrelation of the gradient against depth.
Shattered gradients — the named failure.
The derivative of the output with respect to the first layer is a product along the chain:
\[\frac{\partial \mathbf{y}}{\partial \mathbf{f}_1} = \frac{\partial \mathbf{f}_2}{\partial \mathbf{f}_1}\frac{\partial \mathbf{f}_3}{\partial \mathbf{f}_2}\frac{\partial \mathbf{f}_4}{\partial \mathbf{f}_3}\]
A residual, or skip, connection adds each layer’s input back to its output:
\[\mathbf{h}_k = \mathbf{h}_{k-1} + \mathbf{f}_k\big[\mathbf{h}_{k-1}, \boldsymbol{\phi}_k\big]\]
The branch learns a correction to the identity, not a replacement for it.

56 layers, with skip connections.

The same network, skips removed.
One addition per block, and the surface becomes descendable.

\[\mathbf{y} = \mathbf{x} + \mathbf{f}_1[\mathbf{x}] + \mathbf{f}_2\big[\mathbf{x}+\mathbf{f}_1[\mathbf{x}]\big] + \cdots\]
Eight of the sixteen paths contain \(\mathbf{f}_1\), one of length one:
\[\frac{\partial \mathbf{y}}{\partial \mathbf{f}_1} = \mathbf{I} + \frac{\partial \mathbf{f}_2}{\partial \mathbf{f}_1} + \Big(\frac{\partial \mathbf{f}_3}{\partial \mathbf{f}_1} + \frac{\partial \mathbf{f}_2}{\partial \mathbf{f}_1}\frac{\partial \mathbf{f}_3}{\partial \mathbf{f}_2}\Big) + \cdots\]
Gradients no longer have to survive the whole chain.

Linear then ReLU — the block’s output is non-negative, so it can only add.

ReLU then linear — both signs available.
The block must stay nonlinear, or the whole network collapses to a linear map.

Beyond that, the variance of the forward pass has to be managed.

1, 2, 4, 8 — exponential in the number of blocks.

Rescale by \(1/\sqrt{2}\) after each block: variance held at one.

BatchNorm first in each block: growth becomes linear, 2, 3, 4.
The second is what is used in practice.
Per activation \(h\), over the batch \(\mathcal{B}\):
\[m_h = \frac{1}{|\mathcal{B}|}\sum_{i \in \mathcal{B}} h_i, \qquad s_h = \sqrt{\frac{1}{|\mathcal{B}|}\sum_{i \in \mathcal{B}} (h_i - m_h)^2}\]
Standardize, then apply a learned scale and offset:
\[h_i \leftarrow \frac{h_i - m_h}{s_h + \epsilon}, \qquad h_i \leftarrow \gamma h_i + \delta\]
Mean \(\delta\) and standard deviation \(\gamma\), both learned like any other parameter.
At test time there is no batch to take statistics from.
A rare case where the model at test time is not the model that was fitted.
The cost
The benefit

Standard block
BatchNorm, ReLU, convolution — twice, then added back.
Bottleneck block
\(1\times1\) cuts the channels, \(3\times3\) does the spatial work, \(1\times1\) restores them.
The same \(3\times3\) receptive field, at a fraction of the parameters.
| Network | Layers | Top-5 | Top-1 |
|---|---|---|---|
| AlexNet | 8 | 16.4% | 38.1% |
| VGG | 19 | 6.8% | 23.7% |
| ResNet-200 | 200 | 4.8% | 20.1% |
Human performance on this task is estimated at 5.1% top-5.

Concatenation cannot cross a downsampling, so each resolution starts a new chain.

The long skip is the same idea, at the scale of a whole encoder stage.
The loss surface account
Explains why the network trains at all: a smooth basin is descendable by a finite step, and its minimum is robust.
The ensemble account
Explains why width can substitute for depth, and why gradients never traverse the longest paths.
Neither account says the depth itself is responsible — a plain 56-layer network trains worse.
The inductive bias is the whole of the advantage.
The Computer Vision module takes these architectures as its subject rather than its conclusion.
The same kernel, at the scale of a working system. 🔍