Basics to Machine Learning · v1.1.43
2026-09-08 16:27:15
Three prerequisites are used, and no others.
The three families treated below — the straight line, the shallow network, the deep network — differ in the first item and in nothing else.
A supervised learning model defines a mapping from one or more inputs to one or more outputs.
\[\mathbf{y} = \mathbf{f}[\mathbf{x}].\]
The worked case is the valuation of a second-hand car: the input is the car’s age and mileage, the output is its estimated value in dollars.
Computing \(\mathbf{y}\) from \(\mathbf{x}\) is termed inference.
The model is an equation of fixed form containing parameters \(\boldsymbol{\phi}\), so the honest form carries them explicitly:
\[\mathbf{y} = \mathbf{f}[\mathbf{x}, \boldsymbol{\phi}].\]
The equation names a family of relations; the parameters name one member of that family.
Learning or training means finding parameters that produce sensible predictions, using a training dataset of \(I\) pairs \(\{\mathbf{x}_i, \mathbf{y}_i\}\).
The loss \(L\) summarises how poorly the model with parameters \(\boldsymbol{\phi}\) predicts the training outputs, and training minimizes it:
\[\hat{\boldsymbol{\phi}} = \operatorname*{argmin}_{\boldsymbol{\phi}}\Big[L[\boldsymbol{\phi}]\Big].\]
Inference is a single evaluation of the function with the parameters held fixed.
Training searches the space of parameters for the value minimizing the loss.
That asymmetry holds for every model treated in this unit.
Discriminative
\[\mathbf{y} = \mathbf{f}[\mathbf{x}, \boldsymbol{\phi}]\]
Predicts the output directly from the measurement.
Generative
\[\mathbf{x} = \mathbf{g}[\mathbf{y}, \boldsymbol{\phi}]\]
Computes the measurement from the output; inference needs \(\mathbf{y} = \mathbf{g}^{-1}[\mathbf{x}, \boldsymbol{\phi}]\).
A 1D linear regression model relates a scalar input to a scalar output as a straight line:
\[y = \mathrm{f}[x, \boldsymbol{\phi}] = \phi_0 + \phi_1 x.\]

Three members of the family:
One equation, one family of lines; the intercept and the slope pick out the member.

The data is \(I\) pairs \(\{x_i, y_i\}\), and nothing more.
The mismatch at example \(i\) is the deviation between the height of the line at \(x_i\) and the observed output \(y_i\). The total, termed the training error, is
\[L[\boldsymbol{\phi}] \;=\; \sum_{i=1}^{I}\big(\mathrm{f}[x_i, \boldsymbol{\phi}] - y_i\big)^2 \;=\; \sum_{i=1}^{I}\big(\phi_0 + \phi_1 x_i - y_i\big)^2.\]
Because the best parameters minimize this expression, it is a least-squares loss.


\(L = 7.07\) against \(L = 0.20\).


Two parameters, so the whole surface can be drawn.

This model
\(\partial L/\partial\phi_0\) and \(\partial L/\partial\phi_1\) can be set to zero and solved for the optimum in closed form.
Every model after it
No closed-form solution, and far too many parameters to evaluate the loss for every combination.
Descent is introduced for the second column, not the first.
Training loss
What the minimization returned. Measures performance on the training data alone.
Test loss
The same loss on a separate set of test data. Stands in for performance in the real world.
How well accuracy generalizes depends partly on the training data, partly on the model.
Underfitting
The model family is too restrictive to capture the true relationship.
Overfitting
A very expressive model describes atypical peculiarities of the training sample.
This is why model capacity — for a network, the number of hidden units — is a design decision rather than a quantity to be maximised.
The drawback: a straight line, and nothing else.
A shallow neural network has one hidden layer:
\[\begin{aligned} y = \mathrm{f}[x, \boldsymbol{\phi}] = \;& \phi_0 + \phi_1 a[\theta_{10} + \theta_{11}x] \\ &+ \phi_2 a[\theta_{20} + \theta_{21}x] + \phi_3 a[\theta_{30} + \theta_{31}x]. \end{aligned}\]
\[a[z] = \mathrm{ReLU}[z] = \begin{cases} 0 & z < 0 \\ z & z \ge 0. \end{cases}\]
It returns the input when the input is positive, and zero otherwise — it clips negative values to zero.
Hidden units
\[\begin{aligned} h_1 &= a[\theta_{10} + \theta_{11}x] \\ h_2 &= a[\theta_{20} + \theta_{21}x] \\ h_3 &= a[\theta_{30} + \theta_{31}x] \end{aligned}\]
\(y = \phi_0 + \phi_1 h_1 + \phi_2 h_2 + \phi_3 h_3\)
Joints
Unit \(d\) switches where its linear function crosses zero:
\[x = -\frac{\theta_{d0}}{\theta_{d1}}\]
One joint per unit, so up to four linear regions.



Each unit computes its own \(\theta_{d0} + \theta_{d1}x\) — its own intercept and slope.



Negatives become zero. Each unit gains a joint where its line crossed zero.



The output weights \(\phi_1, \phi_2, \phi_3\) scale and may flip each contribution.

Each linear region corresponds to a different activation pattern. A clipped unit is inactive; an unclipped unit is active.
The slope of a region comes from the slopes \(\theta_{\bullet 1}\) of the units active there, together with the output weights \(\phi_{\bullet}\) applied to them — in the shaded region, \(\theta_{11}\phi_1 + \theta_{31}\phi_3\).
Three units give four regions, but only three of the four slopes are independent.

Intercepts, ReLUs and parameter names are usually omitted — supplied by convention, not by the picture.
\[h_d = a[\theta_{d0} + \theta_{d1}x], \qquad y = \phi_0 + \sum_{d=1}^{D}\phi_d h_d.\]
A 1D function (dashed) approximated by a piecewise linear model with 5, 10 and 20 linear regions.
The theorem holds for the general case, mapping \(\mathbf{x} = [x_1, \ldots, x_{D_i}]^T\) to \(\mathbf{y} = [y_1, \ldots, y_{D_o}]^T\).
Widening the weight matrices leaves the mechanism untouched.


Outputs fan out on the right; inputs fan in on the left. The hidden layer is unchanged.
\[y_1 = \phi_{10} + \sum_{d=1}^{4}\phi_{1d}h_d, \qquad y_2 = \phi_{20} + \sum_{d=1}^{4}\phi_{2d}h_d.\]

One slope parameter per input:
\[\begin{aligned} h_1 &= a[\theta_{10} + \theta_{11}x_1 + \theta_{12}x_2] \\ h_2 &= a[\theta_{20} + \theta_{21}x_1 + \theta_{22}x_2] \\ h_3 &= a[\theta_{30} + \theta_{31}x_1 + \theta_{32}x_2] \end{aligned}\]



Brightness is the unit’s output. Each is a plane tilted its own way.



The cyan line is the joint — a hyperplane. One side is flat zero.



Each contribution is scaled by \(\phi_1, \phi_2, \phi_3\).

Beyond two inputs
Counting them
Shallow networks have more units than inputs, so typically more than \(2^{D_i}\) regions.
\[h_d = a\Big[\theta_{d0} + \sum_{i=1}^{D_i}\theta_{di}x_i\Big], \qquad y_j = \phi_{j0} + \sum_{d=1}^{D}\phi_{jd}h_d.\]
Each of the \(D\) hidden units carries \(D_i\) slopes and one offset; each of the \(D_o\) outputs carries \(D\) slopes and one offset:
\[(D_i + 1)D + (D + 1)D_o.\]
Three inputs, three hidden units, two outputs: fifteen slopes and five offsets, twenty parameters in all.

Why nonlinear
What ReLU costs
Nonlinearity is what depth is built on — and the ReLU pays for it at the flat end.
Two shallow networks of three hidden units each. The first takes \(x\) and returns \(y\):
\[h_d = a[\theta_{d0} + \theta_{d1}x], \qquad y = \phi_0 + \phi_1 h_1 + \phi_2 h_2 + \phi_3 h_3.\]
The second takes \(y\) and returns \(y'\), with its own parameters:
\[h'_d = a[\theta'_{d0} + \theta'_{d1}y], \qquad y' = \phi'_0 + \phi'_1 h'_1 + \phi'_2 h'_2 + \phi'_3 h'_3.\]
This composed model is also piecewise linear — but with potentially more regions than a shallow network of six hidden units.



Three ranges of \(x\) give the same \(y\), so network 2 is applied three times — nine regions, flipped and rescaled.
Shallow: 6 units, one layer
One joint per unit, so \(6 + 1 = 7\) regions.
Deep: 3 units, then 3 units
The second layer’s function repeats on each fold: 9 regions.
The same six units. The arrangement, not the count, decides how many regions they buy.
\[h_d = a[\theta_{d0} + \theta_{d1}x],\]
\[h'_d = a[\psi_{d0} + \psi_{d1}h_1 + \psi_{d2}h_2 + \psi_{d3}h_3],\]
\[y' = \phi'_0 + \phi'_1 h'_1 + \phi'_2 h'_2 + \phi'_3 h'_3.\]
The construction has four steps: linear functions of the input through a ReLU; new linear functions of those units; a second ReLU, clipping them and adding new joints; a linear combination of the second layer.



Already piecewise linear, and all three share their joints.



Clipping breaks the shared pattern — each unit gains joints of its own.



Scaled by \(\phi'_1, \phi'_2, \phi'_3\).

Layer by layer
\[\begin{aligned} \mathbf{h}_1 &= \mathbf{a}[\boldsymbol{\beta}_0 + \boldsymbol{\Omega}_0\mathbf{x}] \\ \mathbf{h}_{k} &= \mathbf{a}[\boldsymbol{\beta}_{k-1} + \boldsymbol{\Omega}_{k-1}\mathbf{h}_{k-1}] \\ \mathbf{y} &= \boldsymbol{\beta}_K + \boldsymbol{\Omega}_K\mathbf{h}_K \end{aligned}\]
Reading it
Every deep network is this: affine maps alternating with an element-wise nonlinearity.

Layer \(k\) has \(D_k\) units, so:
Nothing is chosen here — the widths fix every shape.
Parameters
\(\boldsymbol{\phi} = \{\boldsymbol{\beta}_k, \boldsymbol{\Omega}_k\}\) — the weights and biases, learned by minimizing a loss.
Hyperparameters
The number of layers \(K\) and the widths \(D_1, \ldots, D_K\) — chosen before the parameters are learned.
Fixed hyperparameters give a family of functions; taking them into account as well, a network is a family of families of functions.
Both families approximate anything. The question is at what cost.
Approximation power alone does not decide between the families.
The first is theoretical; the last two are empirical.
Shallow, \(D > 2\) units
up to \(D + 1\) regions on \(3D + 1\) parameters
Deep, \(K\) layers of \(D > 2\) units
up to \((D+1)^K\) regions on \(3D + 1 + (K-1)D(D+1)\) parameters
It is the difference between adding and multiplying. \(K = 5\) layers of \(D = 10\) units: \(471\) parameters, \(161{,}051\) regions — against roughly \(150\) for a shallow network on the same budget.


At a fixed parameter budget, deeper networks produce more regions.
Process local regions in parallel, then integrate over larger ones — which needs multiple layers.
Both are reported as observations, not as results. Neither is well understood.
The least-squares loss was introduced by stipulation, and descend the loss surface was left without an algorithm behind it.
Training Models supplies both — losses derived from maximum likelihood, and the optimizers that minimize them: stochastic gradient descent, momentum, Adam, together with backpropagation.