Preliminaries to Machine Learning

Keywords

ver. 1.0.0, preliminaries_to_machine_learning

Introduces supervised learning, linear and ReLU networks, depth vs width, and how to express and compare network families

This unit introduces supervised learning as fitting a parameterized function to input–output examples, works a complete 1D linear regression example, and explains train vs test error and the bias–variance trade-off. It builds shallow ReLU networks and shows their piecewise linear structure, states the universal approximation theorem (an existence result), extends networks to multivariate inputs/outputs and counts parameters, composes layers into deep networks (explaining how folding multiplies linear regions), writes K-layer networks in matrix form and distinguishes parameters from hyperparameters, and compares shallow and deep architectures for expressivity per parameter and practical trainability. The unit concludes by preparing for the next unit on how to fit these models in practice.

This unit teaches how to view learning as fitting a parameterized function to data and gives the concrete language and examples needed to describe, build, and compare simple model families.

It begins by framing supervised learning: a family of functions f[x, φ] parameterized by φ and a dataset of input–output pairs; “learning” is choosing φ to make outputs match targets. To make this concrete, it develops 1D linear regression end-to-end — specifying the affine model, writing the least-squares loss, and visualizing the loss surface so training appears as a search over parameters.

The unit separates fitting training examples from performing well on new data by defining training error, test error, and generalization. This leads to diagnosing underfitting versus overfitting in terms of model capacity and motivates why capacity control matters.

Moving beyond linear models, it constructs a shallow network with ReLU activations and shows how each hidden unit produces a single joint in the output, so the overall mapping is a continuous piecewise linear function. The universal approximation theorem is stated carefully: sufficiently wide shallow networks can approximate any continuous function arbitrarily well — an existence statement that does not provide a constructive recipe or address efficiency or training.

The shallow ReLU architecture is extended to multivariate inputs and outputs: weight matrices widen the model, the linear regions in input space become convex polytopes, and counting parameters becomes a matter of matrix dimensions. Composing one network into another yields deep networks; composition folds input space and multiplies linear regions, giving an intuitive account of why depth can increase expressivity more than simply adding width.

The general K-layer network is written compactly in matrix notation as alternating affine maps and elementwise nonlinearities; this representation separates learned parameters (weight matrices and biases) from hyperparameters (depth, widths, activation choices). Finally, the unit compares shallow and deep networks: both are universal approximators, but they differ in regions-per-parameter scaling, depth efficiency, and real-world trainability. Practical considerations — computational cost, optimization difficulty, and alignment with structured data — guide architecture choice.

The unit closes by collecting the model families and notation the learner can now build and points forward to the next unit, which will introduce principled loss functions and algorithms for actually training these models.

Materials

Source documents

  • Understanding Deep Learning, Simon Prince, 2026 — Link — Page 31-38
  • Understanding Deep Learning, Simon Prince, 2026 — Link — Page 39-54
  • Understanding Deep Learning, Simon Prince, 2026 — Link — Page 55-69