Training Models

Keywords

ver. 1.0.0, training_models

How to train deep networks end-to-end: principled losses, efficient optimization, stable gradients, and practical regularization

This unit teaches how to turn a neural network into a complete, trainable learning system. It shows how to derive principled loss functions by viewing networks as predictors of output distributions, build common losses (least-squares, binary and multiclass cross-entropy) with a four-step recipe, and minimize those losses reliably using gradient-based methods. You will learn minibatch stochastic gradient descent, momentum and Adam, how to pick learning rates and batch sizes, how to compute gradients by backpropagation, how to initialize weights to avoid vanishing/exploding signals, how to decompose and diagnose test error (noise, bias, variance) including double descent, and how to apply both explicit and implicit regularization (early stopping, dropout, ensembling, augmentation, transfer learning). The unit yields a practical end-to-end training pipeline ready for image models like convolutional nets.

This unit gives a cohesive, practical foundation for training deep neural networks from first principles to deployment-ready practice.

Losses from probability: Networks are recast as predictors of parameters of a probability distribution over outputs; the negative log-likelihood then yields a justified loss rather than an ad hoc choice. Applying a simple four-step recipe (choose a distribution, decide which parameters the model predicts, write the likelihood, take negative log) recovers familiar losses: Gaussian outputs give least-squares for regression; Bernoulli with sigmoid gives binary cross-entropy; categorical with softmax gives multiclass cross-entropy. Viewing losses probabilistically also makes clear how adding penalty terms corresponds to priors over parameters.

Optimization and geometry: Minimizing the loss uses gradient-based descent. Plain gradient descent and the geometry of nonconvex loss surfaces explain why training is an engineering task: plateaus, sharp valleys, and saddle points cause failures. Replacing full-batch gradients with stochastic gradients on random minibatches makes each step cheap and injects noise that helps escape problematic regions.

Practical optimizers: Accumulating gradient history reduces oscillation and adapts step sizes. Momentum damps oscillation in narrow valleys; adaptive methods such as Adam also correct for wildly differing gradient magnitudes across parameters and layers. Choosing the optimizer, learning rate, momentum hyperparameters, batch size and schedules are critical design choices; their interactions determine convergence speed and final performance.

Efficient gradients and stable signals: Backpropagation implements the chain rule in reverse, storing intermediates so gradient computation scales linearly with depth. Proper weight initialization (e.g., He initialization) sets weight variance so activations and gradients neither vanish nor explode as they propagate forward and backward, making deep networks trainable.

Generalization, diagnostics, and double descent: Test error is usefully decomposed into noise, bias, and variance; diagnosing which term dominates suggests whether to increase data, change model capacity, or regularize. Modern models exhibit double descent: beyond the interpolation threshold, test error can decrease again as capacity grows, so classical bias–variance intuition must be extended. Hyperparameters must be selected via a validation split to avoid contaminating the test set.

Regularization, explicit and implicit: Regularization can be added explicitly as penalty terms (interpretable as priors), but optimization itself induces implicit regularization (finite-step SGD biases solutions). Practical regularizers—early stopping, dropout, ensembling, data augmentation, transfer learning, and injecting noise—are surveyed with explanations of the mechanisms by which they improve generalization.

End-to-end training pipeline: Combining principled losses, gradient-based optimizers (with momentum/Adam), minibatching, backpropagation, proper initialization, diagnostic tools for error decomposition and validation, and a toolbox of regularizers yields a complete training pipeline that can be applied to real tasks. With that pipeline in place, model families that exploit structure (for example, convolutional networks for images) become the next focus for building higher-performing systems.

Materials

Source documents

  • Understanding Deep Learning, Simon Prince, 2026 — Link — Page 70-90
  • Understanding Deep Learning, Simon Prince, 2026 — Link — Page 91-109
  • Understanding Deep Learning, Simon Prince, 2026 — Link — Page 110-131
  • Understanding Deep Learning, Simon Prince, 2026 — Link — Page 132-151
  • Understanding Deep Learning, Simon Prince, 2026 — Link — Page 152-174