Basics to Machine Learning

Keywords

ver. 1.0.0

Foundations and practical techniques for building, training, and deploying deep neural networks for vision and general supervised tasks

This node covers supervised learning and neural network families (linear, ReLU, shallow vs deep), principled loss functions and gradient-based optimization (SGD, momentum, Adam, backprop, initialization), regularization and generalization diagnostics (bias–variance, double descent, augmentation, dropout, early stopping, ensembling, transfer learning), and convolutional architectures and modules (convolutions, receptive field, pooling/upsampling, 1×1 convs, residual connections, batch normalization) with guidance on parameter counting and architecture choices for vision models like ResNet, DenseNet, and U‑Net.

This unit provides an end-to-end treatment of modern supervised deep learning, combining theoretical foundations with practical recipes for model design and training.

Main themes - Supervised learning fundamentals: framing learning as fitting a parameterized function to input–output examples, training vs test error, and the bias–variance trade-off illustrated with a complete 1D linear regression example. - Network families and expressivity: construction and parameter counting for linear and ReLU networks, piecewise-linear structure of shallow ReLU nets, universal approximation as an existence result, extension to multivariate inputs/outputs, K-layer matrix-form notation, and comparison of shallow vs deep architectures in expressivity-per-parameter and practical trainability (how depth multiplies linear regions). - Losses and probabilistic view: deriving principled loss functions by treating network outputs as predictive distributions; building least-squares and binary/multiclass cross-entropy losses via a simple recipe. - Optimization and gradient computation: reliable minimization with minibatch stochastic gradient descent, momentum, and Adam; practical choices for learning rate and batch size; efficient gradient computation by backpropagation; weight initialization strategies to avoid vanishing/exploding signals. - Generalization, diagnostics, and regularization: decomposing test error into noise, bias, and variance, accounting for phenomena like double descent, and applying explicit/implicit regularizers such as early stopping, dropout, data augmentation, ensembling, and transfer learning. - Convolutional vision architectures: why fully connected layers are inadequate for images, how convolutions provide translation equivariance and controlled invariance, and how kernel, stride, padding, and dilation determine receptive fields and output size; channels and 2D kernels build feature maps; resolution and channel transforms (pooling, upsampling, 1×1 convolutions) and mapping architectures to vision tasks. - Engineering deep, trainable networks: architectural building blocks and fixes that make very deep models work—residual connections, normalization (batch norm), and block designs—plus parameter counting, receptive-field tracing, and justification of popular vision models such as ResNet, DenseNet, and U‑Net.

Outcome - You will be able to express, compare, and count parameters in network families; choose and derive losses; implement and tune gradient-based optimizers; initialize and monitor training to avoid stability issues; apply regularization and augmentation to improve generalization; and design or analyze convolutional architectures for vision tasks, assembling practical pipelines that train end-to-end for image models.

Units

Preliminaries to Machine Learning

Introduces supervised learning, linear and ReLU networks, depth vs width, and how to express and compare network families

Training Models

How to train deep networks end-to-end: principled losses, efficient optimization, stable gradients, and practical regularization

Convnets

How convolution, resolution changes, skips and normalization make deep vision models trainable and effective