Convnets

Keywords

ver. 1.0.0, convnets

How convolution, resolution changes, skips and normalization make deep vision models trainable and effective

This unit teaches how convolutional networks process images: why fully connected layers fail, how convolutions provide translation equivariance and partial invariance, how to control receptive field and output size with kernel, stride, padding and dilation, and how channels and 2D kernels build feature maps. It explains resolution and channel transforms (pooling, upsampling, 1×1 convolutions), matches architectures to vision tasks, and solves training difficulties in deep nets by introducing residual connections and batch normalization. By the end you can design, analyze, and count parameters in convolutional layers, trace receptive fields, construct residual blocks, and justify architectural choices like ResNet, DenseNet, and U‑Net.

This unit develops a practical and conceptual toolkit for modern convolutional neural networks on images.

It begins by showing why directly applying fully connected models to raw pixels fails: images have local structure and symmetries that dense layers ignore. Convolutional layers replace dense connections with a shared kernel that slides across the input, giving translation equivariance (features shift with the image) and enabling useful invariances through pooling and stacking.

You learn the 1D convolution mathematically and how its hyperparameters—kernel size, stride, padding, dilation—determine what each unit sees and the spatial size of the output. Extending to channels and 2D shows how multiple kernels produce feature maps, how to count layer parameters, and how depth turns local operations into global computation by expanding receptive fields; the unit teaches how to trace a unit’s receptive field through many layers.

Practical plumbing that changes spatial resolution and channel depth is covered next: downsampling (pooling or strided convolution), upsampling (interpolation, transposed convolution), and 1×1 convolutions for channel mixing. Using these building blocks, the unit maps three common vision tasks to architectures: classification needs a global summary, detection requires multi-scale localization, and semantic segmentation needs high-resolution per-pixel predictions; they commonly share a backbone but differ in heads and preserved resolution.

The unit then addresses why naively stacking more layers eventually hurts optimization: very deep plain networks become harder to train, with higher training loss, showing the failure is optimization rather than capacity. Residual connections fix this by adding identity shortcuts around blocks so gradients have unobstructed paths backward; writing layer updates as x_{k+1} = x_k + f(x_k) reframes the network as an ensemble of short routes and eases gradient flow. Because residual addition can increase variance, batch normalization is introduced to stabilize activations, rescale variance, and smooth the loss surface; the treatment explains what BN fixes and what it costs.

Different ways of applying the skip principle are compared: ResNet adds identity shortcuts, DenseNet concatenates earlier features so each layer reuses all previous outputs, and U‑Net carries high-resolution features across long skip connections to preserve spatial detail for segmentation. Finally, competing explanations for residual networks’ success are weighed—smoother optimization landscapes and an implicit ensemble of shallow paths both contribute, beyond the simple claim of “more depth.”

After finishing the unit you can: - Explain why convolutions are preferred on images and state the equivariance/invariance properties they provide. - Define and compute 1D/2D convolutions and predict output sizes from kernel, stride, padding, and dilation. - Build convolutional layers with multiple channels, count their parameters, and trace receptive fields through stacks. - Use downsampling, upsampling, and 1×1 convolutions to change resolution and channel depth appropriately for classification, detection, or segmentation. - Diagnose why very deep plain networks fail to train and construct residual blocks that restore gradient flow. - Apply batch normalization correctly and articulate its benefits and tradeoffs. - Compare ResNet, DenseNet, and U‑Net as different instantiations of skip-based design and justify architectural choices based on task and training dynamics.

The unit closes by showing how these pieces combine into a complete, trainable vision model and by highlighting the unchanging recipe—define a model family suited to the data’s symmetries, assemble operators that respect those symmetries, and use architectural and normalization techniques that make deep models optimizable in practice.

Materials

Source documents

  • Understanding Deep Learning, Simon Prince, 2026 — Link — Page 175-199
  • Understanding Deep Learning, Simon Prince, 2026 — Link — Page 200-220