Residual Convolutional Networks

Keywords

ver. 1.0.0, residual_convolutional_networks

How to make convolutional networks much deeper and much smaller while understanding the tradeoffs

This unit shows how to train very deep convolutional networks and how to design parameter- and latency-efficient CNNs. It introduces residual learning and ResNets to fix the degradation optimization problem, explores dense connectivity for extreme feature reuse, and presents compact architectures (SqueezeNet, MobileNet) together with metaparameter-driven design and width/resolution multipliers to trade accuracy for cost.

The unit explains two complementary problems and their practical solutions: how to make convolutional networks reliably deeper, and how to make them much smaller or faster for constrained budgets.

First, it isolates the degradation problem — deeper plain networks can get worse at both training and test time — and shows this is an optimization failure distinct from overfitting or vanishing gradients. The remedy is residual learning: reformulate a block to learn F(x)=H(x)-x and use an identity shortcut so the block only needs to learn residuals. That change is parameter-free but dramatically alters what gradient descent can find. From that idea the ResNet family is constructed: basic and bottleneck residual blocks yield ResNet-18/34/50/101/152 and make extremely deep models affordable. Empirical studies (including controlled CIFAR experiments up to very large depths) demonstrate where depth helps and where new limits appear.

Taking shortcuts further, the unit introduces dense connectivity: within a block every layer receives the concatenation of all previous feature maps. Dense blocks promote feature reuse, provide implicit deep supervision, and act as a regularizer, producing strong parameter efficiency across CIFAR, SVHN, and ImageNet experiments.

On the flip side of “deeper” is “smaller.” The unit covers squeeze-and-expand Fire modules to meet a hard parameter budget (SqueezeNet), showing how microarchitectural choices yield AlexNet-level accuracy at a fraction of the parameter count. It then treats architecture as a parameterized design space: separate microarchitecture (module internals) from macroarchitecture (how modules are wired) and sweep metaparameters rather than hand-tuning single instances. MobileNet follows by factoring standard convolutions into depthwise and pointwise stages for large computation savings, introducing the width multiplier α and resolution multiplier ρ so a single design becomes a family of models along the accuracy-vs-cost curve.

By the end of the unit a learner can diagnose degradation vs. other training failures, build and justify residual blocks and whole ResNets (basic and bottleneck), read and reason about empirical depth effects, construct dense blocks and explain their effects on gradients and parameters, design compact models under strict budgets (Fire modules, SqueezeNet), factor convolutions for latency-sensitive design (MobileNet), and pick operating points using α and ρ or explore architectures through metaparameters to meet real-world accuracy, size, and latency constraints. These backbones prepare learners for downstream tasks that require stronger spatial reasoning, such as object detection.

Materials

Source documents

  • He, Kaiming, et al. “Deep residual learning for image recognition.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2016.
  • Iandola, Forrest N., et al. “SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and< 0.5 MB model size.” arXiv preprint arXiv:1602.07360 (2016).
  • Huang, Gao, et al. “Densely connected convolutional networks.” 2017 IEEE conference on computer vision and pattern recognition (CVPR). Ieee, 2017.
  • Howard, Andrew G., et al. “Mobilenets: Efficient convolutional neural networks for mobile vision applications.” arXiv preprint arXiv:1704.04861 (2017).