Computer Vision
ver. 1.0.0
How convolutional image classifiers and object detectors evolved, how to design and train them, and how to choose trade-offs between depth, size, speed and accuracy
- Evolution of CNN classifiers from AlexNet to Inception, VGG, ResNet and compact nets; practical design, training and evaluation tools.
- Techniques to train very deep nets (residual/dense connections) and to build efficient models (SqueezeNet, MobileNet, width/resolution multipliers).
- Region-based detectors (R-CNN → Fast R-CNN → Faster R-CNN) and single-shot approaches (YOLO, SSD), including modern multi-scale/feature-pyramid refinements.
- Key trade-offs and training tricks for accuracy, parameter/compute cost, speed and multi-scale detection; how to pick and compare models.
This unit bundle covers the development, design principles, training methods, and evaluation of modern convolutional image classifiers and object detectors, emphasizing practical trade-offs between accuracy, model size, latency and scale.
Classifiers and architectures - Why deep CNNs superseded hand-crafted features and the core building blocks (convolutions, 1×1 convs, pooling, global average pooling). - Historical progression and practical recipes: AlexNet’s training tricks, VGG-style deep nets, Inception modules for multi-path feature extraction, and how to compare models by accuracy, parameter count and compute. - Methods for scaling depth and improving optimization: residual learning (ResNets) to fix degradation, dense connectivity for feature reuse. - Designing compact, efficient networks (SqueezeNet, MobileNet) and metaparameter-driven knobs (width and resolution multipliers) to trade off accuracy versus latency and parameter cost.
Object detection families and trade-offs - Region-based two-stage detectors: formulating detection as classifying/localising region proposals; fine-tuning ImageNet backbones; bounding-box regression; shared convolutional computation via RoI pooling; multi-task losses; and Region Proposal Networks that unify proposal generation with the detection backbone (R-CNN → Fast R-CNN → Faster R-CNN). - Single-shot and grid-based detectors: framing detection as direct regression from pixels to boxes and scores (YOLO) and extending to multi-scale default boxes on feature maps (SSD). - Training and engineering techniques that make detectors work well: weighted losses, responsible-predictor routing, hard negative mining, data augmentation, backbone choices (e.g., Darknet-19), and label hierarchy ideas (YOLO9000). - Restoring multi-scale accuracy: using Feature Pyramid Networks to combine strong proposals with multi-scale Fast R-CNN heads, and how modern detectors balance single-shot speed with two-stage accuracy.
Putting it together: evaluation and model selection - How to evaluate and compare models (accuracy, speed, parameters, compute FLOPs) and reason about trade-offs in practice: choose detectors by expected object scale, latency constraints, and desired accuracy. - Practical outcome: you can design, train and regularize modern classifiers; build and train two-stage detectors; understand single-shot alternatives and their training tricks; and pick or tune architectures that meet specific speed/accuracy/scale requirements.
Units
Residual Convolutional Networks
How to make convolutional networks much deeper and much smaller while understanding the tradeoffs
Early Object Detection Networks
How to build, train and evaluate modern region-based object detectors (R-CNN → Fast R-CNN → Faster R-CNN), their components, and trade-offs
Fast Object Detection Networks
How modern object detectors trade off speed, accuracy and scale: single‑shot grids, multi‑scale defaults, YOLOv2 and SSD training tricks, feature pyramids, and choosing the right detector.