Fast Object Detection Networks

Keywords

ver. 1.0.0, fast_object_detection_networks

How modern object detectors trade off speed, accuracy and scale: single‑shot grids, multi‑scale defaults, YOLOv2 and SSD training tricks, feature pyramids, and choosing the right detector.

This unit explains how contemporary object detection moved from two‑stage proposals to fast single‑shot detectors and back again with powerful feature pyramids. You will learn how grid detectors like YOLO frame detection as a regression from image pixels to boxes and class scores, how SSD extends that idea with default boxes on multiple feature maps, and which training techniques (weighted losses, responsible‑predictor routing, hard negative mining, augmentation) make these designs work. The course also traces YOLO→YOLOv2 improvements (Darknet‑19, ablations, YOLO9000 label hierarchies) and shows how Feature Pyramid Networks restore multi‑scale accuracy for region proposals and Fast R‑CNN heads. Finally you will be able to read benchmarks and pick a detector by reasoning about speed, accuracy and object scale.

This unit teaches the design, training and evaluation of modern object detectors with a focus on the architectural choices that determine speed, accuracy and the scale range of detectable objects.

It begins by reframing detection as a single regression problem: a network predicts all boxes and class probabilities in one forward pass by outputting a fixed S×S grid of B box candidates plus class scores. You will learn how YOLO represents detections as a single tensor and how the training signal is routed to the responsible predictor for each object. The unit explains YOLO’s weighted sum‑squared loss and the role of the responsible‑predictor rule in coping with extreme background dominance, and it quantifies where a coarse grid helps (fewer false positives, high speed) and where it fails (localisation precision and crowded scenes).

The course then shows how SSD keeps the single‑shot advantage while improving scale handling: default (anchor) boxes are tiled over several feature maps at different resolutions, and small convolutional predictors run over each map. You will learn why most default boxes are negatives and how hard negative mining plus aggressive data augmentation stabilise training and unlock SSD’s performance. SSD’s runtime/accuracy tradeoffs are presented and compared to two‑stage detectors so you can see when a single‑shot model is preferable.

YOLOv2 is reconstructed as a sequence of measured changes: a redesigned backbone (Darknet‑19) optimised for detection throughput rather than ImageNet classification, architectural tweaks and training changes whose individual mAP contributions are measured. You’ll study how to combine detection and classification data using a label hierarchy (WordTree) to train a single network that detects thousands of categories (YOLO9000).

To recover accuracy on small objects and across scales, the unit builds the Feature Pyramid Network (FPN): a top‑down pathway with lateral connections that produces feature maps that are both high resolution and semantically strong. With a pyramid available, anchors and RoI pooling are reworked so each pyramid level handles a narrow scale range; an RPN and a Fast R‑CNN head are attached per level and ablations justify these design choices. FPN’s COCO results and broader applicability (e.g., segmentation proposals) are examined.

Throughout, you learn to diagnose error profiles of single‑shot detectors, understand structural limits imposed by grids or feature resolution, and attribute gains to specific interventions (loss weighting, mining, augmentation, backbone redesign). The unit closes by consolidating the design space into actionable reasoning: given a latency budget, accuracy target and dominant object sizes, you will be able to choose or design a detector—single‑shot or two‑stage—with the appropriate backbone, multi‑scale strategy, and training recipe.

Materials

Source documents

  • Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 779-788).
  • Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C. Y., & Berg, A. C. (2016, September). Ssd: Single shot multibox detector. In European conference on computer vision (pp. 21-37). Cham: Springer International Publishing.
  • Redmon, J., & Farhadi, A. (2017, July). YOLO9000: better, faster, stronger. In 2017 IEEE conference on computer vision and pattern recognition (CVPR) (pp. 6517-6525). IEEE.
  • Lin, T. Y., Dollár, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2017, July). Feature pyramid networks for object detection. In 2017 IEEE conference on computer vision and pattern recognition (CVPR) (pp. 936-944). IEEE.