Early Object Detection Networks

Computer Vision · v1.0.13

2026-10-02 02:06:30

Where we are

Where this fits

  • The previous unit built backbones that classify a whole image: ResNet, DenseNet, MobileNet
  • A classifier maps an image to one of \(N\) labels, and the label is the entire output
  • Detection annotation is scarce beside ImageNet, so the classifier that already exists has to be reused

The backbone is given. What changes is the task and the head above it.

Outline of key topics

  1. The detection task, and how a prediction is graded
  2. R-CNN — proposals, per-region features, class-specific SVMs
  3. Fast R-CNN — the RoI pooling layer and one multi-task loss
  4. Faster R-CNN — the region proposal network and shared features
  5. The bottleneck each generation removed

A new learning task

From classification to detection

The list length varies with the image, which is what the classification head cannot produce.

Intersection over union

\[\mathrm{IoU}(A, B) = \frac{|A \cap B|}{|A \cup B|}\]

  • A true positive needs IoU \(\ge 0.5\) with a box of the same class
  • Ranking one class by score gives average precision
  • mAP averages that over classes

Every result in this unit is an mAP.

Two routes that fail at depth

Regression to coordinates

Map the image straight to box coordinates.

  • \(30.5\%\) mAP on VOC 2007
  • against \(58.5\%\) for the region approach

Sliding window

Score a dense grid of positions.

  • Units five convolutional layers deep have \(195 \times 195\) receptive fields
  • at strides of \(32 \times 32\) pixels

Precise localisation inside the sliding-window paradigm is left an open problem.

The R-CNN pipeline

Three modules

  1. A region proposal generator, independent of category
  2. A large convolutional network, one fixed-length vector per region
  3. A set of class-specific linear SVMs

Detection recast as region classification — recognition using regions.

Trial and error, by proposal

  • Selective search in fast mode emits about 2000 candidate boxes per image
  • The method is agnostic to the algorithm — objectness, CPMC and MCG are alternatives
  • The generator is a black box: nothing in it is learned

Every proposal is a guess to be scored, and almost all of them are wrong.

Warping to a fixed size

Proposals warped to \(227 \times 227\).

  • The network requires a fixed \(227 \times 227\) input; proposals are arbitrary rectangles
  • Anisotropic warping with context padding: dilate the box so \(p = 16\) pixels of context survive, then warp
  • Beats square crops by 3 to 5 mAP points

One forward pass per proposal, 2000 per image, each giving a 4096-d vector.

Scoring the candidates

  • Each 4096-d vector is scored by the linear SVM of each class
  • \(N + 1\) classes: \(N\) object types and one background class
  • Class-specific work is one matrix product, \(2000 \times 4096\) against \(4096 \times N\)

CNN parameters are shared across categories, so the 13 s per image is amortised over every class.

Non-maximum suppression

For one class, with \(\mathcal{B} = \{(b_i, s_i)\}\) the scored boxes and threshold \(\tau\):

\[ \begin{aligned} m &\leftarrow \operatorname*{argmax}_{i \,:\, (b_i, s_i) \in \mathcal{B}} s_i,\\ \mathcal{D} &\leftarrow \mathcal{D} \cup \{b_m\},\\ \mathcal{B} &\leftarrow \mathcal{B} \setminus \big(\{(b_m, s_m)\} \cup \{i : \mathrm{IoU}(b_m, b_i) > \tau\}\big) \end{aligned} \]

Repeated until \(\mathcal{B}\) is empty, from \(\mathcal{D} = \emptyset\), independently per class.

The same procedure, as pseudocode

sort B by score, descending
D <- empty
while B not empty:
  b <- B.pop_front()
  D.append(b)
  for c in B:
    if IoU(b, c) > tau:
      B.remove(c)
return D
  • Greedy: a box once kept is never reconsidered
  • Run per class, so boxes of different classes never suppress each other

Boxes that survive are still off target

  • Loc — IoU with the correct class between \(0.1\) and \(0.5\), or a duplicate
  • Significantly more errors are Loc than background or wrong-class
  • The reverse of the pattern DPM shows

The features discriminate; the boundaries are loose.

Bounding-box regression

A surviving proposal \(P = (P_x, P_y, P_w, P_h)\) and its \(\text{pool}_5\) feature give:

\[(d_x, d_y, d_w, d_h) = \mathbf{W}\,\boldsymbol{\phi}_5(P), \qquad \mathbf{W} \in \mathbb{R}^{4 \times 4096}\]

applied as a scale-invariant shift and a log-space rescaling:

\[ \begin{aligned} \hat{G}_x &= P_w d_x + P_x, & \hat{G}_w &= P_w e^{d_w},\\ \hat{G}_y &= P_h d_y + P_y, & \hat{G}_h &= P_h e^{d_h} \end{aligned} \]

Worth 3 to 4 mAP: \(54.2\%\) to \(58.5\%\) on VOC 2007.

Fitting the regressor

The targets

\[ \begin{aligned} t_x &= (G_x - P_x)/P_w\\ t_y &= (G_y - P_y)/P_h\\ t_w &= \log(G_w/P_w)\\ t_h &= \log(G_h/P_h) \end{aligned} \]

Ridge regression, \(\lambda = 1000\).

  • Fitted once per class
  • A proposal is used only if its IoU with the assigned ground truth exceeds \(0.6\)
  • Transforming a distant \(P\) into \(G\) is not well posed

Training, in three stages

  1. Pre-train the CNN on ILSVRC2012 image-level labels
  2. Fine-tune it on warped proposals, with a fresh \((N+1)\)-way layer, at one tenth the learning rate
  3. Fit the SVMs, then the regressors, on cached features

Two definitions of a positive: IoU \(\ge 0.5\) for fine-tuning, ground-truth boxes only for the SVMs.

What it cost

The bottlenecks

  • A full forward pass per proposal, thousands per image
  • Features written to disk between stages
  • Three separate training stages

The numbers, VGG16

  • \(2.5\) GPU-days to extract features for VOC07 trainval
  • Hundreds of gigabytes of storage
  • \(47\) s per image at test time

Nothing is shared between two proposals overlapping by 90%.

Fast R-CNN and the RoI pooling layer

One pass over the image

The whole image through the convolutional layers once; each RoI pooled from the shared map, then two sibling outputs — softmax over \(K+1\) classes, and per-class box offsets.

What an RoI is

An RoI is a rectangular window on the conv feature map,

\[(r, c, h, w)\]

its top-left corner, height and width.

  • One RoI is one proposal, projected onto the map by dividing by the total stride
  • A \(320 \times 480\) proposal becomes about \(20 \times 30\) positions at stride 16

The proposals still come from selective search. Only the forward pass is shared.

What RoI pooling computes

The \(h \times w\) window is divided into an \(H \times W\) grid, the \((i,j)\)-th cell covering

\[ \mathcal{R}_{ij} = \left\{ (y, x) \;:\; \begin{aligned} r + \left\lfloor \tfrac{(i-1)h}{H} \right\rfloor &\le y < r + \left\lceil \tfrac{ih}{H} \right\rceil\\ c + \left\lfloor \tfrac{(j-1)w}{W} \right\rfloor &\le x < c + \left\lceil \tfrac{jw}{W} \right\rceil \end{aligned} \right\} \]

\[\mathbf{Y}_{k,i,j} = \max_{(y,x) \in \mathcal{R}_{ij}} \mathbf{X}_{k,y,x}, \qquad \mathbf{Y} \in \mathbb{R}^{C \times H \times W}\]

Any RoI, whatever its size, yields the same fixed shape — which is what the fully connected layers require.

The grid, drawn

Pooling is per channel, and the output size does not depend on the RoI.

One multi-task loss

\[L(p, u, t^u, v) = L_{\mathrm{cls}}(p, u) + \lambda\,[u \ge 1]\,L_{\mathrm{loc}}(t^u, v)\]

  • \(L_{\mathrm{cls}}(p, u) = -\log p_u\), log loss over the \(K+1\) softmax outputs
  • \([u \ge 1]\) is zero for background, \(u = 0\), which has no ground-truth box
  • \(L_{\mathrm{loc}}\) sums \(\mathrm{smooth}_{L_1}\) over \((t_x, t_y, t_w, t_h)\), with \(\lambda = 1\)

Three training stages collapse into one, and the softmax it trains outperforms post-hoc SVMs.

Smooth \(L_1\)

\[\mathrm{smooth}_{L_1}(x) = \begin{cases} 0.5x^2 & |x| < 1\\ |x| - 0.5 & \text{else}\end{cases}\]

  • Quadratic near zero, linear beyond
  • \(L_2\) with unbounded targets needs careful learning-rate tuning to avoid exploding gradients

What RoI pooling bought

Speed

  • \(47\) s per image falls to \(0.32\) s
  • Training \(9\times\) faster
  • The disk cache disappears

What remains

  • Proposals still come from selective search, about 2 s per image on CPU
  • Low-quality candidates, thousands of them

The convolutional work no longer grows with the proposal count. The proposal step now dominates.

Faster R-CNN and the region proposal network

Proposals from the same features

  • Selective search costs about 2 s per image, an order of magnitude slower than the detector it feeds
  • EdgeBoxes at \(0.2\) s still consumes as much time as detection
  • A GPU reimplementation would ignore the network and miss the chance to share

The change is algorithmic: compute the proposals with a network over the features already there.

The sliding window on the feature map

  • A VGG backbone at stride 16 maps an image to \(H/16 \times W/16\) positions
  • An \(n \times n\) window, \(n = 3\), maps each position to 512-d
  • Then two sibling \(1 \times 1\) convolutions, cls and reg

Shared across all positions, so the head is three convolutions and nothing else.

Anchors

\(k = 9\) anchors per position: three scales \(128^2, 256^2, 512^2\) and three ratios \(1{:}1, 1{:}2, 2{:}1\).

What each position predicts

cls, \(2k\) outputs

Object against background, for each anchor. The class is not decided here.

reg, \(4k\) outputs

\[t_x = \frac{x - x_a}{w_a}, \;\; t_w = \log\frac{w}{w_a}\]

and likewise for \(y\) and \(h\), relative to anchor \(a\).

A pyramid of regression references, replacing pyramids of images or of filters.

The RPN loss

\[ \begin{aligned} L(\{p_i\}, \{t_i\}) = \;& \frac{1}{N_{cls}} \sum_i L_{cls}(p_i, p_i^*)\\ +\;& \lambda \frac{1}{N_{reg}} \sum_i p_i^* L_{reg}(t_i, t_i^*) \end{aligned} \]

  • \(p_i^* = 1\) above IoU \(0.7\), \(0\) below \(0.3\), between contributes nothing
  • \(N_{cls} = 256\), \(N_{reg} \approx 2400\), \(\lambda = 10\)

\(p_i^*\) makes the box loss active only on positive anchors.

Where the features are shared

Shared

The backbone and the feature map it produces, computed once per image.

Per RoI

The RPN head, and the detector head above RoI pooling — and the RPN’s surviving proposals are the detector’s RoIs.

4-step alternating training ties them: RPN, detector, RPN with shared layers fixed, detector head.

What sharing buys

System Proposals Total Rate
Selective search + Fast R-CNN 1510 ms 1830 ms \(0.5\) fps
RPN, shared, VGG-16 10 ms 198 ms 5 fps
RPN, shared, ZF 3 ms 59 ms 17 fps

300 learned proposals also beat 2000 selective search boxes: \(59.9\%\) against \(58.7\%\) mAP.

Summary

Three generations, three bottlenecks

  • R-CNN — about 2000 warped proposals, each with its own forward pass, trained in three stages. Bottleneck: per-proposal convolution.
  • Fast R-CNN — proposals projected onto one feature map and pooled to fixed size, trained in one stage. Bottleneck: the external proposals.
  • Faster R-CNN — proposals from a small network over the same features, at 10 ms. Remaining cost: the region-wise layers.

VOC 2007 with VGG-16: \(66.0\%\) at 47 s, \(66.9\%\) at \(0.32\) s, \(69.9\%\) at \(0.198\) s.

Machinery that recurs

Carried unchanged

  • The box parameterisation \((t_x, t_y, t_w, t_h)\), from R-CNN’s appendix to the RPN
  • The multi-task objective

IoU as a label-maker

  • \(0.5\) fine-tuning positives
  • \(0.3\) SVM negatives
  • \(0.7\) and \(0.3\) anchors
  • \(0.6\) regression pairs

A continuous overlap turned into a training label is the recurring device.

Looking forward

Where next

The next unit keeps the anchors and box regression, and drops the proposal stage.

  • YOLO predicts every box and class from an \(S \times S\) grid in one pass
  • SSD tiles default boxes over feature maps of several resolutions
  • YOLOv2 restores anchors, chosen by \(k\)-means

Dropping the proposal stage cost \(4.8\) mAP in this unit’s matched comparison. Those detectors win it back.