
Computer Vision · v1.0.13
2026-10-02 02:06:30
The backbone is given. What changes is the task and the head above it.

The list length varies with the image, which is what the classification head cannot produce.

\[\mathrm{IoU}(A, B) = \frac{|A \cap B|}{|A \cup B|}\]
Every result in this unit is an mAP.
Regression to coordinates
Map the image straight to box coordinates.
Sliding window
Score a dense grid of positions.
Precise localisation inside the sliding-window paradigm is left an open problem.

Detection recast as region classification — recognition using regions.
Every proposal is a guess to be scored, and almost all of them are wrong.

Proposals warped to \(227 \times 227\).
One forward pass per proposal, 2000 per image, each giving a 4096-d vector.
CNN parameters are shared across categories, so the 13 s per image is amortised over every class.
For one class, with \(\mathcal{B} = \{(b_i, s_i)\}\) the scored boxes and threshold \(\tau\):
\[ \begin{aligned} m &\leftarrow \operatorname*{argmax}_{i \,:\, (b_i, s_i) \in \mathcal{B}} s_i,\\ \mathcal{D} &\leftarrow \mathcal{D} \cup \{b_m\},\\ \mathcal{B} &\leftarrow \mathcal{B} \setminus \big(\{(b_m, s_m)\} \cup \{i : \mathrm{IoU}(b_m, b_i) > \tau\}\big) \end{aligned} \]
Repeated until \(\mathcal{B}\) is empty, from \(\mathcal{D} = \emptyset\), independently per class.

The features discriminate; the boundaries are loose.
A surviving proposal \(P = (P_x, P_y, P_w, P_h)\) and its \(\text{pool}_5\) feature give:
\[(d_x, d_y, d_w, d_h) = \mathbf{W}\,\boldsymbol{\phi}_5(P), \qquad \mathbf{W} \in \mathbb{R}^{4 \times 4096}\]
applied as a scale-invariant shift and a log-space rescaling:
\[ \begin{aligned} \hat{G}_x &= P_w d_x + P_x, & \hat{G}_w &= P_w e^{d_w},\\ \hat{G}_y &= P_h d_y + P_y, & \hat{G}_h &= P_h e^{d_h} \end{aligned} \]
Worth 3 to 4 mAP: \(54.2\%\) to \(58.5\%\) on VOC 2007.
The targets
\[ \begin{aligned} t_x &= (G_x - P_x)/P_w\\ t_y &= (G_y - P_y)/P_h\\ t_w &= \log(G_w/P_w)\\ t_h &= \log(G_h/P_h) \end{aligned} \]
Ridge regression, \(\lambda = 1000\).
Two definitions of a positive: IoU \(\ge 0.5\) for fine-tuning, ground-truth boxes only for the SVMs.
The bottlenecks
The numbers, VGG16
Nothing is shared between two proposals overlapping by 90%.
The whole image through the convolutional layers once; each RoI pooled from the shared map, then two sibling outputs — softmax over \(K+1\) classes, and per-class box offsets.
An RoI is a rectangular window on the conv feature map,
\[(r, c, h, w)\]
its top-left corner, height and width.
The proposals still come from selective search. Only the forward pass is shared.
The \(h \times w\) window is divided into an \(H \times W\) grid, the \((i,j)\)-th cell covering
\[ \mathcal{R}_{ij} = \left\{ (y, x) \;:\; \begin{aligned} r + \left\lfloor \tfrac{(i-1)h}{H} \right\rfloor &\le y < r + \left\lceil \tfrac{ih}{H} \right\rceil\\ c + \left\lfloor \tfrac{(j-1)w}{W} \right\rfloor &\le x < c + \left\lceil \tfrac{jw}{W} \right\rceil \end{aligned} \right\} \]
\[\mathbf{Y}_{k,i,j} = \max_{(y,x) \in \mathcal{R}_{ij}} \mathbf{X}_{k,y,x}, \qquad \mathbf{Y} \in \mathbb{R}^{C \times H \times W}\]
Any RoI, whatever its size, yields the same fixed shape — which is what the fully connected layers require.

Pooling is per channel, and the output size does not depend on the RoI.
\[L(p, u, t^u, v) = L_{\mathrm{cls}}(p, u) + \lambda\,[u \ge 1]\,L_{\mathrm{loc}}(t^u, v)\]
Three training stages collapse into one, and the softmax it trains outperforms post-hoc SVMs.

\[\mathrm{smooth}_{L_1}(x) = \begin{cases} 0.5x^2 & |x| < 1\\ |x| - 0.5 & \text{else}\end{cases}\]
Speed
What remains
The convolutional work no longer grows with the proposal count. The proposal step now dominates.
The change is algorithmic: compute the proposals with a network over the features already there.

cls and regShared across all positions, so the head is three convolutions and nothing else.

\(k = 9\) anchors per position: three scales \(128^2, 256^2, 512^2\) and three ratios \(1{:}1, 1{:}2, 2{:}1\).
cls, \(2k\) outputs
Object against background, for each anchor. The class is not decided here.
reg, \(4k\) outputs
\[t_x = \frac{x - x_a}{w_a}, \;\; t_w = \log\frac{w}{w_a}\]
and likewise for \(y\) and \(h\), relative to anchor \(a\).
A pyramid of regression references, replacing pyramids of images or of filters.
\[ \begin{aligned} L(\{p_i\}, \{t_i\}) = \;& \frac{1}{N_{cls}} \sum_i L_{cls}(p_i, p_i^*)\\ +\;& \lambda \frac{1}{N_{reg}} \sum_i p_i^* L_{reg}(t_i, t_i^*) \end{aligned} \]
\(p_i^*\) makes the box loss active only on positive anchors.
Shared
The backbone and the feature map it produces, computed once per image.
Per RoI
The RPN head, and the detector head above RoI pooling — and the RPN’s surviving proposals are the detector’s RoIs.
4-step alternating training ties them: RPN, detector, RPN with shared layers fixed, detector head.
| System | Proposals | Total | Rate |
|---|---|---|---|
| Selective search + Fast R-CNN | 1510 ms | 1830 ms | \(0.5\) fps |
| RPN, shared, VGG-16 | 10 ms | 198 ms | 5 fps |
| RPN, shared, ZF | 3 ms | 59 ms | 17 fps |
300 learned proposals also beat 2000 selective search boxes: \(59.9\%\) against \(58.7\%\) mAP.
VOC 2007 with VGG-16: \(66.0\%\) at 47 s, \(66.9\%\) at \(0.32\) s, \(69.9\%\) at \(0.198\) s.
Carried unchanged
IoU as a label-maker
A continuous overlap turned into a training label is the recurring device.
The next unit keeps the anchors and box regression, and drops the proposal stage.
Dropping the proposal stage cost \(4.8\) mAP in this unit’s matched comparison. Those detectors win it back.