
Computer Vision · v1.0.6
2026-09-25 22:16:38
Depth is worth having, and it is fragile. Neither paper explains the fragility.

Training error: 56 layers above 20.

Test error: the same ordering.
A 56-layer plain network trains worse than a 20-layer one, on CIFAR-10.
What remains is optimization: a solution exists, and solvers do not find it.
Let \(H(x)\) be the mapping a few stacked layers must fit. Instead, fit
\[F(x) := H(x) - x\]
and recover the original as \(F(x) + x\).
If identity is close to optimal, driving \(F\) toward zero is far easier than fitting an identity through nonlinear layers.

\[y = F(x, \{W_i\}) + x\]
No extra parameter, no meaningful extra computation.
Same dimensions
Dimensions increase
Projections help slightly; identity is the default because the gain does not justify the cost.
VGG-19, the 34-layer plain network, and the 34-layer residual network — shortcuts inserted every two layers, nothing else changed.

Two-layer block, ResNet-34.

Bottleneck block, ResNet-50/101/152.
\(1\times1 \to 3\times3 \to 1\times1\): reduce, then restore.
Identity shortcuts matter most here: a projection on a bottleneck doubles its cost.
| Network | Layers | FLOPs | Top-5 error |
|---|---|---|---|
| VGG-19 | 19 | 19.6 B | — |
| ResNet-34 | 34 | 3.6 B | — |
| ResNet-50 | 50 | 3.8 B | — |
| ResNet-101 | 101 | 7.6 B | — |
| ResNet-152 | 152 | 11.3 B | 4.49% |
152 layers, at lower cost than 19. Depth became cheaper, not more expensive.

Plain nets: higher error as depth grows, 20 to 56 layers.

ResNets: lower error as depth grows, to 110 layers.
The crossover is the whole argument: the same depths, opposite trends.
ResNet-110 against ResNet-1202.
Training error falls under 0.1% — but test error is worse than at 110. Overfitting, not degradation.
A residual layer adds the identity back: \(x_\ell = H_\ell(x_{\ell-1}) + x_{\ell-1}\).
\[x_\ell = H_\ell([x_0, x_1, \dots, x_{\ell-1}])\]
Concatenation keeps the identity and the new features distinct; summation merges them.
A 5-layer dense block, growth rate \(k=4\). Each layer takes every earlier feature map as input.

Transitions — batch norm, a 1x1 convolution, 2x2 average pooling — are the only place channel count is cut back.

Accuracy against parameter count.

Accuracy against FLOPs.
A DenseNet at ResNet-50’s computation performs on par with ResNet-101, at half the FLOPs.
Concatenation lets DenseNet go deeper than ResNet, at a fraction of the parameters.

Standard filters: every output reads every input channel.
A kernel \(D_K \times D_K \times M \times N\) maps an \(M\)-channel input to an \(N\)-channel output:
\[D_K \cdot D_K \cdot M \cdot N \cdot D_F \cdot D_F\]
What if filtering and combining were two cheaper layers instead of one expensive one?

One \(D_K \times D_K\) filter per input channel.
\[\hat{\mathbf{G}}_{k,l,m} = \sum_{i,j} \hat{\mathbf{K}}_{i,j,m} \cdot \mathbf{F}_{k+i-1,\, l+j-1,\, m}\]
It filters each channel and cannot combine them, so it creates no new features.

A \(1\times1\) kernel across all \(M\) channels.
\[\mathbf{G}_{k,l,n} = \sum_{m} \mathbf{K}_{m,n} \cdot \hat{\mathbf{G}}_{k,l,m}\]
It combines channels into the \(N\) new features the depthwise stage could not produce.
Depthwise costs \(D_K \cdot D_K \cdot M \cdot D_F \cdot D_F\); pointwise adds \(M \cdot N \cdot D_F \cdot D_F\). Together, against the standard convolution:
\[\frac{1}{N} + \frac{1}{D_K^2}\]
For a 3x3 kernel: 8 to 9 times less computation, at a small accuracy cost.
A standard convolutional layer with batch norm and ReLU (left), against the depthwise separable pair (right).
Shortcut connections recur across every architecture in this unit.
These backbones are what object detection consumes.
A better backbone lifts every downstream task.