Chapter 12: Transformers

Keywords

ver. 1.0.0, chapter_12_transformers

How transformers turn text into contextual representations and generated text, why they suit language better than CNNs/FC layers, and how they scale and adapt to other modalities

This unit explains the transformer architecture end-to-end: why fully connected and convolutional layers struggle with text, how dot‑product self‑attention (queries, keys, values) is computed in per‑token and matrix form, and the practical additions (positional encodings, scaled dot products, multi‑head attention) that make it usable. It shows how transformer layers are built (multi‑head attention + position‑wise MLP + residuals + LayerNorm), how raw text becomes input via subword tokenization and learned embeddings, and how encoder, decoder and encoder‑decoder variants differ in masking and objective. The unit also covers the O(N^2) cost of attention and approaches to reduce it, and how the same ideas apply to images (patches/pixels as tokens).

Learners will gain a conceptual and practical understanding of transformers sufficient to compute, assemble and adapt them. The unit begins by contrasting text with images: text has variable length, extremely high input dimensionality per token, and long‑range ambiguity that makes local receptive fields and dense fully connected layers poorly suited. Self‑attention supplies a different inductive bias: each token dynamically routes and aggregates information from every other token.

You will learn how to form attention from queries, keys and values — first at the per‑token level (how one token decides how much to take from another) and then collapsed into the standard matrix expression used in implementations. From those formulas you will be able to compute dot‑product self‑attention by hand or in code. The unit then adds the practical fixes that turn the bare mechanism into a robust layer: positional encodings (so order is represented), scaling of dot products (to stabilise gradients), and multi‑head attention (to let the model attend to several relationships in parallel).

With those building blocks you will assemble a full transformer layer: multi‑head self‑attention followed by a position‑wise MLP, wrapped with residual connections and LayerNorm. You will follow raw text through sub‑word tokenization and learned embedding matrices to produce the input matrix X that transformer layers consume. The unit distinguishes encoder, decoder and encoder‑decoder families by their masking patterns and training objectives, showing how BERT uses bidirectional (unmasked) attention with masked language modelling for contextual representations and fine‑tuning, while GPT‑style decoders use causal (masked) attention and autoregressive objectives to generate text and enable few‑shot behaviour at scale. Cross‑attention in encoder‑decoder models is explained as queries coming from one sequence and keys/values from another, demonstrated in machine translation.

The course quantifies the quadratic O(N^2) time and memory bottleneck of full self‑attention and surveys the three main strategies to mitigate it: sparse attention patterns, low‑rank or projected representations, and kernel or algorithmic rewrites that approximate attention more cheaply. Finally, the unit shows how the same transformer machinery adapts to vision by treating patches or pixels as tokens (ImageGPT, ViT, multi‑scale designs), highlighting how reduced inductive bias trades off with scalability and learned structure.

By the end you will be able to: explain why transformers are preferred for text; compute dot‑product self‑attention in token and matrix forms; apply positional encoding, scaling and multi‑head designs; build a transformer block from attention, MLPs, residuals and LayerNorm; convert raw text to transformer inputs via tokenization and embeddings; distinguish encoder/decoder/encoder‑decoder architectures and their objectives; analyse attention’s quadratic cost and describe mitigation approaches; and reason about transformer adaptations to images. The unit also points to next steps for understanding and building large language models.

Materials

Source document

  • Understanding Deep Learning, Simon Prince, 2026 — Link — Page 221-253