Language Models · v1.0.0
2026-08-25 14:45:14
(batch, tokens, embedding dimension).We build attention in four passes, each keeping everything from the pass before.
The attention variants we will code.
Translating German to English requires contextual and grammatical alignment.
The encoder compresses the input into one hidden state; the decoder generates from it.
The decoder can access all input tokens selectively, weighted by attention.
Self-attention lets each position in a sequence consider the relevance of — “attend to” — every other position in the same sequence.
The word “self” marks the contrast: ordinary attention relates elements of two different sequences (e.g. input and output); self-attention relates positions within one.
Each position interacts with and weighs the importance of every other position.
For a query token \(x^{(q)}\) in a sequence of \(T\) tokens:
\[ \omega_{q,i} = x^{(q)} \cdot x^{(i)}, \qquad i = 1, \dots, T. \]
The dot product measures similarity: the higher it is, the more aligned two vectors are.
Attention scores \(\omega\) between the query and every input, computed as dot products.
Normalize the scores so they sum to \(1\), using softmax:
\[ \alpha_{q,i} = \operatorname{softmax}(\omega_{q,\cdot})_i = \frac{e^{\omega_{q,i}}}{\sum_j e^{\omega_{q,j}}}. \]
Softmax is preferred over dividing by the sum: it handles extreme values better and has more favourable gradients, and it guarantees positive, interpretable weights.
Normalizing the scores \(\omega_{2i}\) into weights \(\alpha_{2i}\).
\[ z^{(q)} = \sum_{i=1}^{T} \alpha_{q,i}\, x^{(i)}. \]
The context vector is the weighted sum of all input vectors, weighted by attention.
The context vector \(z^{(2)}\) combines all inputs, weighted by \(\alpha\).
Every token needs the same treatment as a query — not just one.
The second row’s weights, generalized to every row.
\[ \Omega = X X^{\top}, \qquad \Omega \in \mathbb{R}^{T \times T}, \]
where \(\Omega_{i,j}\) is the dot product of token \(i\) with token \(j\).
Softmax must run along the row — the axis over keys, for a fixed query.
Normalizing along the wrong axis is the classic bug here. Checking that every row sums to \(1\) is a cheap way to catch it.
\(\Omega \in \mathbb{R}^{T\times T} \;\to\; \alpha X \in \mathbb{R}^{T \times d}\)
Adding trainable weights to the attention mechanism.
\[ q^{(i)} = W_q x^{(i)}, \qquad k^{(i)} = W_k x^{(i)}, \qquad v^{(i)} = W_v x^{(i)}. \]
Three trainable matrices \(W_q\), \(W_k\), \(W_v\) project each token into a query, key, and value vector.
The query, key, and value vectors, each from its own weight matrix.
\[ \omega_{q,i} = q^{(q)} \cdot k^{(i)}, \qquad \alpha_{q,\cdot} = \operatorname{softmax}\!\left(\frac{\omega_{q,\cdot}}{\sqrt{d_k}}\right). \]
Scaled scores, normalized into attention weights.
\[ z^{(q)} = \sum_i \alpha_{q,i}\, v^{(i)}. \]
The context vector is now a weighted sum over the value vectors, not the raw inputs.
Combining the value vectors via the attention weights.
Given an input tensor, produce query, key, value projections, then run the same three steps as before.
# pseudo-code
class SelfAttention:
def __init__(self, d_in, d_out):
self.W_query, self.W_key, self.W_value = three trainable projections d_in -> d_out
def forward(self, x):
q, k, v = W_query(x), W_key(x), W_value(x)
scores = q @ k.T
weights = softmax(scores / sqrt(d_out), along=rows)
return weights @ v\(X\) is transformed by \(W_q, W_k, W_v\); the attention matrix comes from \(Q\) and \(K\); \(Z\) comes from the weights and \(V\).
A raw trainable parameter matrix and a bias-free linear layer compute the same operation.
A linear layer comes with an optimized weight initialization scheme, which contributes to more stable and effective training — this is why frameworks’ built-in linear layer is carried forward for the rest of the chapter.
Weights above the diagonal are masked out.
Zero, then renormalize. Compute ordinary weights, zero every entry above the diagonal, renormalize each row to sum to \(1\).
Softmax, then zero, then normalize.
Fill with \(-\infty\) before the softmax. Since \(e^{-\infty} \to 0\), mask the scores directly, then apply softmax once.
No renormalization step is needed — each row already sums to \(1\).
Masking the scores with \(-\infty\) before softmax.
Both routes give the same lower-triangular attention matrix. The second is preferred: one softmax pass, not a softmax plus a separate normalization.
In GPT-style transformers, dropout is applied right after the attention weights are computed — the more common of two possible placements.
Dropout with rate \(p\) zeroes a random fraction \(p\) of entries, and scales every surviving entry by \(\dfrac{1}{1-p}\).
The scaling keeps the average influence of the attention mechanism consistent between training and inference, compensating for the reduced number of active elements.
Both are expected effects of the random mask, not errors.
Real inputs arrive as (batch, tokens, embedding dimension), because the data loader from the previous unit produces batched outputs.
The causal attention computation must run per example within the batch, not on a single sequence alone.
# pseudo-code
class CausalAttention:
def __init__(self, d_in, d_out, context_length, dropout):
self.W_query, self.W_key, self.W_value = three projections d_in -> d_out
self.dropout = dropout layer
self.mask = upper-triangular mask, shape (context_length, context_length)
def forward(self, x): # x: (batch, tokens, d_in)
q, k, v = W_query(x), W_key(x), W_value(x)
scores = q @ k.transpose(last two dims) # per-example, not global transpose
scores = fill(scores, mask, -inf)
weights = softmax(scores / sqrt(d_out), along=rows)
weights = dropout(weights)
return weights @ vFrom simplified attention, to trainable weights, to a causal mask. Multi-head is next.
Multi-head attention divides the mechanism into multiple heads, each operating independently, each with its own learned projections.
Two heads, two sets of \(W_q\), \(W_k\), \(W_v\); the two context-vector sets combine into one.
The most direct implementation runs several independent copies of the single-head module and concatenates their outputs along the feature dimension.
d_out * num_heads.Rather than separate classes for a single head and a multi-head wrapper, combine them into one module.
Two heads computed separately (top) vs. one larger projection split into heads (bottom).
# pseudo-code
class MultiHeadAttention:
def __init__(self, d_in, d_out, context_length, dropout, num_heads):
assert d_out % num_heads == 0
head_dim = d_out // num_heads
self.W_query, self.W_key, self.W_value = three projections d_in -> d_out
self.out_proj = projection d_out -> d_out
self.mask = upper-triangular mask
def forward(self, x): # x: (batch, tokens, d_in)
q, k, v = W_query(x), W_key(x), W_value(x) # (batch, tokens, d_out)
split each of q, k, v into (batch, num_heads, tokens, head_dim)
scores = q @ k.transpose(last two dims) # per head, batched
scores = fill(scores, mask, -inf)
weights = dropout(softmax(scores / sqrt(head_dim), along=rows))
context = weights @ v # per head
recombine heads -> (batch, tokens, d_out)
return out_proj(context)head_dim = d_out // num_heads — d_out must divide evenly across heads.d_out width.Only one matrix multiplication per role is needed — one for keys, one for queries, one for values — regardless of the number of heads.
The stacked wrapper repeated that multiplication once per head. This version implements the identical mathematical operation, more efficiently.
Next: 4 Implementing a GPT model from scratch to generate text.
Attention is one sub-layer of a transformer block. The next unit adds the rest of the block, stacks the blocks into a full GPT architecture, and runs the loop that turns the model’s output into generated text.