Language Models · v1.0.0
2026-08-25 14:46:38
Fix the hyperparameters, sketch the whole model in placeholders, then fill each placeholder in.
A GPT model: embedding layers plus one or more transformer blocks containing masked multi-head attention.
The configuration collects every shape the model needs, in one place.
| key | meaning | GPT-2 124M value |
|---|---|---|
vocab_size |
BPE vocabulary size | 50,257 |
context_length |
max tokens via positional embeddings | 1,024 |
emb_dim |
embedding dimension per token | 768 |
n_heads |
attention heads per block | 12 |
n_layers |
number of transformer blocks | 12 |
drop_rate |
dropout probability | 0.1 |
qkv_bias |
bias on Q/K/V projections | False |
A placeholder model establishes the data path before any internals are real.
# pseudo-code
class DummyGPTModel:
def __init__(cfg):
tok_emb, pos_emb = embedding tables over vocab and context length
drop_emb = dropout(cfg.drop_rate)
trf_blocks = n_layers x DummyTransformerBlock # identity, for now
final_norm = DummyLayerNorm # identity, for now
out_head = linear(emb_dim -> vocab_size, no bias)
def forward(token_ids):
x = tok_emb(token_ids) + pos_emb(positions)
x = drop_emb(x)
x = trf_blocks(x)
x = final_norm(x)
return out_head(x) # logitsFeeding two tokenized sentences (four tokens each) through the placeholder gives output shape:
\[[2,\ 4,\ 50257]\]
Two texts, four tokens each, and one logit — an unnormalized score — for every vocabulary entry at every position.
The order in which we code the GPT architecture, starting from the placeholder backbone.
Deep networks are hard to train: vanishing or exploding gradients lead to unstable training dynamics and make it difficult to adjust weights effectively.
Layer normalization adjusts a layer’s activations to have mean \(0\) and unit variance.


\[\hat{x} = \gamma \odot \frac{x - \mu}{\sqrt{\sigma^2 + \varepsilon}} + \beta\]
scale) and \(\beta\) (shift) are trainable vectors of length emb_dim.Attention mixes information across token positions.
The feed-forward sub-layer does the opposite: it transforms each position on its own.
ReLU has long been the default activation, but LLMs generally use something smoother.
\[\mathrm{GELU}(x) = x \cdot \Phi(x)\]
where \(\Phi(x)\) is the standard Gaussian CDF. In practice, a cheaper curve-fit approximation is used — the same one GPT-2 was trained with:
\[\mathrm{GELU}(x) \approx 0.5 x \left(1 + \tanh\!\left[\sqrt{\tfrac{2}{\pi}}\left(x + 0.044715\,x^{3}\right)\right]\right)\]
GELU (left) against ReLU (right).
Two linear projections around a GELU:
With emb_dim = 768, the network widens each token vector to \(3072\) and narrows it back — shape is preserved end to end.
Expansion by a factor of four (768 → 3,072) and contraction back.
Shortcut connections (skip / residual connections) were originally proposed for deep computer-vision networks to mitigate vanishing gradients.
The vanishing gradient problem: gradients become progressively smaller as they propagate backward through layers, making it difficult to train earlier layers.
A shortcut creates an alternative, shorter path for the gradient: it skips one or more layers by adding the output of one layer to the output of a later layer.
Five layers without shortcuts (left) and with them (right), annotated with mean absolute gradient at each layer.
Our model stacks twelve transformer blocks, each with two sub-layers.
Shortcut connections are a core building block of very large models such as LLMs — they help ensure consistent gradient flow across layers. Every sub-layer in the next section is wrapped in one.
The transformer block is repeated a dozen times in GPT-2 124M. It combines multi-head attention, layer normalization, dropout, feed-forward layers, and GELU.
The operation is masked multi-head attention the first time, the feed-forward network the second. Only the layer-norm parameters differ between the two.
Normalization comes before each operation, not after — this is Pre-LayerNorm.
Older architectures, including the original transformer, applied normalization after attention and feed-forward — Post-LayerNorm — which often leads to worse training dynamics.
The block preserves shape: [batch, tokens, emb_dim] in, [batch, tokens, emb_dim] out.
Each row is one token’s 768-dimensional vector; output vectors keep the same dimension as the input.
Everything is now built. The full GPT model is the placeholder skeleton with its two placeholders replaced by real transformer blocks and real layer normalization.
The complete GPT model: token + positional embeddings, twelve transformer blocks, a final layer norm, and a linear output head.
# pseudo-code
class GPTModel:
def __init__(cfg):
tok_emb, pos_emb = embedding tables over vocab and context length
drop_emb = dropout(cfg.drop_rate)
trf_blocks = n_layers x TransformerBlock(cfg)
final_norm = LayerNorm(cfg.emb_dim)
out_head = linear(emb_dim -> vocab_size, no bias)
def forward(token_ids):
x = tok_emb(token_ids) + pos_emb(positions)
x = drop_emb(x)
x = trf_blocks(x)
x = final_norm(x)
return out_head(x) # logitsThanks to TransformerBlock, GPTModel is relatively small and compact.
Feeding the same two-sentence batch through the assembled model still gives output shape [2, 4, 50257].
Summing every parameter tensor in the assembled model gives \(163{,}009{,}536\) — not the 124 million we set out to build.
Where do the extra 39 million come from?
Weight tying was used in the original GPT-2: it reuses the weights from the token embedding layer in the output layer.
[vocab_size, emb_dim], about \(38.6\) million entries each.We do not tie the weights
Weight tying reduces memory footprint and computational complexity, but separate token embedding and output layers give better training and model performance in practice — the same is true for modern LLMs. Our model genuinely holds 163 million parameters; 124M describes OpenAI’s tied version.
At 32-bit floats, 4 bytes per parameter:
\[163{,}009{,}536 \text{ params} \;\approx\; 621.83\text{ MB}\]
That figure covers the weights alone, before optimizer state or activations.
The same architecture produces every GPT-2 size — only the configuration changes.
| Model | emb_dim |
n_layers |
n_heads |
|---|---|---|---|
| GPT-2 small | 768 | 12 | 12 |
| GPT-2 medium | 1,024 | 24 | 16 |
| GPT-2 large | 1,280 | 36 | 20 |
| GPT-2 XL | 1,600 | 48 | 25 |
A model that emits one distribution becomes a text generator by being called repeatedly on its own output.
Given “Hello, I am”, the model predicts a token, appends it, and predicts again — by the sixth iteration, it has constructed a complete sentence.
# pseudo-code
def generate(model, idx, max_new_tokens, context_size):
for _ in range(max_new_tokens):
idx_cond = idx[-context_size:] # crop to the model's context window
logits = model(idx_cond)
logits = logits[last position] # only the next-token prediction matters
probs = softmax(logits)
next_id = argmax(probs) # greedy decoding
idx = concat(idx, next_id)
return idxEncoding “Hello, I am” and running six generation steps through the untrained model produces:
Hello, I am Featureiman Byeswickattribute argue
Six iterations of the token prediction cycle, each appending its prediction to the input.
Gibberish, and expectedly so: we have only implemented the architecture and initialized random weights. Architecture alone does not produce language.
What is missing: nothing in this unit measures how wrong an output is, and nothing changes a weight in response.
Next: 5 Pretraining on unlabeled data.
It defines cross-entropy loss and perplexity, builds a training loop around AdamW, and shows the overfitting a small corpus produces. It replaces greedy decoding with temperature scaling and top-\(k\) sampling, adds checkpointing, and loads OpenAI’s pretrained GPT-2 weights into the architecture we just wrote.