Language Models · v1.0.0
2026-08-25 14:44:41
In 4 Implementing a GPT model from scratch to generate text we assembled a complete GPTModel and a greedy generation loop.
This is stage 2 of building an LLM: pretraining.
We train on a small public-domain story so every experiment runs on a laptop in minutes.
The model is the one built previously, with its context length shortened from 1,024 to 256 tokens so training fits on a laptop.
Generating text encodes text into token IDs, which the LLM processes into logit vectors. The logits are converted back into token IDs and detokenized.
Running the untrained model on "Every effort moves you" with greedy decoding — always taking the highest-probability token — produces gibberish.
Take a batch of tokenized inputs and targets, where each target is the input shifted one position forward.
Before training, the model produces random next-token probability vectors. Training aims to maximise the probability values at the highlighted target token IDs.
Before training, target_probas is close to \(1/|V|\) for a vocabulary of size \(|V|\).
Training is exactly the process of pushing these probabilities up.
Why the logarithm
Working with logarithms of probability scores is more manageable in mathematical optimization than handling the scores directly.
The deep learning convention is not to push the average log probability up toward \(0\), but to bring the negative average log probability down toward \(0\).
This quantity — cross entropy loss — is exactly what PyTorch’s cross_entropy computes in one call from the raw logits and target token IDs.
\[\mathcal{L} = -\frac{1}{N}\sum_{i=1}^{N} \log p(y_i \mid x_i)\]
No need to write the softmax, indexing, or logarithm by hand.
Perplexity is the exponential of the loss:
\[\text{perplexity} = e^{\mathcal{L}}\]
It reads as the effective vocabulary size the model is undecided among.
A loss on a couple of hand-picked sentences is not a measurement of a model.
The cost of pretraining LLMs
Llama 2 7B required 184,320 GPU hours on A100 GPUs over 2 trillion tokens — roughly $690,000. Our corpus has 5,145 tokens.
The corpus is split 90/10 into training and validation text, each fed through the sliding-window loader with window length and stride equal to the context length.
The input text is split into training and validation portions, tokenized, divided into fixed-length chunks, shuffled and organised into batches.
batch_loss measures a single batch.loader_loss averages over a loader, optionally over just a few batches for a cheap estimate.Applied to both loaders before training, both losses come out close to:
\[\log |V| \approx 10.98\]
For each epoch: iterate over training batches, compute the loss, backpropagate, and step the optimizer.
A typical PyTorch training loop iterates over the batches of the training set for several epochs, computing the loss, deriving gradients, and updating the weights.
def pretrain(model, train_loader, val_loader, optimizer, num_epochs):
for epoch in range(num_epochs):
for inputs, targets in train_loader:
loss = batch_loss(model, inputs, targets)
loss.backward()
optimizer.step()
optimizer.zero_grad()
if time_to_evaluate():
log(loader_loss(model, train_loader), loader_loss(model, val_loader))
sample_and_print(model, prompt="Every effort moves you")AdamW
AdamW is a variant of Adam that improves the weight decay approach, which aims to minimise model complexity and prevent overfitting by penalising larger weights.
The adjustment allows more effective regularization and better generalization — AdamW is the standard choice for LLM training.
Training this configuration for 10 epochs on “The Verdict” takes about five minutes on a laptop.
Both losses drop sharply at first. Past epoch 2 the training loss keeps falling while the validation loss stagnates — the model is overfitting.
The divergence past epoch 2, and the size of the gap, tell us the model is overfitting.
Expected: the dataset is small and we trained for multiple epochs. In practice, train on a much larger dataset for only one epoch.
The trained model still disappoints, and the reason is not the weights — it is the decoding.
argmax) is deterministic.Replace argmax with sampling in proportion to each token’s probability — drawing from the softmax distribution rather than taking its peak.
Temperature scaling divides the logits by a number \(T > 0\) before the softmax:
\[p_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}\]
argmax as \(T \to 0\).A temperature of 1 gives the unscaled scores. Lowering it to 0.1 sharpens the distribution; raising it to 5 makes it more uniform.
Creativity and coherence pull in opposite directions, and the temperature is the dial between them.
Higher temperatures explore less likely — and potentially more interesting — paths, but at the cost of sometimes producing nonsensical output.
Top-\(k\) sampling restricts the candidate set before sampling.
With \(k = 3\) we keep the three largest logits, mask the rest with \(-\infty\), and apply the softmax. Every non-top-\(k\) token then has probability 0.
Only the \(k\) largest logits survive; every other logit is set to \(-\infty\), so after softmax those tokens receive probability exactly zero — since \(e^{-\infty} = 0\).
Each step optionally masks to the top-\(k\) logits before deciding how to pick the next token.
Greedy decoding is the special case temperature = 0.0 — the old behaviour is subsumed, not discarded.
Pretraining is computationally expensive even at this scale, so the model must be saved rather than retrained every time it is needed.
Why switch to evaluation mode after loading
Dropout randomly drops neurons during training to prevent overfitting. At inference we do not want to discard learned information, so the reloaded model is switched to evaluation mode, which disables dropout.
Adaptive optimizers such as AdamW store per-parameter statistics used to adjust the learning rate for each weight dynamically.
A checkpoint intended for resuming training holds both the model’s and the optimizer’s state.
Notice the pattern: construct an architecture, then assign parameters into it by name.
That is exactly the mechanism the next section uses, with OpenAI supplying the parameters instead of our own training run.
OpenAI openly released the weights of their GPT-2 models.
OpenAI released four sizes of GPT-2. The core architecture is identical across them; only the embedding size and the number of repeated components differ.
GPT-2 ranges from 124 million to 1,558 million parameters. The architecture is the same; the embedding size and the repetition counts change.
To accept OpenAI’s weights, our configuration must match theirs:
Loading is a walk over every tensor in OpenAI’s checkpoint, assigning each into the corresponding parameter of our model, with a shape check at every step.
def load_weights_into_gpt(gpt, params):
assign(gpt.pos_emb, params.wpe)
assign(gpt.tok_emb, params.wte)
for block, block_params in zip(gpt.trf_blocks, params.blocks):
assign_attention_weights(block.att, block_params.attn) # split combined QKV, transpose
assign_feedforward_weights(block.ff, block_params.mlp)
assign_layernorm_weights(block.norm1, block.norm2, block_params)
assign(gpt.final_norm, params.final_ln)
assign(gpt.out_head, params.wte) # weight tying: reuse the token embedding matrixWith the weights loaded, generating from "Every effort moves you" with temperature and top-\(k\) sampling produces coherent English.
Proof that both the architecture and the weight assignment are correct — a single mismatched tensor would corrupt the output.
A pretrained GPT-2 is a general next-token predictor.
Stage 3 of building an LLM is adapting it.
The next unit, 6 Fine-tuning for classification, takes this pretrained model and specialises it to predict fixed categories.
The loss function, training loop and checkpointing learned here carry over unchanged — only the head and the targets differ.