Language Models · v1.0.0
2026-08-25 14:45:44
In 5 Pretraining on unlabeled data we finished the general-purpose model.
GPTModel.The result completes text fluently and knows nothing about any particular task.
This unit reaps the fruits of that labour — stage 3, step 8: fine-tuning a pretrained LLM as a classifier.
Figure 6.1 The three main stages of coding an LLM. This chapter focuses on stage 3 (step 8): fine-tuning a pretrained LLM as a classifier.
By the end, a 124M-parameter model fine-tuned on roughly a thousand labeled examples over five epochs classifies unseen messages at 95.67% accuracy.
The choice is made before any implementation begins — the two differ in the data they need, the compute they cost, and what the finished model can do.
Figure 6.2 Two different instruction fine-tuning scenarios: spam detection, and English-to-German translation.
The task is not peculiar to language: identifying plant species from images, categorizing news articles, distinguishing benign from malignant tumors.
A classification fine-tuned model is restricted to predicting classes it has encountered during its training.
Figure 6.3 A model fine-tuned for spam classification does not require further instruction alongside the input. It can only respond with “spam” or “not spam.”
We modify and classification fine-tune the GPT model we previously implemented and pretrained.
Figure 6.4 The three-stage process for classification fine-tuning an LLM.
ham or spam.The label counts are far from equal: 4,825 ham against 747 spam. A model that always answered “ham” would be right most of the time and useless.
Dataset undersampling shrinks the majority class to the size of the minority class.
ham \(\to 0\), spam \(\to 1\).The balanced set is shuffled and split 70/10/20 into training, validation and test.
The sliding window from the text-data unit does not apply here — messages have varying lengths, and a batch must be a rectangular tensor.
Sequence padding and truncation: we append the token ID 50256 ("<|endoftext|>"), from the same GPT-2 tokenizer.
Figure 6.6 Shorter sequences are padded with token ID 50256 to match the length of the longest sequence.
Conceptually: encode every message, find the longest encoded sequence in the training split, and pad (or truncate) every sequence to that common length.
max_length=None on the training set discovers 120 tokens as typical — well inside GPT-2’s 1,024-token context limit.
Standard PyTorch data loaders wrap the dataset.
(8, 120) tensor of token IDs, paired with an (8,) tensor of labels.Figure 6.7 A single training batch: eight messages as token IDs, with a class label array.
The targets are class labels, not the next tokens in the text.
We reuse the same GPT-2 configuration and weight-loading utilities from pretraining to obtain a GPTModel populated with pretrained parameters, in evaluation mode.
It struggles to follow instructions, as expected of a model that has only undergone pretraining. Fine-tuning is necessary.
Lower layers generally capture basic language structures applicable across many tasks; upper layers are more task-specific.
Selective parameter freezing disables gradients on every parameter as a starting point:
A small part of the model is re-enabled once the new head is in place.
The trade-off
Fewer trainable parameters means faster training and less overfitting risk on ~1,000 examples. Too few, and the model cannot adapt at all.
We replace the \(768 \rightarrow 50257\) output layer with a smaller classification head: \(768 \rightarrow 2\).
Figure 6.9 Replacing the vocabulary projection with a two-class output layer.
Training the head alone is sufficient, but fine-tuning additional layers noticeably improves performance — so the last transformer block and final LayerNorm are also made trainable.
Figure 6.10 The final LayerNorm and the last transformer block are trainable; the remaining 11 blocks and embeddings stay frozen.
A four-token input produces a \(4 \times 2\) tensor of logits, not \(4 \times 50257\). We need one prediction per example, so we keep only the last position: \(\text{logits}[:, -1, :]\).
Figure 6.11 Only the last row of the output tensor is used for classification.
Because of the causal attention mask: a token’s focus is restricted to itself and the positions before it.
The last token accumulates the most information since it is the only one with access to all previous tokens.
Figure 6.12 The last token, “time”, is the only one that computes attention scores for all preceding tokens.
The two class logits are converted to a label by taking the position of the highest value. Softmax is unnecessary — it does not change which position is largest:
\[ \hat{y} = \arg\max_c \; \text{logits}[-1, c] \]
Figure 6.14 Class labels obtained by looking up the highest-probability index. The model predicts incorrectly because it has not yet been trained.
Classification accuracy: the fraction of examples whose predicted label matches the target label.
Before any fine-tuning, evaluated over ten batches:
Cross-entropy loss is the differentiable proxy we optimize, applied only to the last-position logits:
\[ \mathcal{L} = \text{CrossEntropy}\big(\text{logits}[:, -1, :],\; y\big) \]
Before fine-tuning: 2.453 training, 2.583 validation, 2.322 test.
The training loop is the same overall loop used for pretraining. After each epoch we compute classification accuracy instead of generating a sample text.
Figure 6.15 A typical training loop: iterate over batches, compute loss, derive gradients, update weights.
# pseudo-code
def fine_tune(model, train_loader, val_loader, optimizer, num_epochs):
for epoch in range(num_epochs):
for inputs, targets in train_loader:
loss = cross_entropy(model(inputs)[:, -1, :], targets)
loss.backward(); optimizer.step(); optimizer.zero_grad()
periodically: log(loss on train_loader, val_loader)
log(accuracy on train_loader, val_loader)The optimizer is handed every model parameter; the frozen ones simply receive no gradient.
AdamW, learning rate \(5\times 10^{-5}\), weight decay \(0.1\), five epochs.


Little to no indication of overfitting — no noticeable gap between training and validation losses.
Recomputed over the full loaders:
The slight discrepancy between training and test accuracy suggests minimal overfitting. Validation accuracy is typically somewhat higher than test, because model development tunes hyperparameters against the validation set.
Having fine-tuned and evaluated the model, we are ready to classify new messages.
Figure 6.18 Step 10 — using the fine-tuned model to classify new spam messages.
Inference repeats exactly the preprocessing done at training time: tokenize, truncate to the shorter of max_length and the model’s context length, pad back to max_length, add a batch dimension.
Inference must mirror training. A mismatch here does not raise an error; it silently degrades accuracy.
Only the weights are saved, never the architecture.
The model we have built does exactly one thing.
Next: 7 Fine-tuning to follow instructions.
The dataset preparation, freezing, and training-loop patterns reappear; what changes is the target and how success is measured.