7 Fine-tuning to follow instructions

Language Models · v1.0.0

2026-08-25 14:45:36

Where we are

From a classifier to an assistant

  • Last unit: a spam classifier — one task, one of two labels.
  • This unit: keep the language-modeling head, and train the model to answer arbitrary requests instead.
  • Nothing about the architecture changes. The head stays; the targets change from class labels to full text responses.

What changes from the last unit

Four things distinguish instruction fine-tuning from classification fine-tuning.

  • A prompt template — a fixed layout marking where an instruction ends and its response begins.
  • Dynamic batching — each batch pads to its own longest sequence.
  • More capacity — GPT-2 Medium (355M) instead of GPT-2 Small (124M).
  • A different way to measure success — no accuracy figure for open-ended text.

The three-stage pipeline

Dataset preparation, model setup and fine-tuning, then evaluation — the map for the rest of the unit.

Why base models do not follow instructions

Pretraining rewards completion, not obedience

  • Pretraining optimises one objective: predict the next word.
  • The result “is capable of text completion… given a fragment as input.”
  • Nothing in that objective rewards obeying a request.

Given “Fix the grammar in this text,” a base model just continues the text — it does not fix it.

Examples that expose the gap

Instruction fine-tuning

Instruction fine-tuning continues training the same pretrained model on data where input-output pairs are explicitly provided, so it learns to interpret a task description and emit an appropriate answer.

What actually changes

The objective remains next-token cross-entropy, exactly as in pretraining.

What changes is the data distribution — instruction-response pairs in a fixed layout, instead of arbitrary web text.

“Following instructions” is a consequence of that distribution, not of a new loss term.

Formatting the instruction dataset

The dataset

  • 1,100 instruction-response pairs, created specifically for this book.
  • Each record has three fields: instruction, an optional input, and the reference output.
  • The input field may occasionally be empty — some instructions need no separate input.

Comparing prompt styles

We use the Alpaca style throughout — one of the earliest to publicly detail its instruction fine-tuning process, and one of the most popular.

The Alpaca formatting rule

  • Prepend a fixed preamble, then an ### Instruction: section with the instruction text.
  • Append an ### Input: section only when the input field is non-empty.
  • The reference ### Response: section follows, when building a training example.

The ### Response: marker is the cue the model learns to answer after — and later the marker we slice on to recover a generated answer.

Formatting an entry

# pseudo-code
def format_input(entry):
    text = PREAMBLE + f"### Instruction:\n{entry.instruction}"
    if entry.input:
        text += f"### Input:\n{entry.input}"
    return text
  • Preamble, then instruction, always.
  • Input section only if present.
  • Response appended separately, only for training examples.

Splitting the data

  • 935 training examples (85%)
  • 110 test examples (10%)
  • 55 validation examples (5%)

Batching with a custom collate function

Why we need our own collate function

  • In the classification unit, PyTorch’s default collate function was enough.
  • Here, batching is more involved — we write a custom collate function.

Dynamic padding

Each batch pads to its own longest sequence, using token ID 50256 (<|endoftext|>) — not to a dataset-wide maximum.

Targets: shift by one position

The target vector omits the first input token and appends one end-of-text token at the end.

Masking padding with -100

Padding positions carry no information — replaced with -100. One exception: the first end-of-text token in each target is kept.

Building one batch element

# pseudo-code
def collate_one(ids, pad_id=EOT_ID, ignore_index=-100,
                 batch_max_length):
    ids = ids + [pad_id]                     # one eot token
    ids = pad(ids, batch_max_length, pad_id)  # pad to batch max
    inputs, targets = ids[:-1], ids[1:]       # shift by one
    mask_all_but_first(targets, pad_id,
                        replacement=ignore_index)
    return inputs, targets
  • Append one eot token.
  • Pad to the batch’s longest sequence.
  • Shift to build targets.
  • Keep the first eot, mask the rest.

Why -100

PyTorch’s cross-entropy loss defaults to cross_entropy(..., ignore_index=-100) — targets labeled -100 are simply skipped.

A worked check in the book confirms it directly: appending a target token valued -100 leaves the computed loss numerically unchanged.

Masking the instruction is optional — and contested

  • It is also common to mask the instruction tokens too, so loss is computed on the response alone.
  • Researchers are divided: Shi et al. (2024) found that not masking the instruction benefits performance.
  • This chapter does not mask the instruction — left as an exercise.

Instruction data loaders

Assembling the loaders

The InstructionDataset and the custom collate function plug straight into PyTorch DataLoaders — training, validation and test.

Two configuration choices

  • The collate function moves each batch’s tensors onto the training device itself, as a background step, so it does not block the GPU during training.
  • allowed_max_length is fixed at 1024 — GPT-2’s maximum context length — so any batch is truncated rather than exceeded.

Batch size 8; only the training loader shuffles and drops a trailing partial batch.

Verifying the batches

Printing the shape of each training batch shows the second dimension changing — 61, then 76, then 73, and so on.

Inputs and targets within a batch always share a shape. The varying width is direct evidence padding is dynamic, not global.

Loading GPT-2 Medium

A bigger model

GPT-2 Small GPT-2 Medium
Parameters 124M 355M
Layers 12 24
Attention heads 12 16
Embedding dim 768 1024

Why the bigger model

“The 124-million-parameter model is too limited in capacity to achieve satisfactory results via instruction fine-tuning.”

Smaller models lack the capacity to learn and retain the nuanced behaviours required for high-quality instruction following. The medium checkpoint is roughly 1.42 GB, about three times the size of the small one.

The baseline, before training

Instruction: Convert the active sentence to passive: “The chef cooks the meal every day.”

Untrained model’s response: merely restates the prompt and part of the instruction.

  • Creates a Response section…
  • …but does not actually perform the conversion.
  • This is the failure the whole unit exists to fix.

If the hardware will not cooperate

  • Fallback: use GPT-2 Small instead of Medium.
  • Two epochs on GPT-2 Medium: ~16 minutes on a laptop CPU, under 2 minutes on a datacenter GPU.
  • GPT-2 Small trains faster still.

Fine-tuning on instruction data

Training setup

  • The training loop is reused unchanged from the pretraining unit — everything new lives in the data pipeline, not the loop.
  • Optimiser: AdamW, learning rate \(5\times 10^{-5}\), weight decay \(0.1\).
  • Two epochs, tracking the same validation example throughout.

Loss over training

Training loss (solid) and validation loss (dotted) over two epochs.

Reading the curves

  • Training loss falls sharply within the first epoch, then more slowly — ending near 0.30.
  • Validation loss flattens around 0.66.

Rapid early decrease \(\to\) the model quickly learns meaningful patterns. Slower decrease in epoch two \(\to\) the model is converging to a stable solution.

Why stop at two epochs

  • A deliberate stopping point, not an accident.
  • Extending training further “is not essential and may even be counterproductive” — risk of overfitting.
  • The gap between falling training loss and flattening validation loss is the ordinary signature of a small dataset.

The result

The sample generated after the second epoch: “The meal is cooked every day by the chef.”

A direct contrast with the baseline, which merely echoed the prompt back.

Extracting and saving responses

Stage 3: generate over the test set

The most crucial measure is response quality, not the loss plot alone.

How extraction works

  • Generate from the formatted instruction.
  • Slice the prompt off the front of the output, by its known length.
  • Strip the leftover ### Response: marker \(\to\) model_response.

Nothing here would work if the marker varied between records — the fixed template from stage 1 is what makes this slicing reliable.

Reading a few responses by hand

  • “The car is as fast as a bullet” — a correct simile (reference: “as fast as lightning”).
  • Correctly names Jane Austen, if more wordily than the reference.
  • “cumulus cloud” where the reference says “cumulonimbus” — close, but wrong.

Some responses are right, some nearly right. There is no arithmetic that turns this into a number by itself — motivating the next section.

Evaluating with an LLM judge

Three ways to evaluate

  • Short-answer / multiple-choice benchmarks (e.g. MMLU) — test general knowledge.
  • Human preference comparison — e.g. the LMSYS chatbot arena.
  • Automated conversational benchmarks — another LLM scores the responses (e.g. AlpacaEval).

Human evaluation does not scale. We take the third route, on our own test set, with a judge running locally.

Ollama and Llama 3

Ollama wraps llama.cpp and runs an 8-billion-parameter, instruction-tuned Llama 3 locally. It does inference only — no training or fine-tuning.

The scoring prompt

The judge receives three things: the formatted instruction, the reference output, and our model’s response — and is asked for an integer score from 0 to 100.

“Respond with the integer number only.” Without this, the judge writes an explanatory essay instead of a bare number — scores would not be averageable.

The result

Our fine-tuned model averaged 50.32 / 100 across the 110 test records.

  • Llama 3 8B base model, no fine-tuning: 58.51
  • Llama 3 8B instruct model: 82.6

Treat these numbers as a relative signal, not an absolute measure.

What the number is, and is not

  • Ollama is not entirely deterministic across operating systems — scores vary slightly between runs.
  • Repeating the evaluation and averaging gives more robust results.

Where the field goes next

Four directions extend this work

  • Preference fine-tuning — trains on comparisons between responses, e.g. Direct Preference Optimization (DPO).
  • Parameter-efficient fine-tuning — LoRA, low-rank adaptation.
  • Improving the result at hand — hyperparameters, more data, different prompts, a larger model.
  • Production frameworks — e.g. Axolotl or LitGPT.

Preference tuning and LoRA

Preference fine-tuning

Trains on comparisons between responses rather than single reference answers — the stage separating an instruction-tuned model from a deployed assistant.

LoRA

Freezes the base weights and trains small low-rank matrices instead — the freezing idea from the classification unit, taken further.

Staying current

The field is evolving rapidly — recent arXiv papers, r/LocalLLaMA, discussion on X.

But the durable knowledge is the mechanism — tokenization, attention, transformer blocks, loss, optimization — not any particular model release.

Summary

Key ideas \(^1/_2\)

  • The loss never changes — only the data distribution does, and the instruction-following behaviour follows from it.
  • A fixed prompt template gives the response a learnable boundary: ### Response: is both what the model learns to answer after, and what we slice on afterwards.
  • A custom collate function pads sequences, creates target token IDs, and masks padding tokens.

Key ideas \(^2/_2\)

  • -100 is not a magic number — it is PyTorch’s default ignore_index.
  • One terminal <|endoftext|> is deliberately left unmasked, so the model learns when a response is complete.
  • Capacity is a design decision: 355M parameters, 24 layers, 16 heads, \(\text{emb\_dim} = 1024\), two epochs over 935 examples.
  • Evaluation means extracting responses on a test set and scoring them with another LLM — our fine-tuned model averaged 50.32/100, a benchmark to compare against, not a verdict.

The module as a whole

Tokenization and attention, a GPT architecture from scratch, pretraining, real GPT-2 weights, a spam classifier — and now a model that follows instructions.

The models will change. This mechanism will not.