Language Models · v1.0.0
2026-08-25 14:45:36
Four things distinguish instruction fine-tuning from classification fine-tuning.
Dataset preparation, model setup and fine-tuning, then evaluation — the map for the rest of the unit.
Given “Fix the grammar in this text,” a base model just continues the text — it does not fix it.
Instruction fine-tuning continues training the same pretrained model on data where input-output pairs are explicitly provided, so it learns to interpret a task description and emit an appropriate answer.
The objective remains next-token cross-entropy, exactly as in pretraining.
What changes is the data distribution — instruction-response pairs in a fixed layout, instead of arbitrary web text.
“Following instructions” is a consequence of that distribution, not of a new loss term.
instruction, an optional input, and the reference output.input field may occasionally be empty — some instructions need no separate input.We use the Alpaca style throughout — one of the earliest to publicly detail its instruction fine-tuning process, and one of the most popular.
### Instruction: section with the instruction text.### Input: section only when the input field is non-empty.### Response: section follows, when building a training example.The ### Response: marker is the cue the model learns to answer after — and later the marker we slice on to recover a generated answer.
Each batch pads to its own longest sequence, using token ID 50256 (<|endoftext|>) — not to a dataset-wide maximum.
The target vector omits the first input token and appends one end-of-text token at the end.
-100Padding positions carry no information — replaced with -100. One exception: the first end-of-text token in each target is kept.
# pseudo-code
def collate_one(ids, pad_id=EOT_ID, ignore_index=-100,
batch_max_length):
ids = ids + [pad_id] # one eot token
ids = pad(ids, batch_max_length, pad_id) # pad to batch max
inputs, targets = ids[:-1], ids[1:] # shift by one
mask_all_but_first(targets, pad_id,
replacement=ignore_index)
return inputs, targets-100PyTorch’s cross-entropy loss defaults to cross_entropy(..., ignore_index=-100) — targets labeled -100 are simply skipped.
A worked check in the book confirms it directly: appending a target token valued -100 leaves the computed loss numerically unchanged.
The InstructionDataset and the custom collate function plug straight into PyTorch DataLoaders — training, validation and test.
allowed_max_length is fixed at 1024 — GPT-2’s maximum context length — so any batch is truncated rather than exceeded.Batch size 8; only the training loader shuffles and drops a trailing partial batch.
Printing the shape of each training batch shows the second dimension changing — 61, then 76, then 73, and so on.
Inputs and targets within a batch always share a shape. The varying width is direct evidence padding is dynamic, not global.
| GPT-2 Small | GPT-2 Medium | |
|---|---|---|
| Parameters | 124M | 355M |
| Layers | 12 | 24 |
| Attention heads | 12 | 16 |
| Embedding dim | 768 | 1024 |
“The 124-million-parameter model is too limited in capacity to achieve satisfactory results via instruction fine-tuning.”
Smaller models lack the capacity to learn and retain the nuanced behaviours required for high-quality instruction following. The medium checkpoint is roughly 1.42 GB, about three times the size of the small one.
Instruction: Convert the active sentence to passive: “The chef cooks the meal every day.”
Untrained model’s response: merely restates the prompt and part of the instruction.
Response section…Training loss (solid) and validation loss (dotted) over two epochs.
Rapid early decrease \(\to\) the model quickly learns meaningful patterns. Slower decrease in epoch two \(\to\) the model is converging to a stable solution.
The sample generated after the second epoch: “The meal is cooked every day by the chef.”
A direct contrast with the baseline, which merely echoed the prompt back.
The most crucial measure is response quality, not the loss plot alone.
### Response: marker \(\to\) model_response.Nothing here would work if the marker varied between records — the fixed template from stage 1 is what makes this slicing reliable.
Some responses are right, some nearly right. There is no arithmetic that turns this into a number by itself — motivating the next section.
Human evaluation does not scale. We take the third route, on our own test set, with a judge running locally.


Ollama wraps llama.cpp and runs an 8-billion-parameter, instruction-tuned Llama 3 locally. It does inference only — no training or fine-tuning.
The judge receives three things: the formatted instruction, the reference output, and our model’s response — and is asked for an integer score from 0 to 100.
“Respond with the integer number only.” Without this, the judge writes an explanatory essay instead of a bare number — scores would not be averageable.
Our fine-tuned model averaged 50.32 / 100 across the 110 test records.
Treat these numbers as a relative signal, not an absolute measure.
Preference fine-tuning
Trains on comparisons between responses rather than single reference answers — the stage separating an instruction-tuned model from a deployed assistant.
LoRA
Freezes the base weights and trains small low-rank matrices instead — the freezing idea from the classification unit, taken further.
The field is evolving rapidly — recent arXiv papers, r/LocalLLaMA, discussion on X.
But the durable knowledge is the mechanism — tokenization, attention, transformer blocks, loss, optimization — not any particular model release.
### Response: is both what the model learns to answer after, and what we slice on afterwards.-100 is not a magic number — it is PyTorch’s default ignore_index.<|endoftext|> is deliberately left unmasked, so the model learns when a response is complete.Tokenization and attention, a GPT architecture from scratch, pretraining, real GPT-2 weights, a spam classifier — and now a model that follows instructions.
The models will change. This mechanism will not.