7 Fine-tuning to follow instructions

Keywords

ver. 1.0.0, 7_fine_tuning_to_follow_instructions

Instruction-tune a GPT-2 Medium to follow arbitrary instructions, evaluate it automatically, and understand paths to further improvement.

This unit teaches how to turn a pretrained language model into an instruction-following assistant by supervised instruction fine-tuning. You will prepare and format an instruction dataset, implement dynamic padding and target masking in a PyTorch data pipeline, justify and load GPT-2 Medium, run a training loop and monitor loss, generate and persist responses on a test set, and score those responses automatically with a local LLM judge. It also summarizes next-step techniques that extend simple supervised tuning.

The unit walks through end-to-end supervised instruction fine-tuning of a pretrained autoregressive LLM so it reliably answers arbitrary prompts rather than merely continuing text. You begin by diagnosing the behavior of a base pretrained model on explicit instructions (it typically repeats or continues the input), then define instruction fine-tuning as the corrective approach: training on input–response pairs rendered into a consistent prompt template.

You learn practical data work: downloading an instruction corpus, formatting each record into an Alpaca-style prompt with a format_input helper, and splitting into train/validation/test. The unit covers building a robust PyTorch pipeline: wrapping data in a Dataset, writing a custom collate function that pads each batch to its longest sequence using the model’s padding token, creating shifted targets for language modeling, and replacing padding target tokens with -100 so padded positions do not contribute to gradients. The collate function is parameterized (device, max length) and wired into DataLoaders so you can inspect batch tensors before training.

Model choice and baseline behavior are emphasized: you load GPT-2 Medium (355M parameters) rather than the smaller checkpoint and demonstrate that the base LM cannot yet follow instructions. You then fine-tune the LM with a simple training loop (AdamW) for a couple of epochs, monitor training and validation loss to confirm learning, and observe the model begin to produce proper responses.

The unit covers inference and persistence: generating continuations for every test record, extracting the model’s response text, saving results and model weights, and then scoring open-ended outputs automatically by querying a locally hosted Llama 3 judge via a REST API. You produce per-example numeric scores (0–100) and an overall benchmark average.

Finally, the unit places this supervised instruction-tuning pipeline in context: it names follow-on techniques such as preference optimization / RLHF, parameter-efficient tuning (LoRA), structured hyperparameter search, and higher-level frameworks that package and automate these steps. By the end you can format instruction prompts, implement data and batching logic that correctly masks padding, fine-tune a medium-sized GPT-2 to follow instructions, generate and save responses, run automated evaluation with a local LLM judge, and articulate the next techniques to scale or refine the approach.

Materials

Source document

  • Building a Large Language Model (from scratch), Sebastian Raschka, 2024, Manning Books — Link — Page 226-272