1 Understanding large language models

Keywords

ver. 1.0.0, 1_understanding_large_language_models

Foundational understanding of what LLMs are, how they are built and applied, and the practical roadmap to start building one.

This unit explains what large language models are, how they fit into AI/ML/DL and generative AI, and why scaling a decoder transformer trained on large web corpora produces versatile behaviour. It covers common application areas, the two-stage lifecycle of pretraining then fine-tuning, the composition and compute implications of pretraining data, GPT’s decoder-only autoregressive architecture, emergent zero- and few-shot capabilities, and a three-stage roadmap that points to the hands-on next steps.

Learners finish this unit with a clear conceptual and practical map of large language models. They will be able to place LLMs within the broader fields of AI, machine learning, deep learning and generative AI and contrast them with earlier rule-based NLP approaches; survey concrete application areas where LLMs are effectively deployed and identify the properties that make a task a good fit for a pretrained model; and explain the two-stage lifecycle of building an LLM—a large, unlabeled foundation model trained at scale followed by comparatively small, task-specific adaptation.

The unit makes the transformer-from-the-previous-unit the central engine of LLMs by showing how a decoder-only, autoregressive simplification (as used in GPT) implements next-token prediction and why scaling that setup across massive internet-derived corpora can produce surprisingly general behaviours. It characterizes what a pretraining corpus actually looks like, how many tokens are involved, and the compute costs that implies, so learners understand practical constraints when training or selecting models.

A key conceptual takeaway is why zero-shot and few-shot in-context learning arise: they are emergent behaviours associated with scale and the training objective rather than explicit supervised instruction. The unit also provides a three-stage, actionable roadmap for building an LLM from scratch and explains where subsequent work in the module fits on that roadmap. Finally, learners leave ready to work with text data and to take the next concrete build step in the following unit, equipped to reason about architecture choices, data and compute trade-offs, and how to adapt a foundation model to downstream tasks.

Materials

Source document

  • Building a Large Language Model (from scratch), Sebastian Raschka, 2024, Manning Books — Link — Page 23-38