1 Understanding large language models

Language Models · v1.0.0

2026-08-25 14:44:25

Where we are

From layer to system

In the previous unit we built the transformer layer: dot-product self-attention, positional encoding, multiple heads, the position-wise MLP, residual connections and LayerNorm.

We ended with a glimpse of a GPT-class model carrying out tasks nobody had trained it on.

What this unit covers

  • What a large language model actually is, and how it differs from earlier NLP
  • The lifecycle: pretraining on unlabeled text, then fine-tuning
  • The data and scale behind LLMs, and the zero-shot/few-shot behaviour that arrives with them
  • The roadmap for the model we build over the rest of this module

What is an LLM?

Definition

A large language model (LLM) is a deep neural network trained on massive amounts of text to understand, generate and respond to human-like text.

“Large” does two jobs:

  • The model is large — tens or hundreds of billions of parameters
  • The dataset is large — often large portions of the public internet

On the word “understand”

When we say a language model understands, we mean it processes and generates text that appears coherent and contextually relevant.

We do not mean it possesses human-like consciousness or comprehension.

Where LLMs sit among the fields

LLMs as an application of deep learning.
  • AI — machines performing tasks requiring human-like intelligence
  • Machine learning — algorithms that learn from data
  • Deep learning — neural networks with 3+ layers
  • GenAI and LLMs sit inside deep learning, where it overlaps generative AI

The spam filter, twice

Traditional machine learning

  • Human experts hand-craft features: trigger words, exclamation marks, all-caps, suspicious links
  • Model trained on those features

Deep learning

  • No manual feature extraction
  • Model learns relevant features directly from text

Both still need labels — spam or non-spam has to come from somewhere in either approach.

Where LLMs are used

Four kinds of application

  • Generation and transformation — translation, novel text, summarization, code
  • Understanding — sentiment analysis and other classification of a passage
  • Interaction — chatbots and virtual assistants (ChatGPT, Gemini) augmenting search
  • Knowledge retrieval — sifting, summarizing and answering over specialized text (medicine, law)

What makes a task a good fit

Every application above takes unstructured text as input and produces text as output, posed as a continuation of a prompt.

A task requiring a guaranteed exact computation is not made a good fit by wrapping it in prose.

Pretraining, then fine-tuning

The two-stage pipeline

Pretraining an LLM, then fine-tuning it with a smaller labeled dataset.

Stage one: pretraining

  • Trains on a large, diverse corpus of raw, unlabeled text
  • Uses self-supervised learning — the model generates its own labels from the input
  • Produces a base or foundation model (e.g., GPT-3)

Capable of text completion and limited few-shot behaviour, but that is all — nothing task-specific yet.

Stage two: fine-tuning

Further training the pretrained model on a smaller labeled dataset.

  • Instruction fine-tuning — instruction/answer pairs, e.g. “translate this text” paired with the translation
  • Classification fine-tuning — texts with class labels, e.g. emails labeled spam / not spam

Why build your own

  • Domain performance — custom models can outperform general ones (BloombergGPT for finance)
  • Data privacy — avoid sharing sensitive data with third-party providers
  • Latency and cost — smaller models run locally, on-device
  • Autonomy — full control over updates and modifications

The transformer behind LLMs

The original transformer

Encoder processes input; decoder generates output, one word at a time.

  • Encoder — encodes input text into contextual vectors
  • Decoder — generates output text from those vectors

Two branches: BERT and GPT

Encoder segment (BERT) versus decoder segment (GPT).
  • BERT — encoder only, masked word prediction, strong at classification
  • GPT — decoder only, generative tasks: translation, summarization, code
  • GPT is also adept at zero-shot and few-shot learning

Transformers versus LLMs

The two terms are often used synonymously, and that is imprecise in both directions.

Not all transformers are LLMs (vision transformers). Not all LLMs are transformers (recurrent/convolutional variants exist).

We follow the source: “LLM” means transformer-based, GPT-like. This module follows the GPT branch.

The data behind the scale

The GPT-3 pretraining corpus

Dataset Tokens Proportion
CommonCrawl (filtered) 410 billion 60%
WebText2 19 billion 22%
Books1 12 billion 8%
Books2 55 billion 8%
Wikipedia 3 billion 3%

What the numbers say

  • Web crawl dominates — CommonCrawl alone is 410 billion tokens, about 570 GB
  • Not everything was used — 499B tokens available, 300B trained on
  • Scale and diversity are the point — they drive performance on syntax, semantics, context and general knowledge

Later models widened the mixture further — Meta’s LLaMA added Arxiv papers and StackExchange Q&A.

The cost, and what we will do instead

GPT-3 pretraining is estimated at $4.6 million in cloud compute.

We will implement pretraining and run it for educational purposes on consumer hardware, then reuse openly available weights — the learning is in the mechanism, not the budget.

Inside GPT, and what emerges

The objective is one word

GPT models are pretrained on next-word prediction — a form of self-supervised learning.

  • No labels to collect explicitly
  • The next word in the text is the label
  • Labels are created “on the fly”, so massive unlabeled text becomes usable training data

The architecture is the decoder alone

GPT uses only the decoder: unidirectional, left-to-right, one word at a time.

Because outputs feed back as inputs for future predictions, GPT is an autoregressive model.

What was never trained for

The ability to perform a task the model was never explicitly trained for is called emergent behavior.

Zero-shot and few-shot prompting without retraining.

The roadmap for this module

Three stages

Building an LLM: architecture and data, pretraining, fine-tuning.

Stage 1 — building an LLM

Data preparation and sampling, the attention mechanism, the LLM architecture.

This is the next several units of the module, beginning with 2 Working with text data.

Stage 2 — the foundation model

The training loop, model evaluation, loading pretrained weights.

Pretraining from scratch costs thousands to millions of dollars for GPT-like models — this stage focuses on the mechanism, on a small dataset, plus loading open weights.

Stage 3 — fine-tuning

Take the pretrained model and fine-tune it twice over:

  • Into a personal assistant that follows instructions
  • Into a classifier of texts

Closing

What this unit established

  • LLMs replaced rule-based, narrow NLP with one broadly capable, pretrained model
  • Two-stage lifecycle: pretrain on unlabeled text, fine-tune on labeled data
  • LLMs rest on the transformer; GPT-like LLMs keep only the decoder
  • Pretraining needs billions of tokens and serious compute
  • Next-word prediction alone produces emergent zero-shot and few-shot behaviour

Where next

Next: 2 Working with text data.

Stage 1 of the roadmap begins with the input pipeline: text split into pieces, mapped to integer IDs, cut into input-target pairs, and turned into embedding vectors.