Language Models · v1.0.0
2026-08-25 14:44:25
In the previous unit we built the transformer layer: dot-product self-attention, positional encoding, multiple heads, the position-wise MLP, residual connections and LayerNorm.
We ended with a glimpse of a GPT-class model carrying out tasks nobody had trained it on.
A large language model (LLM) is a deep neural network trained on massive amounts of text to understand, generate and respond to human-like text.
“Large” does two jobs:
When we say a language model understands, we mean it processes and generates text that appears coherent and contextually relevant.
We do not mean it possesses human-like consciousness or comprehension.
Traditional machine learning
Deep learning
Both still need labels — spam or non-spam has to come from somewhere in either approach.
Every application above takes unstructured text as input and produces text as output, posed as a continuation of a prompt.
A task requiring a guaranteed exact computation is not made a good fit by wrapping it in prose.
Pretraining an LLM, then fine-tuning it with a smaller labeled dataset.
Capable of text completion and limited few-shot behaviour, but that is all — nothing task-specific yet.
Further training the pretrained model on a smaller labeled dataset.
Encoder processes input; decoder generates output, one word at a time.
The two terms are often used synonymously, and that is imprecise in both directions.
Not all transformers are LLMs (vision transformers). Not all LLMs are transformers (recurrent/convolutional variants exist).
We follow the source: “LLM” means transformer-based, GPT-like. This module follows the GPT branch.
| Dataset | Tokens | Proportion |
|---|---|---|
| CommonCrawl (filtered) | 410 billion | 60% |
| WebText2 | 19 billion | 22% |
| Books1 | 12 billion | 8% |
| Books2 | 55 billion | 8% |
| Wikipedia | 3 billion | 3% |
Later models widened the mixture further — Meta’s LLaMA added Arxiv papers and StackExchange Q&A.
GPT-3 pretraining is estimated at $4.6 million in cloud compute.
We will implement pretraining and run it for educational purposes on consumer hardware, then reuse openly available weights — the learning is in the mechanism, not the budget.
GPT models are pretrained on next-word prediction — a form of self-supervised learning.
GPT uses only the decoder: unidirectional, left-to-right, one word at a time.
Because outputs feed back as inputs for future predictions, GPT is an autoregressive model.
The ability to perform a task the model was never explicitly trained for is called emergent behavior.
Zero-shot and few-shot prompting without retraining.
Building an LLM: architecture and data, pretraining, fine-tuning.
Data preparation and sampling, the attention mechanism, the LLM architecture.
This is the next several units of the module, beginning with 2 Working with text data.
The training loop, model evaluation, loading pretrained weights.
Pretraining from scratch costs thousands to millions of dollars for GPT-like models — this stage focuses on the mechanism, on a small dataset, plus loading open weights.
Take the pretrained model and fine-tune it twice over:
Next: 2 Working with text data.
Stage 1 of the roadmap begins with the input pipeline: text split into pieces, mapped to integer IDs, cut into input-target pairs, and turned into embedding vectors.