Language Models · v1.0.0
2026-08-25 14:45:28
A transformer accepts a matrix of numbers. Your corpus is a string. This unit is the whole distance between the two.
This unit is step 1 of stage 1: the data sampling pipeline.
Along the way: what does the tokenizer do with an unseen word, and how do we cut one token stream into training examples?
Neural networks operate on continuous values; text is categorical. An embedding maps discrete tokens into a continuous vector space.
\[E : V \to \mathbb{R}^{d}\]
Since \(V\) is finite, \(E\) is fully described by a matrix \(\mathbf{W} \in \mathbb{R}^{|V| \times d}\), one row per token.
LLMs typically produce their own embeddings, as part of the input layer, updated during training — rather than reusing a pretrained model like Word2Vec.
Higher dimensionality captures more nuance, at the cost of computational efficiency.
Tokenization splits input text into individual tokens — words or special characters, including punctuation.
Splitting text into individual tokens.
A tokenizer is a function \(T : \Sigma^{*} \to V^{*}\). The chapter arrives at its rule in three steps.
? : ; " ( ) --"world" and "world." are one entry) while remaining a signalA vocabulary maps each unique token to a unique integer, built by tokenizing the corpus, taking the set, sorting and numbering.
\[V = \operatorname{sorted}(\operatorname{set}(T(\text{corpus}))), \qquad \mathrm{id}(v_i) = i\]
A vocabulary built by sorting unique tokens and numbering them.
Since \(\mathrm{id}\) is a bijection between \(V\) and \(\{0, \dots, |V|-1\}\), one vocabulary supplies both directions.
class SimpleTokenizer:
def __init__(self, vocab):
self.str_to_int = vocab
self.int_to_str = invert(vocab)
def encode(self, text) -> list[int]:
return [self.str_to_int[t] for t in tokenize(text)]
def decode(self, ids) -> str:
text = join_with_space(self.int_to_str[i] for i in ids)
return reattach_punctuation(text)Apply this tokenizer outside the training set and it fails outright.
Encoding "Hello, do you like tea?" raises a lookup error — Hello never occurs in “The Verdict”.
This is not a bug. It is a property of a closed vocabulary built from one small corpus.
Special tokens <|unk|> and <|endoftext|> added to the vocabulary.
<|unk|> — stands in for a word outside the vocabulary, so encode returns an ID instead of raising<|endoftext|> — inserted before each document that follows a previous, unrelated text sourceThe repair is one line in encode: substitute before looking up, so no lookup can miss.
decode is unchanged — the unknown token round-trips as the literal string <|unk|>.
[BOS], [EOS], [PAD] mark sequence start, sequence end, and batch padding.
GPT models use none of these — only <|endoftext|>, which doubles as end-marker and padding token, and no <|unk|>, because GPT uses byte pair encoding instead.
Byte pair encoding (BPE) is the tokenization scheme used to train GPT-2, GPT-3 and the original ChatGPT model.
Rather than a fixed token set, BPE admits sub-word units, so any string can be spelled from pieces the vocabulary already holds.
In practice, one uses an existing implementation — tiktoken, exposing the same encode/decode interface.
<|endoftext|> is assigned token ID 50256 — the vocabulary size is 50,257someunknownPlace encodes and decodes correctly, with no <|unk|> tokenUnknown words broken into sub-words and characters.
"d" + "e" → "de")Input blocks predict the word following each block.
BPE-tokenizing “The Verdict” gives a single stream of 5,145 token IDs. The sampling problem: how to cut it into supervised examples.
Fix a context size \(n\). For a window at position \(i\):
\[\mathbf{x}^{(i)} = (t_i, \dots, t_{i+n-1}), \qquad \mathbf{y}^{(i)} = (t_{i+1}, \dots, t_{i+n})\]
The target at every position is the token that actually followed it — one window of length \(n\) carries \(n\) prediction tasks.
max_length — the context size \(n\)stride — how far the window advances between examplesstride = 1 — each token appears in up to \(n\) examples: maximally large, maximally redundantstride = max_length — windows tile without overlap: each token used once per epochThe chapter settles on the non-overlapping choice to use the dataset fully without skipping words, since more overlap risks overfitting.
An embedding is a weight matrix \(\mathbf{W} \in \mathbb{R}^{|V| \times d}\). Embedding a token ID selects its row:
\[E(t) = \mathbf{W}_{t,:}\]
The embedding layer retrieves a row per token ID.
The embedding layer is a more efficient implementation of one-hot encoding followed by a fully connected layer.
\[\mathbf{e}_t^{\top} \mathbf{W} = \mathbf{W}_{t,:}\]
Same map, different cost — and because it is equivalent to a linear layer, it can be optimized via backpropagation.
The embedding layer maps the same token ID to the same vector, regardless of position.
Token ID 5 gives the same vector in the first or fourth position.
Self-attention itself has no notion of order — position must be injected separately.
OpenAI’s GPT models use absolute positional embeddings, optimized during training. That is what we implement.
| table | rows | indexed by |
|---|---|---|
| token embedding | \(\lvert V \rvert\) = 50,257 | token ID |
| positional embedding | context length \(n\) | position \(0,\dots,n-1\) |
\[\mathbf{Z}_{b,i,:} = E_{\text{tok}}(t_{b,i}) + E_{\text{pos}}(i)\]
<|unk|>, <|endoftext|>) carry information about the text, not part of itThe pipeline hands over a tensor of shape (batch size \(\times\) context length \(\times\) embedding dimension).
Next: 3 Coding attention mechanisms. Self-attention, from a simplified form through trainable weights, causal masking, dropout, and a batched multi-head formulation.