2 Working with text data

Keywords

ver. 1.0.0, 2_working_with_text_data

Turn raw text into the continuous input embeddings an LLM consumes.

This unit teaches how to convert raw text into the numeric tensor inputs a transformer language model requires. You will learn to:

  • Tokenize text with regular expressions and justify splitting choices.
  • Build a vocabulary and a tokenizer with encode/decode that handles unknowns and document boundaries.
  • Use byte-pair encoding (BPE) to avoid out-of-vocabulary failures.
  • Produce next-token prediction examples by sliding a fixed-size window over token streams and batching them.
  • Map token IDs to vectors with torch.nn.Embedding (a learnable lookup table) and add learnable absolute positional embeddings to form the final input tensor.

This unit walks you through the complete input pipeline that transforms raw textual data into the continuous, position-aware embeddings consumed by a transformer-based language model. You will understand why text—discrete symbols—must be represented as continuous vectors so that neural networks can process and optimize over them, and you’ll learn what an embedding actually is: a learnable lookup table that maps token IDs to dense vectors and participates in gradient updates.

You’ll implement tokenization with a regular expression, making explicit the design choices and trade-offs that a splitting rule forces (what counts as a token, how punctuation and whitespace are handled). From those tokens you will construct a vocabulary and a tokenizer class providing matching encode/decode methods, and you will extend the vocabulary with semantic special tokens such as an unknown token and an end-of-text token so your tokenizer survives unseen inputs and marks document boundaries.

To handle arbitrary strings without out-of-vocabulary errors, you’ll replace a brittle word-based approach with byte-pair encoding (BPE), the subword algorithm used in practice (e.g., GPT-2/GPT-3). BPE produces a compact subword vocabulary that can spell any input while preserving frequent multi-character units.

For training, you’ll convert a long token stream into supervised next-token prediction examples by sliding a fixed-length window to form input–target pairs, and pack those examples into batches with a DataSet/DataLoader. This includes the logic for aligning inputs with targets and managing document boundaries and context windows.

On the model-input side you’ll implement torch.nn.Embedding to map token IDs into vectors, and develop an intuitive and technical understanding of that layer as a parameter matrix accessed via index lookup with gradients flowing back into it. Finally you’ll add learnable absolute positional embeddings and combine token and position vectors to produce the final input tensor of shape (batch_size, sequence_length, embedding_dim) that the transformer expects.

By the end you will be able to build a complete, production-quality input pipeline: tokenize and encode text (with robust special tokens or BPE), generate sliding-window next-token training examples and batches, convert token IDs to embeddings, inject positional information, and produce the exact tensors required by attention-based models—preparing you to implement and train transformer attention mechanisms next.

Materials

Source document

  • Building a Large Language Model (from scratch), Sebastian Raschka, 2024, Manning Books — Link — Page 39-71