Lecture notes — 1 Understanding large language models

Published

2026-09-02 00:00

Keywords

ver. 1.2.0, 1_understanding_large_language_models

← 1 Understanding large language models

ver. 1.2.0 · 2026-09-02 16:26:54

Where this fits

The previous unit, Chapter 12: Transformers, established the mechanism: dot-product self-attention with queries, keys and values, positional encodings, scaled dot products, multi-head attention, and the transformer layer assembled from multi-head attention, a position-wise MLP, residual connections and LayerNorm. It also separated the three families by their masking pattern and training objective — bidirectional encoders, causal decoders, and encoder-decoder pairs joined by cross-attention.

This unit takes the decoder family and asks what a working system built from it looks like. The architecture is unchanged. What changes is the scale of the training corpus, the two-stage lifecycle that produces a usable model, and the behaviour that appears at the far end of that scale. No code appears here; the code begins with the input pipeline in the next unit.

Learning outcomes

  1. situate-llms — Place LLMs within AI, machine learning, deep learning and generative AI, and contrast them with earlier rule-based NLP.
  2. describe-llm-applications — Survey the tasks LLMs are deployed on and identify what makes a task a good fit.
  3. explain-pretraining-finetuning — Explain the two-stage lifecycle of pretraining a foundation model and fine-tuning it for a task.
  4. characterize-pretraining-data — Characterize the scale and composition of an LLM pretraining corpus, and the compute it implies.
  5. relate-gpt-to-transformer — Relate the GPT architecture to the transformer of the previous unit as a decoder-only, autoregressive simplification.
  6. explain-emergent-abilities — Explain zero-shot and few-shot in-context learning as behaviour that emerges from scale rather than from explicit training.
  7. navigate-the-build-roadmap — Navigate the three-stage roadmap for building an LLM from scratch and locate each unit of this module on it.

Concepts introduced

  • Large Language Model (LLM) — a deep neural network with tens or hundreds of billions of parameters, trained on a corpus of comparable scale, designed to understand, generate and respond to human-like text.
  • Transformer architecture — the deep network architecture of the 2017 paper “Attention Is All You Need”, built from an encoder submodule and a decoder submodule connected by self-attention.
  • Self-attention mechanism — the component that weighs the importance of different words or tokens in a sequence relative to each other.
  • Next-word prediction — the self-supervised objective of predicting the upcoming word in a sequence from the words preceding it.
  • Pretraining — the first training stage, on a large and diverse corpus of raw unlabeled text, which yields a base or foundation model.
  • Fine-tuning — the second training stage, on a smaller labeled dataset narrower in task or domain.
  • Encoder versus decoder architectures — the split between BERT-style models built on the encoder submodule and GPT-style models built on the decoder submodule.
  • Emergent behavior — a capability the model was never explicitly trained to perform, arising as a consequence of exposure to vast and diverse data.
  • Zero-shot and few-shot learning — carrying out a task from the prompt alone, with no examples or with a handful of examples supplied in the input.

What an LLM is

An LLM is a neural network designed to understand, generate and respond to human-like text. It is a deep neural network trained on massive amounts of text data, sometimes encompassing large portions of the entire publicly available text on the internet.

The word “large” carries two claims at once, and both are load-bearing.

  • Model size. Models of this kind often have tens or even hundreds of billions of parameters — the adjustable weights in the network, optimized during training.

  • Dataset size. The immense corpus on which the model is trained.

The parameters are optimized on a single task: predicting the next word in a sequence. Next-word prediction is a sensible objective because it harnesses the inherent sequential nature of language, training the model on context, structure and relationships within text. It is a very simple task, and it surprised many researchers that it can produce such capable models.

ImportantWhat “understand” means here

When a language model is said to understand language, the claim is that it can process and generate text in ways that appear coherent and contextually relevant. It is not a claim that the model possesses human-like consciousness or comprehension. The word is used in the first sense throughout.

The nested fields

Figure 1.1. LLMs as a specific application of deep learning, nested inside machine learning and artificial intelligence, and overlapping generative AI.
  • Artificial intelligence covers the creation of machines that can perform tasks requiring human-like intelligence — understanding language, recognizing patterns, making decisions. Machine learning and deep learning are subfields of it, but not the whole of it: AI also includes rule-based systems, genetic algorithms, expert systems, fuzzy logic and symbolic reasoning.
  • Machine learning develops algorithms that learn from data and make predictions or decisions without being explicitly programmed.
  • Deep learning is the subset of machine learning that uses neural networks with three or more layers to model complex patterns and abstractions in data.
  • Generative AI uses deep neural networks to create new content, such as text, images or other media. LLMs, being capable of generating text, sit in this region as well.

Because LLMs use the transformer architecture, they can pay selective attention to different parts of the input when making predictions, which is what suits them to the nuances and complexities of human language.

Where the boundary between machine learning and deep learning falls

A spam filter makes the distinction concrete. In traditional machine learning, human experts manually extract features from the email text: the frequency of certain trigger words such as “prize”, “win” or “free”, the number of exclamation marks, the use of all uppercase words, or the presence of suspicious links. The dataset created from those expert-defined features is what trains the model.

Deep learning requires no such manual feature extraction. Human experts do not identify and select the relevant features. Both approaches still require labels — spam or non-spam — gathered from an expert or from users; the difference is confined to who chooses the features.

Earlier NLP models were typically designed for specific tasks: text categorization, language translation and so on. They excelled in their narrow applications. Before LLMs, traditional methods handled categorization tasks such as email spam classification and straightforward pattern recognition well, but underperformed on language tasks demanding complex understanding and generation — parsing detailed instructions, conducting contextual analysis, producing coherent original text. Previous generations of language models could not write an email from a list of keywords, a task that is trivial for contemporary LLMs.

Learning outcomes

  • situate-llms Place LLMs within AI, machine learning, deep learning and generative AI, and contrast them with earlier rule-based NLP.

Concepts

  • large-language-model a deep neural network of tens or hundreds of billions of parameters, trained on a comparably large corpus, for understanding and generating human-like text
  • transformer-architecture the architecture LLMs are built on, which lets them attend selectively to parts of the input when making a prediction
  • next-word-prediction the task on which the network’s parameters are optimized, exploiting the sequential nature of language

Where LLMs are used

The applications follow from one capability: parsing and understanding unstructured text data. The list below is drawn from the source, and its breadth is the evidence for how general a single pretrained model is.

  • Generation and transformation of text. Machine translation, generation of novel texts, and text summarization. Content creation extends this to writing fiction, articles, and computer code.

    Each of these was once served by a model built for it alone. They are now downstream uses of one pretrained model.

  • Understanding. Sentiment analysis, and the categorization tasks that traditional methods already handled.

  • Interaction. Chatbots and virtual assistants such as OpenAI’s ChatGPT or Google’s Gemini, formerly called Bard, which answer user queries and augment traditional search engines such as Google Search or Microsoft Bing.

  • Knowledge retrieval. Retrieval from vast volumes of text in specialized areas such as medicine or law — sifting through documents, summarizing lengthy passages, and answering technical questions.

The common shape is a task that involves parsing and generating text. LLMs are applicable to almost any such task, which is what makes their range of application so wide.

Figure 1.2. A chat interface: the user supplies the instruction “Write a 4-line poem containing the words Wisconsin, AI, and pizza”, and the model returns the poem.

Figure 1.2 is worth reading closely, because it shows the interface an LLM presents rather than the model itself. The user’s instruction and the model’s output occupy the same channel — natural language. Nothing in the exchange names a task, selects a model, or configures an output format. The specification “4-line poem containing the words Wisconsin, AI, and pizza” is itself part of the input the network reads.

This is the property that makes one model serve many applications. A task is specified in the same medium as the data, so switching tasks requires no change to the model.

Learning outcomes

  • describe-llm-applications Survey the tasks LLMs are deployed on and identify what makes a task a good fit.

Concepts

  • large-language-model one pretrained model serves translation, summarization, sentiment analysis, code generation, conversation and domain knowledge retrieval, because all of them reduce to parsing and generating text

Pretraining, then fine-tuning

The general process of creating an LLM has two stages, and the distinction between them accounts for most of what follows in this module.

Figure 1.3. Pretraining on raw unlabeled text produces a foundation model with text completion and few-shot capabilities; a smaller labeled dataset then fine-tunes it for a specific task.

Stage 1: pretraining

The “pre” in pretraining refers to the initial phase, in which a model is trained on a large and diverse dataset to develop a broad understanding of language. The corpus is raw text: regular text without any labeling information. Filtering may be applied — removing formatting characters, or documents in unknown languages — but no annotation is added. The sources in figure 1.3 are internet texts, books, Wikipedia and research articles, on the order of trillions of words.

NoteWhere the labels come from

Traditional machine learning models and deep neural networks trained under the conventional supervised paradigm require labeling information. Pretraining does not. It uses self-supervised learning, in which the model generates its own labels from the input data: the next word in the sentence is the label for the position before it.

The output of stage 1 is an initial pretrained LLM, called a base model or foundation model. A typical example is the GPT-3 model, the precursor of the original model offered in ChatGPT. It is capable of text completion — finishing a half-written sentence provided by a user — and it has limited few-shot capabilities, which means it can learn to perform new tasks from only a few examples instead of needing extensive training data.

Stage 2: fine-tuning

Fine-tuning continues training the pretrained model on labeled data. The two most popular categories differ only in what the labels are.

  • Instruction fine-tuning. The labeled dataset consists of instruction and answer pairs — for example, a query to translate a text accompanied by the correctly translated text.

  • Classification fine-tuning. The labeled dataset consists of texts and associated class labels — for example, emails associated with “spam” and “not spam” labels.

Why build a custom model

Research has shown that custom-built LLMs, tailored for specific tasks or domains, can outperform general-purpose LLMs such as ChatGPT, which are designed for a wide array of applications. BloombergGPT, specialized for finance, and LLMs tailored for medical question answering are the source’s examples. Three further reasons are practical rather than about accuracy.

  • Data privacy. Companies may prefer not to share sensitive data with third-party LLM providers.
  • On-device deployment. Smaller custom LLMs can be deployed directly on customer devices such as laptops and smartphones, which decreases latency and reduces server-related costs.
  • Autonomy. Custom LLMs grant developers complete control over updates and modifications to the model.

Learning outcomes

  • explain-pretraining-finetuning Explain the two-stage lifecycle of pretraining a foundation model and fine-tuning it for a task.

Concepts

  • pretraining training on a large diverse corpus of raw unlabeled text by self-supervised learning, which yields a base or foundation model
  • fine-tuning continued training of the pretrained model on a smaller labeled dataset, either instruction and answer pairs or texts with class labels
  • large-language-model a custom model tailored to a domain can outperform a general-purpose one, and can be run locally for privacy, latency and cost

The transformer behind LLMs

Most modern LLMs rely on the transformer architecture, introduced in the 2017 paper “Attention Is All You Need”. It was developed for machine translation — translating English texts to German and French.

The architecture consists of two submodules.

  • The encoder processes the input text and encodes it into a series of numerical representations, or vectors, that capture the contextual information of the input.
  • The decoder takes these encoded vectors and generates the output text.

In a translation task the encoder encodes the text from the source language into vectors, and the decoder decodes these vectors to generate text in the target language. Both the encoder and the decoder consist of many layers connected by a self-attention mechanism, which allows the model to weigh the importance of different words or tokens in a sequence relative to each other. This is what enables the model to capture long-range dependencies and contextual relationships within the input data.

Figure 1.4. The original transformer for translation. The encoder reads “This is an example”; the decoder, given the partial translation “Das ist ein”, completes it to “Das ist ein Beispiel”.

Figure 1.5. The encoder submodule (left, BERT-like) fills in randomly masked words; the decoder submodule (right, GPT-like) receives incomplete text and learns to generate one word at a time.

The two branches

Later variants built on this concept to adapt the architecture for different tasks.

  • BERT — short for bidirectional encoder representations from transformers — is built upon the original transformer’s encoder submodule. It specializes in masked word prediction: the model predicts masked or hidden words in a given sentence. In figure 1.5 the input reads This is an __ of how concise I __ be, and the model’s task is to recover the original sentence. This training strategy equips BERT with strengths in text classification tasks, including sentiment prediction and document categorization. As an application, X (formerly Twitter) uses BERT to detect toxic content.

  • GPT — short for generative pretrained transformers — focuses on the decoder portion. It receives incomplete text, as in This is an example of how concise I can, and learns to generate one word at a time. It is designed for tasks that require generating texts: machine translation, text summarization, fiction writing, writing computer code, and more.

NoteTransformers and LLMs are not the same set

Today’s LLMs are based on the transformer architecture, so the two terms are often used synonymously. Neither inclusion holds. Not all transformers are LLMs, since transformers can also be used for computer vision — the vision transformers of the previous unit are the case in point. Not all LLMs are transformers either, as there are LLMs based on recurrent and convolutional architectures; the main motivation behind these alternative approaches is to improve the computational efficiency of LLMs. Whether they can compete with the capabilities of transformer-based LLMs remains to be seen.

The remainder of this module follows the GPT branch.

Learning outcomes

  • relate-gpt-to-transformer Relate the GPT architecture to the transformer of the previous unit as a decoder-only, autoregressive simplification.

Concepts

  • transformer-architecture the 2017 encoder-decoder design for machine translation, in which the encoder produces contextual vectors and the decoder generates the target text from them
  • self-attention-mechanism weighs the importance of tokens in a sequence relative to each other, which is how long-range dependencies within the input are captured
  • encoder-vs-decoder-llms BERT builds on the encoder submodule and is trained by masked word prediction for classification; GPT builds on the decoder submodule and generates text one word at a time

The data behind the scale

The training datasets for popular GPT- and BERT-like models are diverse and comprehensive text corpora encompassing billions of words, covering a vast array of topics and both natural and computer languages. Table 1.1 of the source gives the dataset used for pretraining GPT-3, which served as the base model for the first version of ChatGPT.

Dataset name Dataset description Number of tokens Proportion in training data
CommonCrawl (filtered) Web crawl data 410 billion 60%
WebText2 Web crawl data 19 billion 22%
Books1 Internet-based book corpus 12 billion 8%
Books2 Internet-based book corpus 55 billion 8%
Wikipedia High-quality text 3 billion 3%

A token is a unit of text that a model reads. The number of tokens in a dataset is roughly equivalent to the number of words and punctuation characters in the text.

Two features of the table repay attention, and both are easy to misread.

  • The proportions are proportions of the sampled data, not of the corpus. They sum to 100% of the sampled data, adjusted for rounding errors. WebText2 contributes 22% of the training data from 19 billion tokens, while Books2 contributes 8% from 55 billion; the column is a sampling weight, not a share of the token counts.

  • The subsets total 499 billion tokens, but the model was trained on only 300 billion. The authors of the GPT-3 paper did not specify why the model was not trained on all 499 billion tokens.

For a sense of the physical scale: CommonCrawl alone consists of 410 billion tokens and requires about 570 GB of storage. Later iterations of models, such as Meta’s LLaMA, have expanded their training scope to include additional data sources like Arxiv research papers (92 GB) and StackExchange’s code-related Q&As (78 GB).

The scale and diversity of this dataset are what allow these models to perform well on diverse tasks, including language syntax, semantics, and context — and even some requiring general knowledge.

The authors of the GPT-3 paper did not share the training dataset. A comparable publicly available one is Dolma: An Open Corpus of Three Trillion Tokens for LLM Pretraining Research by Soldaini et al. 2024. The collection may contain copyrighted works, and the exact usage terms may depend on the intended use case and country.

What this implies for the course

Pretraining an LLM requires access to significant resources and is very expensive. The GPT-3 pretraining cost is estimated at $4.6 million in terms of cloud computing credits. That figure puts reproduction beyond the reach of any course, and the source’s response is a division of labour.

  • The pretraining code is implemented and used to pretrain an LLM for educational purposes, with all computations executable on consumer hardware.
  • After the pretraining code is implemented, openly available model weights are reused and loaded into the same architecture, which skips the expensive pretraining stage when the model is fine-tuned.

Many pretrained LLMs are available as open source models and can be used as general-purpose tools to write, extract and edit texts that were not part of the training data. Fine-tuning them on specific tasks requires relatively smaller datasets, which reduces the computational resources needed.

Learning outcomes

  • characterize-pretraining-data Characterize the scale and composition of an LLM pretraining corpus, and the compute it implies.

Concepts

  • pretraining GPT-3 sampled 300 billion tokens from a 499-billion-token corpus dominated by filtered web crawl, at an estimated $4.6 million in cloud computing credits
  • large-language-model the scale and diversity of a billions-of-tokens corpus is what produces competence across syntax, semantics, context and general knowledge rather than in one domain

Inside GPT, and what emerges

GPT was originally introduced in the paper “Improving Language Understanding by Generative Pre-Training” by Radford et al. from OpenAI. GPT-3 is a scaled-up version of this model, with more parameters and trained on a larger dataset. The original model offered in ChatGPT was created by fine-tuning GPT-3 on a large instruction dataset using the method of OpenAI’s InstructGPT paper.

The objective and the architecture

The model is simply trained to predict the next word. Next-word prediction is a form of self-supervised learning, which is a form of self-labeling: labels for the training data are not collected explicitly, but the structure of the data itself is used — the next word in a sentence or document is the label the model is supposed to predict. Because this task allows labels to be created “on the fly”, it is possible to use massive unlabeled text datasets to train LLMs.

Compared with the original transformer architecture, the general GPT architecture is relatively simple. It is just the decoder part without the encoder. Since decoder-style models generate text by predicting text one word at a time, they are a type of autoregressive model: autoregressive models incorporate their previous outputs as inputs for future predictions. In GPT, each new word is chosen based on the sequence that precedes it, which improves the coherence of the resulting text.

Figure 1.8. Three iterations of generation. The input “This” produces “This is”, which becomes the input of iteration 2, and so on; the output of the previous round serves as input to the next round.

Architectures such as GPT-3 are also significantly larger than the original transformer model. The original transformer repeated the encoder and decoder blocks six times. GPT-3 has 96 transformer layers and 175 billion parameters in total.

GPT-3 was introduced in 2020, which by the standards of deep learning and large language model development is a long time ago. More recent architectures, such as Meta’s Llama models, are still based on the same underlying concepts, introducing only minor modifications.

Emergent behavior

The original transformer, consisting of encoder and decoder blocks, was explicitly designed for language translation. GPT models — despite their larger yet simpler decoder-only architecture aimed at next-word prediction — are also capable of performing translation tasks. This capability was initially unexpected to researchers, as it emerged from a model primarily trained on a next-word prediction task, which is a task that did not specifically target translation.

The ability to perform tasks that the model was not explicitly trained to perform is called an emergent behavior. This capability is not explicitly taught during training, but emerges as a natural consequence of the model’s exposure to vast quantities of multilingual data in diverse contexts.

Figure 1.6. Text completion, zero-shot and few-shot behaviour of a GPT-like model, distinguished only by what the input contains.

Figure 1.6 separates the three cases by the content of the input alone. No retraining, fine-tuning, or task-specific model architecture change accompanies the switch between them.

  • Text completion — the input Breakfast is the yields most important meal of the day.

    This is the pretraining objective, applied at inference time. The model creates plausible text given a partial input text.

  • Zero-shot — the input Translate English to German: breakfast => yields Frühstück. Zero-shot learning is the ability to generalize to completely unseen tasks without any prior specific examples.

    The task is stated in the prompt and no example of it is given. Nothing in the next-word objective asked for translation.

  • Few-shot — the input gaot => goat, sheo => shoe, pohne => yields phone. Few-shot learning involves learning from a minimal number of examples the user provides as input.

    The task is never named. The three lines are the specification, and the model infers both the operation and the output format from them.

Learning outcomes

  • relate-gpt-to-transformer Relate the GPT architecture to the transformer of the previous unit as a decoder-only, autoregressive simplification.
  • explain-emergent-abilities Explain zero-shot and few-shot in-context learning as behaviour that emerges from scale rather than from explicit training.

Concepts

  • encoder-vs-decoder-llms GPT is the decoder part of the original transformer without the encoder, processing text unidirectionally from left to right
  • next-word-prediction the label is the next word in the text, so labels are created on the fly and each predicted token is appended to the input for the following prediction
  • emergent-behavior translation and other untargeted capabilities arise from exposure to vast multilingual data rather than from any training objective that asked for them

The roadmap for this module

Building an LLM from scratch takes the fundamental idea behind GPT as a blueprint and proceeds in three stages.

Figure 1.9. The three stages of coding an LLM: implementing the architecture and data preparation, pretraining to obtain a foundation model, and fine-tuning that model into a classifier or a personal assistant.
  • Stage 1 — building an LLM. Three steps: data preparation and sampling, the attention mechanism, and the LLM architecture. This stage implements the data sampling and establishes the basic mechanism.

    2 Working with text data covers step 1, 3 Coding attention mechanisms covers step 2, and 4 Implementing a GPT model from scratch to generate text covers step 3.

  • Stage 2 — the foundation model. Pretraining is step 4, connecting stage 1 to the foundation model; the stage then comprises the training loop, model evaluation, and loading pretrained weights. It pretrains the LLM on unlabeled data to obtain a foundation model for further fine-tuning.

    5 Pretraining on unlabeled data covers this stage. Pretraining an LLM from scratch demands thousands to millions of dollars in computing costs for GPT-like models, so the emphasis falls on training for educational purposes using a small dataset, together with code for loading openly available model weights.

  • Stage 3 — fine-tuning. Two branches from the same foundation model. Step 8 fine-tunes it with a dataset with class labels to create a classifier; step 9 fine-tunes it with an instruction dataset to create a personal assistant or chat model.

    6 Fine-tuning for classification covers the classifier and 7 Fine-tuning to follow instructions covers the personal assistant. Following instructions such as answering queries, and classifying texts, are the most common tasks in applications and research.

Stage 2 also covers the fundamentals of evaluating LLMs. Coding an LLM from the ground up is the exercise that yields an understanding of its mechanics and limitations, and it supplies the knowledge required for pretraining or fine-tuning existing open source LLM architectures on a specific domain’s datasets or tasks.

Most LLMs today are implemented using the PyTorch deep learning library, which is what this module uses.

Learning outcomes

  • navigate-the-build-roadmap Navigate the three-stage roadmap for building an LLM from scratch and locate each unit of this module on it.
  • explain-pretraining-finetuning Explain the two-stage lifecycle of pretraining a foundation model and fine-tuning it for a task.

Concepts

  • transformer-architecture stage 1 implements the attention mechanism and the transformer LLM architecture in code
  • pretraining stage 2 implements the training loop for next-word prediction and produces the foundation model, or loads openly available weights in its place
  • fine-tuning stage 3 adapts the foundation model twice, into a text classifier and into an instruction-following assistant

Summary

  • LLMs have transformed the field of natural language processing, which previously mostly relied on explicit rule-based systems and simpler statistical methods.

    The change is the substitution of learned representations for handcrafted rules, which is what removed the one-model-per-task constraint.

  • Modern LLMs are trained in two main steps: pretrained on a large corpus of unlabeled text using the prediction of the next word in a sentence as a label, then fine-tuned on a smaller labeled target dataset to follow instructions or perform classification tasks.

    Stage 1 needs no annotation, which is what makes a corpus the size of the web usable at all. Stage 2 is where a labeled dataset is required, and it is small.

  • LLMs are based on the transformer architecture, whose key idea is an attention mechanism that gives the LLM selective access to the whole input sequence when generating the output one word at a time.

    Selective access to the whole sequence is the property that a fixed-width receptive field cannot provide.

  • The original transformer architecture consists of an encoder for parsing text and a decoder for generating text. LLMs for generating text and following instructions, such as GPT-3 and ChatGPT, only implement decoder modules.

    The encoder exists to represent a second sequence. Next-word prediction has only one sequence, so there is nothing for the encoder to read.

  • Large datasets consisting of billions of words are essential for pretraining LLMs.

    GPT-3 was trained on 300 billion tokens at an estimated $4.6 million in cloud computing credits, which is why openly available weights are loaded rather than reproduced.

  • While the general pretraining task for GPT-like models is to predict the next word in a sentence, these LLMs exhibit emergent properties, such as capabilities to classify, translate, or summarize texts.

    The objective is narrow and the resulting competence is not, which is the single most consequential observation in this unit.

  • Once an LLM is pretrained, the resulting foundation model can be fine-tuned more efficiently for various downstream tasks, and LLMs fine-tuned on custom datasets can outperform general LLMs on specific tasks.

    This is the asymmetry that the three-stage roadmap is organised around: one expensive stage, many cheap ones.

Stage 1 of that roadmap begins with the input pipeline. Text is discrete and a network optimizes over continuous vectors, so before the next word can be predicted the text has to become tokens: split by a rule, assigned integer IDs, extended with special tokens for unknown words and document boundaries, made robust with byte pair encoding, cut into input-target pairs by a sliding window, and mapped through a learnable embedding table with positional embeddings added. 2 Working with text data builds that pipeline, and it is the first code in this module.

References

  • Building a Large Language Model (from scratch), Sebastian Raschka, 2024, Manning Books — Link — Page 23-38