I Foundations

Keywords

ver. 1.0.0, i_foundations

Introduces core reinforcement learning concepts, formalism, common algorithm families, practical challenges, and how to formulate and train an RL agent for a control task.

This unit teaches the foundations of reinforcement learning: how learning from evaluative rewards differs from supervised labels, the agent–environment interaction loop and trajectories, the Markov decision process and policy formalism, and the collect-then-update training loop with its shifting data distribution. It names four central RL challenges (credit assignment, optimisation instability, reward design, exploration vs exploitation), classifies algorithms by the object they learn (policy, value, or model), compares algorithm families and selection criteria, and demonstrates a complete formulation and training pipeline on a rocket-landing control example.

The unit equips a learner to understand and formulate reinforcement learning problems and to choose and justify practical algorithmic approaches.

It explains how RL’s feedback is evaluative rather than prescriptive: rewards indicate how well the agent performed but not which actions were correct, creating the credit assignment problem over time. The agent–environment interaction loop is defined step-by-step (observe S_t, take action A_t, receive reward R_{t+1}, observe S_{t+1}), and sequences of such steps are described as trajectories. The Markov property is introduced, tasks are formalised as Markov Decision Processes (MDPs), and policies are written as π(A_t | S_t); strategies for handling non‑Markov observations are covered.

Training is presented as a two-level process: a fast interaction loop that generates data and a slower outer loop that turns experience into updated parameters. This collect-then-update paradigm explains why the agent’s own behaviour continually shifts the data distribution and why that matters for learning.

Four recurring practical difficulties are identified and diagnosed: credit assignment across delayed rewards, optimisation instability from nonstationary data and bootstrapped targets, designing reward functions that induce desired behaviour without unintended incentives, and the exploration–exploitation tradeoff. Understanding these challenges guides evaluation and debugging of RL projects.

Algorithms are classified by the object they learn: direct policy methods that parameterise and optimise π, value-based methods that estimate V(s) or Q(s,a) and derive policies from those estimates, and model-based methods that learn transition and reward dynamics and plan within a learned model. The unit explains how these objects are combined in hybrids such as actor–critic methods and gives practical criteria to choose an algorithm family based on action space, sample efficiency needs and stability concerns.

A full worked example demonstrates the end-to-end process: formulating a control task (starship/rocket landing), designing rewards, using curriculum learning and demonstration bootstrapping, and training with a practical algorithm (PPO) while observing how the four core challenges manifest and are mitigated.

By the end of the unit, a learner can write the agent–environment loop and describe trajectories, test and formalise tasks as MDPs with policies, articulate why rewards differ from labels and how credit assignment arises, explain the collect‑then‑update loop and distribution shift, diagnose the four core RL challenges in a task, classify algorithms by what they learn and choose an appropriate family for a given problem, and design a complete RL formulation (reward, curriculum, demonstrations) for a concrete control task. The unit also prepares the learner to proceed to deeper algorithmic details in the next unit.

Materials

Source document

  • The Little Book of Reinforcement Learning, Alexandre Torres Leguet, 2026 — Link — Page 10-39