II Diving deeper

Keywords

ver. 1.0.0, ii_diving_deeper

From Bellman backups to PPO: value methods, temporal-difference learning, function approximation, policy gradients and modern stabilisation techniques.

This unit teaches practical reinforcement learning foundations: value functions and Bellman equations; dynamic programming, Monte Carlo and temporal-difference methods; on-policy (SARSA) and off-policy (Q‑learning) control; multi-step returns and TD(λ); scaling Q‑learning with neural networks, replay buffers and target networks; direct policy optimisation via policy gradients including variance reduction (reward‑to‑go, baselines, GAE); and trust‑region style stabilisation culminating in PPO’s clipped objective. By the end, a learner can derive and implement core update rules, choose tradeoffs between bias and variance, and apply stabilisation techniques for deep RL.

This unit turns RL theory into a toolbox of algorithms and design principles a practitioner can apply and reason about.

It starts by formalising state-value and action-value functions and their Bellman equations, and uses a backup-depth vs width perspective to organise methods. Learners gain hands-on capability with exact dynamic programming when a model is available: they can perform policy evaluation and improvement, write Bellman operators, and explain contraction-based convergence.

Dropping the model, learners implement Monte Carlo estimation from full sampled episodes and convert those estimates into control with ε‑greedy policies, understanding the unbiased-but-high-variance nature of MC. Moving toward online updates, they derive and apply single-step bootstrapping updates: SARSA for on‑policy TD control and Q‑learning by replacing the sampled next action with a max to obtain off‑policy control. They can predict behavioral differences between on‑ and off‑policy learners (e.g., safe vs edge-following exploration) and understand implications for data collection and reuse.

The unit generalises temporality by introducing n‑step returns and TD(λ), showing how to trade bias and variance along a continuous depth axis. Learners learn to choose and implement intermediate returns and combine them into eligibility-trace style updates.

To scale beyond tabular methods, the course replaces Q‑tables with neural networks and teaches the two stabilising engineering fixes that make deep Q‑learning work in practice: experience replay and a slowly updated target network. Students will be able to implement and justify these mechanisms and recognise why they prevent divergence.

Switching to policy optimisation, learners derive the policy gradient via the score-function trick and implement REINFORCE. They progressively reduce variance using reward‑to‑go, learned baselines, and Generalized Advantage Estimation (GAE). Finally, the performance difference lemma is used to motivate a surrogate objective that remains valid across multiple update steps; learners implement PPO’s clipped surrogate to enforce a trust‑region style constraint in a simple, practical way.

By the end of the unit a learner can: - Define V^π and Q^π, write their Bellman equations, and place algorithms on the depth/width backup spectrum. - Run policy evaluation and policy improvement with a known model and explain operator contraction and convergence. - Estimate values from sampled episodes and turn them into ε‑greedy control. - Write and apply SARSA and Q‑learning single‑transition updates and predict behavioural differences between on‑ and off‑policy methods. - Trade bias and variance with n‑step returns, TD(λ), and GAE. - Replace tabular Q with neural networks and stabilise training using replay buffers and target networks. - Derive the policy gradient, reduce its variance with baselines and reward‑to‑go, and implement GAE. - Justify and implement a trust‑region style surrogate, and use PPO’s clipped objective for stable policy optimisation.

These capabilities prepare learners to move on to large‑scale RL applications — applying PPO and related methods to complex environments such as language models or game‑playing agents.

Materials

Source document

  • The Little Book of Reinforcement Learning, Alexandre Torres Leguet, 2026 — Link — Page 40-111