III RL at scale

Keywords

ver. 1.0.0, iii_rl_at_scale

How modern RL is applied to large language models and games, trading critics for sampled baselines, engineering inference-trainer systems, and reusing search via self-play.

This unit explains how post-training regimes for large models work in practice: supervised fine-tuning (SFT), preference-based RL (RLHF) and verifier-based RL (RLVR); how token generation maps to an MDP; how GRPO removes a learned value by using group-sampled baselines; practical fixes (entropy targeting, length control, dynamic sampling); the distributed trainer/inference architecture and importance-sampling corrections; the relationship and hybrids between SFT and RL; and how policy iteration with Monte‑Carlo Tree Search and self-play produces AlphaZero-style learning and why zero-sum self-play converges. By the end you’ll be able to reason about credit assignment, bias–variance trade-offs, moving distributions and stale data across these algorithms.

This unit develops practical mastery of reinforcement-style post-training on large models and the analogous loop used in game-playing agents. You will: distinguish three post-training paradigms for foundation models — imitation (SFT), preference optimisation (RLHF), and verifier-driven optimisation (RLVR) — and explain what objective each one optimises and where each approach breaks down. You will formalise autoregressive token generation as an MDP (states, actions, transitions and extremely sparse rewards) and use that formulation to reason about RLVR runs.

You will learn a production-capable variant of policy optimisation (GRPO) that removes the learned critic by using a group of sampled completions as a baseline while retaining clipped probability ratios. You will evaluate the practical refinements required to make GRPO work at LLM scale: entropy targeting to maintain exploration and calibration, length penalties and sequence-level objectives to shape generation, dynamic sampling strategies that trade variance and compute, and the sequence-versus-token trade-offs in credit assignment.

The unit covers the systems side required for correctness: the separation of trainer and inference fleets, asynchronous rollout collection, and the importance-sampling corrections needed to compensate for distributional drift between collectors and trainers. You will be able to explain how SFT and RL objectives relate (SFT as a limiting case of certain RL losses) and describe hybrid methods that interpolate between imitation and preference optimisation.

Switching domain, you will see policy iteration reappear in AlphaZero: MCTS acts as the improvement operator and self-play rollouts provide evaluation; a single network is trained to reproduce both the search policy and the resulting value. You will understand why two-player zero-sum self-play converges to a minimax equilibrium rather than collapsing, and what properties a task must satisfy for test‑time search to transfer usefully (e.g., a strong forward model, tractable branching, and alignment between search-time objective and task reward).

Finally, you will be able to identify recurring design patterns across algorithms: credit assignment difficulties, bias–variance trade-offs, problems caused by moving data distributions and stale experiences, and the engineering choices that tip correctness and performance. Armed with these concepts, you can evaluate and extend RL-style post-training pipelines for language models and assess when search-and-self-play techniques are applicable to new domains.

Materials

Source document

  • The Little Book of Reinforcement Learning, Alexandre Torres Leguet, 2026 — Link — Page 112-154