A visual map of how reinforcement learning organizes itself in 2025: value functions, direct policy optimization, actor-critic compromises, world models, offline constraints, multi-agent games, and LLM-era alignment loops. The page emphasizes structure, tensions, and practical tradeoffs more than prose.
The survey’s chapter structure acts like a gravity map. Larger nodes mark families that the discipline treats as foundational rather than peripheral.
Mock scores summarize where each branch sits along theory anchor, deployment fit, planning leverage, and dependence on fresh interaction.
Bellman-style estimates first, behavior second.
Behavior is optimized directly.
Actor-critic became the practical equilibrium.
Switch the optimization family. The same environment loop changes character depending on whether value estimates, policy gradients, or actor-critic structure dominate the update.
Mock placement of classic families. The survey’s organizing logic makes sense only if you remember that elegant objectives and survivable training are not the same thing.
Offline RL lives or dies on what the logged dataset actually covers. Bright cells are well-supported state-action pairs. Dim cells are where optimistic value estimates become dangerous.
Behavior cloning is safest but capped by the log. Conservative critics trade ambition for caution. Sequence models gain flexibility but still inherit dataset limits.
Datasets can be expert, medium, replay-like, or mixed.
Lower values on unsupported actions before they become fantasy plans.
Decision Transformer made offline RL legible to the transformer era.
Toggle between the two expansion directions the transcript highlights: RL for LLMs, and LLMs or world models used inside RL pipelines.
Which branch leans on planning, human preference data, multi-step simulation, and fixed logged trajectories?
Search over futures instead of reacting one step at a time.
Preference data changed RL’s public face.
Reusable predictive assets matter more than a neat loop diagram suggests.
Selected papers and related episode references used for this visual companion.