Interactive visualization companion arXiv: 2412.05265 viz permalink extra arXiv IDs in transcript: none detected

Reinforcement Learning in 2025: An Overview

A visual map of how reinforcement learning organizes itself in 2025: value functions, direct policy optimization, actor-critic compromises, world models, offline constraints, multi-agent games, and LLM-era alignment loops. The page emphasizes structure, tensions, and practical tradeoffs more than prose.

Survey spine
6 branches
Core tension
theory vs stability
Data regime split
online / offline
Modern center
actor-critic
Episode lens

The transcript treats the survey as an atlas that also makes an argument: what gets a full chapter reads as central, while offline RL and data realism feel comparatively sidelined.

Visual guide

Use the tabs to move from taxonomy to training mechanics, then to offline support mismatch, and finally to world models and LLM-related RL.

Key primitives
MDP
States, actions, transitions, rewards.
Policy
Maps state to action distribution.
Value
Expected future return from here.

Field Map

The survey’s chapter structure acts like a gravity map. Larger nodes mark families that the discipline treats as foundational rather than peripheral.

hover nodes
direct policy optimization
value estimation
planning / models
data / alignment expansions

2025 Taxonomy Heatmap

Mock scores summarize where each branch sits along theory anchor, deployment fit, planning leverage, and dependence on fresh interaction.

6×6 matrix
Blue cells are low weight, orange is rising importance, red is structural pressure.

Value-Based

Bellman-style estimates first, behavior second.

Canonical ladder
Value iteration → TD learning → SARSA → Q-learning → deep Q nets.
Strength
Clear control logic once action values are reliable.
Failure mode
Approximation error and exploding Q-values under nonlinear training.

Policy-Based

Behavior is optimized directly.

Canonical ladder
REINFORCE → trust regions → PPO → off-policy actor variants.
Strength
Natural fit for continuous actions and stochastic policies.
Failure mode
High gradient variance and brittle update geometry.

Modern compromise

Actor-critic became the practical equilibrium.

Actor
Directly changes the policy.
Critic
Reduces variance with a learned evaluator.
Why central
Cleaner than pure value methods in action spaces, steadier than pure policy gradients in practice.

Credit Assignment Loop

Switch the optimization family. The same environment loop changes character depending on whether value estimates, policy gradients, or actor-critic structure dominate the update.

The redder the reward trace, the harder the delayed-credit problem becomes for that step.

Stability vs Sample Efficiency

Mock placement of classic families. The survey’s organizing logic makes sense only if you remember that elegant objectives and survivable training are not the same thing.

bubble chart
value-heavy
policy-heavy
model-based / planning

Support Mismatch Heatmap

Offline RL lives or dies on what the logged dataset actually covers. Bright cells are well-supported state-action pairs. Dim cells are where optimistic value estimates become dangerous.

Static-Data Method Tradeoffs

Behavior cloning is safest but capped by the log. Conservative critics trade ambition for caution. Sequence models gain flexibility but still inherit dataset limits.

normalized return
OOD risk

D4RL intuition

Datasets can be expert, medium, replay-like, or mixed.

Narrow expert log
High reward but poor coverage for recovery behavior.
Broad replay log
Messier trajectories, better support for counterfactuals.

CQL intuition

Lower values on unsupported actions before they become fantasy plans.

Goal
Bias the critic toward caution outside the behavior distribution.
Cost
Some conservatism remains even when a better action exists.

Sequence modeling bridge

Decision Transformer made offline RL legible to the transformer era.

View
Trajectory as conditional sequence generation, not only Bellman backup.
Limit
The model still cannot invent reliable support from unseen behavior.

World Models and LLM Loops

Toggle between the two expansion directions the transcript highlights: RL for LLMs, and LLMs or world models used inside RL pipelines.

Expansion Matrix

Which branch leans on planning, human preference data, multi-step simulation, and fixed logged trajectories?

capability grid

Model-based RL

Search over futures instead of reacting one step at a time.

Decision-time planning
Use the model while acting, AlphaGo or MuZero style.
Background planning
Improve the policy offline with imagined rollouts.

LLM post-training

Preference data changed RL’s public face.

RLHF pipeline
Samples → comparisons → reward signal → policy update.
DPO tension
Simplifies the stack, but implicit reward generalization can be brittle under shift.

World-model pretraining

Reusable predictive assets matter more than a neat loop diagram suggests.

Trend
Video and robot data pretrain latent dynamics before task-specific control.
Tension
Imagined rollouts look coherent before reality exposes model bias.

References

Selected papers and related episode references used for this visual companion.