The foundation of the paper's experiments: two agents choose Cooperate (C) or Defect (D) simultaneously. Mutual cooperation yields higher total reward, but defection dominates individually.
In standard RL, the environment is fixed. In multi-agent settings, other agents' policies change during training, creating a moving target.
Self-interested optimization leads to mutual defection (Nash equilibrium), leaving both agents worse off than mutual cooperation (Pareto optimum).
From Meulemans et al. (2025): adaptive agents become exploitable, mutual exploitation creates cooperative pressure, heterogeneous training unlocks the full cycle.
Press & Dyson (2012): ZD strategies unilaterally fix linear payoff relationships. Extortion strategies exploit adaptive learners by channeling their gradient updates.
When both agents run ZD strategies, neither sustains unilateral control. The equilibrium converges toward mutual cooperation.
Sequence models trained on diverse episodes learn to run learning algorithms in their forward pass, adapting within episodes without weight updates.
Training against mixed populations—naive learners, learning-aware agents, adversarial policies—is critical for emergent cooperation.
Within-episode trajectory showing how in-context best-response converges toward optimal strategy against the inferred co-player type.
Trained agents achieve stable cooperation in iterated Prisoner's Dilemma, comparable to LOLA-class methods but without explicit opponent modeling.
The sequence model's ability to classify opponent type from early-episode behavior, enabling strategic adaptation.
Performance across different opponent configurations demonstrates robust in-context best-response.
Evolution of learning-aware agent architectures, showing progressive removal of rigid assumptions.
Quantifying architectural constraints: what each method assumes about opponent learning dynamics.
Current evidence versus claimed applicability. The gap between IPD demonstrations and complex multi-agent settings.