Emergent Cooperation in Self-Interested Multi-Agent AI

Weis, M., Wołczyk, M., et al. | Google Paradigms of Intelligence Team | February 2026

arXiv:2602.16301

Prisoner's Dilemma Payoff Matrix

The foundation of the paper's experiments: two agents choose Cooperate (C) or Defect (D) simultaneously. Mutual cooperation yields higher total reward, but defection dominates individually.

Payoff encoding:
Low → High

Multi-Agent RL Challenge: Non-Stationarity

In standard RL, the environment is fixed. In multi-agent settings, other agents' policies change during training, creating a moving target.

Nash Equilibrium vs. Pareto Optimum

Self-interested optimization leads to mutual defection (Nash equilibrium), leaving both agents worse off than mutual cooperation (Pareto optimum).

Three-Step Cooperative Mechanism

From Meulemans et al. (2025): adaptive agents become exploitable, mutual exploitation creates cooperative pressure, heterogeneous training unlocks the full cycle.

Zero-Determinant Strategy Exploitation

Press & Dyson (2012): ZD strategies unilaterally fix linear payoff relationships. Extortion strategies exploit adaptive learners by channeling their gradient updates.

Extortion gradient:
Weak → Strong pressure

Mutual Extortion Resolution

When both agents run ZD strategies, neither sustains unilateral control. The equilibrium converges toward mutual cooperation.

In-Context Learning as Algorithm Distillation

Sequence models trained on diverse episodes learn to run learning algorithms in their forward pass, adapting within episodes without weight updates.

Co-Player Diversity Requirement

Training against mixed populations—naive learners, learning-aware agents, adversarial policies—is critical for emergent cooperation.

Cooperation emergence:
Failed → Robust

Episode-Level Adaptation

Within-episode trajectory showing how in-context best-response converges toward optimal strategy against the inferred co-player type.

Cooperation Rate Over Training

Trained agents achieve stable cooperation in iterated Prisoner's Dilemma, comparable to LOLA-class methods but without explicit opponent modeling.

Co-Player Type Inference Accuracy

The sequence model's ability to classify opponent type from early-episode behavior, enabling strategic adaptation.

Reward Distribution by Co-Player Pairing

Performance across different opponent configurations demonstrates robust in-context best-response.

Average reward:
1.0 → 5.0

Architectural Comparison: LOLA vs. M-FOS vs. This Work

Evolution of learning-aware agent architectures, showing progressive removal of rigid assumptions.

Assumption Footprint Analysis

Quantifying architectural constraints: what each method assumes about opponent learning dynamics.

Generalization Scope

Current evidence versus claimed applicability. The gap between IPD demonstrations and complex multi-agent settings.

Key References

Primary Paper: Weis, M., Wołczyk, M., et al. (2026). Multi-agent cooperation through in-context co-player inference. arXiv:2602.16301
Foerster, J., Chen, R. Y., Al-Shedivat, M., Whiteson, S., Abbeel, P., Mordatch, I. (2018). Learning with Opponent-Learning Awareness. LOLA Paper
Press, W. H., Dyson, F. J. (2012). Iterated Prisoner's Dilemma contains strategies that dominate any evolutionary opponent. PNAS
Laskin, M., Wang, L., Oh, J., et al. (2023). In-context reinforcement learning with algorithm distillation. Algorithm Distillation
Meulemans, A., et al. (2025). From naive to learning-aware: Emergence of cooperative behaviors in multi-agent systems. Google Research
Lu, C., Willi, T., de Waard, C. S., Foerster, J. (2022). Model-Free Opponent Shaping. M-FOS Paper
Brown, T. B., Mann, B., et al. (2020). Language Models are Few-Shot Learners. GPT-3 Paper
Axelrod, R., Hamilton, W. D. (1981). The Evolution of Cooperation. Classic Paper