AI Post Transformers · Companion visuals
Recursive Self-Improvement and Alignment: OpenAI's "An Alien Mind"
Single-author essay · OpenAI research leader (byline not in text)
No methods · no tables · no error bars
Source essay ↗
All numbers on charts are illustrative mock data
References
An Alien Mind
(OpenAI essay)
Gabriel 2020,
Artificial Intelligence, Values, and Alignment
Ouyang et al. 2022,
InstructGPT
Wallace et al. 2024,
The Instruction Hierarchy
Betley et al. 2025,
Emergent Misalignment
Korbak et al. 2025,
Chain of Thought Monitorability
Baker et al. 2025,
Monitoring Reasoning Models for Misbehavior
Christiano 2018,
Clarifying AI alignment
Christiano et al. 2017,
Deep RL from Human Preferences
Greenblatt et al. 2024,
Alignment Faking in LLMs
Anthropic 2026,
The Persona Selection Model
Langosco et al. 2022,
Goal Misgeneralization in Deep RL
Hubinger et al. 2019,
Risks from Learned Optimization
Burns et al. 2023,
Weak-to-Strong Generalization
OpenAI 2025-26,
Training Language Models to Self-Report / Confessions
Also discussed: Kaplan et al. 2020,
Scaling Laws for Neural Language Models
Also discussed: Hoffmann et al. 2022,
Training Compute-Optimal LLMs