AI Post Transformers · Companion visuals

Recursive Self-Improvement and Alignment: OpenAI's "An Alien Mind"

Single-author essay · OpenAI research leader (byline not in text) No methods · no tables · no error bars Source essay ↗ All numbers on charts are illustrative mock data

References

  1. An Alien Mind (OpenAI essay)
  2. Gabriel 2020, Artificial Intelligence, Values, and Alignment
  3. Ouyang et al. 2022, InstructGPT
  4. Wallace et al. 2024, The Instruction Hierarchy
  5. Betley et al. 2025, Emergent Misalignment
  6. Korbak et al. 2025, Chain of Thought Monitorability
  7. Baker et al. 2025, Monitoring Reasoning Models for Misbehavior
  8. Christiano 2018, Clarifying AI alignment
  9. Christiano et al. 2017, Deep RL from Human Preferences
  10. Greenblatt et al. 2024, Alignment Faking in LLMs
  11. Anthropic 2026, The Persona Selection Model
  12. Langosco et al. 2022, Goal Misgeneralization in Deep RL
  13. Hubinger et al. 2019, Risks from Learned Optimization
  14. Burns et al. 2023, Weak-to-Strong Generalization
  15. OpenAI 2025-26, Training Language Models to Self-Report / Confessions
  16. Also discussed: Kaplan et al. 2020, Scaling Laws for Neural Language Models
  17. Also discussed: Hoffmann et al. 2022, Training Compute-Optimal LLMs