← All episodes Technical AGI Safety and Security Framework

Technical AGI Safety and Security Framework

Jun 5, 2026
This episode explores DeepMind’s paper on technical AGI safety and security, focusing on how labs might prevent severe, humanity-scale harm before highly capable systems are deployed. It breaks down the paper’s core distinctions between misuse and misalignment, explains what the authors mean by Exceptional AGI and the no-human-ceiling assumption, and examines dangerous capability evaluations in areas like cyber, biology, persuasion, and self-proliferation. The discussion highlights the paper’s main argument that safety measures such as refusal training, jailbreak hardening, access controls, monitoring, anomaly detection, and model-weight security only matter if they are explicitly tied to capability thresholds that trigger real deployment restrictions. Listeners would find it interesting because it turns abstract AGI risk debates into a concrete governance and engineering framework for deciding when a model is too dangerous to release under normal conditions.
Sources:
1. Technical AGI Safety and Security Framework
https://arxiv.org/pdf/2504.01849
2. The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation — Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, et al., 2018
https://scholar.google.com/scholar?q=The+Malicious+Use+of+Artificial+Intelligence%3A+Forecasting%2C+Prevention%2C+and+Mitigation
3. Model evaluation for extreme risks — Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, et al., 2023
https://scholar.google.com/scholar?q=Model+evaluation+for+extreme+risks
4. Frontier AI Regulation: Managing Emerging Risks to Public Safety — Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, et al., 2023
https://scholar.google.com/scholar?q=Frontier+AI+Regulation%3A+Managing+Emerging+Risks+to+Public+Safety
5. Evaluating Frontier Models for Dangerous Capabilities — Mary Phuong, Matthew Aitchison, Elliot Catt, Victoria Krakovna, et al., 2024
https://scholar.google.com/scholar?q=Evaluating+Frontier+Models+for+Dangerous+Capabilities
6. Guidance on the Assurance of Machine Learning in Autonomous Systems (AMLAS) — Richard Hawkins, Colin Paterson, Chiara Picardi, Ibrahim Habli, et al., 2021
https://scholar.google.com/scholar?q=Guidance+on+the+Assurance+of+Machine+Learning+in+Autonomous+Systems+%28AMLAS%29
7. Safety Cases: How to Justify the Safety of Advanced AI Systems — Joshua Clymer, Nick Gabrieli, David Krueger, Thomas Larsen, 2024
https://scholar.google.com/scholar?q=Safety+Cases%3A+How+to+Justify+the+Safety+of+Advanced+AI+Systems
8. Safety case template for frontier AI: A cyber inability argument — Arthur Goemans, Marie Davidsen Buhl, Jonas Schuett, Geoffrey Irving, et al., 2024
https://scholar.google.com/scholar?q=Safety+case+template+for+frontier+AI%3A+A+cyber+inability+argument
9. The BIG Argument for AI Safety Cases — Ibrahim Habli, Richard Hawkins, Colin Paterson, Mark Sujan, et al., 2025
https://scholar.google.com/scholar?q=The+BIG+Argument+for+AI+Safety+Cases
10. Safety cases: Justifying the safety of advanced AI systems — J. Clymer, N. Gabrieli, D. Krueger, and T. Larsen, 2024
https://scholar.google.com/scholar?q=Safety+cases%3A+Justifying+the+safety+of+advanced+AI+systems
11. AI control: Improving safety despite intentional subversion — R. Greenblatt, B. Shlegeris, K. Sachan, and F. Roger, 2024
https://scholar.google.com/scholar?q=AI+control%3A+Improving+safety+despite+intentional+subversion
12. Alignment faking in large language models — R. Greenblatt, C. Denison, B. Wright, F. Roger, M. MacDiarmid, S. Marks, J. Treutlein, T. Belonax, J. Chen, D. Duvenaud, et al., 2024
https://scholar.google.com/scholar?q=Alignment+faking+in+large+language+models
13. Towards evaluations-based safety cases for AI scheming — M. Balesni, M. Hobbhahn, D. Lindner, A. Meinke, T. Korbak, J. Clymer, B. Shlegeris, J. Scheurer, C. Stix, R. Shah, et al., 2024
https://scholar.google.com/scholar?q=Towards+evaluations-based+safety+cases+for+AI+scheming
14. Generative AI misuse: A taxonomy of tactics and insights from real-world data — N. Marchal, R. Xu, R. Elasmar, I. Gabriel, B. Goldberg, and W. Isaac, 2024
https://scholar.google.com/scholar?q=Generative+AI+misuse%3A+A+taxonomy+of+tactics+and+insights+from+real-world+data
15. Stress-Testing Capability Elicitation With Password-Locked Models — Ryan Greenblatt et al., 2024
https://scholar.google.com/scholar?q=Stress-Testing+Capability+Elicitation+With+Password-Locked+Models
16. Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models — Cameron Tice et al., 2024
https://scholar.google.com/scholar?q=Noise+Injection+Reveals+Hidden+Capabilities+of+Sandbagging+Language+Models
17. Benchmarking Misuse Mitigation Against Covert Adversaries — Davis Brown et al., 2025
https://scholar.google.com/scholar?q=Benchmarking+Misuse+Mitigation+Against+Covert+Adversaries
18. On scalable oversight with weak LLMs judging strong LLMs — Zachary Kenton et al., 2024
https://scholar.google.com/scholar?q=On+scalable+oversight+with+weak+LLMs+judging+strong+LLMs
19. Scaling Laws For Scalable Oversight — Joshua Engels et al., 2025
https://scholar.google.com/scholar?q=Scaling+Laws+For+Scalable+Oversight
20. Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs — Kyle O'Brien et al., 2025
https://scholar.google.com/scholar?q=Deep+Ignorance%3A+Filtering+Pretraining+Data+Builds+Tamper-Resistant+Safeguards+into+Open-Weight+LLMs
21. AI Post Transformers: When LLM Judges Become Coin Flips — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-when-llm-judges-become-coin-flips-8b43ef.mp3
22. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3
23. AI Post Transformers: Split Personality Training Reveals Latent Knowledge — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-08-split-personality-training-reveals-laten-c84616.mp3
24. AI Post Transformers: Trajectory Summaries for Long-Horizon Coding Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-24-trajectory-summaries-for-long-horizon-co-0194be.mp3