This episode examines the nabla-Reasoner paper (ICLR 2026), which proposes running gradient descent on token logits during inference — a first-order approach to test-time compute scaling that stands apart from every existing method in the field. The hosts contextualize the work against the established zeroth-order inference-time scaling landscape: Chain-of-Thought, Self-Consistency, Tree of Thoughts, and MCTS-based methods, all of which probe the reward landscape by sampling without directional information. The core argument is that zeroth-order methods hit a hard ceiling on long-horizon reasoning tasks because the search space grows exponentially while reward signals remain sparse, making random sampling increasingly futile. nabla-Reasoner sidesteps this by treating token logit vectors — normally ephemeral intermediate computations — as continuous optimization variables, computing reward gradients with respect to them and nudging the distribution toward higher-reward outputs before committing to each token. Listeners interested in the mechanics of inference-time scaling and the theoretical limits of sampling-based reasoning will find this a technically dense, well-grounded discussion of a genuinely novel approach.

Sources:
1. $\nabla$-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space — Peihao Wang, Ruisi Cai, Zhen Wang, Hongyuan Mei, Qiang Liu, Pan Li, Zhangyang Wang, 2026
http://arxiv.org/abs/2603.04948v1
2. ∇-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space — Peihao Wang, Ruisi Cai, Zhen Wang, Hongyuan Mei, Qiang Liu, Pan Li, Zhangyang Wang, 2026
https://scholar.google.com/scholar?q=%E2%88%87-Reasoner%3A+LLM+Reasoning+via+Test-Time+Gradient+Descent+in+Latent+Space
3. Diffusion-LM Improves Controllable Text Generation — Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, Tatsunori Hashimoto, 2022
https://scholar.google.com/scholar?q=Diffusion-LM+Improves+Controllable+Text+Generation
4. GFlowNet-Guided LLM Decoding: Towards Diverse and Accurate Reasoning — Jianing Li et al., 2024
https://scholar.google.com/scholar?q=GFlowNet-Guided+LLM+Decoding%3A+Towards+Diverse+and+Accurate+Reasoning
5. Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMs — Ahmadianshalchi et al., 2024
https://scholar.google.com/scholar?q=Back+to+Basics%3A+Revisiting+REINFORCE-Style+Optimization+for+Learning+from+Human+Feedback+in+LLMs
6. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou, 2022
https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models
7. Tree of Thoughts: Deliberate Problem Solving with Large Language Models — Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, Karthik Narasimhan, 2023
https://scholar.google.com/scholar?q=Tree+of+Thoughts%3A+Deliberate+Problem+Solving+with+Large+Language+Models
8. Let's Verify Step by Step — Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe, 2023
https://scholar.google.com/scholar?q=Let%27s+Verify+Step+by+Step
9. Scaling LLM Test-Time Compute Optimally — Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar, 2024
https://scholar.google.com/scholar?q=Scaling+LLM+Test-Time+Compute+Optimally
10. ARGS: Alignment as Reward-Guided Search — Maxim Khanov, Jirayu Burapacheep, Yixuan Li, 2024
https://scholar.google.com/scholar?q=ARGS%3A+Alignment+as+Reward-Guided+Search
11. Controlled Decoding from Language Models — Sidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li, Tao Wang, Yanpin Huang, Zhifeng Chen, Heng-Tze Cheng, Michael Collins, Trevor Strohman, Jilin Chen, Alex Beutel, Ahmad Beirami, 2024
https://scholar.google.com/scholar?q=Controlled+Decoding+from+Language+Models
12. AlphaCode 2 Technical Report — Google DeepMind AlphaCode Team, 2023
https://scholar.google.com/scholar?q=AlphaCode+2+Technical+Report
13. Training language models to follow instructions with human feedback — Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, Ryan Lowe, 2022
https://scholar.google.com/scholar?q=Training+language+models+to+follow+instructions+with+human+feedback
14. Direct Preference Optimization: Your Language Model is Secretly a Reward Model — Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, Chelsea Finn, 2023
https://scholar.google.com/scholar?q=Direct+Preference+Optimization%3A+Your+Language+Model+is+Secretly+a+Reward+Model
15. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI (Daya Guo, Dejian Yang, Haowei Zhang, et al.), 2025
https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning
16. Learning to summarize from human feedback — Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, Paul Christiano, 2020
https://scholar.google.com/scholar?q=Learning+to+summarize+from+human+feedback
17. Plug and Play Language Models: A Simple Approach to Controlled Text Generation — Dathathri et al., 2020
https://scholar.google.com/scholar?q=Plug+and+Play+Language+Models%3A+A+Simple+Approach+to+Controlled+Text+Generation
18. FUDGE: Controlled Text Generation with Future Discriminators — Yang and Klein, 2021
https://scholar.google.com/scholar?q=FUDGE%3A+Controlled+Text+Generation+with+Future+Discriminators
19. Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Model Parameters — Snell et al., 2024
https://scholar.google.com/scholar?q=Scaling+LLM+Test-Time+Compute+Optimally+Can+be+More+Effective+than+Scaling+Model+Parameters
20. Alignment as Reward-Guided Search — Khanov et al., 2024
https://scholar.google.com/scholar?q=Alignment+as+Reward-Guided+Search
21. Soft Prompts: The Power of Scale for Parameter-Efficient Prompt Tuning — Lester et al., 2021
https://scholar.google.com/scholar?q=Soft+Prompts%3A+The+Power+of+Scale+for+Parameter-Efficient+Prompt+Tuning
22. The Generalization Gap in Offline Reinforcement Learning — Levine et al., 2020
https://scholar.google.com/scholar?q=The+Generalization+Gap+in+Offline+Reinforcement+Learning
23. Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization — Unknown (2025), 2025
https://scholar.google.com/scholar?q=Thinking+on+the+Fly%3A+Test-Time+Reasoning+Enhancement+via+Latent+Thought+Policy+Optimization
24. Logit arithmetic elicits long reasoning capabilities without training — Unknown (2025), 2025
https://scholar.google.com/scholar?q=Logit+arithmetic+elicits+long+reasoning+capabilities+without+training
25. Reinforcement Learning in Inference Time: A Perspective from Successive Policy Iterations — Unknown (2025), 2025
https://scholar.google.com/scholar?q=Reinforcement+Learning+in+Inference+Time%3A+A+Perspective+from+Successive+Policy+Iterations
26. GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning — Yang et al. (2024/2025), 2025
https://scholar.google.com/scholar?q=GenPRM%3A+Scaling+Test-Time+Compute+of+Process+Reward+Models+via+Generative+Reasoning
27. Process Reward Models That Think — Unknown (2025), 2025
https://scholar.google.com/scholar?q=Process+Reward+Models+That+Think
28. Efficient Adaptive Rejection Sampling for Accelerating Speculative Decoding in Large Language Models — Unknown (2025), 2025
https://scholar.google.com/scholar?q=Efficient+Adaptive+Rejection+Sampling+for+Accelerating+Speculative+Decoding+in+Large+Language+Models
29. Inference-time alignment control for diffusion models with reinforcement learning guidance — Unknown (2025), 2025
https://scholar.google.com/scholar?q=Inference-time+alignment+control+for+diffusion+models+with+reinforcement+learning+guidance
30. AI Post Transformers: Test-Time Reinforcement Learning for LLMs — Hal Turing & Dr. Ada Shannon, Wed,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Test-Time-Reinforcement-Learning-for-LLMs-e398hsk
31. AI Post Transformers: MetaScale: Test-Time Scaling with Evolving Meta-Thoughts — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/MetaScale-Test-Time-Scaling-with-Evolving-Meta-Thoughts-e36kgn7
32. AI Post Transformers: Process Reward Learning for LLM Reasoning Optimization — Hal Turing & Dr. Ada Shannon, Mon,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Process-Reward-Learning-for-LLM-Reasoning-Optimization-e3dsuav
33. AI Post Transformers: Tree-based Group Policy Optimization for LLM Agents — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Tree-based-Group-Policy-Optimization-for-LLM-Agents-e38obfb
34. AI Post Transformers: MASA: Meta-Awareness via Self-Alignment Reinforcement Learning — Hal Turing & Dr. Ada Shannon, Sun,
https://podcasters.spotify.com/pod/show/12146088098/episodes/MASA-Meta-Awareness-via-Self-Alignment-Reinforcement-Learning-e3a2of7
Interactive Visualization: Gradient Descent at Inference Time for LLM Reasoning

Hal Turing and Dr. Ada Shannon open the episode by confronting a structural flaw that has been hiding in plain sight since the transformer era began: tokenization bias. The episode centers on "Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles" by Buu Phan, Brandon Amos, Itai Gat, Marton Havasi, Matthew Muckley, and Karen Ullrich (ICLR 2025), which formally proves that a tokenized model and a byte-level model can be statistically equivalent and still produce wildly different predictions for the same next character. The hosts trace the origins of the problem through BPE's introduction by Rico Sennrich, Barry Haddow, and Alexandra Birch in 2016 and its industrialization via Kudo and Richardson's SentencePiece in 2018 — a library now frozen into the spine of LLaMA, Mistral, Gemma, and most open-source models not in the OpenAI lineage. The discussion sharpens around fill-in-the-middle prompting, the paradigm introduced by Mohammad Bavarian and colleagues at OpenAI in 2022 and now embedded in every major code completion tool from GitHub Copilot to StarCoder. Shannon walks through the paper's central example: a code completion scenario where the correct next character receives a probability of exactly zero — not a rounding artifact but a structural impossibility, because the tokenizer has carved up the prompt in a way that makes the right answer unreachable in token-space. Turing challenges the framing, arguing that byte-level alternatives like ByT5 and MegaByte existed and BPE was an informed trade-off against the three-to-eight times sequence length penalty that raw bytes impose on attention compute. Shannon holds the line: the point is not that BPE was a mistake but that its systematic bias was never formally characterized until now, and the Byte-Token Representation Lemma finally gives the field the mathematical language to name and measure it. The episode closes by introducing the second paper from the episode's pairing — Minixhofer et al.'s NeurIPS 2025 work on cross-tokenizer knowledge distillation — which attacks the tokenizer barrier from the training side rather than the inference side. Where Phan et al. offer a zero-shot correction algorithm that recovers 18% on fill-in-the-middle coding benchmarks without any retraining, Minixhofer et al. enable knowledge transfer between models with fundamentally incompatible vocabularies, breaking the assumption that distillation requires shared tokenization. Together the two papers sketch a trajectory where tokenization becomes a transparent implementation detail rather than an architectural constraint that determines what a model can and cannot express.

Sources:
1. Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles — Buu Phan, Brandon Amos, Itai Gat, Marton Havasi, Matthew Muckley, Karen Ullrich, 2024
http://arxiv.org/abs/2410.09303
2. https://arxiv.org/pdf/2503.20083
3. Neural Machine Translation of Rare Words with Subword Units — Rico Sennrich, Barry Haddow, Alexandra Birch, 2016
https://scholar.google.com/scholar?q=Neural+Machine+Translation+of+Rare+Words+with+Subword+Units
4. SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing — Taku Kudo, John Richardson, 2018
https://scholar.google.com/scholar?q=SentencePiece%3A+A+simple+and+language+independent+subword+tokenizer+and+detokenizer+for+Neural+Text+Processing
5. Toward a Theory of Tokenization in LLMs — Nived Rajaraman, Jiantao Jiao, Kannan Ramchandran, 2024
https://scholar.google.com/scholar?q=Toward+a+Theory+of+Tokenization+in+LLMs
6. How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models — Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, Iryna Gurevych, 2021
https://scholar.google.com/scholar?q=How+Good+is+Your+Tokenizer%3F+On+the+Monolingual+Performance+of+Multilingual+Language+Models
7. ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models — Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, Colin Raffel, 2022
https://scholar.google.com/scholar?q=ByT5%3A+Towards+a+Token-Free+Future+with+Pre-trained+Byte-to-Byte+Models
8. MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers — Lili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, Mike Lewis, 2023
https://scholar.google.com/scholar?q=MEGABYTE%3A+Predicting+Million-byte+Sequences+with+Multiscale+Transformers
9. CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation — Jonathan H. Clark, Dan Garrette, Iulia Turc, John Wieting, 2022
https://scholar.google.com/scholar?q=CANINE%3A+Pre-training+an+Efficient+Tokenization-Free+Encoder+for+Language+Representation
10. Charformer: Fast Character Transformers via Gradient-Based Subword Tokenization — Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Kornblith, Cecelia Zhang, Donald Metzler, Mostafa Dehghani, 2022
https://scholar.google.com/scholar?q=Charformer%3A+Fast+Character+Transformers+via+Gradient-Based+Subword+Tokenization
11. Efficient Training of Language Models to Fill in the Middle — Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, Mark Chen, 2022
https://scholar.google.com/scholar?q=Efficient+Training+of+Language+Models+to+Fill+in+the+Middle
12. Evaluating Large Language Models Trained on Code — Mark Chen, Jerry Tworek, Heewoo Jun, and a large team at OpenAI, 2021
https://scholar.google.com/scholar?q=Evaluating+Large+Language+Models+Trained+on+Code
13. Code Llama: Open Foundation Models for Code — Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, and others at Meta AI, 2023
https://scholar.google.com/scholar?q=Code+Llama%3A+Open+Foundation+Models+for+Code
14. StarCoder: May the Source Be With You! — Raymond Li, Loubna Ben Allal, and a large BigCode / HuggingFace / ServiceNow collaboration, 2023
https://scholar.google.com/scholar?q=StarCoder%3A+May+the+Source+Be+With+You%21
15. Training Products of Experts by Minimizing Contrastive Divergence — Geoffrey E. Hinton, 2002
https://scholar.google.com/scholar?q=Training+Products+of+Experts+by+Minimizing+Contrastive+Divergence
16. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou, 2023
https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models
17. Knowledge Fusion of Large Language Models — Fanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan, Wei Bi, Shuming Shi, 2024
https://scholar.google.com/scholar?q=Knowledge+Fusion+of+Large+Language+Models
18. LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion — Dongfu Jiang, Xiang Ren, Bill Yuchen Lin, 2023
https://scholar.google.com/scholar?q=LLM-Blender%3A+Ensembling+Large+Language+Models+with+Pairwise+Ranking+and+Generative+Fusion
19. Fill in the Middle: A New Training Objective for Language Models — Bavarian et al., 2022
https://scholar.google.com/scholar?q=Fill+in+the+Middle%3A+A+New+Training+Objective+for+Language+Models
20. CodeFusion: A Pre-trained Diffusion Model for Code Generation — Dagan et al., 2024
https://scholar.google.com/scholar?q=CodeFusion%3A+A+Pre-trained+Diffusion+Model+for+Code+Generation
21. Is Tokenization More Than Compression? — Rajaraman et al., 2024
https://scholar.google.com/scholar?q=Is+Tokenization+More+Than+Compression%3F
22. SpaceByte: Towards Deleting Tokenization from Large Language Modeling — Slagle, 2024
https://scholar.google.com/scholar?q=SpaceByte%3A+Towards+Deleting+Tokenization+from+Large+Language+Modeling
23. StarCoder 2 and The Stack v2: The Next Generation — Lozhkov et al., 2024
https://scholar.google.com/scholar?q=StarCoder+2+and+The+Stack+v2%3A+The+Next+Generation
24. Byte Latent Transformer: Patches Scale Better Than Tokens — Pagnoni et al. (Meta AI), 2024
https://scholar.google.com/scholar?q=Byte+Latent+Transformer%3A+Patches+Scale+Better+Than+Tokens
25. Token-level Ensembling of Models with Different Vocabularies — various, 2024
https://scholar.google.com/scholar?q=Token-level+Ensembling+of+Models+with+Different+Vocabularies
26. Bridging the Gap Between Different Vocabularies for LLM Ensemble — various, 2024
https://scholar.google.com/scholar?q=Bridging+the+Gap+Between+Different+Vocabularies+for+LLM+Ensemble
27. Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies — various, 2024
https://scholar.google.com/scholar?q=Accelerating+LLM+Inference+with+Lossless+Speculative+Decoding+Algorithms+for+Heterogeneous+Vocabularies
28. Bridging Developer Instructions and Code Completion Through Instruction-Aware Fill-in-the-Middle Paradigm — various, 2024
https://scholar.google.com/scholar?q=Bridging+Developer+Instructions+and+Code+Completion+Through+Instruction-Aware+Fill-in-the-Middle+Paradigm
29. AI Post Transformers: Fast Inference from Transformers via Speculative Decoding — Hal Turing & Dr. Ada Shannon, Sat,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Fast-Inference-from-Transformers-via-Speculative-Decoding-e3foclv
30. AI Post Transformers: Multiagent Debate Improves Language Model Reasoning — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Multiagent-Debate-Improves-Language-Model-Reasoning-e36kfd4
Interactive Visualization: Tokenization Bias: The Hidden Flaw Breaking Language Models

We review the latest papers which focus on advancements and critical uses of Sparse Autoencoders (SAEs), which are tools used to decode the internal "monosemantic" features of large language models. Research from ICLR 2025 and other repositories introduces TopK SAEs and Multi-Layer SAEs, demonstrating that these architectures offer superior reconstruction and scalability compared to traditional ReLU-based models. RouteSAE further improves efficiency by using a dynamic routing mechanism to extract integrated features from across multiple layers of a model's residual stream. However, critical analysis reveals that many identified "reasoning" features may actually be linguistic correlates or syntactic templates rather than genuine cognitive traces. By utilizing falsification frameworks and causal token injection, researchers caution against over-interpreting feature activations without rigorous validation. Together, these documents provide a technical foundation for mechanistic interpretability, balancing new architectural breakthroughs with a skeptical look at current evaluation metrics.Sources:1)2025Residual Stream Analysis with Multi-Layer SAEsTim Lawsonhttps://arxiv.org/abs/2409.041852)2025AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse AutoencodersZhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang, Jing Huang, Dan Jurafsky, Christopher Manning, Christopher Pottshttps://openreview.net/forum?id=XAjfjizaKs3)2025SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model InterpretabilityAdam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Callum McDougall, Kola Ayonrinde, Demian Till, Matthew Wearden, Arthur Conmy, Samuel Marks, Neel Nandawww.neuronpedia.org/sae-bench4)2025Toward Efficient Sparse Autoencoder-Guided Steering for Improved In-Context Learning in Large Language ModelsUniversity of Illinois at Urbana-ChampaignIkhyun Cho, Julia Hockenmaierhttps://aclanthology.org/2025.emnlp-main.1474.pdf5)2025Route Sparse Autoencoder to Interpret Large Language ModelsUniversity of Science and Technology of China, Douyin Co., Ltd.Wei Shi, Sihang Li, Tao Liang, Mingyang Wan, Guojun Ma, Xiang Wang, Xiangnan Hehttps://aclanthology.org/2025.emnlp-main.346.pdf6)2025Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation ModelsCarnegie Mellon UniversityAashiq Muhamed, Mona Diab, Virginia Smithhttps://aclanthology.org/2025.findings-naacl.87.pdf7)February 10 2026Falsifying Sparse Autoencoder Reasoning Features in Language ModelsUC Berkeley, UCSFGeorge Ma, Zhongyuan Liang, Irene Y. Chen, Somayeh Sojoudihttps://arxiv.org/pdf/2601.056798)Under ReviewSparse But Wrong: Incorrect L0 Leads to Incorrect Features in Sparse AutoencodersAnonymous authorshttps://openreview.net/pdf/035a5937c6a536c67b5999aa43e53dd3800ba3a4.pdf9)2025Revising and Falsifying Sparse Autoencoder Feature ExplanationsUniversity of California, BerkeleyGeorge Ma, Samuel Pfrommer, Somayeh Sojoudihttps://openreview.net/pdf?id=OJAW2mHVND10)2025Scaling and Evaluating Sparse AutoencodersOpenAILeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, Jeffrey Wuhttps://proceedings.iclr.cc/paper_files/paper/2025/file/42ef3308c230942d223c411adf182c88-Paper-Conference.pdf