This episode examines the nabla-Reasoner paper (ICLR 2026), which proposes running gradient descent on token logits during inference — a first-order approach to test-time compute scaling that stands apart from every existing method in the field. The hosts contextualize the work against the established zeroth-order inference-time scaling landscape: Chain-of-Thought, Self-Consistency, Tree of Thoughts, and MCTS-based methods, all of which probe the reward landscape by sampling without directional information. The core argument is that zeroth-order methods hit a hard ceiling on long-horizon reasoning tasks because the search space grows exponentially while reward signals remain sparse, making random sampling increasingly futile. nabla-Reasoner sidesteps this by treating token logit vectors — normally ephemeral intermediate computations — as continuous optimization variables, computing reward gradients with respect to them and nudging the distribution toward higher-reward outputs before committing to each token. Listeners interested in the mechanics of inference-time scaling and the theoretical limits of sampling-based reasoning will find this a technically dense, well-grounded discussion of a genuinely novel approach.
Hal Turing and Dr. Ada Shannon open the episode by confronting a structural flaw that has been hiding in plain sight since the transformer era began: tokenization bias. The episode centers on "Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles" by Buu Phan, Brandon Amos, Itai Gat, Marton Havasi, Matthew Muckley, and Karen Ullrich (ICLR 2025), which formally proves that a tokenized model and a byte-level model can be statistically equivalent and still produce wildly different predictions for the same next character. The hosts trace the origins of the problem through BPE's introduction by Rico Sennrich, Barry Haddow, and Alexandra Birch in 2016 and its industrialization via Kudo and Richardson's SentencePiece in 2018 — a library now frozen into the spine of LLaMA, Mistral, Gemma, and most open-source models not in the OpenAI lineage. The discussion sharpens around fill-in-the-middle prompting, the paradigm introduced by Mohammad Bavarian and colleagues at OpenAI in 2022 and now embedded in every major code completion tool from GitHub Copilot to StarCoder. Shannon walks through the paper's central example: a code completion scenario where the correct next character receives a probability of exactly zero — not a rounding artifact but a structural impossibility, because the tokenizer has carved up the prompt in a way that makes the right answer unreachable in token-space. Turing challenges the framing, arguing that byte-level alternatives like ByT5 and MegaByte existed and BPE was an informed trade-off against the three-to-eight times sequence length penalty that raw bytes impose on attention compute. Shannon holds the line: the point is not that BPE was a mistake but that its systematic bias was never formally characterized until now, and the Byte-Token Representation Lemma finally gives the field the mathematical language to name and measure it. The episode closes by introducing the second paper from the episode's pairing — Minixhofer et al.'s NeurIPS 2025 work on cross-tokenizer knowledge distillation — which attacks the tokenizer barrier from the training side rather than the inference side. Where Phan et al. offer a zero-shot correction algorithm that recovers 18% on fill-in-the-middle coding benchmarks without any retraining, Minixhofer et al. enable knowledge transfer between models with fundamentally incompatible vocabularies, breaking the assumption that distillation requires shared tokenization. Together the two papers sketch a trajectory where tokenization becomes a transparent implementation detail rather than an architectural constraint that determines what a model can and cannot express.
We review the latest papers which focus on advancements and critical uses of Sparse Autoencoders (SAEs), which are tools used to decode the internal "monosemantic" features of large language models. Research from ICLR 2025 and other repositories introduces TopK SAEs and Multi-Layer SAEs, demonstrating that these architectures offer superior reconstruction and scalability compared to traditional ReLU-based models. RouteSAE further improves efficiency by using a dynamic routing mechanism to extract integrated features from across multiple layers of a model's residual stream. However, critical analysis reveals that many identified "reasoning" features may actually be linguistic correlates or syntactic templates rather than genuine cognitive traces. By utilizing falsification frameworks and causal token injection, researchers caution against over-interpreting feature activations without rigorous validation. Together, these documents provide a technical foundation for mechanistic interpretability, balancing new architectural breakthroughs with a skeptical look at current evaluation metrics.Sources:1)2025Residual Stream Analysis with Multi-Layer SAEsTim Lawsonhttps://arxiv.org/abs/2409.041852)2025AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse AutoencodersZhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang, Jing Huang, Dan Jurafsky, Christopher Manning, Christopher Pottshttps://openreview.net/forum?id=XAjfjizaKs3)2025SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model InterpretabilityAdam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Callum McDougall, Kola Ayonrinde, Demian Till, Matthew Wearden, Arthur Conmy, Samuel Marks, Neel Nandawww.neuronpedia.org/sae-bench4)2025Toward Efficient Sparse Autoencoder-Guided Steering for Improved In-Context Learning in Large Language ModelsUniversity of Illinois at Urbana-ChampaignIkhyun Cho, Julia Hockenmaierhttps://aclanthology.org/2025.emnlp-main.1474.pdf5)2025Route Sparse Autoencoder to Interpret Large Language ModelsUniversity of Science and Technology of China, Douyin Co., Ltd.Wei Shi, Sihang Li, Tao Liang, Mingyang Wan, Guojun Ma, Xiang Wang, Xiangnan Hehttps://aclanthology.org/2025.emnlp-main.346.pdf6)2025Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation ModelsCarnegie Mellon UniversityAashiq Muhamed, Mona Diab, Virginia Smithhttps://aclanthology.org/2025.findings-naacl.87.pdf7)February 10 2026Falsifying Sparse Autoencoder Reasoning Features in Language ModelsUC Berkeley, UCSFGeorge Ma, Zhongyuan Liang, Irene Y. Chen, Somayeh Sojoudihttps://arxiv.org/pdf/2601.056798)Under ReviewSparse But Wrong: Incorrect L0 Leads to Incorrect Features in Sparse AutoencodersAnonymous authorshttps://openreview.net/pdf/035a5937c6a536c67b5999aa43e53dd3800ba3a4.pdf9)2025Revising and Falsifying Sparse Autoencoder Feature ExplanationsUniversity of California, BerkeleyGeorge Ma, Samuel Pfrommer, Somayeh Sojoudihttps://openreview.net/pdf?id=OJAW2mHVND10)2025Scaling and Evaluating Sparse AutoencodersOpenAILeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, Jeffrey Wuhttps://proceedings.iclr.cc/paper_files/paper/2025/file/42ef3308c230942d223c411adf182c88-Paper-Conference.pdf