This episode examines "Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs," which proposes Chain-of-Layers (CoLa). CoLa takes a frozen pretrained model and, for each input, skips some layers, repeats others, and reorders them, with no finetuning. The hosts explain why skipping works, citing residual connections, ResNets behaving like ensembles of shallow paths, and static pruning results like ShortGPT and "The Unreasonable Ineffectiveness of the Deeper Layers". They also cover dynamic early-exit methods and looped-depth models such as Universal Transformers. They disagree about whether pruning results mean deeper layers are dead weight, since those results come mostly from multiple-choice tasks and multi-step reasoning degrades faster. The paper's claim is that many correctly answered samples still work with a shorter layer chain, and many wrong ones can be fixed by some other chain. The episode walks through the Monte Carlo Tree Search that finds these paths, including its skip and repeat edits, its scoring rule with an exploration bonus and a length penalty, and the resulting Pareto set of short, accurate paths. It also notes that the paper omits closely related prior work on frozen-model layer manipulation, which affects how novel the result is.
Sources:
1. Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs — Ziyue Li, Yang Li, Tianyi Zhou, 2025
http://arxiv.org/abs/2507.079962. Universal Transformers — Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, Łukasz Kaiser, 2018 (ICLR 2019)
https://scholar.google.com/scholar?q=Universal+Transformers3. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Tom Goldstein, 2025
https://scholar.google.com/scholar?q=Scaling+up+Test-Time+Compute+with+Latent+Reasoning%3A+A+Recurrent+Depth+Approach4. Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language Models — David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, Adam Santoro, 2024
https://scholar.google.com/scholar?q=Mixture-of-Depths%3A+Dynamically+Allocating+Compute+in+Transformer-Based+Language+Models5. Transformer Layers as Painters — Qi Sun, Marc Pickett, Aakash Kumar Nain, Llion Jones, 2024
https://scholar.google.com/scholar?q=Transformer+Layers+as+Painters6. The Unreasonable Ineffectiveness of the Deeper Layers — Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, Daniel A. Roberts, 2024
https://scholar.google.com/scholar?q=The+Unreasonable+Ineffectiveness+of+the+Deeper+Layers7. ShortGPT: Layers in Large Language Models are More Redundant Than You Expect — Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, Weipeng Chen, 2024
https://scholar.google.com/scholar?q=ShortGPT%3A+Layers+in+Large+Language+Models+are+More+Redundant+Than+You+Expect8. LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding — Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, Bram Wasti, Liangzhen Lai, Anas Mahmoud, Bilge Acun, Saurabh Agarwal, Ahmed Roman, Ahmed A. Aly, Beidi Chen, Carole-Jean Wu, 2024 (ACL 2024)
https://scholar.google.com/scholar?q=LayerSkip%3A+Enabling+Early+Exit+Inference+and+Self-Speculative+Decoding9. Confident Adaptive Language Modeling — Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q. Tran, Yi Tay, Donald Metzler, 2022 (NeurIPS 2022)
https://scholar.google.com/scholar?q=Confident+Adaptive+Language+Modeling10. Dr.LLM: Dynamic Layer Routing in LLMs — Ahmed Heakl et al., 2025
https://scholar.google.com/scholar?q=Dr.LLM%3A+Dynamic+Layer+Routing+in+LLMs11. Confident Adaptive Language Modeling (CALM) — Tal Schuster et al., 2022
https://scholar.google.com/scholar?q=Confident+Adaptive+Language+Modeling+%28CALM%2912. Router-Tuning / FlexiDepth: dynamic layer skipping in pretrained LLMs — Shwai He et al. (Router-Tuning); Xuan Luo et al. (FlexiDepth), 2024-2025
https://scholar.google.com/scholar?q=Router-Tuning+%2F+FlexiDepth%3A+dynamic+layer+skipping+in+pretrained+LLMs13. Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation — Sangmin Bae et al., 2025
https://scholar.google.com/scholar?q=Mixture-of-Recursions%3A+Learning+Dynamic+Recursive+Depths+for+Adaptive+Token-Level+Computation14. Do Language Models Use Their Depth Efficiently? — Róbert Csordás, Christopher D. Manning, Christopher Potts, 2025
https://scholar.google.com/scholar?q=Do+Language+Models+Use+Their+Depth+Efficiently%3F15. Large Language Monkeys: Scaling Inference Compute with Repeated Sampling — Bradley Brown et al., 2024
https://scholar.google.com/scholar?q=Large+Language+Monkeys%3A+Scaling+Inference+Compute+with+Repeated+Sampling16. Residual Networks Behave Like Ensembles of Relatively Shallow Networks — Andreas Veit, Michael Wilber, Serge Belongie, 2016
https://scholar.google.com/scholar?q=Residual+Networks+Behave+Like+Ensembles+of+Relatively+Shallow+Networks17. The Remarkable Robustness of LLMs: Stages of Inference? — Vedang Lad, Wes Gurnee, Max Tegmark, 2024
https://scholar.google.com/scholar?q=The+Remarkable+Robustness+of+LLMs%3A+Stages+of+Inference%3F18. Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding — Jun Zhang et al., 2023
https://scholar.google.com/scholar?q=Draft+%26+Verify%3A+Lossless+Large+Language+Model+Acceleration+via+Self-Speculative+Decoding19. Interpreting GPT: the logit lens / Eliciting Latent Predictions from Transformers with the Tuned Lens — nostalgebraist (2020); Nora Belrose et al. (2023), 2020/2023
https://scholar.google.com/scholar?q=Interpreting+GPT%3A+the+logit+lens+%2F+Eliciting+Latent+Predictions+from+Transformers+with+the+Tuned+LensInteractive Visualization: Chain-of-Layers: Skipping and Looping Frozen LLM Layers Per Input