← All episodes Transformer Layers as Painters: Skipping and Reordering Frozen LLM Layers

Transformer Layers as Painters: Skipping and Reordering Frozen LLM Layers

Sep 20, 2026
This episode examines "Transformer Layers as Painters," which asks whether a frozen, ordinarily trained Llama 2 can tolerate having its layers skipped, swapped, reordered, or run in parallel without any retraining. The discussion uses a painter-on-an-assembly-line analogy: the residual stream serves as a shared canvas, so layers read and write the same space. It places the paper alongside related work on residual networks, logit and tuned lens, depth pruning, and the "stages of inference" study. Testing on Llama2-7B, 13B, and 70B plus BERT-Large across five benchmarks, the first results show a sharp split. Removing or swapping the first and last layers collapses performance, while the uniform middle layers barely register a change, and accuracy falls off gradually instead of breaking. Listeners interested in layer pruning, conditional computation, and the latency cost of depth get a look at what a pretrained model can absorb without being trained for it.
Sources:
1. Transformer Layers as Painters — Qi Sun, Marc Pickett, Aakash Kumar Nain, Llion Jones, 2024
http://arxiv.org/abs/2407.09298
2. Eliciting Latent Predictions from Transformers with the Tuned Lens — Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, Jacob Steinhardt (building on nostalgebraist's 2020 'interpreting GPT: the logit lens'), 2023
https://scholar.google.com/scholar?q=Eliciting+Latent+Predictions+from+Transformers+with+the+Tuned+Lens
3. The Remarkable Robustness of LLMs: Stages of Inference? — Vedang Lad, Wes Gurnee, Max Tegmark, 2024
https://scholar.google.com/scholar?q=The+Remarkable+Robustness+of+LLMs%3A+Stages+of+Inference%3F
4. The Unreasonable Ineffectiveness of the Deeper Layers — Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, Daniel A. Roberts, 2024
https://scholar.google.com/scholar?q=The+Unreasonable+Ineffectiveness+of+the+Deeper+Layers
5. Residual Networks Behave Like Ensembles of Relatively Shallow Networks — Andreas Veit, Michael Wilber, Serge Belongie, 2016
https://scholar.google.com/scholar?q=Residual+Networks+Behave+Like+Ensembles+of+Relatively+Shallow+Networks
6. Reducing Transformer Depth on Demand with Structured Dropout (LayerDrop) — Angela Fan, Edouard Grave, Armand Joulin, 2019 (ICLR 2020)
https://scholar.google.com/scholar?q=Reducing+Transformer+Depth+on+Demand+with+Structured+Dropout+%28LayerDrop%29
7. LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding — Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, Bram Wasti, et al. (Meta), 2024
https://scholar.google.com/scholar?q=LayerSkip%3A+Enabling+Early+Exit+Inference+and+Self-Speculative+Decoding
8. Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language Models — David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, Adam Santoro (Google DeepMind), 2024
https://scholar.google.com/scholar?q=Mixture-of-Depths%3A+Dynamically+Allocating+Compute+in+Transformer-Based+Language+Models
9. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Tom Goldstein, 2025
https://scholar.google.com/scholar?q=Scaling+up+Test-Time+Compute+with+Latent+Reasoning%3A+A+Recurrent+Depth+Approach
10. Interpreting GPT: the logit lens (blog post) and Eliciting Latent Predictions from Transformers with the Tuned Lens — nostalgebraist; Nora Belrose et al., 2020 / 2023
https://scholar.google.com/scholar?q=Interpreting+GPT%3A+the+logit+lens+%28blog+post%29+and+Eliciting+Latent+Predictions+from+Transformers+with+the+Tuned+Lens
11. Do Language Models Use Their Depth Efficiently? — Róbert Csordás, Christopher D. Manning, Christopher Potts, 2025
https://scholar.google.com/scholar?q=Do+Language+Models+Use+Their+Depth+Efficiently%3F
12. The Curse of Depth in Large Language Models — Wenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin, Yefeng Zheng, Shiwei Liu, 2025
https://scholar.google.com/scholar?q=The+Curse+of+Depth+in+Large+Language+Models
13. Massive Activations in Large Language Models — Mingjie Sun, Xinlei Chen, J. Zico Kolter, Zhuang Liu, 2024
https://scholar.google.com/scholar?q=Massive+Activations+in+Large+Language+Models
14. Highway and Residual Networks Learn Unrolled Iterative Estimation — Klaus Greff, Rupesh K. Srivastava, Jürgen Schmidhuber, 2016
https://scholar.google.com/scholar?q=Highway+and+Residual+Networks+Learn+Unrolled+Iterative+Estimation
15. Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding — Jun Zhang et al., 2023
https://scholar.google.com/scholar?q=Draft+%26+Verify%3A+Lossless+Large+Language+Model+Acceleration+via+Self-Speculative+Decoding
16. Universal Transformers / Deep Equilibrium Models — Mostafa Dehghani et al.; Shaojie Bai, J. Zico Kolter, Vladlen Koltun, 2019
https://scholar.google.com/scholar?q=Universal+Transformers+%2F+Deep+Equilibrium+Models
17. Your Transformer is Secretly Linear — Anton Razzhigaev et al., 2024
https://scholar.google.com/scholar?q=Your+Transformer+is+Secretly+Linear
18. Layer by Layer: Uncovering Hidden Representations in Language Models — Oscar Skean et al., 2025
https://scholar.google.com/scholar?q=Layer+by+Layer%3A+Uncovering+Hidden+Representations+in+Language+Models
19. Transformer Feed-Forward Layers Are Key-Value Memories — Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy, 2021
https://scholar.google.com/scholar?q=Transformer+Feed-Forward+Layers+Are+Key-Value+Memories
20. BERT Rediscovers the Classical NLP Pipeline — Ian Tenney, Dipanjan Das, Ellie Pavlick, 2019
https://scholar.google.com/scholar?q=BERT+Rediscovers+the+Classical+NLP+Pipeline