Every layer reads the current state off the stream, computes attention + feed-forward, and adds its contribution back on — it never overwrites. Bar height below each block is that layer's mock contribution magnitude. Deep-middle layers add almost nothing; the final layer is the one hard exception.
For every starting layer l and block size n, the paper measures how similar the input to layer l is to the input at layer l+n, using only the final token of the sequence, averaged over 10,000 C4 samples. Hover a cell for its value. The coolest (bluest) cell is l* — the safest place to cut.
Mock curves illustrating the paper's reported pattern across four models. Toggle between metric families: QA benchmarks show a flat region followed by a sharp collapse; reasoning benchmarks degrade immediately; C4 loss climbs smoothly straight through the point where QA accuracy falls off a cliff.
After a block is removed, the two remaining halves are spliced directly and fine-tuned to smooth over the mismatch — using QLoRA so the whole operation fits on one GPU instead of a training cluster.
Hosts push back on the paper's own framing throughout the episode. Bar length is a rough confidence read on each claim after that back-and-forth — hover a bar for the reasoning.
References
- The Unreasonable Ineffectiveness of the Deeper Layers — Gromov, Tirumala, Shapourian, Glorioso, Roberts (2024)
arxiv.org/abs/2403.17887 - ShortGPT: Layers in Large Language Models are More Redundant Than You Expect — Men, Xu, Zhang, Wang, Lin, Lu, Han, Chen (2024)
Google Scholar ↗ - Locating and Editing Factual Associations in GPT (ROME) — Meng, Bau, Andonian, Belinkov (2022)
Google Scholar ↗ - Transformer Feed-Forward Layers Are Key-Value Memories — Geva, Schuster, Berant, Levy (2021)
Google Scholar ↗ - Compact Language Models via Pruning and Knowledge Distillation (Minitron) — Muralidharan et al., NVIDIA (2024)
Google Scholar ↗ - The Truth is in There: Improving Reasoning with Layer-Selective Rank Reduction (LASER) — Sharma, Ash, Misra (2023)
Google Scholar ↗ - LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding — Elhoushi, Shrivastava, Liskovich, Hosmer et al. (2024)
Google Scholar ↗ - Are Emergent Abilities of Large Language Models a Mirage? — Schaeffer, Miranda, Koyejo (2023)
Google Scholar ↗ - SliceGPT: Compress Large Language Models by Deleting Rows and Columns — Ashkboos, Croci, Gennari do Nascimento, Hoefler, Hensman (2024)
Google Scholar ↗ - Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks — Yang, Yu, Zhu, Hayou (2023)
Google Scholar ↗