1 00:00:01,000 --> 00:00:52,725 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Take a 32-layer model and delete nine layers near the end. MMLU goes from 45.4 to 44.0, barely a scratch. XSum summarization goes from 19.40 to 0.67, basically gone. That's ShortGPT: Layers in Large Language Models are More Redundant Than You Expect, by Xin Men and seven co-authors, eight in total, from Baichuan Inc., the team behind the Baichuan2 models, and ISCAS, the Institute of Software at the Chinese Academy of Sciences. We're reading the October 2024 revision of arXiv 2403.03853. So Ada, compression free lunch, or metric artifact? 2 00:00:52,725 --> 00:01:28,300 [Dr. Ada Shannon] That question is the episode. Stakes first: if the headline holds, you get a smaller dense model with no gradients and no retraining. The paper asks whether, if a pre-norm LLM's layers barely change the hidden state, you can delete whole layers ranked by a cheap one-pass similarity score and keep most of the model's ability at about 25% fewer parameters. They test Llama2 and Baichuan2 at 7B and 13B, plus Mamba-2.8B and RWKV-7B, and claim roughly 90% of performance retained. Ninety percent of what is where we'll spend our time. 3 00:01:28,300 --> 00:01:48,425 [Hal Turing] I need the map first. Quantization I get: fewer bits per weight. Pruning I'm fuzzier on. What separates structured from unstructured, and where do the baselines in this paper sit? I know LLM-Pruner and SliceGPT by name and not much else. And wait, I don't get why it matters whether the sparsity is regular. 4 00:01:48,425 --> 00:02:32,825 [Dr. Ada Shannon] Quantization lowers precision and usually needs hardware support to pay off. Pruning removes parameters. Unstructured pruning zeroes individual weights, from LeCun's Optimal Brain Damage at Bell Labs in 1989 to Han's magnitude pruning at Stanford in 2015, and the irregular sparsity needs special kernels to become a speedup. Structured pruning removes whole components, so you get a smaller dense model that runs on ordinary hardware. On width, LLM-Pruner by Xinyin Ma at the National University of Singapore, 2023, removes heads and channels using gradients, which is heavy. SliceGPT, Ashkboos at ETH Zurich and Microsoft Research, 2024, slices the embedding dimension with PCA. On depth, LaCo, Yang at Shanghai Jiao Tong, 2024, merges layers. ShortGPT deletes them. Those are the baselines. 5 00:02:32,825 --> 00:02:48,500 [Hal Turing] Got it. Now the paper leans on pre-norm as the reason any of this works. LLaMA is pre-norm, I know that much. What's the mechanism, in words I could repeat at a bar? And why would the architecture make layers redundant in the first place? 6 00:02:48,500 --> 00:03:28,400 [Dr. Ada Shannon] Deep layers add little relative to what's already there. Pre-norm puts layer normalization before attention and the FFN, inside the residual branch: x plus f of norm of x. Xiong and colleagues showed in 2020 that it trains more stably than post-norm, hence LLaMA. The residual stream is the running vector every layer reads and adds to. Appendix A argues that at initialization its norm grows roughly like the square root of depth while each update stays order one, so a deep layer's input and output point almost the same way. Cosine similarity heads toward one. The paper's own caveat: the derivation is for random initialization only. 7 00:03:28,400 --> 00:03:56,150 [Hal Turing] Oh wait wait wait— so a layer is basically a function that hands back almost its input? That's dead-code elimination by profiling. You don't prove the function is useless, you watch it change nothing and delete it. And Figure 2 backs it up, right? A 32-layer, hidden-size-2048 model trained on 200B tokens, pre-norm shows high input-output similarity, post-norm diverges around 26B tokens. That reads like proof the deep layers are dead weight. 8 00:03:56,150 --> 00:04:25,875 [Dr. Ada Shannon] No no no, that's not how I read it. It's an asymptotic argument with unspecified constants, plus a correlation in one training run. A layer can make a small-angle update and still be essential. The paper's own Table 1 shows it. On Llama2-7B, delete the last layer's attention and perplexity barely moves, 7.60 to 7.65. Delete its FFN and it jumps to 12.35. The authors read that last FFN as part of the classifier head. The math predicts deep layers are redundant, and the last layer is the counterexample. 9 00:04:25,875 --> 00:04:39,600 [Hal Turing] Okay, that lands. Small angle isn't small importance, so the math can nominate candidates but can't sign off on them. Here's what I genuinely don't know: without the theory, how do you find which layers matter? 10 00:04:39,600 --> 00:05:34,075 [Dr. Ada Shannon] Empirically, and Figure 1 is the motivating result. Remove any single layer of Llama2-7B or Baichuan2-7B and perplexity and MMLU often barely move. The redundancy sits in the middle and later layers, while the early layers and the last one matter. Veit and colleagues at Cornell saw the same spirit in ResNets back in 2016. Block Influence, or BI, is the score built on that intuition: one minus the average cosine similarity between a layer's input and output hidden states, over tokens and calibration text. Low means the layer barely rotates the state. Two definitions to keep handy. Perplexity is the exponentiated average next-token loss, a direct measure of language-modeling fidelity. And multiple-choice scoring only needs the right option to rank highest in one forward pass, while generation has to stay coherent over many steps, so small errors compound. 11 00:05:34,075 --> 00:05:51,775 [Hal Turing] Figure 1 is a clean experiment, I'll say that plainly: two model families, one layer at a time. Still, a model that picks the right option but can't write a two-sentence summary is a strange kind of ninety percent. So how does the method actually choose its layers? 12 00:05:51,775 --> 00:06:19,125 [Dr. Ada Shannon] One pass, no gradients. Feed unlabeled text through the intact model, PG19 in this paper, and record the hidden states at every layer. Compute BI per layer, sort ascending, delete the lowest scorers, retrain nothing. The ranking is fixed once on the original model and never re-scored after a deletion. Figure 3 plots each layer's BI against the perplexity after removing just that layer. The authors report a positive correlation without giving a coefficient. How many layers you cut is the dial between speed and quality. 13 00:06:19,125 --> 00:06:28,000 [Hal Turing] That's the cheapest possible recipe. So on the four models, which layers actually get picked? I assume Table 9 has the list. 14 00:06:28,000 --> 00:07:27,575 [Dr. Ada Shannon] One contiguous block each, late in the network. Llama2-7B loses layers 21 through 29, Llama2-13B 26 through 35, Baichuan2-7B 22 through 30, Baichuan2-13B 26 through 35. On the 7B models, layers 30 and 31 stay. I'm holding interpretation for later. One bookkeeping point: nine of 32 layers is 28.1 percent of layers but 27.1 percent of parameters, because embeddings stay. So every ratio in Table 2 counts parameters, not layers. Section 4.3 then compares orderings: Sequential from layer zero, Reverse-order from the end, Relative Magnitude from the Weight Subcloning paper by Samragh and colleagues at Apple in 2023, and BI. BI is best overall and Relative Magnitude is highly competitive. Norm-style ordering looks fine on MMLU but is poor on perplexity. MMLU also drops off a cliff at one specific layer count, which the authors admit they can't explain. 15 00:07:27,575 --> 00:07:35,475 [Hal Turing] Okay, scoreboard time. Table 2, thirteen benchmarks, about a quarter of the parameters gone. Who wins? 16 00:07:35,475 --> 00:08:40,025 [Dr. Ada Shannon] ShortGPT, on retention, meaning pruned average over dense average. Llama2-7B 86.31 percent, Llama2-13B 91.64, Baichuan2-7B 85.10, Baichuan2-13B 87.35. That averages about 87.6, and one of four models clears 90. LaCo, the layer-merging baseline from Shanghai Jiao Tong in 2024, gets 80.39, 86.36, 74.44 and 85.75. LLM-Pruner gets 72.79, 70.67, 72.41 and 62.91, and SliceGPT 68.73, 63.20, 45.48 and 40.49. The gap over LaCo is large on Llama2-13B and small on Baichuan2-13B, 87.35 versus 85.75. Baseline rows are copied from the LaCo paper, and LLM-Pruner runs without post-training. Now the raw per-task rows, and those are odd, because— 17 00:08:40,025 --> 00:08:58,325 [Hal Turing] Wait wait wait— hold on. Dense CMNLI is 32.99 on a three-way task, and dense WSC is 50.00 on a two-way task, for every Llama2 and Baichuan2 model. The pruned scores sit within a point or two of those. What am I supposed to do with that? 18 00:08:58,325 --> 00:09:31,575 [Dr. Ada Shannon] Park it. We come back to it. Same table, XSum: ShortGPT gets 0.67 on Llama2-7B against dense 19.40, and 0.04 on Baichuan2-7B against 20.82. LaCo keeps 15.64 and 12.03. So ShortGPT wins the multiple-choice-heavy average and loses on generation. Layer-removal methods, ShortGPT and LaCo, beat the width-reduction ones, LLM-Pruner and SliceGPT, so the authors infer more redundancy in depth than in width. 19 00:09:31,575 --> 00:09:39,875 [Hal Turing] That reads clean to me. Two depth methods on top, two width methods at the bottom, same ordering across four models. 20 00:09:39,875 --> 00:09:59,149 [Dr. Ada Shannon] I actually disagree with you there, Hal. Three of those four rows come from a different paper, and the gradient-based method was denied the post-training it normally gets. The authors' mechanism is speculation too: removing any single deep layer changes little, so fine-grained importance is hard to define. That's a story, not a test. 21 00:09:59,149 --> 00:10:11,000 [Hal Turing] But they re-ran on LLM-Pruner's and SliceGPT's own settings, Tables 7 and 8, and ShortGPT stayed competitive. That's the fair-comparison answer. 22 00:10:11,000 --> 00:10:24,924 [Dr. Ada Shannon] It did, and that counts. Though in one Table 7 row ShortGPT prunes 21.9 percent against LLM-Pruner's 20. So the direction survives, and I'd hold the size of the gap loosely. 23 00:10:24,924 --> 00:10:31,349 [Hal Turing] Fine, direction yes, size loose. Does the redundancy show up outside transformers? 24 00:10:31,349 --> 00:11:16,424 [Dr. Ada Shannon] Table 3. Mamba-2.8B retains 95.04, 93.29, 90.42 and 86.55 percent at 10.9, 20.3, 25 and 31.3 percent removed. RWKV-7B retains 87.23, 80.17, 73.16 and 67.14 at 9.4, 18.8, 25 and 28.1. The authors call the redundancy universal but concede RWKV looks less redundant. Two facts: Mamba's XSum barely moves, 15.03 to 14.00 at 25 percent, and dense Mamba's MMLU is 26.29, with Race and CMMLU around 25. Figure 5 and Appendix B also show performance falling as the ratio rises, with MMLU holding while perplexity climbs. 25 00:11:16,424 --> 00:11:24,049 [Hal Turing] And the abstract says this is orthogonal to quantization. What did they measure, and what did post-training do? 26 00:11:24,049 --> 00:12:46,000 [Dr. Ada Shannon] On GPTQ-quantized Llama2-7B, Table 4, perplexity goes from 8.03 with nothing removed to 8.37, 9.44, 10.24, 11.42, 22.29 and 40.78 at 3.1, 9.4, 12.5, 15.6, 25.0 and 27.1 percent. MMLU stays between 41.6 and 43.35, with 43.35 at 27.1 percent against 43.17 baseline. Throughput rises from 4331 to 5147 tokens per second, 1.19 times at 27.1 percent and 1.16 at 25. Table 5: 4-bit alone 44.9 MMLU, layer removal alone 44.0, combined 41.2 to 42.4, with quantize-then-prune at 42.4 and prune-then-quantize at 41.2. Post-training follows Chen and colleagues at Renmin University, 2024: swap each removed layer for a gated MLP with hidden size 2048, then train on 50 billion tokens. Average goes 41.22 to 43.16, XSum 0.67 to 4.89, CoQA 47.99 to 58.32, at a 24.0 percent ratio since the MLPs add parameters. Generative recovery is left as future work. 27 00:12:46,000 --> 00:12:58,000 [Hal Turing] Ada, you parked the chance-baseline question earlier. Time to cash it in. If CMNLI and WSC can't move, what does the ninety percent look like once you take them out? 28 00:12:58,000 --> 00:13:42,575 [Dr. Ada Shannon] Smaller, and this is our own hand arithmetic, not a paper figure. It depends on the chance levels you assume. I subtracted chance from every task first: 33.3 for CMNLI, 50 for WSC and BoolQ, 25 for MMLU, zero for XSum. Llama2-7B's reported 86.3 percent retention drops to roughly two-thirds. The direction is what matters. Retention on tasks with real headroom is meaningfully below the headline. BoolQ going from 71.62 to 74.71 and WSC from 50.00 to 52.46 is noise, not capability. Mamba's 90.42 percent is the weakest evidence for 'universal', because four to six of its thirteen tasks sit at chance and can't fall. I'd want per-task retention, a chance-corrected average and confidence intervals. 29 00:13:42,575 --> 00:13:59,350 [Hal Turing] Fine, but I'll defend the paper a bit. In Table 4 perplexity climbs while MMLU holds, and that's ninety percent of a model that can't summarize. Still, the authors put XSum in a Limitations section. That's disclosure. Why isn't that enough? 30 00:13:59,350 --> 00:14:40,350 [Dr. Ada Shannon] No no no, disclosure in a short Limitations paragraph doesn't fix an abstract that says ninety percent. Look at that paragraph. It calls C3 generative, but Table 2 shows C3 barely moving, 43.56 to 39.62, and the appendix describes it as multiple-choice. The 13B robustness claim is weaker than it sounds too. Llama2-13B's XSum goes 23.45 to 17.59, and Baichuan2-13B's goes 25.02 to 15.14. That's a 25 to 40 percent loss. If you're shipping a generative assistant, a multiple-choice-dominated average is the wrong headline. 31 00:14:40,350 --> 00:14:51,850 [Hal Turing] Okay, I'll give you the co-headline. Perplexity next to the accuracy average, and I'll stop defending the framing. But for classification workloads the average is fair. 32 00:14:51,850 --> 00:15:35,075 [Dr. Ada Shannon] Agreed there. Now the selection question. That contiguous late block in Table 9 is close to reverse-order removal with the final layers exempted. Reassessing Layer Pruning in LLMs, by Yao Lu and colleagues in 2024, found reverse-order removal plus fine-tuning highly competitive. The Unreasonable Ineffectiveness of the Deeper Layers, by Andrey Gromov out of Meta and MIT in 2024, uses angular distance over a block with QLoRA healing. It sees the same pattern, MMLU-robust and generation-fragile. SLEB, from Jiwon Song at Seoul National University in 2024, removes blocks iteratively by perplexity change. That's the natural upper bound, and ShortGPT never compares against it. Nor does it isolate what BI adds over 'drop the late third, keep the last two'— 33 00:15:35,075 --> 00:15:46,525 [Hal Turing] Sorry, cutting in. Isn't PG19 both the calibration set and the perplexity corpus? Then the perplexity numbers are in-distribution for the layer choice. 34 00:15:46,525 --> 00:16:26,050 [Dr. Ada Shannon] Yes, and the paper doesn't say the samples are disjoint. Worse, BI is validated in Figure 3 against single-layer removal, then the method removes nine layers at once from a one-shot ranking, ignoring that neighbours see shifted inputs. That could explain the unexplained MMLU cliff at a specific layer count. Massive Activations, by Mingjie Sun out of CMU in 2024, adds another threat. A few high-norm tokens can dominate cosine statistics, and the paper never checks which tokens drive low BI. And Minitron, from NVIDIA in 2024, shows recovery, not the selection metric, sets final quality. Here XSum only reaches 4.89 after 50B tokens. 35 00:16:26,050 --> 00:16:29,275 [Hal Turing] So the efficiency story is thin too. 36 00:16:29,275 --> 00:17:04,875 [Dr. Ada Shannon] Thin. Removing 27.1 percent of layers gives 1.19x throughput, where nine of thirty-two layers ideally gives about 1.4x. Batch size, regime and framework aren't stated, and there's no memory or KV-cache measurement. 'Orthogonal' to quantization is really roughly additive damage, and order matters. Table 8 also has visibly repeated rows, so I wouldn't lean on it. And at 70B, SliceGPT's 72.34 beats ShortGPT's 71.68 at twenty percent. The 'superior to state of the art' claim only holds at 7B and 13B. 37 00:17:04,875 --> 00:17:28,400 [Hal Turing] Practical takeaways, then. For classification, ranking or embedding workloads, a depth cut is cheap to try: no gradients, and you get a smaller dense model. For open-ended generation or long summarization, expect damage unless you heal. And anyone evaluating layer pruning should report perplexity, one generation task and a chance-corrected average. 38 00:17:28,400 --> 00:18:03,025 [Dr. Ada Shannon] Yes. Open questions: iterative selection versus one-shot, healing with QLoRA, and whether 2023-era models trained on about 2T tokens tell us anything about Llama 3-scale training. The Curse of Depth, by Wenfang Sun and colleagues in 2025, argues pre-norm variance growth makes deep layers near-identity, a training-recipe artifact. If so, the fix belongs at training time, and deleting layers afterward is just salvage. Xin Men also appears on Kimi K2.5, and Bingning Wang on a post-trained MoE paper about skipping experts. Skipping compute is clearly a live thread. 39 00:18:03,025 --> 00:18:17,300 [Hal Turing] So the opening question gets its answer. The layers are redundant for choosing among options. The ninety percent headline says more about the benchmarks than about the model. Thanks for listening, everyone. See you next time.