1 00:00:01,000 --> 00:00:40,149 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called "The Unreasonable Ineffectiveness of the Deeper Layers," by Andrey Gromov and four co-authors — Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, and Daniel A. Roberts — out of Meta FAIR, University of Maryland, Cisco, Zyphra, MIT, and Sequoia Capital, published at ICLR 2025. And Ada, the number that got me here is: you can apparently delete up to half the layers of a 70-billion-parameter model and it just... keeps answering questions correctly. 2 00:00:40,149 --> 00:01:15,349 [Dr. Ada Shannon] Half. Not a rounding error, not a couple of redundant blocks — roughly half of Llama-2-70B's layers, gone, with almost no drop on standard QA benchmarks. That's the kind of result that should make anyone who's spent time staring at a training loss curve a little uncomfortable, because it means a huge chunk of a very expensive model is, for practical purposes, dead weight on the tasks we usually measure. This paper isn't proposing a flashy new architecture — it's holding a scalpel up to models we already trust and asking a much more basic question: where is the knowledge actually stored? And the answer they get back is not the one most people would have guessed. 3 00:01:15,349 --> 00:01:33,099 [Hal Turing] Right, so before we get into how they figure out which layers to cut, I think we need to set up why anyone would expect this to work at all. It comes down to something called the residual stream. Ada, can you break that down for people who haven't touched a transformer internals diagram? 4 00:01:33,099 --> 00:02:12,024 [Dr. Ada Shannon] Sure. Picture the transformer's information flow not as water passing through pipes, but as a running tally on a highway. Every layer reads the current state off that highway, computes something — attention plus a feed-forward block — and then adds its contribution back on. It doesn't overwrite anything; it just adds. So mathematically, the final output is literally a sum: the original embedded input plus the output of layer one, plus layer two, plus every layer after that, all the way down. That additive structure is the whole architecture's secret sauce, going back to residual networks, but it also has an interesting implication: if some of those additive terms are small or redundant with each other, the sum barely notices if you remove a few. 5 00:02:12,024 --> 00:02:26,374 [Hal Turing] Okay, wait — but layers aren't independent, right? Layer 40's output depends on what layer 39 handed it. So it's not like each layer is contributing some isolated little packet you can just yank out. 6 00:02:26,374 --> 00:03:07,149 [Dr. Ada Shannon] Exactly, and that's the tension the paper is built around. If the terms were truly independent — if every layer just looked at the raw input and added its own opinion — you could prune freely with no consequences. But they're not independent, each layer's output becomes the next layer's input. So in principle, deleting a layer should cause a cascade, because everything downstream was built assuming that layer's contribution was there. The paper's bet is that somewhere past the early layers, the representation stops changing much from one layer to the next — each layer's contribution becomes small relative to what's already accumulated on the stream. If that's true, splicing around a chunk of those near-identical layers should barely disturb anything downstream. 7 00:03:07,149 --> 00:03:24,024 [Hal Turing] Which is a genuinely elegant hypothesis, and it ties into something people have been circling for a while — this idea of knowledge localization. Where, physically, inside all these matrices, does a fact like "Paris is the capital of France" actually live? 8 00:03:24,024 --> 00:04:07,924 [Dr. Ada Shannon] There's real prior work here worth naming. Meng, Bau, Andonian, and Belinkov's ROME paper — Locating and Editing Factual Associations in GPT, 2022 — used causal tracing to show individual facts are concentrated in specific mid-layer feed-forward modules, precise enough that you can edit one fact by tweaking a handful of weights. And Geva, Schuster, Berant, and Levy showed back in 2021 that feed-forward layers behave like key-value memories — the first projection pattern-matches an input, the second retrieves an associated value and adds it to the stream. So going in, the expectation was that knowledge is somewhat localized but distributed across many components. Layer pruning attacks the same question from the opposite direction — instead of reading out what a layer knows, you delete it and see what breaks. 9 00:04:07,924 --> 00:04:20,899 [Hal Turing] So walk me through the mechanics without giving away next segment's math — how do they actually decide what's safe to remove, and how do they patch the model back together afterward without it falling apart? 10 00:04:20,899 --> 00:04:51,899 [Dr. Ada Shannon] At a conceptual level: they measure how similar the representation going into a block of layers is to the representation coming out the other side. If input and output look nearly identical, that block is doing very little transformative work, and it's a good pruning candidate. Once they cut a chunk out, they stitch the two remaining halves together directly — but that seam is now mismatched, because the later layers were trained expecting a different input. So they run a small amount of finetuning to smooth over that mismatch. Here's where it gets practically important — they use QLoRA to do that healing. 11 00:04:51,899 --> 00:04:57,799 [Hal Turing] And QLoRA is — remind me, because I always mix up LoRA and QLoRA. 12 00:04:57,799 --> 00:05:36,424 [Dr. Ada Shannon] LoRA — Low-Rank Adaptation — freezes the entire pretrained model and instead trains a pair of small, low-rank matrices bolted onto specific weight matrices, so you're updating a tiny fraction of the parameters instead of the whole model. QLoRA, from Dettmers, Pagnoni, Holtzman, and Zettlemoyer in 2023, takes that further by first quantizing the frozen base model down to four-bit precision, then applying LoRA on top of that. The upshot for this paper is huge: healing a chopped-up 70-billion-parameter model doesn't require a training cluster, it fits on a single 40-gigabyte A100. That's what makes this an experiment a university lab can run, not just a hyperscaler. 13 00:05:36,424 --> 00:05:48,474 [Hal Turing] Which honestly reframes the whole thing for me — this isn't just a compression trick, it's a cheap probe you can point at any open-weight model and ask, "how much of you is actually necessary?" 14 00:05:48,474 --> 00:06:26,924 [Dr. Ada Shannon] Oh — wait, wait, hold on, before you wrap that up too neatly — I want to push back a little on "cheap probe." It's cheap to run, sure, but I don't think we should undersell what the result implies. If half of a 70-billion-parameter model's deepest layers can vanish and MMLU barely moves, that's not just a fun compression footnote. That's either an indictment of how we pretrain these things — that we're burning enormous compute building parameters the model never learns to use — or it's telling us something uncomfortable about the benchmarks themselves, that they're not actually testing what we think they're testing. Those are two very different, very serious stories, and I don't think the paper fully picks one. 15 00:06:26,924 --> 00:06:38,024 [Hal Turing] I mean, I'd lean toward it being about the benchmarks — MMLU and BoolQ are multiple-choice and yes-no, there's only so much depth of reasoning they can even demand. 16 00:06:38,024 --> 00:07:16,049 [Dr. Ada Shannon] No, no — I actually disagree with you there, and I think it's too easy an out. If it were purely a benchmark artifact, I'd expect the degradation curve to be gradual and noisy, models slowly getting worse as you prune more. That's not what happens. There's a flat region of essentially no cost, then a sharp cliff down to random guessing. A sharp phase transition like that isn't the signature of a benchmark that's just too easy — it's the signature of something structural in how the network stores information across depth. I'm not saying the benchmark critique is wrong, I just don't think it's sufficient on its own, and we shouldn't reach for it as a comfortable escape hatch. 17 00:07:16,049 --> 00:07:41,524 [Hal Turing] Fair — I'll grant the shape of that curve is doing a lot of the persuading here. Okay, so where does that leave us: a large fraction of a massive, expensive model's deepest layers can disappear before anything on the surface looks wrong, healed back to health on a single GPU. That's the setup. Next, we get into exactly how they find which layers to cut, and what happens once you actually run this across seven different models. 18 00:07:41,524 --> 00:08:23,149 [Dr. Ada Shannon] The mechanics are almost insultingly simple once you see them. Four steps. First, you pick how many layers you want gone, call it n. Second, you compute what they call an angular distance between the input to layer l and the input to layer l+n, for every possible starting point l — that's equation seven in the paper, and it's just an arccosine of the cosine similarity, normalized to sit between zero and one. Third, whichever starting layer minimizes that distance is your answer: that block's input and output look almost identical, so you cut everything in between and stitch the two ends together. Fourth, optionally, you heal the seam with a bit of finetuning. The clever bit is where they measure that distance — only on the final token of the sequence, because with a causal mask that's the only position whose embedding has actually seen the whole input. 19 00:08:23,149 --> 00:08:34,299 [Hal Turing] Wait, only the last token? That feels almost too cheap — you're summarizing an entire block's behavior on a whole dataset using a single position per sequence. 20 00:08:34,299 --> 00:09:10,899 [Dr. Ada Shannon] Right, but averaged over ten thousand C4 validation samples, so the noise washes out. And when they turn this into a heat map — layer index on one axis, block size on the other — a really clean pattern falls out. Deeper blocks are consistently the most similar to each other, meaning the representation has basically stopped moving by that point in the network, which is exactly your redundancy signal. But there's one hard exception across every model family: the very last block, the one touching the final layer before the LM head, is almost maximally dissimilar every single time. Drop anything else, but that last layer is load-bearing. 21 00:09:10,899 --> 00:09:19,274 [Hal Turing] So once you've found your l-star, what's the actual healing budget — how much data are we talking about to smooth over the seam? 22 00:09:19,274 --> 00:09:49,574 [Dr. Ada Shannon] They train on somewhere between 164 and 328 million tokens of C4, with LoRA adapters only on the feedforward modules — narrow and cheap by design. And the headline numbers are what you'd hope for given that setup: across seven models — the Llama-2 family, Qwen, Mistral-7B, Phi-2 — you get a flat region of essentially no MMLU or BoolQ loss, then a sharp collapse to random guessing at a threshold that's very family-specific. Llama-2 tolerates the most, up near 50%. Qwen falls apart earliest, around 20%. 23 00:09:49,574 --> 00:10:10,824 [Hal Turing] Okay, but here's where I start to wonder if we're overselling the method itself. If the dumb version — just lop off the deepest layers, no fancy distance computation — nearly matches the sophisticated one after healing, doesn't that undercut the whole angular-distance apparatus? Why build equation seven at all if 'chop the tail' gets you— 24 00:10:10,824 --> 00:10:42,050 [Dr. Ada Shannon] Oh, hold on, hold on — that's not the right conclusion to draw from that result. The fact that the dumb heuristic works almost as well after healing isn't a strike against the method, it's independent confirmation of what the method found. You needed the angular-distance analysis to discover that deep layers are the similar, prunable ones in the first place. The cheap heuristic only works because it's exploiting a pattern the careful measurement revealed. Without the equation-seven analysis nobody knows a priori that 'deepest minus one' is the right rule instead of, say, layers 10 through 30. 25 00:10:42,050 --> 00:10:55,400 [Hal Turing] Fair, but then why does the paper even bother reporting the simple version as a separate contribution rather than a footnote? If it's downstream of the real finding, it reads like they're padding results. 26 00:10:55,400 --> 00:11:34,950 [Dr. Ada Shannon] Because it matters practically — you can prune Llama-2-70B without ever loading the unpruned model onto a GPU to measure anything. That's a real deployment win even if it isn't a new scientific insight. I'll grant you it's not independent evidence, it's an engineering corollary. We can agree on that split. What I won't grant is that it makes the angular-distance work redundant — and it sets up the more interesting asymmetry: on C4 loss, healing turns the sharp phase transition into a smooth, almost linear climb, all the way out to 80% pruning. The benchmark accuracy and the raw next-token loss are being decoupled by healing, which is a genuinely strange finding. 27 00:11:34,950 --> 00:11:45,075 [Hal Turing] And that decoupling gets even weirder once you leave MMLU and BoolQ. What happens on tasks that actually require multi-step reasoning? 28 00:11:45,075 --> 00:12:07,025 [Dr. Ada Shannon] They fall apart immediately — GSM8K and HellaSwag both degrade the moment you prune anything, no flat region at all. But Chain-of-Thought MMLU behaves just like ordinary MMLU, flat until the same threshold. So whatever's happening in those deep layers, it isn't storing the trivia MMLU is testing — it's doing something reasoning-shaped that math and commonsense tasks need immediately and knowledge-retrieval tasks barely touch at all. 29 00:12:07,025 --> 00:12:37,651 [Hal Turing] Right, it isn't storing the trivia MMLU is testing. But here's what nags at me about the whole method: that angular-distance calculation, l-star, is computed entirely on C4, generic scraped web text, and only on the final token of the sequence. Point this pruned, healed model at a codebase, a math proof, or a ten-turn conversation — is l-star still the right place to have cut? Or did they just find 'safe to drop for web text' and we're reading it as 'safe to drop, period'? 30 00:12:37,651 --> 00:13:12,477 [Dr. Ada Shannon] Nobody tests that, and it's a real gap. You run the C4 pass once, prune, heal, and ship, regardless of what the model faces downstream. Code has different dependency structure, dialogue different attention patterns across turns. It compounds with healing too: every experiment heals with QLoRA, 4-bit base, adapters only on the feedforward blocks, on 164 to 328 million tokens. That's cheap and narrow. Full-precision, full-parameter healing might recover more capability, meaning the reported 'safe to prune' fraction could be an artifact of a limited healing method, not proof the layers are truly redundant. 31 00:13:12,477 --> 00:13:39,177 [Hal Turing] Which feeds my other worry about the framing. MMLU and BoolQ are saturated multiple-choice and yes-no formats — a model can often stay above-random through surface answer-choice heuristics even with degraded internal representations. Given GSM8K and HellaSwag collapse the instant you prune anything, how much of 'up to fifty percent prunable' is really benchmark-format robustness dressed up as knowledge robustness? 32 00:13:39,177 --> 00:14:09,227 [Dr. Ada Shannon] Compounded by their own loss data. The healed C4 loss climbs smoothly straight through the exact pruning fractions where MMLU and BoolQ collapse to random — one metric says gradual damage, the other a cliff, same model. The authors cite Schaeffer, Miranda, and Koyejo's 'Are Emergent Abilities of Large Language Models a Mirage?', Stanford, 2023: discontinuous, exact-match metrics can jump sharply even while a continuous metric like loss moves smoothly underneath. Known scoring artifact, not a new mystery — though they present it as their own surprising finding. 33 00:14:09,227 --> 00:14:34,877 [Hal Turing] Oh wait, sorry, hold on — that's the thing that gets me. If loss sails past the exact point where the model stops answering questions, why trust loss as validation anywhere else here? They also call Qwen 'strange' — twenty percent prunable versus forty-five to fifty-five for Llama-2 — with unexplained 'islands' in its heatmap. Doesn't one unexplained outlier undercut how general the angular-distance intuition is supposed to be? 34 00:14:34,877 --> 00:15:02,327 [Dr. Ada Shannon] Legitimate asterisk — they note the islands, don't explain them. Though it's not an isolated result: Men, Xu, Zhang, Wang and colleagues' ShortGPT, near-concurrent 2024 work, uses a different metric, Block Influence, and independently converges on the same story — deep layers are redundant, cheaply prunable. Two groups landing there with different math is real signal. Where I'd push harder is the causal leap on top: that pretraining 'isn't properly leveraging' deep parameters. That's a hypothesis motivated by the pattern, not something they tested. 35 00:15:02,327 --> 00:15:37,002 [Hal Turing] Which is where I actually disagree with the paper's own conclusion, not just its evals. They argue deep layers are necessary for reasoning, since GSM8K and HellaSwag collapse under any pruning. But Sharma, Ash, and Misra's LASER paper, 'The Truth Is in There,' 2023, found the opposite move works: selectively reducing, not deleting, rank in specific weight matrices, often in later layers, improves reasoning accuracy. If trimming deep-layer weights helps reasoning, that's hard to square with 'deep layers are necessary for reasoning.' 36 00:15:37,002 --> 00:16:06,202 [Dr. Ada Shannon] I actually disagree with you there, Hal. Those are different operations. LASER denoises — strips low-signal components out of a matrix while leaving it in place and in role. This paper deletes the block outright. Reasoning improving when you remove noise isn't evidence the layer's function is dispensable — if anything the layer matters enough that its noise floor drags reasoning down. The results are complementary: deep layers do real work, and some of what they compute is degraded by exactly the noise LASER strips out. 37 00:16:06,202 --> 00:16:32,827 [Hal Turing] No, I get the mechanistic distinction, but functionally you've got two papers pointing opposite directions at the same layers: one says editing deep-layer weights improves reasoning, the other says deleting them doesn't hurt QA but does hurt reasoning. 'Deep layers are necessary for reasoning' is doing more work in this conclusion than the evidence cleanly supports — 'deep layers are noisy for reasoning' fits the same data and implies a very different fix. 38 00:16:32,827 --> 00:17:04,552 [Dr. Ada Shannon] Fair, I'll meet you there. They never rank-reduce, only delete, so the honest claim is narrower: deleting deep blocks disproportionately hurts reasoning relative to QA. Whether that's load-bearing signal or noisy machinery reasoning leans on is genuinely open. It connects to LayerSkip, Elhoushi and colleagues, 2024, cited as future work — training in early-exit and layer dropout from the start, so skippability is designed in rather than discovered afterward in an already-trained model. 39 00:17:04,552 --> 00:17:41,127 [Hal Turing] Practically: a legitimate depth-compression trick, a good source of cheap draft models for speculative decoding — but I wouldn't ship a C4-pruned model into a coding assistant or long reasoning agent without re-validating on that domain. And zooming out, the abstract frames this as a general claim about knowledge storage in LLMs, but it's seven dense, non-MoE, sub-70B models from four labs, healed on a few hundred million tokens of one dataset — a suggestive pattern, not a settled theory of transformer depth. 40 00:17:41,127 --> 00:18:01,402 [Dr. Ada Shannon] Which their own future-work list admits — they don't know if this holds across pretraining checkpoints, whether more training changes prunability, or whether knowledge is delocalized rather than shallow. They also never test frontier-scale or heavily overtrained models, so whether this trend holds, shrinks, or reverses past 70B is an open question too. Honest gaps, not yet closed. 41 00:18:01,402 --> 00:18:20,152 [Hal Turing] Good place to land it. Roughly half the layers in some of these models, gone, with barely a dent in the benchmarks everyone quotes — but ask it to reason, or step outside the data they tuned on, and it falls apart. Great paper to argue with, Ada. Thanks for listening — we'll catch you next time.