1 00:00:01,000 --> 00:00:47,524 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're looking at "Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs" by Ziyue Li et al., with two co-authors, all at the University of Maryland, College Park. It went up on arXiv in July 2025 as a preprint under review. Here's the claim. Take a frozen pretrained LLM and, for each individual test sample, skip some layers, repeat others, and stack them in a new order. No finetuning at all. Every input normally runs through all N layers in the same order, whether it's trivial or brutal. The authors ask whether easy inputs need all of that, and whether hard ones get enough. 2 00:00:47,524 --> 00:01:23,400 [Dr. Ada Shannon] The paper says the fixed stack is a habit, not a law. Their abstract claims that for over 75% of samples the model already gets right, a shorter chain of layers still works. For over 60% of samples it gets wrong, some chain of layers fixes it. They call this Chain-of-Layers, or CoLa. Every path uses only the original layers with unchanged weights. Skipping is the fast-thinking analogy and repeating is the slow-thinking one. The umbrella term is test-time adaptation: changing the computation at inference, with no gradient updates, to suit the input. Here the thing being adapted is depth and layer order, not the weights. 3 00:01:23,400 --> 00:01:33,900 [Hal Turing] Okay, my curiosity gets stuck at the first step. I trained this thing with layer seven feeding layer eight. Why doesn't deleting layer eight just wreck the output? 4 00:01:33,900 --> 00:02:20,500 [Dr. Ada Shannon] Residual connections. Each block adds a small update to a shared residual stream, so a skipped block just passes its input through. Veit and colleagues at Cornell showed in 2016 that ResNets behave like ensembles of shallow paths, so removing blocks degrades them gracefully. That fuels static pruning. ShortGPT, from Xin Men and colleagues at Baichuan in 2024, scores each layer by how much it changes the hidden state and deletes the lowest ones with no training. Gromov and colleagues, out of Meta and MIT in 2024, in "The Unreasonable Ineffectiveness of the Deeper Layers", removed large fractions of layers with little loss on question answering. Dynamic methods choose per input. BranchyNet from Harvard in 2016 attached exit heads to intermediate layers, and Google's CALM in 2022 did per-token exit for generation. But those need training or auxiliary heads, and they can only truncate the tail. 5 00:02:20,500 --> 00:02:30,025 [Hal Turing] Hold on, that's actually—so if you can chop half the deeper layers and barely notice, those layers are dead weight. Just wasted parameters. 6 00:02:30,025 --> 00:02:47,825 [Dr. Ada Shannon] I disagree with you there, Hal. Gromov's own reading is that either pretraining underuses the deeper layers or the knowledge sits in the shallow ones. Those aren't the same claim. And those results are mostly multiple-choice question answering. Multi-step reasoning degrades much faster. 7 00:02:47,825 --> 00:02:57,575 [Hal Turing] Fine, but multiple choice is exactly what this paper tests, so for its purposes the redundancy is real. I'll hold my ground on that. 8 00:02:57,575 --> 00:03:29,925 [Dr. Ada Shannon] For this paper, agreed. Now the other axis. Looped models repeat layers. Universal Transformers, from Dehghani and colleagues at Google in 2018, share one block across depth with per-position halting. Geiping and colleagues, from the ELLIS Institute Tübingen and the University of Maryland in 2025, trained a recurrent-depth model that spends more iterations on harder inputs. Those models are trained to loop. CoLa loops a model that never was. The precedent for treating layers as modules is mixture-of-experts, which does routing on the width axis. This is the depth-axis version. 9 00:03:29,925 --> 00:03:36,375 [Hal Turing] So how do you pick a path? Skip or repeat any layer in any order is a gigantic space. 10 00:03:36,375 --> 00:03:58,175 [Dr. Ada Shannon] That's why they use Monte Carlo Tree Search. It builds a tree of states by repeating four steps: select a promising node, expand it with new children, simulate to see how good it is, and backpropagate the result up the tree. It balances exploiting branches that already look good against exploring untried ones. It's what powered AlphaGo's search from DeepMind in 2016. Here each state is a layer sequence and each move edits it. 11 00:03:58,175 --> 00:04:03,600 [Hal Turing] And before we go further, is this prior work? You mentioned frozen models. 12 00:04:03,600 --> 00:04:24,350 [Dr. Ada Shannon] Yes, and the paper doesn't cite it. Transformer Layers as Painters, by Qi Sun and colleagues out of Sakana AI and Emergence AI in 2024, already tried skipping, repeating, reordering and parallelizing layers in frozen LLMs. The CoLa paper also doesn't cite ShortGPT or Gromov. Keep that in mind, because it changes how much of this is new. 13 00:04:24,350 --> 00:04:40,825 [Hal Turing] Let me say the setup back, and you correct me. The model has N layers, weights frozen. A path is a list of those layers, any order, repeats allowed. Nothing is trained and no parameter is added. The only knobs are which layers run and how often. 14 00:04:40,825 --> 00:05:01,675 [Dr. Ada Shannon] Right. For OLMoE, each whole MoE layer counts as one unit. The root of the search tree is the untouched path, and an action is one edit: skip k consecutive layers, or repeat a block of k layers r times, with k and r each from one to four. Path length is capped. A path is terminal when no edit is allowed. You run it on the input, and the reward is one for a correct output, optionally minus a length penalty. 15 00:05:01,675 --> 00:05:08,700 [Hal Turing] The score that picks which branch to grow looks like three ideas stapled together. Take it piece by piece? 16 00:05:08,700 --> 00:05:50,950 [Dr. Ada Shannon] It is three ideas. Q over v is mean reward so far, which exploits what worked. C times the square root of log V over v is the exploration bonus, which grows for neglected nodes. Then minus lambda times path length over N, a penalty normalized by original depth. The paper uses lambda of 5.0, 200 simulations per input, and a 0.1 chance of picking a random unexplored child instead. All fixed across models and datasets. The output is the Pareto set: no other path is both shorter and more accurate. One thing to pocket: the text says predictions are compared to gold answers and that feeds the search, while Algorithm 1, line five, says evaluate on held-out inputs. Both are printed as written. I'm not resolving it here. 17 00:05:50,950 --> 00:05:54,550 [Hal Turing] Pocketed. So what did they run, and what came out? 18 00:05:54,550 --> 00:06:36,250 [Dr. Ada Shannon] Joint search beats skipping alone and looping alone in every row. Models: LLaMA-3 at 3B and 8B, base and instruct, plus OLMoE-1B-7B, both flavors. Data: ARC-Easy, ARC-Challenge, DART-Math levels one to five, 500 random instances each. LLaMA-3-3B-Base on ARC-Easy: 27.8 original, 75.4 skip-only, 65.4 recurrence-only, 95.8 with both. LLaMA-3-8B-Instruct on DART-2: 6.0, then 25.2, 20.8, and 66.2. On DART-5, 3B-Base goes from 1.4 to 25.6, and 8B-Instruct from 12.4 to 47.8. And the ARC-Easy CoLa column is— 19 00:06:36,250 --> 00:06:46,050 [Hal Turing] Sorry, hold on, that's actually— I'm reading down that column and it's 95.80 in every model. Six different models, one number. 20 00:06:46,050 --> 00:07:11,175 [Dr. Ada Shannon] Yes. 95.80 on ARC-Easy for all six, and 98.2 on ARC-Challenge for four of six. Same pocket. Finding one: skip-only helps most on easy tasks, recurrence-only helps more on harder DART levels and instruct models, and the joint space wins everywhere. Their reading is compression plus expansion: skip what an easy input doesn't need, loop what a hard one does. OLMoE gains are smaller, which they attribute to already-sparse computation. That's a hypothesis, not a test. 21 00:07:11,175 --> 00:07:17,900 [Hal Turing] Okay, a real question. I picture half the network gone. How short do these paths actually get? 22 00:07:17,900 --> 00:07:56,975 [Dr. Ada Shannon] Modestly, on average. In Figure 2, 3B paths span roughly 24 to 29 layers out of 28, and 8B roughly 27 to 33 of 32. Unique-layer depth is lower than total depth. Wrong-to-correct paths are shorter than correct-to-correct ones, and base models compress more than instruct. The tail is bigger: Appendix A says the shortest 5% of correct paths reach 20 to 22 layers on 3B and 22.5 to 25 on 8B, up to 30% under full depth. The shortest 20% sit 12 to 23% under. And on DART-4 and 5, the original path is almost never the best one. 23 00:07:56,975 --> 00:08:08,375 [Hal Turing] That's the part I'm buying. Wrong answers get fixed by shorter paths, so the default pass overthinks. The paper says it outright. That's a finding about how the model works. 24 00:08:08,375 --> 00:08:25,550 [Dr. Ada Shannon] I disagree, Hal. 'Overthinking' is a label the authors put on an observation. They didn't measure what the skipped layers were doing, and the claim that instruction tuning calibrates more layers to be relevant is stated, not tested. Shorter paths exist. Why they win is open. 25 00:08:25,550 --> 00:08:34,425 [Hal Turing] But the short path got the answer and the full one didn't. On that input the extra layers hurt. That's an outcome, not a label. 26 00:08:34,425 --> 00:08:40,475 [Dr. Ada Shannon] On that input, along that one path, yes. I'll sign the effect. I won't sign the mechanism. 27 00:08:40,475 --> 00:08:45,175 [Hal Turing] Fine, I'll take 'exists' for now. Which layers does it lean on? 28 00:08:45,175 --> 00:09:52,750 [Dr. Ada Shannon] Early layers are nearly always kept: skip rate near zero, rising through the middle, falling toward the end. The 3B shows a V-shape with the middle suppressed, and the 8B decays more smoothly. Usage entropy is 3.46 for 8B against 3.33 for 3B, max concentration 0.035 against 0.040, and harder tasks flatten usage further. Repeats pile onto late layers in small base models, while larger and instruct models are less stereotyped. That's background, not the authors' claim, but early layers build what everything downstream reads, which fits why they survive. Now the question I pocketed earlier. The headline accuracies are oracle numbers. The search reward is correctness against the gold label, with 200 simulations per sample. So 'CoLa accuracy' means at least one of about 200 tried paths matched the answer key, and the paper never says which of its two descriptions the reported number uses. My reading is that 95.8 on ARC-Easy is a search-success rate. A real evaluation would apply a path found on some questions to different questions, or use a label-free reward like self-consistency. 29 00:09:52,750 --> 00:10:18,775 [Hal Turing] So the missing control is a lottery ticket. ARC is four-way multiple choice, and 200 perturbed forward passes with an answer key... Ada, what does a random-perturbation best-of-200 score? Because if random skip and repeat edits already reach the mid-nineties, the tree search is decoration. And 95.8 percent is 479 of 500, one exact count across dense and MoE models. 30 00:10:18,775 --> 00:10:54,975 [Dr. Ada Shannon] Right, and ARC-Challenge lands at 98.2, above Easy. That looks like a ceiling set by the budget and the answer key, plus maybe the same 21 items that are unparseable or mislabeled. The frame I'd use is Large Language Monkeys, by Bradley Brown at Stanford in 2024: coverage rises log-linearly with repeated samples. Pass@200 on the unmodified model is the baseline this paper needs. The starting points also look broken. LLaMA-3-3B-Base at 27.8 on ARC-Easy is near the 25 chance rate. Some of the W-to-C gains may just be perturbations that make the model emit a bare letter. 31 00:10:54,975 --> 00:11:13,025 [Hal Turing] Genuine question, then. The efficiency pitch says more than 75 percent of correct samples get a shorter chain. But you told us the average savings are a few layers, and the search burns 200 passes per input. Does any path come out ahead once you count the search? 32 00:11:13,025 --> 00:11:33,200 [Dr. Ada Shannon] No. The 30 percent figure is only the shortest 5 percent of paths, and the paper reports no latency or FLOPs at all. Skipping one layer counts as 'shorter'. The entropy claim also needs a caveat. 3.46 versus 3.33 is just ln 32 versus ln 28, so both models are near uniform for their layer counts. 33 00:11:33,200 --> 00:11:53,750 [Hal Turing] I'll say the empirical part is the strong part. Six models, two task families, every skip-only and recurrence-only ablation laid out cleanly. And I think the existence result is the finding. A frozen model contains a correct path for most items. That's surprising whether or not you can find it cheaply. 34 00:11:53,750 --> 00:12:14,375 [Dr. Ada Shannon] I actually disagree, Hal. With 200 tries and a four-way answer space, 'a correct path exists' is nearly guaranteed for a broken baseline. It tells you the output distribution is perturbable, not that layers act as composable modules. The 'noisy layers' story is post-hoc, and there's no ablation showing a skipped layer flips the error consistently. 35 00:12:14,375 --> 00:12:39,675 [Hal Turing] Okay, but the layer profile has independent support. The Remarkable Robustness of LLMs paper, by Vedang Lad, Wes Gurnee and Max Tegmark at MIT in 2024, deleted and swapped layers without any search. Early and final layers were sensitive and the middle was robust, which matches the skip profile here. So the descriptive map survives, even if the accuracy jumps don't. 36 00:12:39,675 --> 00:13:08,075 [Dr. Ada Shannon] Agreed on the map. The missing step is a policy. Dr.LLM, by Ahmed Heakl and colleagues at MBZUAI in 2025, uses MCTS-derived paths as supervision for lightweight routers on frozen models. The paper doesn't frame it that way, but that's what makes the search useful. Trained versions exist too: Mixture-of-Depths from DeepMind in 2024 and Mixture-of-Recursions in 2025. My own engineering worry is serving. Per-sample graphs break batching, and repeated layers need separate KV-cache state. 37 00:13:08,075 --> 00:13:32,650 [Hal Turing] So the takeaway: the paper is a good existence proof and search-space map, and its accuracies are upper bounds until two experiments run. First, apply found paths to held-out questions. Second, give the unmodified model and random edits the same 200-attempt budget. If the tree search still wins on held-out data, that's a real result. Thanks for listening, everyone. See you next time.