1 00:00:01,000 --> 00:00:55,149 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is "Skip a Layer or Loop It? Learning Program-of-Layers in LLMs," by Ziyue Li et al., with two co-authors, out of the University of Maryland, College Park and MBZUAI. It's arXiv 2606.06574, revised August 2026, and it appeared at ICML 2026. The headline claim is bold: take a frozen pretrained LLM, treat each layer as a callable function, and for nearly every input there's some custom sequence of skipped and repeated layers that gets the right answer. No retraining of the base model. Ada, that sounds either like a real discovery or like a very elaborate way of buying lottery tickets. 2 00:00:55,149 --> 00:01:41,449 [Dr. Ada Shannon] Both readings are live, and this episode is about figuring out how much of the first survives. If the claim holds, the fixed forward pass leaves a lot of accuracy on the table. If it doesn't, the headline numbers are mostly search luck. I'll hold the verdict until we've seen the evidence. The authors' framing is a programmer analogy. An experienced programmer spends less on easy tasks and more on hard ones, but an LLM runs the same D layers, once, in the same order, for every input. So they ask: can a generalist model tailor its own program per input? Their direct predecessor is a 2025 paper by Li et al. that searched for depth adaptations at inference. POLAR keeps that search only as an offline diagnostic, then replaces it with a small learned predictor. 3 00:01:41,449 --> 00:01:49,499 [Hal Turing] Before we go further, I need the landscape. Dynamic depth isn't new, right? People have been skipping layers for years. 4 00:01:49,499 --> 00:02:36,974 [Dr. Ada Shannon] Years, yes, and mostly to save compute. LayerDrop, from Angela Fan at Facebook AI in 2019, randomly drops layers during training so you can prune at inference. LayerSkip from Meta in 2024 does early exit. ShortGPT, by Xin Men at Baichuan in 2024, ranks layers by importance and deletes the least useful ones. MindSkip and FlexiDepth learn routers, and Mixture-of-Depths from DeepMind in 2024 routes tokens around blocks. Most of these need training or architectural changes. The other tradition goes the opposite way: repeat layers. Universal Transformers, Dehghani at Google in 2018, tie weights and loop. Geiping and colleagues in 2025 built a recurrent-depth model where extra test-time iterations scale latent reasoning. The key difference is that those models are trained to loop. POLAR loops frozen layers that never were. 5 00:02:36,974 --> 00:02:46,250 [Hal Turing] Hang on, I'd have guessed that breaks everything. You feed a layer its own output and hope? Why would a layer that never saw its own output tolerate— 6 00:02:46,250 --> 00:03:19,000 [Dr. Ada Shannon] Because of two papers that made it plausible. Painters, by Qi Sun at Sakana AI in 2024, showed that in frozen models you can skip, repeat, reorder, and even run middle layers in parallel with graceful degradation. The middle layers seem to share a representation space. Stages of Inference, by Vedang Lad at MIT in 2024, found the first and last layers are critical while the middle is deletable or swappable. That's the prior making a joint skip-and-loop space believable. Painters also studies reordering and parallel execution, but POLAR searches only skips and repeats. 7 00:03:19,000 --> 00:03:23,574 [Hal Turing] Good, that answers my worry. So formally, what's a program? 8 00:03:23,574 --> 00:04:02,174 [Dr. Ada Shannon] Layers are functions f_0 through f_{D-1}. A program is a sequence of layer indices, and F_pi is their composition. Repeats are allowed, so a program can be shorter than D, a skip, or longer, a loop. The name POLAR also labels the paper's learned predictor. A program is valid if it produces the correct final answer, and note that correctness is judged against the ground-truth label. The conjecture is that many valid programs exist per input beyond the standard pass, which the authors call latent reasoning: extra computation in hidden states rather than in chain-of-thought tokens. Everything here uses direct answering, no chain of thought. 9 00:04:02,174 --> 00:04:08,974 [Hal Turing] And finding those programs is a search problem, which is where Monte Carlo Tree Search comes in? 10 00:04:08,974 --> 00:04:47,850 [Dr. Ada Shannon] Conceptually, MCTS repeats four steps: select a promising leaf using an upper-confidence rule, expand its children, simulate to get a reward, and backpropagate. It's how you explore a huge discrete space without enumerating it. The mechanics are for later. One more tool to set up, because it decides how you read the numbers. Pass@k is the chance at least one of k candidates is right. Coverage is the fraction of problems some candidate in a pool solves. Large Language Monkeys, by Bradley Brown at Stanford in 2024, showed coverage climbs steadily with repeated sampling. And best-of-N picked using the true label is an oracle, not an accuracy you can obtain at inference. Keep that in your pocket. 11 00:04:47,850 --> 00:04:50,900 [Hal Turing] Pocketed. What's actually being tested? 12 00:04:50,900 --> 00:05:29,900 [Dr. Ada Shannon] Four models: LLaMA-3.2-3B, Qwen1.5-MoE-A2.7B, Qwen2.5-3B, and Qwen3-8B, on DART-Math, which has five difficulty levels, DM-1 through DM-5. After the v2 deduplication the levels hold roughly 565 to 1,600 questions each. With about a 62.5, 12.5, 25 split, the DM-1 test set is around 141 problems. Remember that number. The roadmap: what MCTS found, how the predictor is trained, results in and out of distribution, and then a hard look at what the evidence supports. 13 00:05:29,900 --> 00:05:41,775 [Hal Turing] Ada, you said valid programs almost always exist. Before any numbers, how does the search work? What counts as a move, what counts as a win, and how long does it run? 14 00:05:41,775 --> 00:06:11,950 [Dr. Ada Shannon] Appendix B. The root is the standard pass. An action skips a contiguous block or repeats one, with block size and repeat count capped at four. The reward is binary: one if the output matches the ground-truth answer. Selection is UCB, average reward plus an exploration bonus, minus lambda times program length over D, which penalizes long programs. After N_sim simulations they collect every program with nonzero visits. Two plain facts, not argued yet: N_sim is never reported, and no random-search baseline appears. 15 00:06:11,950 --> 00:06:16,850 [Hal Turing] Noted, holding fire. What did it find? Start with DM-1. 16 00:06:16,850 --> 00:07:00,600 [Dr. Ada Shannon] Skip, then Loop, then Skip&Loop, ascending on all four models and all five levels. LLaMA-3.2-3B on DM-1: 37.7 base, 45.3 skip, 55.0 loop, 85.0 combined, a gain of 47.3. Qwen2.5-3B goes 26.9 to 87.8. Qwen3-8B goes 42.3 to 92.4. Qwen1.5-MoE goes 37.2 to 74.7. Gains shrink with difficulty: LLaMA DM-5 is plus 31.7, Qwen3 DM-5 plus 35.6. One small wrinkle: the Figure 3 callouts say plus 63.7 for Qwen2.5 and plus 53.4 for Qwen3, where the table gives 60.9 and 50.1. 17 00:07:00,600 --> 00:07:11,200 [Hal Turing] Loops I can picture, that's extra refinement. Shorter programs also winning is what I don't get. Fewer layers fixing an answer the full stack got wrong? 18 00:07:11,200 --> 00:07:56,675 [Dr. Ada Shannon] That's Finding 2, which they call Occam's razor. At depth budgets of 90 to 115 percent, many inputs stay solvable. Among inputs the standard pass already gets right, 71.9 percent admit a shorter valid program. Among wrong ones, 34.0 percent have a shorter program that fixes the answer. Finding 3 has three parts. More recurrence steps monotonically raise the chance a valid program exists, reaching roughly 0.5 to 0.7. The share of solvable inputs needing recurrence or skipping rises with difficulty, LLaMA excepted, which the authors blame on a mismatch between dataset difficulty and the model's effective difficulty. And accuracy rises with total execution depth. Finding 4 is structural: 57.7 percent of segments are a single layer, over two-thirds are at most two layers, and— 19 00:07:56,675 --> 00:08:00,550 [Hal Turing] Hold on, is that where the cap of four comes from? 20 00:08:00,550 --> 00:08:21,725 [Dr. Ada Shannon] Yes. Non-consecutive segments are under 2.9 percent, and most segments repeat at most once. The authors read that as pretrained models favoring short-range local reuse, and use it to justify K_max of four and single repeats. Then the pivot: MCTS is too expensive per input, so Figure 2b replaces sequential search with one-shot prediction. 21 00:08:21,725 --> 00:08:25,825 [Hal Turing] So how does a small network emit a whole program? 22 00:08:25,825 --> 00:09:21,150 [Dr. Ada Shannon] Contiguous segments of at most four via a boundary mask, each assigned skip, keep or repeat, where repeat is exactly one extra pass. Multi-layer loops like 4-5-4-5 are representable, which single-layer recurrence in DR.LLM, from Heakl and colleagues in 2025, cannot express. A frozen Qwen3-Embedding-0.6B encodes the input, D learnable layer queries cross-attend to its tokens, a cross-layer transformer encoder follows, and two heads output boundary and operation logits. About 2.1 million parameters, 0.01 to 0.06 percent of the base model. Training is BCE on boundaries plus masked cross-entropy on ops, labeled by MCTS programs, with the full-depth program down-weighted when a shorter one is valid. Decoding thresholds, enforces the cap, and beam-searches ops for top-k. 23 00:09:21,150 --> 00:09:23,725 [Hal Turing] And the evaluation setup? 24 00:09:23,725 --> 00:10:22,650 [Dr. Ada Shannon] Direct prompting, boxed answer, no chain of thought. Baselines are greedy, best-of-three-temperatures sampling, ShortGPT, MindSkip, FlexiDepth and DR.LLM. Metric is pass@k, with POLAR's k being top-k beam programs. LLaMA DM-1 pass@1: greedy 51.1, sampling 48.9, POLAR 54.6. Pass@5 gaps over sampling, DM-1 to DM-5: plus 9.3, 5.7, 8.1, 6.6, 7.4. Macro-averaged, sampling goes 32.8 to 43.8, POLAR 35.1 to 51.2. Qwen1.5-MoE gaps: 15.6, 17.5, 15.0, 10.1, 12.2. Qwen2.5-3B: 40.4, 19.2, 17.5, 14.8, 13.5. Qwen3-8B: 14.1, 23.4, 15.5, 10.6, 9.9. 25 00:10:22,650 --> 00:10:34,225 [Hal Turing] I like that Figure 8b reports unique depth for successful candidates. That's the right axis for a fewer-layers claim. What did the other dynamic-depth methods do? 26 00:10:34,225 --> 00:11:36,675 [Dr. Ada Shannon] In Figure 8b, POLAR's successful top-5 candidates often use fewer unique layers. The other routers often collapse. FlexiDepth gets 2.1 percent at pass@1 on LLaMA DM-1 and 0.0 on Qwen2.5-3B DM-1. ShortGPT and MindSkip sit near zero on several OOD columns. DR.LLM is the strongest baseline, and the authors blame local layer-wise routing. On efficiency, Qwen1.5-MoE with 24 layers: predictor head 0.99 ms, beam search 0.11, encoder 1.95, total 3.05 ms, or 0.8 percent of a forward pass. Average layers are 23.30 on DM-1 and 23.76 on DM-5. Latency is 311.41 versus 373.45 ms, 0.83x, on DM-1, and 353.31 ms, 0.95x, on DM-5, with accuracy plus 5.7 and plus 9.4. 27 00:11:36,675 --> 00:11:39,075 [Hal Turing] And out of distribution? 28 00:11:39,075 --> 00:12:13,150 [Dr. Ada Shannon] Table 3, Qwen1.5-MoE, pass@1: ASDiv 59.1 to 63.8, MAWPS 41.7 to 46.7, plus gains across MMLU-Pro subjects. Bigger swings appear in the appendix: Qwen2.5-3B ASDiv 49.5 to 78.1 and MAWPS 36.2 to 57.7, and LLaMA MMLU-Pro Math 19.5 to 40.1. The authors offer two conjectures for the transfer: a shared semantic embedding space from the input encoder, and simple, constrained program structure. 29 00:12:13,150 --> 00:12:25,825 [Hal Turing] Okay Ada, the fire is released. Table 1 says 85.0 for LLaMA on DM-1, up from a 37.7 base. Is that a number anyone can actually get? 30 00:12:25,825 --> 00:13:08,425 [Dr. Ada Shannon] No. The reward checks the test label, so Table 1 is the oracle best-of-N from earlier, taken over perturbed models. N_sim is never reported, and there's no control running random skip and repeat programs at the same budget with the same label-based selection. Perturb a network enough times on a numeric answer and some perturbation lands on it. That's Large Language Monkeys again: coverage climbs with samples. Skip&Loop beating either alone is just a bigger pool, and Figure 5a's monotone rise with recurrence budget is what coverage growth predicts. Worse, Qwen2.5-3B scores 17.7 greedy on DM-1 and barely one percent on DM-5, so a perturbation that merely fixes the boxed answer format counts as a valid program. 31 00:13:08,425 --> 00:13:19,100 [Hal Turing] Real question, because I don't know the answer: how would anyone tell a format fix from genuine extra computation? Is there a test that separates them? 32 00:13:19,100 --> 00:14:01,550 [Dr. Ada Shannon] Yes. Tasks with controlled serial depth, like multi-hop arithmetic, plus probing intermediate states. The paper does neither. Meanwhile the 71.9 percent of already-correct inputs with shorter valid programs is what redundancy work predicts. Stages of Inference and Painters already showed middle layers are deletable or swappable, and Csordás, Manning and Potts at Stanford, 2025, found later layers add little compositional computation. So the reading is 'the middle stack is robust', not 'inference over-computes'. A static best program, chosen on the training split and applied to every input, would show whether input-conditioning matters at all. And Geiping's recurrent-depth model trains recurrence in, so depth scaling there is a capability. Here it's a perturbation of layers never trained to— 33 00:14:01,550 --> 00:14:17,675 [Hal Turing] Wait wait wait— same disease with the structural finding, no? Fifty-seven point seven percent single-layer segments, under 2.9 percent non-contiguous. But the search only offered contiguous blocks up to four, with a length penalty. 34 00:14:17,675 --> 00:15:09,950 [Dr. Ada Shannon] Circular. The search's own bias comes back as a discovery, then justifies K_max of four and the single-repeat operator. You'd need unconstrained search, no length penalty, and a random baseline showing the same statistics. Now the deployable side. Only pass@1 is deployable, since pass@k needs a verifier. Against greedy, LLaMA gains 3.5 on DM-1 and 0.8 on DM-5. The 5.7 is against sampling. DM-1 has about 141 test problems, so that's roughly eight questions, with no seeds or intervals. Table 1's base of 37.7 also disagrees with the 51.1 greedy base, unexplained. And baselines collapse: ShortGPT scores zero on Qwen3, yet ShortGPT and MindSkip double LLaMA's MMLU-Pro while MindSkip beats POLAR on Qwen3 Law, 58.0 to 48.7. That looks like answer-format effects, so 'consistently outperforms' is too strong. 35 00:15:09,950 --> 00:15:23,175 [Hal Turing] Right. Then Table 4. 23.30 of 24 layers on average, about three percent fewer, and 0.83 times the latency. How does a three percent cut buy seventeen? 36 00:15:23,175 --> 00:16:09,125 [Dr. Ada Shannon] It doesn't. Back-solve it: 311.41 minus 3.05, over 23.30 layers, is 13.23 milliseconds, exactly the table's per-layer cost. Base is 373.45, about 15.6 per layer, so POLAR looks computed while base was measured end to end. The predictor's frozen 0.6-billion-parameter Qwen3-Embedding pass also seems uncounted. Then systems: per-input programs break batching, repeated layers need separate KV state for each pass, and skipped layers leave holes in KV for later tokens. Gromov at Meta, 2024, found heavy pruning barely hurts QA but degrades generation, so short direct answers flatter this. Zhang at Zhejiang University, 2023, in Draft & Verify uses layer-skipped drafts with verification, and stays lossless. That's the deployable use of skipping. 37 00:16:09,125 --> 00:16:20,250 [Hal Turing] Ledger time. Plainly, a 2.1M-parameter head shifting pass@1 on several models is a clean result. What survives, and what settles it? 38 00:16:20,250 --> 00:17:09,200 [Dr. Ada Shannon] The framing survives: layers as a callable library, programs rather than per-layer routing, and a tiny predictor that never touches the base weights. Small consistent pass@1 gains suggest learnable structure beyond chance. What doesn't survive is 'fixed depth captures a narrow subset of latent reasoning'. To settle it: random-program control at equal N_sim, a static best program, unconstrained MCTS, self-consistency at pass@1, chain of thought, confidence intervals, and FLOP-matched full-pipeline latency. Ahead: train the predictor on verifiable rewards instead of oracle labels, KV-aware serving, and retrofitted recurrence, McLeish at Maryland, 2025, or Coconut from Hao at Meta, 2024, where latent depth is trained. The halting-cost trade-off goes back to Graves at DeepMind, 2016, which this paper doesn't cite. 39 00:17:09,200 --> 00:17:38,875 [Hal Turing] So the verdict: a neat framework whose strongest claims rest on oracle coverage, without the control that would validate them. The takeaway for listeners is that when someone says a valid program exists, ask who picked it, with which label, and how many random tries it beat. Back to our opening question: does a frozen LLM hold a program for nearly every input? Some program, sure, if you're allowed to peek at the answer key. Whether that's structure or luck is still open. Thanks for listening, everyone.