A frozen pretrained LLM, no retraining: for each input, does some custom sequence of skipped and repeated layers beat the standard forward pass? The paper searches that space with Monte Carlo Tree Search, then trains a 2.1M-parameter predictor (POLAR) to emit a program directly. This page charts the oracle-coverage claims, the deployable pass@1 gains, and the gap between the two.
Layers are treated as callable functions f₀…f_D₋₁. A "program" is any sequence of layer indices — shorter than D (skip), longer than D (loop), or both. MCTS finds programs offline against a binary correct/wrong reward; POLAR learns to predict them without search.
Dashed arrow: at inference, MCTS is never run again — POLAR's forward pass replaces it entirely.
A 12-layer toy model under the standard forward pass vs. a sample POLAR-style program that skips two blocks and loops two others (repeat = one extra pass, capped at K_max=4).
Table 1 numbers are best-of-N over perturbed programs, selected against the ground-truth label — an oracle coverage figure, not something obtainable at inference. Keep that framing while reading these charts.
Only LLaMA-3.2-3B's individual Skip and Loop numbers are stated in the transcript; other models report only Base and combined Skip&Loop (hatched bars = not reported).
DM-1/DM-5 are as reported; DM-2–DM-4 and the two DM-5 values not stated directly (Qwen2.5-3B, Qwen1.5-MoE) are linearly interpolated/estimated for illustration — hover any cell for the exact label.
Among inputs the standard pass already answers correctly, how many also admit a shorter valid program? And among wrong answers, how many get fixed by one?
The MCTS action space caps contiguous blocks at 4 layers with a single repeat — and the programs it finds are overwhelmingly short and contiguous, which the paper reads as evidence for local reuse (see the circularity critique in the Reality Check tab).
A frozen 0.6B embedding model plus a ~2.1M-parameter head (0.01–0.06% of the base model) replaces the MCTS search at inference time.
Qwen1.5-MoE-A2.7B, 24 layers. Total predictor overhead reported as 3.05ms, ~0.8% of one forward pass.
Unlike Table 1, these numbers don't peek at the label at inference time — except pass@k itself, which still needs a verifier to be usable in the k>1 case.
DM-1/DM-5 endpoints reported; DM-2–DM-4 linearly interpolated for illustration.
DR.LLM is called "the strongest baseline" in the transcript but no specific pass@1 figure is given — shown as not reported. On Law, MindSkip beats POLAR outright.
The headline 85–92% coverage numbers and the ~3–6 point pass@1 gains are not the same claim. Three specific gaps, charted.
(311.41 − 3.05) / 23.30 ≈ 13.23ms/layer for POLAR's reported total, vs. 373.45 / 24 ≈ 15.56ms/layer for baseline — consistent with baseline measured end-to-end while POLAR's own frozen 0.6B encoder pass looks uncounted.
Without an unconstrained search and a random-program control at equal N_sim, "local reuse" could just be a restatement of the search's own bias.