AI Post Transformers · Episode Companion

Chain-of-Layers: Skipping and Looping Frozen LLM Layers Per Input

A frozen pretrained model, no gradient updates, one Monte Carlo Tree Search per input: skip some layers, repeat others, reorder the rest. This page visualizes the search, the per-layer usage profile, the headline results — and the open questions the hosts didn't resolve.

arXiv:2507.07996 Li, Li, Zhou · Univ. of Maryland · 2025 Interactive viz (this page)

What CoLa does to a frozen stack

The base model always runs layers 1…N in the same order. CoLa searches, per input, for an edited chain: skip a run of layers, repeat a block several times, and let the order shift — using only the original, unchanged weights.

Example search outcome

The kind of path MCTS finds differs by input difficulty: easy inputs tend to compress, hard ones tend to loop. Toggle between a representative easy-input path and a hard-input path.

Headline claims

>75%
of already-correct samples still work with a shorter layer chain
>60%
of wrong samples get fixed by some other chain
0
gradient updates — weights are never touched

Monte Carlo Tree Search over layer edits

Each state is a layer sequence. Each move edits it: skip k consecutive layers or repeat a block of k layers r times, k,r ∈ [1,4]. Click through the four stages.

The node score, decomposed

Selection balances three terms: mean reward (exploit), an unexplored-node bonus (explore), and a length penalty normalized by original depth N. λ = 5.0, 200 simulations per input, ε = 0.1 random-child chance. Values below are an illustrative example, not paper-reported numbers.

Which layers survive the skip?

Early layers are almost never skipped; usage dips through the middle, then the pattern diverges by model size. Reconstructed qualitatively from the paper's description — exact per-layer values are not published.

Skip-rate heatmap, same data

How short do paths actually get?

Figure 2 and Appendix A: average CoLa paths save only a handful of layers. The shortest 5% of correct paths cut up to ~30% of depth — but that's the tail, not the norm.

Skip-only vs. recurrence-only vs. joint (CoLa)

Joint search beats either axis alone in every reported row. Select a task/model pair.

DART-5: original vs. CoLa

Harder multi-step math reasoning, levels 1–5. Original accuracy collapses; CoLa recovers a large fraction of it.

One number, six models

ARC-Easy CoLa accuracy is 95.80 for all six models tested — dense and MoE, base and instruct, 3B and 8B alike (479/500). ARC-Challenge reads 98.2 for four of six.

Flagged in the episode: a search budget of 200 tries against a 4-way answer key can itself produce near-ceiling "pass@200" numbers. The hosts wanted this compared against a random-edit baseline with the same budget — the paper doesn't report one.

Where CoLa sits in the landscape

Two axes: does the method touch weights (trained) or not (frozen), and does it adapt per-input (dynamic) or apply the same edit to every input (static)? CoLa's corner — frozen weights, dynamic per-input search — is the one Transformer Layers as Painters already explored, uncredited by this paper.

Dead weight, or overthinking?

Hal: deep layers barely move accuracy when pruned, so they're wasted parameters.

Ada: that's mostly multiple-choice evidence; multi-step reasoning degrades much faster when the same layers are cut. Redundancy for one task type isn't redundancy in general.

Independent support for the shape of the effect: Lad, Gurnee & Tegmark (2024) found early/final layers sensitive and the middle robust to deletion/swap — without any search.

Existence proof, or search-success rate?

The paper's own text says predictions are checked against gold labels and that result feeds the search — while Algorithm 1, line 5, says evaluate on held-out inputs. Both are printed as written; the episode doesn't resolve which one produced the headline numbers.

Two missing controls the hosts proposed: (1) apply a path found on some questions to different held-out questions, and (2) give the unmodified model the same 200-try budget with random edits, à la Large Language Monkeys' pass@k scaling.

References