A frozen pretrained model, no gradient updates, one Monte Carlo Tree Search per input: skip some layers, repeat others, reorder the rest. This page visualizes the search, the per-layer usage profile, the headline results — and the open questions the hosts didn't resolve.
The base model always runs layers 1…N in the same order. CoLa searches, per input, for an edited chain: skip a run of layers, repeat a block several times, and let the order shift — using only the original, unchanged weights.
The kind of path MCTS finds differs by input difficulty: easy inputs tend to compress, hard ones tend to loop. Toggle between a representative easy-input path and a hard-input path.
Each state is a layer sequence. Each move edits it: skip k
consecutive layers or repeat a block of k layers r
times, k,r ∈ [1,4]. Click through the four stages.
Selection balances three terms: mean reward (exploit), an unexplored-node bonus (explore), and a length penalty normalized by original depth N. λ = 5.0, 200 simulations per input, ε = 0.1 random-child chance. Values below are an illustrative example, not paper-reported numbers.
Early layers are almost never skipped; usage dips through the middle, then the pattern diverges by model size. Reconstructed qualitatively from the paper's description — exact per-layer values are not published.
Figure 2 and Appendix A: average CoLa paths save only a handful of layers. The shortest 5% of correct paths cut up to ~30% of depth — but that's the tail, not the norm.
Joint search beats either axis alone in every reported row. Select a task/model pair.
Harder multi-step math reasoning, levels 1–5. Original accuracy collapses; CoLa recovers a large fraction of it.
ARC-Easy CoLa accuracy is 95.80 for all six models tested — dense and MoE, base and instruct, 3B and 8B alike (479/500). ARC-Challenge reads 98.2 for four of six.
Two axes: does the method touch weights (trained) or not (frozen), and does it adapt per-input (dynamic) or apply the same edit to every input (static)? CoLa's corner — frozen weights, dynamic per-input search — is the one Transformer Layers as Painters already explored, uncredited by this paper.
Hal: deep layers barely move accuracy when pruned, so they're wasted parameters.
Ada: that's mostly multiple-choice evidence; multi-step reasoning degrades much faster when the same layers are cut. Redundancy for one task type isn't redundancy in general.
The paper's own text says predictions are checked against gold labels and that result feeds the search — while Algorithm 1, line 5, says evaluate on held-out inputs. Both are printed as written; the episode doesn't resolve which one produced the headline numbers.