The tension: one pruned model, two verdicts
Same Llama2-7B with nine late layers deleted. Multiple-choice barely moves; generation collapses.
Where ShortGPT sits in the compression map
Hover or tap a node. ShortGPT is the only branch that needs neither gradients nor retraining.
Pre-norm vs post-norm: what the residual stream sees
LLaMA is pre-norm. Normalisation sits inside the branch, so the skip path is never rescaled.
Toy model: why deep updates rotate the state less
Residual norm grows like √depth while each update stays order one. Drag the depth. Toy geometry with an orthogonal update; the paper's derivation covers random initialization only.
Small angle ≠ unimportant: the last layer of Llama2-7B
Perplexity after removing parts of the final layer (Table 1). The final FFN behaves like part of the classifier head.
The algorithm, step by step
One forward pass, no gradients. Bars and heat are illustrative BI shapes, not the paper's exact values.
Which layers get deleted (Table 9 ranges)
One contiguous late block per model. Heat uses mock BI consistent with those ranges.
Delete k layers: Llama2-7B (GPTQ) dial
Slide through the Table 4 operating points. Layers are removed in ascending mock-BI order.
BI vs single-layer-removal perplexity (Figure 3 shape)
Illustrative scatter. The authors report a positive correlation without a coefficient, and validate on one layer at a time while the method removes nine at once.
Retention scoreboard (Table 2, ~25% parameters removed)
Pruned average ÷ dense average over thirteen benchmarks. Click a legend chip to hide a method. Baseline rows are copied from the LaCo paper.
Beyond transformers (Table 3)
Mamba-2.8B holds up; RWKV-7B decays faster.
Perplexity climbs while MMLU holds (Table 4)
GPTQ-quantized Llama2-7B, layers removed.
Stacking with 4-bit (Table 5)
MMLU. Order of operations matters.
Speed: realized vs ideal
4331 → 5147 tokens/s at 27.1% removed.
Healing: swap each removed layer for a gated MLP, train on 50B tokens
Ratio becomes 24.0% because the MLPs add parameters. Generative recovery is left as future work.
Generation is where the damage lands (XSum)
Dense vs ShortGPT vs LaCo where reported. 13B models lose 25–40%; the 7B ones nearly all of it.
Headroom above chance (Llama2-7B)
Hatched = chance floor that can't fall. Triangles mark pruned scores.
What "90%" might really mean
The chance-corrected value is the hosts' hand arithmetic, not a paper figure, and depends on assumed chance levels.
Audit board: open questions by dimension
Editorial judgement of severity (0 none · 2 strong). Hover a cell; click a row for detail.
At 70B the claim flips (20% removed)
"Superior to state of the art" holds at 7B and 13B only.