AI Post Transformers · Episode Companion

ShortGPT: Deleting Redundant LLM Layers, Free Lunch or Artifact?

Men, Xu, Zhang, Wang, Lin, Lu, Han, Chen — Baichuan Inc. & ISCAS · Layers in Large Language Models are More Redundant Than You Expect · October 2024 revision

arXiv 2403.03853 Hal Turing & Dr. Ada Shannon
9 / 32layers deleted (Llama2-7B), 27.1% of parameters
0gradients, 0 retraining steps
45.4→44.0MMLU after pruning
19.40→0.67XSum after pruning

The tension: one pruned model, two verdicts

Same Llama2-7B with nine late layers deleted. Multiple-choice barely moves; generation collapses.

Where ShortGPT sits in the compression map

Hover or tap a node. ShortGPT is the only branch that needs neither gradients nor retraining.

Pre-norm vs post-norm: what the residual stream sees

LLaMA is pre-norm. Normalisation sits inside the branch, so the skip path is never rescaled.

Toy model: why deep updates rotate the state less

Residual norm grows like √depth while each update stays order one. Drag the depth. Toy geometry with an orthogonal update; the paper's derivation covers random initialization only.

Small angle ≠ unimportant: the last layer of Llama2-7B

Perplexity after removing parts of the final layer (Table 1). The final FFN behaves like part of the classifier head.

The algorithm, step by step

One forward pass, no gradients. Bars and heat are illustrative BI shapes, not the paper's exact values.

Which layers get deleted (Table 9 ranges)

One contiguous late block per model. Heat uses mock BI consistent with those ranges.

Delete k layers: Llama2-7B (GPTQ) dial

Slide through the Table 4 operating points. Layers are removed in ascending mock-BI order.

BI vs single-layer-removal perplexity (Figure 3 shape)

Illustrative scatter. The authors report a positive correlation without a coefficient, and validate on one layer at a time while the method removes nine at once.

Retention scoreboard (Table 2, ~25% parameters removed)

Pruned average ÷ dense average over thirteen benchmarks. Click a legend chip to hide a method. Baseline rows are copied from the LaCo paper.

Beyond transformers (Table 3)

Mamba-2.8B holds up; RWKV-7B decays faster.

Perplexity climbs while MMLU holds (Table 4)

GPTQ-quantized Llama2-7B, layers removed.

Stacking with 4-bit (Table 5)

MMLU. Order of operations matters.

Speed: realized vs ideal

4331 → 5147 tokens/s at 27.1% removed.

Healing: swap each removed layer for a gated MLP, train on 50B tokens

Ratio becomes 24.0% because the MLPs add parameters. Generative recovery is left as future work.

Generation is where the damage lands (XSum)

Dense vs ShortGPT vs LaCo where reported. 13B models lose 25–40%; the 7B ones nearly all of it.

Headroom above chance (Llama2-7B)

Hatched = chance floor that can't fall. Triangles mark pruned scores.

What "90%" might really mean

The chance-corrected value is the hosts' hand arithmetic, not a paper figure, and depends on assumed chance levels.

Audit board: open questions by dimension

Editorial judgement of severity (0 none · 2 strong). Hover a cell; click a row for detail.

Click a row.

At 70B the claim flips (20% removed)

"Superior to state of the art" holds at 7B and 13B only.

References