One Backbone, Five Adaptation Regimes
Same FLAN-T5-XL, same stopping rule, same task mix. Click a method and the diagram rewires the benchmark around the patch it trains.
The footprint figures are illustrative scale markers for the visual only; the benchmark story is the comparison shape, not one universal percent for all implementations.
Quality vs Patch Footprint
Vertical position is average within-task score retention across the benchmark. Horizontal position compresses trainable footprint on a log scale.
Why PEFT Exists At All
Training cost is not the whole story. Storage duplication, task swapping, and adapter management are part of the deployment math.
Where The Gradients Are Allowed To Land
The forward pass still runs through the giant transformer. What changes is the tiny region marked trainable.
Trainable State Snapshot
Each mode compresses the editable region differently: full-grid movement, low-rank deltas, scaling vectors, bias nudges, or prompt embeddings.
Task-Swap Surface
The system motivation is multi-task reuse: one frozen base, many small task payloads, and much less copy churn than full fine-tuning.
Messy Winner Table, Clean Scope Warning
The transcript gives exact cells for the headline wins. Unquoted cells here are interpolated only to preserve the reported ordering and make the grid navigable.
Placement Is Its Own Optimization Problem
Once a family wins the horse race, the next question is where to spend its tiny parameter budget. The ablation message is selective placement, not blanket insertion.