MiCA freezes LoRA's up-projection matrix to the minor singular directions of the pretrained weight — the low-energy corners classical compression discards — and trains only the down-projection into that subspace. This page visualizes the mechanism, the spectrum it exploits, and the results (and gaps) the episode debates.
Full fine-tuning updates every weight. LoRA freezes W and learns a low-rank correction with both factors free to drift. MiCA keeps the same low-rank shape but fixes one factor to the pretrained weight's own minor singular directions.
MiCA's best BLOGS-MC result used rank 16 vs. LoRA's rank 128 — and one whole factor is fixed geometry, not learned — so the trainable-parameter gap compounds on top of the rank gap.
Sorted singular values of a (mock) weight matrix. The classical move — Eckart–Young–Mirsky, 1936 — keeps the big ones and discards the tail as noise. MiCA's bet: that discarded tail is exactly where new knowledge has room to live. Hover a bar for its value.
A schematic decomposition — U (left singular vectors) × Σ (singular values, diagonal) × VT (right singular vectors). Cell color = magnitude via the blue→orange→red heat scale.
The abstract's 5.9× figure is 11.8‑point MiCA gain ÷ 2.0‑point LoRA gain on HISTORY‑MC (102 items). Small denominators make ratios swing wildly — a chart makes that visceral.
Same rank, same learning rate, same fixed‑B/trained‑only‑A architecture — only which subspace B is frozen to changes. Qwen2.5‑7B‑Instruct, instruct baseline = 72.91%.