MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA

arXiv:2604.01694 Rüdiger et al. · 2026-04-02 PEFT / Knowledge Editing / SVD

MiCA freezes LoRA's up-projection matrix to the minor singular directions of the pretrained weight — the low-energy corners classical compression discards — and trains only the down-projection into that subspace. This page visualizes the mechanism, the spectrum it exploits, and the results (and gaps) the episode debates.

Three ways to adapt a model

Full fine-tuning updates every weight. LoRA freezes W and learns a low-rank correction with both factors free to drift. MiCA keeps the same low-rank shape but fixes one factor to the pretrained weight's own minor singular directions.

all weights trainable trainable low-rank factor frozen (fixed geometry)

Trainable parameter footprint (illustrative, Llama‑2‑7B attention proj.)

MiCA's best BLOGS-MC result used rank 16 vs. LoRA's rank 128 — and one whole factor is fixed geometry, not learned — so the trainable-parameter gap compounds on top of the rank gap.

W = U Σ VT — the ranked structure inside every weight matrix

Sorted singular values of a (mock) weight matrix. The classical move — Eckart–Young–Mirsky, 1936 — keeps the big ones and discards the tail as noise. MiCA's bet: that discarded tail is exactly where new knowledge has room to live. Hover a bar for its value.

major / dominant subspace (top‑r) unused middle spectrum minor subspace (bottom‑r) — MiCA's target

Factorizing W: three matrices, color-coded by magnitude

A schematic decomposition — U (left singular vectors) × Σ (singular values, diagonal) × VT (right singular vectors). Cell color = magnitude via the blue→orange→red heat scale.

LoRA vs. MiCA, side by side

Step through the MiCA recipe

1
2
3
4

Knowledge-injection accuracy

Baseline LoRA MiCA

The headline ratio, and why it's fragile

The abstract's 5.9× figure is 11.8‑point MiCA gain ÷ 2.0‑point LoRA gain on HISTORY‑MC (102 items). Small denominators make ratios swing wildly — a chart makes that visceral.

Which subspace matters? Major vs. Random vs. Minor

Same rank, same learning rate, same fixed‑B/trained‑only‑A architecture — only which subspace B is frozen to changes. Qwen2.5‑7B‑Instruct, instruct baseline = 72.91%.

Training dynamics: minor keeps climbing, random plateaus

Major‑r Random Minor‑r (MiCA)

Open cracks in the paper

No held-out split found. Hyperparameters (rank, LR, epochs) appear tuned directly against the same 300 / 102 items used for reported accuracy — a thin buffer against overfitting the search itself.
SOMA (2025) is barely discussed. A near-identically-named prior method — "Singular value decomposed Minor components Adaptation" — gets grouped as a footnote with no head-to-head comparison.
LoRA baseline restricted to Q/V projections only ("following the original LoRA study"), when current practice usually adds K, O, and MLP layers — echoing the "Learning Rate Matters" critique that PEFT-variant gains over LoRA often shrink once LoRA is tuned properly.
The ablation itself survives scrutiny. Minor‑r beats both Major‑r and Random at identical rank/params, and the training curve shows a genuinely different trajectory, not just a higher endpoint — this part doesn't depend on the LoRA baseline being fair.

References

  1. [1] MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA — arXiv:2604.01694
  2. [2] Eckart, C. & Young, G. — The Approximation of One Matrix by Another of Lower Rank, 1936 — Google Scholar
  3. [3] Hu, E. J. et al. — LoRA: Low-Rank Adaptation of Large Language Models, 2021 — Google Scholar
  4. [4] Meng, F., Wang, Z., Zhang, M. — PiSSA, 2024 — Google Scholar
  5. [5] Zhang, Q. et al. — AdaLoRA, 2023 — Google Scholar
  6. [6] Yun, S., Chae, S., Lee, D., Ro, Y. — SOMA, 2025 — Google Scholar
  7. [7] Lee, Y.-A., Ko, C.-Y., Chen, P.-Y., Yeh, M.-Y. — Learning Rate Matters, 2026 — Google Scholar
  8. [8] Liu, S.-Y. et al. — DoRA, 2024 — Google Scholar
  9. [9] Meng, K., Bau, D., Andonian, A., Belinkov, Y. — ROME, 2022 — Google Scholar
  10. [10] Ilharco, G. et al. — Editing Models with Task Arithmetic, 2023 — Google Scholar