AI Post Transformers β€” Episode Companion

Naive Test-Time Adaptation Destabilizes LLM Predictions

πŸ“„ arXiv:2602.09719 Longhuan Xu, Cunjian Chen, Feng Yin β€” 2026 SCALENET Β· Hypernetwork LR Control Llama-3.3-70B Β· Qwen3-32B

SCALENET is a hypernetwork fix for unsupervised test-time adaptation in LLMs. A 70B Llama model's negative log-likelihood balloons from 2.21 to 11.49 after five naive adaptation steps β€” this page visualizes why single-sample gradients blow up, how a per-layer per-step learned control signal stabilizes them, and where the fix still fails on open-ended reasoning tasks.

Unsupervised, sample-specific TTA: the adapt-and-reset loop

Each prompt gets its own gradient episode, on the prompt's own NLL β€” no gold label, no memory carried to the next prompt.
Frozen backbone stays untouched. Only LoRA Q/V matrices in attention layers absorb the K-step update, then it is discarded before the next prompt arrives.

The gradient the update actually has to serve

You only ever compute βˆ‡NLL(x). The bet is that it correlates with βˆ‡NLL(y | x). SCALENET's whole job is keeping that inner product non-negative β€” without ever seeing y.

Llama-3.3-70B negative log-likelihood across 5 TTA steps

A single global learning rate cannot serve both regimes: too small and nothing moves, too large and it explodes.
Fixed-rate baseline Step-wise learned LR Layer-wise learned LR (SCALENET)

Why one prompt can't average out its own noise

Batch training cancels idiosyncratic gradient directions across thousands of examples. A single-sample TTA step has no such averaging β€” hover the bars.

SCALENET: a per-layer, per-step learning-rate controller

A 2-layer MLP (hidden=128) reads a compressed prompt representation and outputs a non-negative multiplier for every layer Γ— step Γ— {Q, V} pair.

Training vs. deployment

Supervision moves upstream β€” it doesn't disappear.

TTA lineage

Where SCALENET sits relative to prior unsupervised TTA methods.

NLL at step 5 β€” Llama-3.3-70B

Layer-wise control isn't a marginal fix β€” it's the difference between collapse and stability.

ROUGE across datasets β€” where it holds, where it breaks

Extractive tasks (prompt already contains the answer) climb steadily. Open-ended reasoning tasks regress below the no-TTA baseline (dashed line). Points below the baseline are flagged in orange/red.
Fixed-rate Step-wise Layer-wise (SCALENET) No-TTA baseline (dashed)

Learned per-layer, per-step multiplier magnitude

No clean monotonic story β€” neighboring layers can differ by orders of magnitude. The one consistent pattern: magnitude peaks at step 1, then decays fast.
Hover any cell for its exact learned multiplier. Rows are representative attention layers; columns are adaptation steps 1–5.

References

1
Unsupervised Layer-Wise Dynamic Test Time Adaptation for LLMs β€” Longhuan Xu, Cunjian Chen, Feng Yin (2026)
2
Test-Time Training with Self-Supervision for Generalization under Distribution Shifts β€” Sun, Wang, Liu, Miller, Efros, Hardt (2020)
3
Tent: Fully Test-Time Adaptation by Entropy Minimization β€” Wang, Shelhamer, Liu, Olshausen, Darrell (2021)
4
Test-Time Training on Nearest Neighbors for Large Language Models β€” Hardt, Sun (2024)
5
The Surprising Effectiveness of Test-Time Training for Abstract Reasoning β€” AkyΓΌrek, Damani, Qiu, Guo, Kim, Andreas (2024)
6
Test-time Learning for Large Language Models β€” Hu, Zhang, Chen, Wen, Shuai, Luo, Xiao, Li, Tan (2025)
7
SLOT: Sample-specific Language Model Optimization at Test-time β€” Hu, Zhang, Fang, Chen, Wang, Zhang, Qi (2025)
8
COME: Test-time Adaption by Conservatively Minimizing Entropy β€” Zhang, Bian, Kong, Zhao, Zhang (2024)
9
Revisiting Dynamic Evaluation: Online Adaptation for Large Language Models β€” Rannen-Triki, Bornschein, Pascanu, Hutter, et al. (2024)
10
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks (MAML) β€” Finn, Abbeel, Levine (2017)