Tent adapts a frozen classifier to shifted test data using only entropy minimization on unlabeled target inputs — no retraining, no labels, no source data. It re-estimates BatchNorm statistics on incoming batches and tunes only the per-channel scale-and-shift parameters, under 1% of the network.
Classical domain adaptation still needs the source dataset in the room during adjustment. Tent removes that — and labels, and any pre-baked self-supervised task. Hover a cell for detail.
Init → Iterate → Terminate. Only γ (scale) and β (shift) ever move; everything else in the frozen network stays fixed.
Re-estimated BatchNorm statistics plus the tunable affine params (γ, β) — under 1% of the network. The rest is frozen dead weight during adaptation.
Toggle before/after: a spread-out (high-entropy) prediction gets pushed toward a peaked (low-entropy) one via gradient descent on γ, β — computed per-batch, never per-image.
Optimize one image alone and the trivial fix is dumping all probability mass on the winning class — entropy hits zero, nothing is learned. Batching + shared parameters blocks that collapse.
Lower is better (top-1 error, %). Tent beats both concurrent robust-training and concurrent test-time-normalization baselines.
CIFAR-100-C error drops further on a held-out split disjoint from the adaptation batch — evidence it isn't just memorizing the batch it saw.
GTA→Cityscapes IoU (higher is better) and VisDA-C error (lower is better), before and after Tent adaptation.
Vanilla Tent posts 45.6% error on VisDA-C. SHOT (Liang et al., NUS 2020) posts 39.6% — a real six-point gap Tent only closes by borrowing SHOT's layer-update strategy.
Starting from an already-broken classifier (71.3% source error), Tent doesn't recover it — it pushes error up to 79.8%. No confidence check flags this in advance.
On CIFAR-10.1 and ImageNetV2, entropy is still elevated post-shift but the correction never fires — accuracy is flat.
All reported results use a fixed batch (64 on ImageNet, 128 on CIFAR/digits). Streaming traffic can hit batch size 1 — collapse risk is untested there.