AI Post Transformers · Episode Companion

Hilbert Operator Reveals What Networks Actually Learn

arXiv:2607.21366 Mobahi & Bartlett · Google DeepMind / UC Berkeley Posted 2026-07-23 Structured Pruning · Interpretability

HOPE treats network compression as a measurement instrument, not a deployment trick — representing neurons as objects in a Hilbert space so that pruning, merging, and residual-block eviction all fall out of one data-free, hyperparameter-free capacity score.

From checkpoint to compressed network — no data, no tuning

Click any stage in the pipeline to see what it does. The loop at the bottom is the receding-horizon control strategy: execute one action, rescore only what changed, replan.

Click a stage above to read what it does.

Two neurons, one function

Scale a neuron's incoming weights up, and BatchNorm divides the scaling back out downstream. Magnitude alone can't tell these two neurons apart — toggle to see BatchNorm cancel the difference.

low mid high

Magnitude rank vs. Hilbert-norm rank

10 mock neurons, ranked two ways. Rank 1 = kept first, rank 10 = pruned first. Lines that cross a lot show why a magnitude rule and a functional-size rule disagree about what's safe to remove. Hover a neuron for its ranks.

Neuron capacity — Hilbert–Schmidt norm

18 neurons across three depths, sorted by capacity. Everything left of the dashed line survives the prune; everything right of it doesn't. Hover a bar for details.

early layer mid layer deep layer

Gaussian surrogate distribution

Built purely from each neuron's BatchNorm mean/variance — no forward pass on real data.

Hilbert–Schmidt inner products

Pairwise similarity between 8 neurons. Bright cells off the diagonal are merge candidates — near-duplicate functions. Hover a cell.

Receding-horizon action selection

Each step, three candidate actions get scored as distortion ÷ parameters freed. Lowest ratio wins, executes, and only the touched neurons get rescored before the next step.

    Accuracy vs. surviving-neuron density

    Density = fraction of neurons kept. HOPE (structured, data-free) vs. three magnitude/BN-based 2017-era baselines. Hover a point for its value.

    Reading note: Section 11.1 in the paper reports this trend as a density-vs-accuracy plot only — no numeric table, no confidence intervals, and no comparison against SynFlow, the data-free baseline it cites but never runs. Values here are an illustrative reconstruction of the described trend, not a transcription of exact reported numbers.

    References