Linear Classifier Probes for Intermediate Layers

A visualization-first companion to the episode: how tiny frozen readouts measure what information is linearly accessible at each layer, why separability often rises with depth, and why that still does not prove where computation happens.

Episode metadata

Primary paper: Guillaume Alain & Yoshua Bengio (2016, updated 2018)
Core theme: representation-facing measurement rather than causal explanation
Interactive focus: probe pipeline, depth trends, confounds, training diagnostics
Frozen hidden states Linear probe only CNN layers → labels
Additional arXiv IDs found in transcript/sources: 1610.01644, 2102.12452, 2309.16042
4
interactive tabs
6
custom SVG views
2
mode toggles
∞
caveats per probe

Frozen-network probe pipeline

Click a layer tap to attach the probe. The base model never receives probe gradients.
base network forward path selected probe readout classification loss on labels

Representation geometry sketch

Same label set, different depth. Early layers are mixed; later layers often become easier for a linear boundary.

Probe interface size

Early convolutional maps are huge. Pooling/subsampling keeps the probe from becoming a parameter monster.
Important design choice: if one layer gets a gigantic probe and another gets a tiny one, cross-layer “separability” can partly become a probe-capacity contest.

Linear probe accuracy across depth

Toggle architecture and task framing. Mock data follows the paper’s central pattern: original-label linear decode tends to improve with depth in supervised CNNs.
Inception-v3 ResNet-50 MNIST toy CNN

Layer × property heatmap

Hover cells. Not every “better” trend points in the same direction: global context rises, transfer utility may peak mid-layer, and causal certainty stays low.

Best layer depends on the question

Original training label often favors deeper layers. Other tasks can peak earlier or mid-network.

Probe dashboard during training

Switch between healthy and stalled optimization. Probes can localize where progress stops becoming linearly accessible.

Progressive disclosure

Step through the probe workflow: freeze, tap, train readout, compare curves.

Where the dashboard helps

Operationally useful use-case: diagnostics rather than philosophy.

Accessibility is not causal use

A feature can be decodable from a hidden state without being the mechanism the model actually uses. Hover the paths.
Probe success answers “can a simple decoder recover it?” not “did this layer compute it?” This is why modern interpretability pairs probing with interventions, patching, ablations, or mechanistic evidence.

Confound matrix

Each row is a source of distortion in cross-layer comparison.

Probe toolkit map

Different interpretability tools answer different questions.

References

Yosinski et al. (2014). On the Transferability of Features in Deep Neural Networks
Kornblith et al. (2019). Do Better ImageNet Models Transfer Better?
Kim et al. (2022). A Survey on Probing Methods for Linguistic Information in Neural Language Models
Tenney et al. (2019). What do you learn from context? Probing for sentence structure in contextualized word representations
Related interpretability context: saliency, feature visualization, transfer learning, activation patching