AI Post Transformers · Interactive Viz

Selective Classification with Deep Neural Networks

A visual companion to the episode: how a standard classifier gets a calibrated reject option, why coverage and selective risk must be read together, and how SGR turns confidence ranking into an abstention policy.

arXiv 1705.08500 arXiv 1901.09192
Transcript IDs found: 1705.08500, 1901.09192
Focus: SGR, softmax response, MC-dropout
Closed-world calibration
i.i.d. assumption
Risk-coverage tradeoff
Shift breaks guarantees
Live Episode Viz Link

Post-Hoc Selective Classification

Trained model + confidence score + threshold gate
accepted prediction abstained sample confidence ranking

Key Quantities

Judging the subset the model chooses to answer

Forced prediction hides the decision policy. Selective classification changes the question from “How often is the model right?” to “When the model decides to answer, how risky is that answered set?”

SGR: Threshold Selection with a Calibration Set

Finite-sample upper bound drives the chosen operating point
Target Risk 0.05

What SGR Is Actually Doing

No new backbone. Just a calibrated rejection rule.

The guarantee is narrow but useful: given an i.i.d. calibration set, a target risk, and a confidence level, select a threshold whose true selective risk is below target with high probability.

Risk-Coverage Tradeoff

A safer model usually answers less
Confidence Threshold 0.72
softmax response MC-dropout current threshold

Operating Point

Mocked to match the paper’s qualitative behavior

The paper’s main empirical lesson is blunt: simple maximum softmax probability often performs about as well as, and sometimes better than, heavier uncertainty machinery for selective ranking on in-distribution data.

Confidence Score Anatomy

Hover cells to inspect class probability structure

Hover Readout

Accepted examples should cluster high score + low error
Hover a heatmap cell or confidence bar to inspect one sample. The strongest selective score is not “highest confidence ever,” but “best ordering of easy-right above hard-wrong.”

This page stays inside the paper’s world. Under dataset shift, open-set inputs, or subgroup drift, these score structures can become misleading even if the in-distribution calibration looked clean.

References

Compact trail from reject-option theory to later uncertainty work