A visual companion to the episode: how a standard classifier gets a calibrated reject option, why coverage and selective risk must be read together, and how SGR turns confidence ranking into an abstention policy.
Forced prediction hides the decision policy. Selective classification changes the question from “How often is the model right?” to “When the model decides to answer, how risky is that answered set?”
The guarantee is narrow but useful: given an i.i.d. calibration set, a target risk, and a confidence level, select a threshold whose true selective risk is below target with high probability.
The paper’s main empirical lesson is blunt: simple maximum softmax probability often performs about as well as, and sometimes better than, heavier uncertainty machinery for selective ranking on in-distribution data.
This page stays inside the paper’s world. Under dataset shift, open-set inputs, or subgroup drift, these score structures can become misleading even if the in-distribution calibration looked clean.