The paper frames “self-knowledge” as a calibration problem: can a model detect when it lacks an answer, or has it only learned polished uncertainty language? This page turns that argument into diagrams, heatmaps, and interactive failure surfaces.
The episode’s central move is to separate raw correctness from the way a model expresses certainty. The diagram combines the paper’s quadrant with the product pipeline it actually tests.
Unanswerable prompts are grouped into several causes. Then semantically similar answerable questions are retrieved from standard QA sets. Hover the matrix to see how semantic closeness does not guarantee model answerability.
The episode’s skeptical reading is that better benchmark performance may come from better uncertainty language rather than cleaner epistemic calibration. Switch views to compare behavioral improvement against reliability shape.
Unsupported, ambiguous, subjective, and future-contingent prompts should not all collapse into one binary uncertainty label. The more the product stakes rise, the more routing policy matters.