From Question to Trust Outcome
Step through the paper’s control loop. Each stage changes where confidence can separate helpful truth from polished error.
The central claim is visualized here as a control problem: trust breaks less from raw error than from the region where false answers still look certain. The page maps knowledge boundaries, calibration gaps, discrimination failure, and the mismatch between intrinsic uncertainty and what the model says out loud.
Step through the paper’s control loop. Each stage changes where confidence can separate helpful truth from polished error.
A stylized phase map: models often expand the blue region of things they know, but the risky frontier remains the band where internal uncertainty and verbal confidence diverge.
Mock instance-level data across factoid categories. Hover each cell to see how often a confidence bin is actually right.
This scatterplot tracks whether the words the model uses match its latent doubt. Hover points to inspect examples.
The paper’s practical dilemma: abstain too often and you lose helpful answers; answer too freely and confident falsehoods survive.
Metacognition matters because uncertainty is not only for display. It should trigger different actions at different risk levels.
The episode sits at the intersection of self-knowledge, verbalized uncertainty, truthful generation, and selective abstention.
Not just better answers, but better boundaries: evidence that a model can tell when to answer, hedge, abstain, or seek evidence on a case-by-case basis.
Compact citation map. arXiv links are included where explicit IDs are available; other items point to the provided source URLs.