Visualization Companion

Metacognition Against Confident Hallucinations

The central claim is visualized here as a control problem: trust breaks less from raw error than from the region where false answers still look certain. The page maps knowledge boundaries, calibration gaps, discrimination failure, and the mismatch between intrinsic uncertainty and what the model says out loud.

arXiv: 2605.01428 Gal Yona · Mor Geva · Yossi Matias Confident error is the danger zone Transcript IDs found: 2605.01428 Additional IDs: none

From Question to Trust Outcome

Step through the paper’s control loop. Each stage changes where confidence can separate helpful truth from polished error.

The highlighted stage changes with the buttons above. The orange-red zone is where wrong answers still carry enough confidence to damage trust.

Knowledge Boundary vs Expressed Uncertainty

A stylized phase map: models often expand the blue region of things they know, but the risky frontier remains the band where internal uncertainty and verbal confidence diverge.

reliable answers hedge or retrieve confident hallucination

Confidence Heatmap by Question Type

Mock instance-level data across factoid categories. Hover each cell to see how often a confidence bin is actually right.

Rows are categories. Columns are verbal confidence bands. Good systems concentrate red only where confidence is justified, not where prose is merely assertive.

Verbal vs Intrinsic Uncertainty

This scatterplot tracks whether the words the model uses match its latent doubt. Hover points to inspect examples.

faithful caution honest confidence overconfident error

Utility Tax vs Hallucination Risk

The paper’s practical dilemma: abstain too often and you lose helpful answers; answer too freely and confident falsehoods survive.

The sweet spot shifts when retrieval exists, but the control problem does not disappear. It moves upward into search quality, verifier reliability, and tool trust.

Action Policy Ladder

Metacognition matters because uncertainty is not only for display. It should trigger different actions at different risk levels.

Research Constellation

The episode sits at the intersection of self-knowledge, verbalized uncertainty, truthful generation, and selective abstention.

What the Episode Argues Is Missing

Not just better answers, but better boundaries: evidence that a model can tell when to answer, hedge, abstain, or seek evidence on a case-by-case basis.

References

Compact citation map. arXiv links are included where explicit IDs are available; other items point to the provided source URLs.

Hallucinations Undermine Trust; Metacognition is a Way Forward Gal Yona, Mor Geva, Yossi Matias, 2026 · arXiv:2605.01428
Language Models (Mostly) Know What They Know Kadavath et al., 2022 · source
Teaching Models to Express Their Uncertainty in Words Lin, Hilton, Evans, 2022 · source
What Large Language Models Know and What People Think They Know Steyvers et al., 2025 · source
Can LLMs Express Their Uncertainty in Their Generated Responses? Yona et al., 2024 · source
When Can LLMs Actually Correct Their Own Mistakes? Kamoi et al., 2024 · source
The Geometry of Truth Marks and Tegmark, 2023 · source
The Internal State of an LLM Knows When It’s Lying Levinstein and Herrmann, 2023 · source
Selective-LAMA Yoshikawa and Okazaki, 2023 · source
UncertaintyRAG Li et al., 2024 · source