← All episodes Teaching Language Models to Bet Honestly on Their Own Answers

Teaching Language Models to Bet Honestly on Their Own Answers

Sep 23, 2026
This episode examines RewardingDoubt, a reinforcement-learning method for training large language models to express calibrated confidence rather than defaulting to near-maximal certainty regardless of correctness. The discussion covers why miscalibrated confidence is dangerous in high-stakes domains like medicine, legal consultation, and customer service, where deferring to a human reviewer only works if the model's stated confidence is trustworthy. It walks through the technical core of the approach: a proper scoring rule (specifically the logarithmic scoring rule) that mathematically rewards models for reporting their true beliefs rather than gaming their confidence scores, framed intuitively as a betting game where lying about certainty costs real "money." The conversation contrasts this RL-based approach with prior black-box (output-only) and white-box (internals-probing) calibration methods, including Kadavath et al.'s self-evaluation technique and Lin, Hilton, and Evans' supervised fine-tuning approach, arguing RL sidesteps the quality ceiling imposed by fixed ground-truth labels. It also details the paper's MDP formulation, where the model's answer is frozen before a separate PPO-trained pass generates the confidence sequence, and explains the two evaluation metrics—Expected Calibration Error and AUROC—used to judge whether the method actually works.
Sources:
1. Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models — David Bani-Harouni, Chantal Pellegrini, Paul Stangel, Ege Özsoy, Kamilia Zaripova, Nassir Navab, Matthias Keicher, 2025
http://arxiv.org/abs/2503.02623v6
2. Language Models (Mostly) Know What They Know — Kadavath, Conerly, Askell, Henighan, Drain, Perez, Schiefer, Hatfield-Dodds, DasSarma, Tran-Johnson, et al. (Anthropic), 2022
https://scholar.google.com/scholar?q=Language+Models+%28Mostly%29+Know+What+They+Know
3. Teaching Models to Express Their Uncertainty in Words — Stephanie Lin, Jacob Hilton, Owain Evans, 2022
https://scholar.google.com/scholar?q=Teaching+Models+to+Express+Their+Uncertainty+in+Words
4. LACIE: Listener-Aware Finetuning for Confidence Calibration in Large Language Models — Elias Stengel-Eskin, Peter Hase, Mohit Bansal, 2024
https://scholar.google.com/scholar?q=LACIE%3A+Listener-Aware+Finetuning+for+Confidence+Calibration+in+Large+Language+Models
5. Taming Overconfidence in LLMs: Reward Calibration in RLHF — Jixuan Leng, Chengsong Huang, Banghua Zhu, Jiaxin Huang, 2024
https://scholar.google.com/scholar?q=Taming+Overconfidence+in+LLMs%3A+Reward+Calibration+in+RLHF
6. Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs — Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, Bryan Hooi, 2024
https://scholar.google.com/scholar?q=Can+LLMs+Express+Their+Uncertainty%3F+An+Empirical+Evaluation+of+Confidence+Elicitation+in+LLMs
7. Calibration-Tuning: Teaching Large Language Models to Know What They Don't Know — Sanyam Kapoor, Nate Gruver, Manley Roberts, Arka Pal, Samuel Dooley, Micah Goldblum, Andrew Wilson, 2024
https://scholar.google.com/scholar?q=Calibration-Tuning%3A+Teaching+Large+Language+Models+to+Know+What+They+Don%27t+Know
Interactive Visualization: Teaching Language Models to Bet Honestly on Their Own Answers