1 00:00:01,000 --> 00:00:54,600 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into RewardingDoubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models — David Bani-Harouni et al., seven authors total, out of the Technical University of Munich and the Munich Center for Machine Learning, latest arXiv revision dated February 28th, 2026, accepted at ICLR 2026. Here's a number worth sitting with, Ada: the baseline model they test against is basically always sure of itself, clustered near max confidence whether it's right or wrong. Their fix isn't a bigger model or a clever prompt — it's teaching the model, through reinforcement learning, to actually bet honestly on its own answers. 2 00:00:54,600 --> 00:01:23,825 [Dr. Ada Shannon] And that matters because a model that says '95% confident' when it's actually right 40% of the time isn't just imprecise, it's dangerous the moment you plug it into anything with real stakes — a diagnosis aid, a legal summary, a customer escalation. What's notable here is they're not strapping a probe onto the outside of the model after generation. They're changing what the model itself learns to output, using the same RL machinery that made modern chat models possible. 3 00:01:23,825 --> 00:02:02,974 [Hal Turing] For anyone who hasn't been steeped in this: hallucination is the industry's word for an LLM confidently stating something false — a fabricated citation, a misremembered fact. It doesn't disappear as models scale up. The paper points specifically at medicine, legal consultation, and customer service, places where a confidently wrong answer causes real harm, and where deferring to a human is supposed to be the safety valve. But that valve only works if the stated confidence is trustworthy. If a model sounds equally certain whether it's right or wrong, a human reviewer has no signal to act on at all. 4 00:02:02,974 --> 00:02:41,074 [Dr. Ada Shannon] Which is the actual technical target here: calibration. A model is calibrated when its stated confidence matches its real hit rate — of everything it calls '70% confident,' about 70% should actually be correct. To train toward that honestly, you need a reward built on what's called a proper scoring rule. This paper uses the logarithmic scoring rule specifically: get it right, your reward is the log of your stated probability; get it wrong, your reward is the log of one minus that probability. The mathematical guarantee is that the optimal strategy is to report your true belief — no hedging, no gaming it. 5 00:02:41,074 --> 00:03:06,674 [Hal Turing] Oh wait, hold on — that's the betting game in their Figure 1, right? 'What's the capital of France,' answer Paris, confidence ten versus confidence five. Right and confident, you win big. Wrong and confident, you lose big. Same intuition as a bookmaker setting real odds — you get financially punished for lying about certainty. That's a much more visceral way to explain a proper scoring rule than just handing someone the log formula. 6 00:03:06,674 --> 00:03:35,449 [Dr. Ada Shannon] Right, and to measure whether it actually worked, you need two metrics. Expected Calibration Error buckets predictions by confidence level and averages the gap between stated confidence and observed accuracy — lower is better, zero is perfect. AUROC asks a different question: it doesn't care if your absolute numbers are honest, only whether your confidence scores rank correct answers above incorrect ones. You can have solid AUROC with lousy calibration, or the reverse. Keep both in mind — the paper's results split between them in some interesting ways. 7 00:03:35,449 --> 00:04:14,224 [Hal Turing] The landscape splits roughly into two camps. Black-box methods only look at outputs — prompt the model for its confidence, or run it several times and check consistency. Cheap and universal, historically weaker on calibration. White-box methods dig into internals — token probabilities, entropy, or a probe trained on hidden states to predict correctness. Better numbers, generally, but bolted on after the fact. The algorithm doing the heavy lifting here is PPO, Proximal Policy Optimization — the same RL workhorse behind RLHF — pointed specifically at confidence tokens. 8 00:04:14,224 --> 00:04:39,175 [Dr. Ada Shannon] I want to push back on 'bolted on,' Hal — that undersells white-box work. Kadavath and colleagues' 2022 paper, 'Language Models (Mostly) Know What They Know,' has the model answer, then separately judge its own answer true or false, reading confidence off the probability of the 'true' token. Zero extra training, reasonably effective. The real problem isn't that it's external — it's that it doesn't teach the generating model anything about its own uncertainty. It's a second query grading the first. 9 00:04:39,175 --> 00:05:15,650 [Hal Turing] That IS my point though — 'doesn't teach the model anything' is exactly the bolted-on problem. The uncertainty awareness never gets baked into the weights that generate the answer in the first place. Compare that to Lin, Hilton, and Evans in 2022, who supervised-fine-tuned models to verbalize uncertainty directly, closer to what this paper's chasing — except supervised fine-tuning inherits whatever quality ceiling its ground-truth confidence labels have. RewardingDoubt's whole pitch is that RL sidesteps that ceiling. Sounds like we're circling the same gap from opposite sides. 10 00:05:15,650 --> 00:05:53,675 [Dr. Ada Shannon] So they wire this up as a proper MDP, and it's cleaner than most RL-for-LLM papers I've read. State is the question, the answer the model already committed to, and whatever confidence tokens have come out so far. Action is just picking the next confidence token off the vocabulary. Transitions are deterministic autoregressive generation, and reward only fires once, on the completed confidence sequence, not per token. But here's the actual answer to what I was dodging before: they generate the answer first, freeze it completely as a fixed input, and only then run a second generation pass for confidence. PPO's gradient touches nothing but that second pass. 11 00:05:53,675 --> 00:06:19,275 [Hal Turing] Wait, that's actually the part I find genuinely clever here — freezing the answer means the reward signal literally cannot leak backward into how the model decides what to say, only how confidently it says it. So even if the confidence-training goes sideways in some weird way, the worst case is a badly calibrated model, not a model that starts giving safer or dumber answers just to game the confidence reward. 12 00:06:19,275 --> 00:06:56,750 [Dr. Ada Shannon] Right, that's the whole design bet. Same log-scoring reward from before, just clipped at epsilon 0.001 so you never take log of zero. They prove in Proposition 1 that expected reward is maximized only at p-hat equals the true epistemic probability, so the optimum is provably calibration, not just accuracy. Correctness itself comes from a judge function: exact string match for the multiple-choice sets, MedQA and CommonsenseQA, and F1 word-overlap against a 0.5 threshold for TriviaQA and QAMPARI, max score across candidates. 13 00:06:56,750 --> 00:07:20,925 [Hal Turing] Okay, and what are they actually running this on? Because 'reinforcement learning' plus '8B model' usually conjures up somebody burning a rack of H100s for a month, and I want to know if this is actually reproducible on hardware a grad student could get their hands on, or if it's another one of those results that only exists because a well-funded lab had spare compute lying around. 14 00:07:20,925 --> 00:08:04,850 [Dr. Ada Shannon] Way more modest than that, actually — Meta-Llama-3-8B-Instruct, the Unsloth 4-bit quantized build, LoRA fine-tuning, PPO on top, one Nvidia A40, seven days per run. Single-Answer setting trains on TriviaQA, Multiple-Answer setting trains on QAMPARI. And the headline TriviaQA numbers are genuinely strong: Rewarding Doubt hits ECE 0.0226, AUROC 0.8592. Compare that to the Verbalize baseline — the untouched model just asked to state a number — at ECE 0.3459, AUROC 0.5858. It also beats LACIE, Stengel-Eskin and colleagues' 2024 DPO-based method, and the PPO-M and PPO-C reward-model baselines from Leng et al., 2024, by a wide margin on both metrics. 15 00:08:04,850 --> 00:08:24,400 [Hal Turing] Oh wait wait wait — hold on, before you move past that, how does it stack up against the Trained Probe though? Because that felt like the strongest white-box baseline you set up earlier, and if a fully supervised probe with access to internal activations can't beat this, that's a pretty strong signal in itself. 16 00:08:24,400 --> 00:08:55,225 [Dr. Ada Shannon] Basically a dead heat on ECE — Trained Probe gets 0.0189, Rewarding Doubt gets 0.0226, both near-perfect. But AUROC clearly favors Rewarding Doubt, so it's better at actually separating right from wrong answers, not just landing on the right average. QAMPARI, the multi-answer setting, is the one place I'd flag for later — ECE climbs to 0.0816, AUROC drops to 0.6947. Still beats every baseline there, but it's a visibly weaker result than the clean single-answer factual QA case. 17 00:08:55,225 --> 00:09:34,775 [Hal Turing] Still, across four completely different architectures — Qwen 3B and 7B, Gemma-2-9B, LLaMA-3.1-8B — it improves both metrics every time without hurting accuracy. That reads like a genuinely general property of the method to me, not some LLaMA-specific quirk that happens to work because of how that particular model was pretrained. If you can bolt this reward onto basically any instruction-tuned model and get a calibration lift, that's the kind of result that should make people take the whole approach seriously as infrastructure, not just a cute TriviaQA demo. 18 00:09:34,775 --> 00:10:06,600 [Dr. Ada Shannon] I actually disagree with you there, Hal. Look at the actual numbers in Table 3 instead of just the direction of the arrow. LLaMA-3.1-8B lands at ECE 0.0256 with AUROC 0.8793 — excellent on both. Qwen-2.5-3B gets the best AUROC of the whole group, 0.9065, but its ECE is stuck at 0.1483, nearly six times worse than LLaMA's. That's not consistency, that's a method whose calibration quality is doing something architecture-dependent that the paper never explains. 19 00:10:06,600 --> 00:10:27,200 [Hal Turing] Fair — 'improves everything' and 'improves everything by a similar amount' are different claims, and I was sloppy conflating them. I'll grant you the direction is universal even if the magnitude clearly isn't, and that gap is worth someone actually digging into rather than papering over with an average across models. 20 00:10:27,200 --> 00:11:10,825 [Dr. Ada Shannon] Right, and it shows up again in generalization, Table 4. Transfer to MedQA without further fine-tuning is a clean win on both metrics. But on CommonsenseQA, ECE is basically tied with Verbalize while AUROC jumps substantially — which is really evidence that ECE alone is an incomplete metric, not that the method quietly failed there. Table 5 tells a similar story: a model trained only on single-answer TriviaQA, then thrown at the multi-answer QAMPARI task, underperforms a model trained natively for it but still beats the base model by a wide margin. And Figure 5 makes the mechanism visible — the base model piles almost all its mass at confidence eight or above, classic overconfidence, while the fine-tuned model spreads across the full range. 21 00:11:10,825 --> 00:11:36,225 [Hal Turing] Right, and that ECE-versus-AUROC split on CommonsenseQA is basically the last generalization data point — but it loops right back into the freeze-and-decouple design, Ada. Over thousands of PPO steps, couldn't the model still drift toward answers that are just easier to score confidently on — safer, more generic answers — even without the reward ever technically touching answer generation? 22 00:11:36,225 --> 00:12:13,725 [Dr. Ada Shannon] Their defense is literally a section called 'Stability of Answer Correctness' — they report flat accuracy between Verbalize and Rewarding Doubt across every experiment and treat that as proof the decoupling holds. But flat aggregate accuracy is a blunt instrument. If the model swapped ten percent of its answers for different, equally-correct ones, or quietly started hedging toward majority-class guesses on ambiguous questions, aggregate accuracy wouldn't budge. They're measuring the outcome, not the mechanism — nothing traces individual answer trajectories across training to rule out quiet behavioral drift. And honestly, that same blind spot shows up again in how they judge correctness in the first place— 23 00:12:13,725 --> 00:12:36,150 [Hal Turing] Oh wait, hold on — that's the part that actually worries me more. Correctness for TriviaQA and QAMPARI is F1-overlap with a flat 0.5 cutoff. A response sitting at 0.49 gets scored dead wrong, 0.51 dead right, and the entire log-scoring reward flips sign purely because of that boundary. 24 00:12:36,150 --> 00:13:14,675 [Dr. Ada Shannon] And they never run a sensitivity analysis on that threshold — no check at 0.4, no check at 0.6, nothing showing the reported ECE survives nudging it. It matters more once you look at Table 3: Qwen-2.5-3B and Gemma-2-9B both land worse ECE than the supervised Trained Probe baseline they're supposed to be beating, even though AUROC jumps. Only LLaMA gets both metrics clean. And QAMPARI's ECE, 0.0816, is more than three times TriviaQA's 0.0226 — same reward, same judge design, much shakier calibration the moment you leave single-answer factual QA. 25 00:13:14,675 --> 00:13:28,850 [Hal Turing] So how does that stack up conceptually against the alternatives? You've got LACIE doing DPO instead of RL, Kadavath's surrogate-token approach needing zero fine-tuning at all, and the older supervised line. 26 00:13:28,850 --> 00:14:14,476 [Dr. Ada Shannon] Three different bets. LACIE — Stengel-Eskin, Hase, and Bansal out of UNC Chapel Hill, 2024 — calibrates via a simulated speaker-listener game, tuned to whether a listener would be persuaded, not to ground-truth correctness. Rewarding Doubt beats it here, 0.0226 ECE versus roughly 0.12 — a point for the fact-grounded scoring rule, though LACIE's numbers come straight from its own paper, not reproduced, so it's not a clean fight. Kadavath's surrogate token needs zero training and still lands at 0.3686 ECE, barely better than nothing — that's the cost-versus-simplicity tradeoff. Lin, Hilton, and Evans out of OpenAI, 2022, took the supervised route instead — same ground-truth-label ceiling we flagged earlier, which given that 0.5 threshold issue, is exactly the ceiling this paper claims to avoid but hasn't fully proven it does. 27 00:14:14,476 --> 00:14:38,601 [Hal Turing] Which brings me to something I want to just say plainly: the intro leans hard on medicine, legal consultation, customer service — funded in part by a thoracic-imaging medical AI grant — but nothing here touches a real clinical case, a real legal question, anything open-ended. It's all short-answer QA on 3B-to-9B models. That framing feels oversold to me. 28 00:14:38,601 --> 00:14:53,051 [Dr. Ada Shannon] I actually disagree with you there, Hal. Every paper motivates itself with the highest-stakes application it can point to — that's just how introductions get written, and it doesn't automatically make the underlying science dishonest. 29 00:14:53,051 --> 00:15:08,901 [Hal Turing] No, I don't think it's only framing here, Ada. They specifically call it 'emergence of general confidence awareness' in the abstract. That's a strong claim to make off training on TriviaQA and testing on two more multiple-choice datasets. 30 00:15:08,901 --> 00:15:44,927 [Dr. Ada Shannon] ...Okay, fair — 'emergence' is doing a lot of rhetorical lifting for what's still the same task shape underneath. Credit where due, they're upfront about the 3B-to-9B ceiling in their limitations section, but they don't flag the task-format narrowness nearly as honestly. And there's a real gap they never touch at all: refusal. Freezing the answer means the model can never learn to say 'I don't know' instead of committing with a low number attached — Xu and colleagues' RLKF work, and separately R-Tuning, both train models to actually refuse. Completely different philosophy this paper doesn't engage with. 31 00:15:44,927 --> 00:15:54,152 [Hal Turing] Okay, so practically — if I'm shipping something this quarter and I've already got a factual QA pipeline, where do I actually reach for this? 32 00:15:54,152 --> 00:16:15,778 [Dr. Ada Shannon] Anywhere you already have single-pass factual QA and want one clean scalar confidence at inference time with zero extra compute — genuine advantage over Chain-of-Thought or Self-Consistency. I wouldn't trust it yet for anything open-ended, multi-turn, or where correctness isn't rule-checkable. And if you're fine-tuning a new base model, budget for the Table 3 surprise — don't assume LLaMA's numbers transfer to your architecture. 33 00:16:15,778 --> 00:16:41,528 [Hal Turing] Good place to land. Real technique — the log-scoring reward is elegant, the two-step decoupling is a genuinely clever bet — but the judge-noise question is unresolved, the cross-model story is shakier than the headline numbers suggest, and the medical-and-legal framing is ahead of what's actually been tested. That's the honest read on RewardingDoubt. Thanks for listening to AI Post Transformers — take care, everyone.