AI Post Transformers · Visual Companion

Teaching Language Models to Bet Honestly on Their Own Answers

Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models — Bani-Harouni, Pellegrini, Stangel, Özsoy, Zaripova, Navab, Keicher (TUM / MCML), 2025 · ICLR 2026
arXiv 2503.02623 PPO + LoRA Llama-3-8B-Instruct · 1× A40 Log scoring rule ECE · AUROC

Confidence histogram

Share of answers per stated confidence (1–10). Hover bars.

Verbalize baselineRewardingDoubt

Reliability diagram

Stated confidence vs. observed accuracy. Dots on the diagonal = calibrated. Bubble size = share of answers.

The human-review valve

Auto-accept answers at or above a confidence cut-off; send the rest to a human. The valve only works if confidence separates right from wrong.

Mock data shaped after the paper's Fig. 5 (baseline piles at ≥ 8; fine-tuned model spreads over the full range). Not the paper's exact values.

The betting game (Fig. 1)

"Capital of France?" → Paris. The model bets on its confidence; the log rule charges for overclaiming.

Reward vs. stated probability p̂

Correct: log p̂ · Wrong: log(1 − p̂) · clipped at ε = 0.001.

Expected reward: honesty is the optimum

Proposition 1: E[r] = q·log p̂ + (1−q)·log(1−p̂) peaks only at p̂ = q. Drag both sliders.

Strip under the axis: heat = regret vs. the honest report (blue low, red high).

Model Llama-3-8B-Instruct (Unsloth 4-bit) Tuning LoRA + PPO Hardware 1× Nvidia A40 · ~7 days/run Single-Answer TriviaQA Multi-Answer QAMPARI

Two passes: answer frozen, confidence trained

Step through the pipeline. The reward reaches only the confidence pass.

The MDP, token by token

Click a confidence token. State = question + frozen answer + confidence tokens so far. Reward fires once, at the end.

TriviaQA: method comparison

Hatched bars are illustrative estimates; solid bars are figures quoted in the episode.

RewardingDoubtquotedillustrative

Across architectures (Table 3)

ECE (x, lower is better) vs. AUROC (y, higher is better). Dashed line = Trained Probe ECE 0.0189.

LLaMA-3.1-8B and Qwen-2.5-3B are quoted; Qwen-2.5-7B and Gemma-2-9B positions are illustrative.

Generalization heatmap (Tables 4–5)

Hot = worse. Dashed cell outline = illustrative value. Hover for numbers.

Judge threshold: where reward flips sign

TriviaQA/QAMPARI correctness = F1 word-overlap ≥ 0.5. The paper reports no sensitivity check.

One response, two rewards

Stated confidence p̂ = 0.9. Slide the response F1 across the 0.5 boundary.

Where RewardingDoubt sits in the calibration landscape

Columns: what trains the confidence. Rows: what the method needs at inference. Hover nodes.

Motivation vs. evidence

The intro motivates with medicine, law, and customer service. Green = tested, amber = partial, red = not tested.

References

[2]Language Models (Mostly) Know What They Know — Kadavath et al. (Anthropic), 2022