Confidence histogram
Share of answers per stated confidence (1–10). Hover bars.
Reliability diagram
Stated confidence vs. observed accuracy. Dots on the diagonal = calibrated. Bubble size = share of answers.
The human-review valve
Auto-accept answers at or above a confidence cut-off; send the rest to a human. The valve only works if confidence separates right from wrong.
Mock data shaped after the paper's Fig. 5 (baseline piles at ≥ 8; fine-tuned model spreads over the full range). Not the paper's exact values.
The betting game (Fig. 1)
"Capital of France?" → Paris. The model bets on its confidence; the log rule charges for overclaiming.
Reward vs. stated probability p̂
Correct: log p̂ · Wrong: log(1 − p̂) · clipped at ε = 0.001.
Expected reward: honesty is the optimum
Proposition 1: E[r] = q·log p̂ + (1−q)·log(1−p̂) peaks only at p̂ = q. Drag both sliders.
Strip under the axis: heat = regret vs. the honest report (blue low, red high).
Two passes: answer frozen, confidence trained
Step through the pipeline. The reward reaches only the confidence pass.
The MDP, token by token
Click a confidence token. State = question + frozen answer + confidence tokens so far. Reward fires once, at the end.
TriviaQA: method comparison
Hatched bars are illustrative estimates; solid bars are figures quoted in the episode.
Across architectures (Table 3)
ECE (x, lower is better) vs. AUROC (y, higher is better). Dashed line = Trained Probe ECE 0.0189.
LLaMA-3.1-8B and Qwen-2.5-3B are quoted; Qwen-2.5-7B and Gemma-2-9B positions are illustrative.
Generalization heatmap (Tables 4–5)
Hot = worse. Dashed cell outline = illustrative value. Hover for numbers.
Judge threshold: where reward flips sign
TriviaQA/QAMPARI correctness = F1 word-overlap ≥ 0.5. The paper reports no sensitivity check.
One response, two rewards
Stated confidence p̂ = 0.9. Slide the response F1 across the 0.5 boundary.
Where RewardingDoubt sits in the calibration landscape
Columns: what trains the confidence. Rows: what the method needs at inference. Hover nodes.
Motivation vs. evidence
The intro motivates with medicine, law, and customer service. Green = tested, amber = partial, red = not tested.