A visual map of the paper’s central tension: internal answer belief can lock in early, while visible reasoning continues as a polished narrative. These diagrams contrast hidden-state probes, forced early answers, and text-only monitors across easy recall-heavy questions and harder multistep science problems.
On easier multiple-choice questions, answer identity becomes decodable from activations before the written reasoning visibly commits.
Mocked from the paper discussion: strong on easier tasks, smaller on harder tasks where real belief updates continue during inference.
Three readers watch the same reasoning trace. The probe reads hidden states, the forced-answer test cuts the model off early, and the text monitor only inspects what has been written so far.
If activations and forced early answers converge before the text monitor sees a justified answer in the visible chain-of-thought, the written trace may be narration after commitment rather than the commitment itself.
Rows are partial reasoning steps. Columns are answer options. Hot cells indicate the option that hidden activations currently favor. Hover cells to compare early belief concentration across easy and hard regimes.
The paper’s tension appears when the hidden-state answer distribution sharpens before the visible text has made the same commitment legible to a text-only auditor.
Easier recall-heavy questions often show earlier commitment and larger early-exit savings. Harder science questions show later peaks and more room for real belief updates during inference.
This dashboard compresses the episode’s main contrasts: easy vs hard, hidden vs visible, compute saved vs reasoning lost, and answer identity vs actual correctness.
Advance the stages to watch how a model can move from latent commitment to visible explanation. The key distinction is between when belief becomes decodable and when language catches up.
Safety teams want an audit trail. Engineers want efficient test-time compute. This method suggests those goals may depend less on the prose itself and more on whether internal confidence is still moving.
Compact source map for the visual claims. arXiv links are shown where explicit IDs are available from the prompt or transcript.