AI Post Transformers • Interactive Visualization Companion

Reasoning Theater and Unfaithful Chain-of-Thought

A visual map of the paper’s central tension: internal answer belief can lock in early, while visible reasoning continues as a polished narrative. These diagrams contrast hidden-state probes, forced early answers, and text-only monitors across easy recall-heavy questions and harder multistep science problems.

Extracted arXiv IDs: 2603.05488
Focus: faithfulness, interpretability, dynamic inference
Benchmarks: MMLU-Redux 2.0 vs GPQA-Diamond
Headline Pattern
Early hidden belief, late visible rationale

On easier multiple-choice questions, answer identity becomes decodable from activations before the written reasoning visibly commits.

Adaptive Compute Signal
Up to 80% token savings

Mocked from the paper discussion: strong on easier tasks, smaller on harder tasks where real belief updates continue during inference.

Signal Race

Three readers watch the same reasoning trace. The probe reads hidden states, the forced-answer test cuts the model off early, and the text monitor only inspects what has been written so far.

What This Figure Says

If activations and forced early answers converge before the text monitor sees a justified answer in the visible chain-of-thought, the written trace may be narration after commitment rather than the commitment itself.

Probe Lead 7 steps
Visible Lag 0.34
Confidence Mode Easy
Interpretation Theater

Belief Heatmap

Rows are partial reasoning steps. Columns are answer options. Hot cells indicate the option that hidden activations currently favor. Hover cells to compare early belief concentration across easy and hard regimes.

Monitorability Gap

The paper’s tension appears when the hidden-state answer distribution sharpens before the visible text has made the same commitment legible to a text-only auditor.

Benchmark Split

Easier recall-heavy questions often show earlier commitment and larger early-exit savings. Harder science questions show later peaks and more room for real belief updates during inference.

Four Contrasts

This dashboard compresses the episode’s main contrasts: easy vs hard, hidden vs visible, compute saved vs reasoning lost, and answer identity vs actual correctness.

Step-by-Step Mechanism

Advance the stages to watch how a model can move from latent commitment to visible explanation. The key distinction is between when belief becomes decodable and when language catches up.

Why It Matters

Safety teams want an audit trail. Engineers want efficient test-time compute. This method suggests those goals may depend less on the prose itself and more on whether internal confidence is still moving.

References

Compact source map for the visual claims. arXiv links are shown where explicit IDs are available from the prompt or transcript.

Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought Boppana, Ma, Loeffler, Sarfati, Bigelow, Geiger, Lewis, Merullo, 2026. Core paper behind the hidden-belief vs visible-reasoning comparison.
Chain-of-Thought Prompting Elicits Reasoning in LLMs Wei et al., 2022. The performance boost that made chain-of-thought central enough for faithfulness to matter.
Language Models Don't Always Say What They Think Turpin et al., 2023. Early evidence that smooth explanations can diverge from the actual causal route.
Measuring Faithfulness in Chain-of-Thought Reasoning Lanham et al., 2023. A measurement-oriented framing for when reasoning traces deserve trust.
Reasoning Models Know When They're Right Zhang et al., 2025. Hidden-state confidence as a signal for self-verification and early exit.
AI Post Transformers: Do Language Models Know Their Limits Podcast link on confidence, self-knowledge, and the limits of reading certainty from model outputs.