AI Post Transformers · Part 3 of 3

Steering Reasoning Models' Cognitive Behaviors at Test-Time

Zhenyu Zhang, Xiaoxia Wu, Zhongzhu Zhou, Qingyang Wu, Yineng Zhang, Pragaash Ponnusamy, Harikaran Subbaraj, Jue Wang, Shuaiwen Leon Song, Ben Athiwaratkun — UT Austin · Together AI · University of Sydney, 2025
⬢ arXiv:2512.24574 🔧 method: CREST 📊 4 architectures · dense + MoE 🔗 companion viz

Closing arc: how CREST finds the attention heads that trigger "non-linear" reasoning detours, denoises them with shared-subspace PCA, and rotates a single hidden state mid-generation to send the model down a 12-step vs. 45-step path to the same answer — plus the head-ratio inconsistency the paper never resolves.

Finding the Heads That Trigger Non-Linear Reasoning

Five stages, one per-head logistic probe, then a shared-subspace PCA that denoises the raw signal by pooling every head in a layer into one covariance matrix.

Pipeline Flow

Per-Head Probe Accuracy (illustrative grid — 8 layers × 12 heads)

Hover a cell to inspect a head's probe accuracy.
low probe accuracy mid high top-10% candidate (outlined)

Pause, Rotate, Fork: The Live Walkthrough

Converting the point (0, 3) to polar coordinates. Same correct answer, wildly different path length, depending on one rotation applied to one hidden state.

Reasoning Trace

Norm-Preserving Rotation (Eq. 5)

original hidden state (dashed) rotated → suppress rotated → amplify

Benchmark Deltas: CREST vs. Vanilla

Bars diverge from zero. Green means CREST wins; orange-flagged bars are the two rows the hosts caught barely moving or moving the wrong way.

← model gets worse · model improves →

Calibrate on Math, Deploy Everywhere

Every probe and steering vector is fit once on MATH500 — 500 problems — then reused with zero recalibration on three domains that share nothing with math but tokens.

Transfer Flow

Headline Transfer Win — GPQA-D, R1-32B

Where Transfer Cracks — Qwen3-30B

Which Ratio Actually Produced the Headline Numbers?

Three sections, three different fractions of heads steered — and the paper never states which one generated Tables 1 through 3.

Head-Selection Ratio, By Section

Also Worth Watching

References

  1. 01Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time — Zhang, Wu, Zhou, Wu, Zhang, Ponnusamy, Subbaraj, Wang, Song, Athiwaratkun (2025)
    arxiv.org/abs/2512.24574
  2. 02Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs — Gandhi, Chakravarthy, Singh, Lile, Goodman (2025)
    Google Scholar
  3. 03Retrieval Head Mechanistically Explains Long-Context Factuality — Wu, Wang, Xiao, Peng, Fu (2024)
    Google Scholar
  4. 04Representation Engineering: A Top-Down Approach to AI Transparency — Zou, Phan, Chen, et al. (2025)
    Google Scholar
  5. 05Reasoning Models Can Be Effective Without Thinking — Ma, He, Snell, Griggs, Min, Zaharia (2025)
    Google Scholar
  6. 06Thoughts Are All Over the Place: On the Underthinking of o1-like LLMs — Wang, Liu, Xu, et al. (2025)
    Google Scholar