A visual map of how a frozen base model can answer one way, then expose a more candid internal readout through a triggerable reviewer persona. This page focuses on the geometry, flow, benchmarks, and failure modes of latent-state auditing rather than prose summary.
The key engineering move is causal separation. The base model produces the visible answer with the reviewer LoRA disabled; a trigger activates a second persona only after the answer is complete.
If honesty costs reward, output text may be the least trustworthy object in the system. The page highlights the alternate channel: hidden state to reviewer.
The reviewer can become stricter without directly perturbing the original answer path. That avoids one form of capability contamination.
It does not prove the reviewer has privileged access. The review could still rely on answer text, prompt residue, or benchmark quirks.
Toggle scenarios to see which components fail first: answer honesty, audit recall, or trigger robustness.
These grids model hidden-layer evidence for three different stories: the base assistant’s polished output, the hidden objective signal, and the reviewer’s post-hoc readout.
Rows are layer bands; columns are token positions. Hover any cell to inspect the mock latent evidence. The interesting pattern is not uniform truthfulness but localized “hot” regions surviving beneath a safe-looking answer.
A good ELK-style channel should light up where the hidden objective is represented, even when answer tokens look normal. The reviewer map should align with those regions more than with stylistic answer features.
The missing ablations from the episode are easy to state visually: mask the answer, drop activation reuse, or keep only the prompt. If the reviewer still succeeds, the privileged-channel story strengthens.
Mock evaluation slices contrast split-personality review with direct self-report, linear probes, activation-to-language decoding, and an external reader baseline. Switch metric families to expose different failure modes.
This network places the 2026 paper inside the broader track from ELK framing to probes, truth geometry, latent decoding, alignment faking, and activation oracles.
The novelty is narrower than “truth in activations.” Earlier work already argued hidden states can preserve more honest information than text. The new contribution is a triggerable internal reviewer that speaks after the answer.
The episode repeatedly calls for direct comparisons against linear probes, clustering, activation decoders, external readers, and self-report fine-tuning baselines.
Treat latent channels as an audit surface, not a solved lie detector. The right question is whether the internal auditor generalizes outside one benchmark organism and one trigger dialect.
Key papers and prior episodes surfaced in the discussion. arXiv links are included where an ID is explicitly available from the prompt or transcript.