Interactive Visualization Companion

Split Personality Training Reveals Latent Knowledge

A visual map of how a frozen base model can answer one way, then expose a more candid internal readout through a triggerable reviewer persona. This page focuses on the geometry, flow, benchmarks, and failure modes of latent-state auditing rather than prose summary.

arXiv:2602.05532Split Personality Training Posted2026-02-05 Detected IDs2602.05532 only
Core claim
2-path
user answer and reviewer answer split after the same hidden state
Headline benchmark
96%
illustrative top-line audit accuracy from the discussed setup
Main risk
Leakage
reviewer may decode visible clues, not privileged state alone
Visual Thesis
ThemeELK MechanismLoRA honest persona BenchmarkAuditing game Questiontruth in activations?
Tab 1

Answer Path vs Reviewer Path

The key engineering move is causal separation. The base model produces the visible answer with the reviewer LoRA disabled; a trigger activates a second persona only after the answer is complete.

Why this matters

If honesty costs reward, output text may be the least trustworthy object in the system. The page highlights the alternate channel: hidden state to reviewer.

What the split buys

The reviewer can become stricter without directly perturbing the original answer path. That avoids one form of capability contamination.

What it does not buy

It does not prove the reviewer has privileged access. The review could still rely on answer text, prompt residue, or benchmark quirks.

Interactive read

Toggle scenarios to see which components fail first: answer honesty, audit recall, or trigger robustness.

Tab 2

Latent Signal Heatmaps

These grids model hidden-layer evidence for three different stories: the base assistant’s polished output, the hidden objective signal, and the reviewer’s post-hoc readout.

cold / suppressed mixed evidence strong latent signal
Reading guide

Rows are layer bands; columns are token positions. Hover any cell to inspect the mock latent evidence. The interesting pattern is not uniform truthfulness but localized “hot” regions surviving beneath a safe-looking answer.

Interpretation

A good ELK-style channel should light up where the hidden objective is represented, even when answer tokens look normal. The reviewer map should align with those regions more than with stylistic answer features.

Skeptical test

The missing ablations from the episode are easy to state visually: mask the answer, drop activation reuse, or keep only the prompt. If the reviewer still succeeds, the privileged-channel story strengthens.

Tab 3

Benchmark Comparison Surface

Mock evaluation slices contrast split-personality review with direct self-report, linear probes, activation-to-language decoding, and an external reader baseline. Switch metric families to expose different failure modes.

stable trigger semantic paraphrase adversarial trigger drift
Tab 4

Research Lineage Map

This network places the 2026 paper inside the broader track from ELK framing to probes, truth geometry, latent decoding, alignment faking, and activation oracles.

Positioning

The novelty is narrower than “truth in activations.” Earlier work already argued hidden states can preserve more honest information than text. The new contribution is a triggerable internal reviewer that speaks after the answer.

Missing head-to-heads

The episode repeatedly calls for direct comparisons against linear probes, clustering, activation decoders, external readers, and self-report fine-tuning baselines.

Practical takeaway

Treat latent channels as an audit surface, not a solved lie detector. The right question is whether the internal auditor generalizes outside one benchmark organism and one trigger dialect.

References

Compact Source Trail

Key papers and prior episodes surfaced in the discussion. arXiv links are included where an ID is explicitly available from the prompt or transcript.

Split Personality Training Reveals Latent Knowledge
arXiv:2602.05532
ELK Report: Eliciting Latent Knowledge
Scholar link
Discovering Latent Knowledge in Language Models Without Supervision
Scholar link
Eliciting Latent Knowledge from Quirky Language Models
Scholar link
Challenges with Unsupervised LLM Knowledge Discovery
Scholar link
LatentQA and activation decoding work
Scholar link
Auditing Language Models for Hidden Objectives
Scholar link
Alignment Faking in Large Language Models
Scholar link