Transformer Circuits Source arXiv Search Anthropic · July 6 2026 Interactive SVG companion

Verbalizable Representations and the Global Workspace

This page treats the episode as a mechanics problem, not a philosophy essay: where does a model keep the tiny slice of state it can report on, reason over, and deliberately steer? Hover the visuals. Click the tabs. Most numbers here are synthetic, but shaped to match the paper’s qualitative findings and intervention style.

Paper
Wes Gurnee, Nicholas Sofroniew, Jack Lindsey et al. on a workspace-like verbalizable subspace in Claude-family models.
Mechanism
A Jacobian-based J-space that targets directions causally poised to become language later.
Evidence
Concept swaps, partial and full ablations, plus examples in code, safety, biology, and multi-step reasoning.
Boundary
The claim is about access-like report and control, not phenomenal consciousness or subjective experience.
Mid-layer verbalizable band
A quick layer profile: weak early, strongest in the middle, then more output-tied late.

Overview: a small corridor with outsized influence

Click the paper’s five functional claims and watch the same mid-layer corridor change role.

Verbal report
Layer × concept heatmap
low readiness mid high

Technical deep-dive: from readout to future verbalization

The key move is averaging Jacobians over many contexts so the probe finds what is poised for report, not only what is easy to decode in one prompt.

Token-linked directions across layers
Top directions at the selected layer
Context-averaged coherence by layer
This synthetic curve emphasizes the paper’s main contrast: J-space is most useful where abstract, reusable state is still active, not only right before output.
logit lens tuned lens J-lens

Interventions: if it is causal, edits should reroute behavior

Choose a case. Then compare a clean steering edit against partial and full ablations.

France ↔ China
Normalized intervention metrics

Selectivity and broadcast: many readers, fewer strong writers

The paper’s stronger claim is not only that hidden meaning exists, but that one narrow subspace is especially reusable and especially vulnerable to targeted edits.

Read/write asymmetry sketch
Toggle whether to emphasize who writes into the corridor, who reads from it, or the full broadcast picture.
Task retention under J-space ablation
Capacity view: only a few active slots at once
A stylized slot-by-step heatmap of the kind of unspoken words the paper argues can occupy the workspace during a reasoning episode.

References

Compact pointers to the papers and adjacent threads most relevant to the episode’s visual argument.

Source Paper

Verbalizable Representations Form a Global Workspace in Language Models

Gurnee, Sofroniew, Pearce, Lindsey et al. · 2026
Readout Baseline

Eliciting Latent Predictions from Transformers with the Tuned Lens

Belrose, Ostrovsky, McKinney, Steinhardt et al. · 2023
Feature Ecology

Sparse Autoencoders Find Highly Interpretable Features in Language Models

Cunningham, Ewart, Riggs, Huben, Sharkey · 2023
Scaling Features

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

Templeton, Conerly, Marcus, Lindsey et al. · 2026
Meaning Dynamics

Implicit Representations of Meaning in Neural Language Models

Belinda Z. Li, Maxwell Nye, Jacob Andreas · 2021
Steering

Steering Language Models With Activation Engineering

Turner, Thiergart, Leech, Udell et al. · 2023
Steering Follow-up

Improving Instruction-Following in Language Models through Activation Steering

Stolfo, Balachandran, Yousefi, Horvitz, Nushi · 2024
Faithfulness

Measuring Faithfulness in Chain-of-Thought Reasoning

Lanham, Chen, Radhakrishnan, Perez et al. · 2023
Latent Reasoning

Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding

Chen, Feng, Liu, Yao et al. · 2024
Latent Refinement

Efficient Post-Training Refinement of Latent Reasoning in Large Language Models

Wang, Wang, Ying, Bai et al. · 2025
Content Steering

Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering

Valentino, Kim, Dalal, Zhao, Freitas · 2025
Classic Background

A Neuronal Model of a Global Workspace in Effortful Cognitive Tasks

Dehaene, Kerszberg, Changeux · 1998