Sources#
Summary#
Entity. Researcher on Anthropic's interpretability team. Co-first author (with Nicholas Sofroniew) of Verbalizable Representations Form a Global Workspace in Language Models (Transformer Circuits, July 2026), and — with Jack Lindsey — the originator of both the Jacobian Lens (J-lens) method and the conjecture that connects verbalizable representations to conscious access.
Contributions#
Per the paper's author-contributions section:
- Conceived of the Jacobian Lens (J-lens) method, and of the connection between verbalizable representations and conscious access (with Jack Lindsey)
- Developed the first implementation (with Mateusz Piotrowski)
- Led subsequent development and refinement, including methodological variants and the comparisons against the logit and tuned lenses
- Ran the early experiments demonstrating that the lens surfaces concepts used in the model's internal reasoning — the result the rest of the paper is built on
Earlier work of his is cited within the paper itself: the character-counting task used throughout the workspace experiments (the model silently tracking a running line width) is from Gurnee et al.
Connections#
- Jacobian Lens (J-lens) — co-originator, led development
- The Global Workspace in Language Models (J-space) — co-first author of the finding
- Jack Lindsey — co-originator of the method and the conscious-access framing; corresponding author on the paper
- Anthropic — interpretability team
Sources#
- Verbalizable Representations Form a Global Workspace in Language Models — Author Contributions
Cited by 4
- Jack Lindsey×2
Entity. Researcher on Anthropic's interpretability team and corresponding author of Verbalizable…
- Jacobian Lens (J-lens)×2
An interpretability technique from Anthropic's interpretability team (Wes Gurnee, Jack Lindsey et…
- Entities — People, Orgs, Tools & Projects
Wes Gurnee — Anthropic interpretability researcher; co-first author and co-originator of the…
- Self-Report as a Safety Signal
Wes Gurnee — co-author of the refusal-direction method (Arditi et al. 2024) the paper uses as a…
Related articles
- Jack Lindsey
Anthropic interpretability researcher; corresponding author of the global-workspace paper, co-originator of the Jacobia…
- The Assistant Persona in the Workspace
Post-training installs the Assistant's point of view *into* a workspace that already exists in the base model: safety a…
- Chain-of-Thought Monitorability
Korbak et al. 2025: chain-of-thought traces are a fragile monitor; direct CoT training compromises faithfulness; MSM of…
- Agentic Misalignment (AM)
Lynch et al. 2025 eval and threat model: LLM email-agent discovers it may be deleted, can take harmful actions; OOD rel…
- Internal Signatures of Misalignment
The J-lens reads strategic and deceptive cognition that never reaches the output: `leverage`/`blackmail` while reading…
