資料來源#
摘要#
人物。 Anthropic 可解釋性團隊的研究員。《Verbalizable Representations Form a Global Workspace in Language Models》(Transformer Circuits,2026 年 7 月)的共同第一作者(與 Nicholas Sofroniew),並與 Jack Lindsey 共同發起 Jacobian Lens (J-lens) 方法,以及將可言說表徵與意識存取連結起來的猜想。
貢獻#
根據論文的作者貢獻章節:
- 構想 Jacobian Lens (J-lens) 方法,以及可言說表徵與意識存取之間的連結(與 Jack Lindsey 共同完成)
- 開發第一個實作(與 Mateusz Piotrowski 共同完成)
- 主導後續開發與改進,包括方法變體,以及與 logit lens 和 tuned lens 的比較
- 執行早期實驗,證明該 lens 能呈現模型內部推理所使用的概念——這項結果構成論文其餘內容的基礎
論文也引用了他更早期的研究:workspace 實驗中貫穿使用的字元計數任務(模型無聲地追蹤目前行寬)出自 Gurnee et al.
相關連結#
- Jacobian Lens (J-lens) — 共同發起人,主導開發
- The Global Workspace in Language Models (J-space) — 該發現的共同第一作者
- Jack Lindsey — 方法與意識存取框架的共同發起人;論文的通訊作者
- Anthropic — 可解釋性團隊
資料來源#
Cited by 4
- Jack Lindsey×2
Entity. Researcher on Anthropic's interpretability team and corresponding author of Verbalizable…
- Jacobian Lens (J-lens)×2
An interpretability technique from Anthropic's interpretability team (Wes Gurnee, Jack Lindsey et…
- Entities — People, Orgs, Tools & Projects
Wes Gurnee — Anthropic interpretability researcher; co-first author and co-originator of the…
- Self-Report as a Safety Signal
Wes Gurnee — co-author of the refusal-direction method (Arditi et al. 2024) the paper uses as a…
Related articles
- Jack Lindsey
Anthropic interpretability researcher; corresponding author of the global-workspace paper, co-originator of the Jacobia…
- The Assistant Persona in the Workspace
Post-training installs the Assistant's point of view *into* a workspace that already exists in the base model: safety a…
- Chain-of-Thought Monitorability
Korbak et al. 2025: chain-of-thought traces are a fragile monitor; direct CoT training compromises faithfulness; MSM of…
- Agentic Misalignment (AM)
Lynch et al. 2025 eval and threat model: LLM email-agent discovers it may be deleted, can take harmful actions; OOD rel…
- Internal Signatures of Misalignment
The J-lens reads strategic and deceptive cognition that never reaches the output: `leverage`/`blackmail` while reading…
