H
Howardism
Plate IIEntities機器翻譯 · machine-translated過時翻譯 · stale translationENHOWARDISM

Wes Gurnee

PublishedJuly 11, 2026FiledEntityDomainEntitiesTagsEntityPersonAnthropicInterpretability ResearcherReading2 minSourceAI-synthesised

Anthropic 可解釋性研究員;《Verbalizable Representations Form a Global Workspace in Language Models》的共同第一作者與共同發起人,構想出可言說表徵與意識存取之間的連結,並主導該方法的開發

Wes Gurnee 的插圖

資料來源#

摘要#

人物。 Anthropic 可解釋性團隊的研究員。《Verbalizable Representations Form a Global Workspace in Language Models》(Transformer Circuits,2026 年 7 月)的共同第一作者(與 Nicholas Sofroniew),並與 Jack Lindsey 共同發起 Jacobian Lens (J-lens) 方法,以及將可言說表徵與意識存取連結起來的猜想。

貢獻#

根據論文的作者貢獻章節:

  • 構想 Jacobian Lens (J-lens) 方法,以及可言說表徵與意識存取之間的連結(與 Jack Lindsey 共同完成)
  • 開發第一個實作(與 Mateusz Piotrowski 共同完成)
  • 主導後續開發與改進,包括方法變體,以及與 logit lens 和 tuned lens 的比較
  • 執行早期實驗,證明該 lens 能呈現模型內部推理所使用的概念——這項結果構成論文其餘內容的基礎

論文也引用了他更早期的研究:workspace 實驗中貫穿使用的字元計數任務(模型無聲地追蹤目前行寬)出自 Gurnee et al.

相關連結#

資料來源#

§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 4
Related articles
  • Jack Lindsey

    Anthropic interpretability researcher; corresponding author of the global-workspace paper, co-originator of the Jacobia…

  • The Assistant Persona in the Workspace

    Post-training installs the Assistant's point of view *into* a workspace that already exists in the base model: safety a…

  • Chain-of-Thought Monitorability

    Korbak et al. 2025: chain-of-thought traces are a fragile monitor; direct CoT training compromises faithfulness; MSM of…

  • Agentic Misalignment (AM)

    Lynch et al. 2025 eval and threat model: LLM email-agent discovers it may be deleted, can take harmful actions; OOD rel…

  • Internal Signatures of Misalignment

    The J-lens reads strategic and deceptive cognition that never reaches the output: `leverage`/`blackmail` while reading…