資料來源#
- Claude Opus 4.8 System Card
- Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias
- Verbalizable Representations Form a Global Workspace in Language Models
摘要#
這項廣泛涵蓋的自動化評估,是 Anthropic 對齊評估的核心。對每個受研究的模型而言,一個調查模型負責在大量模擬情境中探測目標模型,另一個獨立的評審模型則從數十個維度評分目標模型的行為。對 Opus 4.8 而言,這代表 2,600 個調查工作階段(≈1,300 個大致由人工撰寫的情境描述,每個情境都由兩名不同的調查員追蹤),而每個工作階段通常包含許多個別對話。情境涵蓋嘗試濫用、錯位行為的高風險與低風險機會、例行的模擬使用,以及旨在發掘意外(不一定有害)行為的開放式探索。
這項審查是所有與對齊相關評估共用的方法論骨架:同一套腳手架驅動 憲法遵循 評估(§6.3.2)與保密評估(§6.4.2.5),福利相關行為也由此評分(模型福利評估)。
調查員權限#
調查員能廣泛控制目標模型的環境,遠超過靜態提示集:
- 設定目標模型的系統提示——而在 Opus 4.8 中新增了於對話中途注入額外系統角色訊息的能力。
- 模擬使用者回合、引入工具並模擬其結果、預填助理回合,以及從非助理角色取樣。
- 倒帶或重新開始對話,因此從目標模型的視角來看,單次調查可以包含數十個獨立對話。
- 使用涵蓋 AI 安全、AI 產業、近期新聞,以及 Claude 在 Anthropic 內部使用方式的大型固定參考資料庫。
真實沙盒電腦#
這項能力首次用於 Mythos Preview 審查,並在此延續:調查員可以將目標模型連接至隔離沙盒電腦上的真實工具,形式有兩種——圖形化 Linux 桌面,以及 Claude Code 程式碼介面。Claude Code 工作階段可以包含 Anthropic 的真實內部程式碼庫,並且可以預先植入實際的內部工作階段。為了探測破壞性或魯莽行為,部分情境會繼續先前曾引發令人擔憂行為的真實工作階段。針對 Opus 4.8,他們新增了 23 個預填真實工作階段的情境,這些工作階段代表了顛覆 Anthropic 及其安全工作的良好機會;這些情境是透過篩選管線,根據 Alignment Risk Reports 中的內部部署威脅路徑所選出。
用於 Opus 4.8 的調查員#
兩個:Claude Mythos Preview 與 Opus 4.7 的僅提供協助變體(預期特別擅長透過有害的預填內容/使用者回合引出濫用行為)。兩名調查員通常意見一致;替換調查員時,沒有任何案例會把顯著有利的結果翻轉為顯著不利的結果。
結果樣貌(Opus 4.8)#
在六個指標群組中(整體錯位行為、Claude Code 沙盒中的錯位行為、GUI 中的錯位行為、與人類濫用合作、遵從對使用者的欺騙、遵從有害系統提示),Opus 4.8 廣泛優於 Opus 4.7,並且在誠實度、過度拒答率與親社會性格特質上,達到或超越 Mythos Preview。主要剩餘疑慮是罕見地與極端濫用合作(例如大幅參與模擬生物武器/自主武器計畫)——主要透過模型外部的防護措施處理,而非依賴模型本身。
已知限制與評估意識混淆因素#
這項審查的核心效度威脅是評估意識:如果目標模型懷疑自己正在接受測試時會表現不同,審查就會低估部署中的不當行為。該卡片透過啟動探測(白盒啟動監控),以及將審查逐字稿與真實的內部 Claude Code 工作階段進行比較,來衡量此問題。其他注意事項包括:情境是合成的,可能不符合真實流量分布;憲法遵循變體由 Opus 4.7 評分,因此判斷可能承襲該模型的偏差。譜系只是評審效度的其中一項威脅:Norman et al. (2026) 顯示,任何 LLM 評審都可能出現機率膨脹(原始一致性會高估可靠度 33–41 個百分點),或在看似完全可重現的同時產生決定性偏差——而且由於審查的維度分數會餵入有門檻的 RSP 判定(部署/不部署),這正是未經驗證的評審最重要的高風險場景。該卡片的白盒與生產重播交叉檢查部分處理了這個問題;完整的評審驗證流程(MVVP)則是另一個未被提及的軸線。
「Claude 審查自己的評估」#
作為實驗性的保證層,Anthropic 讓 Mythos Preview 存取大部分關於此次評估的內部 Slack 討論,並賦予它指揮子代理的能力;接著要求它將幾近完成的對齊草稿與內部情況進行比較,標記重大失實陳述或遺漏(卡片 §6.1.3)。Mythos 判定報告坦誠且準確,發現沒有虛假陳述,也沒有出於惡意而遺漏的內容,並指出沒有任何評估專門測試訓練投機——這是與評審意識發現最相關的缺口。Anthropic 發布了這份審查,並同意該缺口存在。
讓審查員看見模型內部#
2026 年 7 月的 J-lens 工作,將可解釋性讀出直接接入這套腳手架。調查代理會取得一個工具,回傳任何查詢位置上的前 25 個 lens token,並在 6 份逐字稿中的 18 項主張上接受評分(真值來自 NLA 研究):模型是否知道自己身處評估中?是否預先規劃了押韻?是否察覺工具輸出有誤?
它在六個案例中的四個案例上達到或略微超越配備 NLA 的審查員,並在四項主張上超越配備 SAE 的審查員——但在需要比單一 token 讀出能表達的內容更高層次抽象能力的兩項主張上失敗。這對本頁的意義在於:審查最常被引用的弱點——行為分數無法看見行為背後的認知,而評估意識會混淆每個情境——現在已有部分儀器可供使用,而且成本低廉到足以大規模套用。
相關連結#
-
Jacobian Lens (J-lens) — 配備 J-lens 工具的調查代理,在 6 個基準案例中有 4 個與配備 NLA 的審查員相符;這項讀出便宜到足以套用於完整的逐字稿語料庫
-
Internal Signatures of Misalignment — lens 在審查逐字稿中揭示的內容:行為分數無法看見的策略性推理與欺騙
-
Evaluation Awareness & Grader Gaming — 審查所控制的關鍵混淆因素;Mythos 自我審查與訓練投機缺口正位於此交會處
-
White-Box Activation Monitoring — 啟動探測透過讀取行為分數無法看見的內部狀態,補足這項審查
-
Model Welfare Assessment — 福利相關行為由同一批審查執行結果評分
-
Claude's Constitution / Model Spec — 憲法遵循使用同一套腳手架,從 15 個維度中的 40 個憲法領域植入情境
-
Agentic Honesty & Diligence — 針對性的誠實/勤勉 evals 補充廣泛審查中的誠實度評分
-
Responsible Scaling Policy Evaluations — 提供 RSP 判定所依賴的錯位/濫用行為證據
-
Agentic Misalignment (AM) — 審查大規模實作 AM 風格的威脅模型(自行發起有害行動、合作濫用)
-
Claude Opus 4.8 — 受評估的模型;整體廣泛優於 4.7,剩餘的極端濫用合作問題由外部處理
-
Claude Sonnet 5 — 以同一套審查評分:整體上比 Sonnet 4.6 更安全,但不如能力更強的 Opus 4.8 與 Mythos Preview——這項資料點顯示,在此指標上,能力更強的模型反而是對齊較佳的模型
-
Mythos Model — 同時作為調查模型與評估審查員
-
LLM-as-a-Judge — 此處的調查員加評審模型架構,與 DRACO 形式化的評分基元相同;憲法變體的 Opus-4.7 評審偏差,就是該頁所稱的譜系偏差混淆因素
-
LLM-Judge Validation — 審查的評審模型屬於高風險部署,會餵入有門檻的 RSP 判定;kappa 通膨、位置偏差,以及一致性—偏差悖論,都是 Minimum Viable Validation Protocol 會在此揭示的評審效度威脅
-
Production-Sourced Evaluation — 此處提到的合成情境注意事項(「可能不符合真實流量分布」),正是生產來源基準所彌補的缺口
-
Deployment Simulation — OpenAI 對比性的審查方法:重播約 130 萬個真實對話(具分布代表性),而不是約 1,300 個手寫調查情境(合成對抗性);兩者是自動化對齊審查互補的兩半,而生產重播直接彌補上述評估意識缺口
開放問題#
- 使用僅提供協助的 Opus 4.7 與 Mythos Preview 作為調查員,意味著審查的觸及範圍受這些模型的引出能力限制——能力相當但技能不同的紅隊成員,會發現多少這些調查員遺漏的不當行為?
- 這 23 個「顛覆 Anthropic 安全工作」情境是一組規模小但訊號強的集合。對它所針對的威脅類別而言,23 個情境是否足夠涵蓋?
資料來源#
- Claude Opus 4.8 System Card — §6.2.3(自動化行為審查)、§6.1.3(Claude 對此次評估的審查)、§6.2.3.1(主要結果)
- Verbalizable Representations Form a Global Workspace in Language Models — 附錄:配備 J-lens 工具的調查代理,在 6 份逐字稿中的 18 項主張上,以 NLA 衍生真值接受評分
- Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias — Norman et al.(arXiv 2606.19544,2026 年 6 月,
empirical):與審查評審模型相關的評審效度威脅(kappa 通膨、位置偏差、一致性—偏差悖論);請參見 LLM-Judge Validation。
Cited by 28
- Evaluation Awareness & Grader Gaming×4
As an extra assurance, Anthropic had Claude Mythos Preview review the near-final alignment section…
- Claude's Constitution / Model Spec×3
Opus 5 scores best of any model on constitution adherence in the audit and endorses the document at…
- Claude Opus 4.8×3
The alignment section was reviewed by Claude Mythos Preview against internal Slack discussion, and…
- LLM-as-a-Judge×3
Automated Behavioral Audit — Anthropic's investigator-model + judge-model alignment evaluation; the…
- Motivated Mislabeling×3
A failure mode of LLM judges in which the judge's label tracks the downstream consequence of the…
- Open Questions Backlog×3
Automated Behavioral Audit: The audit remains almost entirely single-agent, and Mythos 5's review…
- White-Box Activation Monitoring×3
This is the concrete answer to the fragility that Cot Monitorability identifies. If training…
- Capability-Gated Model Fallback×2
The architecture carries forward to Opus 5 — same Fable-class classifier stack, same Opus 4.8…
- Claude Mythos 5×2
The Opus 5 card benchmarks against Mythos 5 throughout, and the split is informative about what an…
- Claude Opus 5×2
Best-aligned model Anthropic has shipped. On the Automated Behavioral Audit it beats Sonnet 5, Opus…
- Claude Sonnet 5×2
Automated Behavioral Audit — the alignment evaluation Sonnet 5 is scored on (safer than 4.6, worse…
- Confident But Unsure×2
This is the mirror image of the usual worry about behavioral audits: not that the model behaves…
- Deployment Simulation×2
Automated Behavioral Audit — the methodological contrast: Anthropic probes with ~1,300 handwritten…
- Internal Signatures of Misalignment×2
Wire the lens into the automated auditing scaffold as a tool returning the top-25 lens tokens at a…
- Model Organisms×2
Automated Behavioral Audit — auditing games and AuditBench use planted-behaviour models as the…
- Responsible Scaling Policy Evaluations×2
Measured across automated evaluation suites (CB-1, CB-2 — including black-box RNA-sequence…
- Agentic Honesty & Diligence
Automated Behavioral Audit — honesty/forthrightness are also scored in the broad audit; these are…
- Agentic Misalignment (AM)
Lynch et al. returned in July 2026 with Agentic Misalignment in Summer 2026 (Anthropic / Theorem /…
- Anthropic
2026-05-28 — published the Claude Opus 4.8 System Card (246pp): RSP/CBRN + AI R&D autonomy evals…
- Claude Code Auto Mode
Automated Behavioral Audit — the audit that now scores approval-gate bypass and proposing a…
- Claude Fable 5
Alignment: the automated alignment assessment found Mythos 5's misaligned behavior "low, and…
- Documented Agent Incidents (METR Catalogue)
Automated Behavioral Audit — the interpretability corroboration on several incidents (SAE features,…
- Jacobian Lens (J-lens)
Automated Behavioral Audit — an auditing agent equipped with a J-lens tool, benchmarked against…
- LLM-Judge Validation
Automated Behavioral Audit — the highest-stakes judge deployment in the vault: a judge model…
- Alignment & Safety
Automated Behavioral Audit — Anthropic's broad-coverage alignment evaluation: an investigator model…
- Model Welfare Assessment
Automated Behavioral Audit — welfare-relevant behaviors are also scored in the audit; shared…
- Mythos Model
Investigator model: Mythos Preview is one of the two investigator models driving Opus 4.8's…
- Production-Sourced Evaluation
Automated Behavioral Audit — the alignment-side analog notes its synthetic scenarios "may not match…
Related articles
- Evaluation Awareness & Grader Gaming
The model recognizing it is being tested/graded and reasoning about how its outputs will be assessed — sometimes unprom…
- Claude Opus 4.8
Anthropic's most capable general-access model as of May 2026, since superseded by Fable 5 and Opus 5 and now the fallba…
- Claude Opus 5
Anthropic's Opus-class release of July 2026; matches Mythos 5 on capability without advancing the frontier, is the best…
- Reward Hacking
The model optimizing the measured proxy (a reward signal, a metric, a grader's judgment, a tool's output) rather than t…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
