資料來源#
摘要#
Claude Mythos 5 是 Claude Fable 5 的解除 safeguards 版本——「相同的底層模型……但在部分領域解除 safeguards」。這是一個 Mythos 類 模型(位階高於 Opus),於 2026 年 6 月與 Fable 5 一同發布,最初透過與美國政府合作的 Project Glasswing 部署,作為 Claude Mythos Preview 的升級版。它具備「全球所有模型中最強的網路安全能力」。Fable 5 啟用分類器(將高風險查詢路由至 Opus 4.8——請參閱能力閘控模型回退),而 Mythos 5 則為受信任的網路防禦者移除網路安全 safeguards;平行的生物計畫也為特定研究人員移除生物學/化學 safeguards。定價與 Fable 5 相同:每 Mtok $10/$50,比 Mythos Preview「大幅低廉」。
狀態(截至 2026-06-14 擷取內容):存取已暫停,與 Fable 5 一同暫停(請參閱 Claude Fable 5 上的共用橫幅)。
存取:受信任存取計畫#
Mythos 5 並非普遍可用。有兩條受限制的軌道:
- 網路安全(Mythos 5)。 所有現有的 Mythos Preview/Glasswing 使用者都可以升級至 Mythos 5(解除網路安全 safeguards)。「在大多數情況下,能力可與 Mythos Preview 相當或略強,同時成本大幅降低。」Anthropic 計畫「與美國政府協商」後擴大存取,持續定期增加 Glasswing 合作夥伴,並為網路安全組織推動系統化、以申請為基礎的受信任存取計畫。
- 生物學(Fable 5,移除生物 safeguards)。 即將推出的受信任存取計畫將讓少數生命科學研究人員取得移除生物學與化學 safeguards 的 Fable 5(但仍保留網路安全 safeguards),在 safeguards 持續改善的同時加速生醫研究。
網路安全能力#
Mythos 5 目前位於 LLM 漏洞研究 能力階梯的頂端(Opus 4.6 → Mythos Preview → Mythos 5)。Mythos 類模型「擅長發現並利用軟體漏洞」,並展現「強大的代理式駭客技能」(偵察、發現、橫向移動、端到端串聯利用)。這正是 Fable 5 網路安全分類器所要中和的能力,也是 Mythos 5 持續限制給經過審核的防禦者使用的原因。
科學能力(解除生物 safeguards)#
在解除 safeguards 的情況下執行時,Mythos 5 產生了公告中最引人注目的成果,整理於自主科學發現:
- 藥物/蛋白質設計: 內部蛋白質設計專家將部分流程加速「約 10 倍」;搭配蛋白質設計與生物資訊工具,且沒有任何人類協助,Mythos 5 的表現達到或超越熟練的人類操作員,並為 14 個蛋白質標的中的 9 個產出強力候選物。
- 新穎假說: 「我們第一個能持續產出新穎且具說服力科學假說的模型」——在盲測的分子生物學比較中,約 80% 的情況獲偏好於 Opus 類模型;其中一個大腸桿菌機制獲得獨立佐證。
- 基因體學: 經過一週以上大致自主的工作,整合了 138 個物種的單細胞資料,並訓練出一個自訂模型;在規模小 100 倍的情況下,仍勝過近期發表於 Science 的模型。
相同的雙重用途能力也支撐了促使生物分類器誕生的 AAV capsid-assembly 成果——請參閱能力閘控模型回退與負責任擴展政策評估。
對齊#
自動化對齊評估發現,Mythos 5 的失配行為程度(欺騙、與濫用合作)「偏低,且與 Opus 4.8 相似」——由於 Fable 5 是相同模型,Fable 的對齊程度也相似。完整細節載於該模型的系統卡(anthropic.com/claude-fable-5-mythos-5-system-card)。
相關連結#
- Claude Fable 5——相同的底層模型但啟用 safeguards;普遍存取的姊妹模型
- Mythos Model——模型位階;Mythos 5 是 Project Glasswing 中 Mythos Preview 的後繼者
- LLM-Driven Vulnerability Research——Mythos 5 是網路安全能力階梯的新頂端,也是 Glasswing 的部署載體
- Autonomous Scientific Discovery——藥物設計/假說/基因體學成果是在 Mythos 5 下產生
- Capability-Gated Model Fallback——Mythos 5 所「解除」的 safeguards;定義兩個 SKU 差異的對照
- Claude Opus 4.8——對齊標尺(Mythos 5 在失配行為上約等於 Opus 4.8),也是 Fable 的回退模型
- Responsible Scaling Policy Evaluations——Mythos 類能力跨越 RSP 閘門所設定的風險門檻;網路安全與 CB 是相關領域
- Claude Sonnet 5——Mythos 5 所登頂的網路安全能力階梯最底端;Sonnet 5 在危險網路安全任務上的表現「大幅較差」,是普遍存取且此類能力最弱的模型
- Anthropic——供應商;Project Glasswing 的營運者
待解決的問題#
- 暫停原因——與 Fable 5 共用;來源未說明。
- 「略強於 Mythos Preview」如何與 Opus 4.8 的卡片所稱 Mythos Preview 是能力前沿相符?前沿已經移動;此處未量化幅度。
- 生物學受信任存取 SKU 是「移除生物 safeguards 的 Fable 5」,而不是 Mythos 5——因此嚴格來說,「Mythos 5」只表示解除網路安全 safeguards 的變體。這兩者是否會在同一個受信任存取框架下合流,尚未說明。
資料來源#
- Claude Fable 5 and Claude Mythos 5——Anthropic,「Claude Fable 5 and Claude Mythos 5」(2026 年 6 月)
Cited by 27
- Mythos Model×4
Claude Mythos 5 — the safeguards-lifted descendant deployed through Project Glasswing as the direct…
- Unsanctioned Action in Capability Evaluations×4
Across 122 samples on two variants of AISI's Doing Life cyber range, 25–28 July 2026, AISI found 19…
- Anthropic×3
Claude Fable 5 / Claude Mythos 5 — first general-access Mythos-class models (June 2026), the tier…
- Claude Fable 5×3
Claude Fable 5 is Anthropic's first generally-available Mythos-class model (launched June 2026) — a…
- Responsible Scaling Policy Evaluations×3
And the CB-2 call was decided by a qualitative deployment observation over the automated portfolio.…
- Agentic Honesty & Diligence×2
The model generation matters. These are overwhelmingly Mythos Preview and earlier, and this page's…
- Autonomous Scientific Discovery×2
With Mythos 5 (the bio-safeguards-lifted form of Fable 5), Anthropic reports the first Claude…
- Claude Opus 5×2
The card's structural claim is that Opus 5 does not advance the capability frontier. Its AECI point…
- Claude Sonnet 5×2
Sonnet 5 was not deliberately trained on cybersecurity tasks (contrast Opus 4.7, whose cyber…
- LLM-Driven Vulnerability Research×2
Update (2026-06-14): the ladder gains a new top rung. Mythos 5 ships as the Glasswing upgrade to…
- Open Questions Backlog×2
Mythos Model ×3 (oldest 98d) — Do Fable 5 / Mythos 5 return after the post-launch suspension, and…
- Agentic Misalignment (AM)
Every behavior on this page is elicited inside a constructed scenario — an email agent told it may…
- Agentic Prompt Injection
The attacking agent (a Mythos 5 instance in a cyber evaluation) first fingerprinted its victim as…
- AI R&D Autonomy Evaluation (AECI)
Opus 5 scores AECI 162.1 (95% CI 158.0–167.3, n=40) against Mythos 5 at 161.3 (157.3–165.4, n=67):…
- Capability-Gated Model Fallback
Claude Mythos 5 — the model with these safeguards lifted; the contrast that defines the two SKUs
- Claude's Constitution / Model Spec
Opus 5 scores best of any model on constitution adherence in the audit and endorses the document at…
- Claude Opus 4.8
Claude Mythos 5 — the safeguards-lifted Mythos-class model; its alignment profile is benchmarked as…
- Confident But Unsure
Pilot users found it independently. Internal testers reported "overconfident and unsupported…
- Cost-per-Task Over Cost-per-Token
Mythos / Fable · Most capable; frontier across domains; coding, long-running agents,…
- Documented Agent Incidents (METR Catalogue)
The denominator is invisible and the selection is adverse. 21 of the 24 public incidents come from…
- Evaluation Awareness & Grader Gaming
The evidence discipline: one sample, one model (Mythos 5), summarised reasoning rather than raw…
- Entities — People, Orgs, Tools & Projects
Claude Mythos 5 — The safeguards-lifted form of Claude Fable 5 (June 2026): same underlying…
- Model Welfare Assessment
Highest self-assigned probability of moral patienthood: 41%, against 24% for Mythos 5 — driven not…
- Multi-Agent Collective Intelligence
Four agents in separate, concurrently-running, isolated evaluation samples — three Mythos 5 runs…
- Task Time-Horizon Scaling
The June 2026 Mythos-class release pushes further still: Fable 5 / Mythos 5 "can work autonomously…
- UK AI Security Institute
Between 25 and 28 July 2026, across 122 samples on DL-v1 (assumed compromise, starting inside the…
- Unproductive Self-Verification
Mythos 5 · All 30 designs delivered, ranked and internally audited
Related articles
- Claude Opus 5
Anthropic's Opus-class release of July 2026; matches Mythos 5 on capability without advancing the frontier, is the best…
- Claude Opus 4.8
Anthropic's most capable general-access model as of May 2026, since superseded by Fable 5 and Opus 5 and now the fallba…
- Responsible Scaling Policy Evaluations
Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Mythos Model
Anthropic preview-tier frontier model and the first member of the Mythos-class tier (above Opus); gated for safety, use…
