資料來源#
摘要#
METR(Model Evaluation & Threat Research)是一個評估前沿 AI 能力的獨立組織,最廣為人知的是其 時間範圍測量:模型能可靠地獨立完成任務的長度。其資料是 Anthropic Institute〈When AI builds itself〉文章所依據的外部基準核心,也構成了本 Wiki 任務時間範圍擴展頁面的基礎。
它的工作#
- 時間範圍。 報告模型在一組任務中達到 50% 可靠度時所對應的任務持續時間(在 80% 可靠度下也呈現相同趨勢)。METR 的主要發現是,這個範圍大約每四個月翻倍,高於早期約七個月翻倍的速度——這是能力正在加速而不只是提升的量化證據。
- 前沿模型的長任務測量。 METR 發現,Claude Mythos Preview 能工作「至少」16 小時,並且「已達 [METR] 在不新增任務的情況下所能測量的上限」——也就是說,前沿模型已開始超越基準測試自身的天花板。
- 獨立的第三方訊號。 由於 METR 位於各實驗室之外,其數據可作為 Anthropic 內部加速主張的外部佐證,例如約 8 倍的程式碼產出量(AI 加速 AI 開發)。
- 被其他評估者重複使用。 UK AI Security Institute 於 2026 年 7 月進行的 test-time-compute 研究,使用 METR 的 211 項軟體工程任務集(以及 AISI 自有的網路安全任務),並延伸了時間範圍的框架:研究顯示,時間範圍及其翻倍速度都取決於預算(見任務時間範圍擴展)。
相關連結#
- 任務時間範圍擴展 — 建立在 METR 時間範圍指標上的概念頁面
- AI 加速 AI 開發 — METR 的外部趨勢線為 Anthropic 的內部產出量證據提供佐證
- 遞迴自我改進 — 將翻倍曲線外推後,它成為 RSI 比預期更早到來的量化依據
- Mythos Model — METR 評估為「至少 16 小時」、超出其目前測量上限的模型
- UK AI Security Institute — 重複使用 METR 任務集,並顯示時間範圍指標取決於預算的姊妹獨立評估者
- Researcher Uplift from Code Output — METR 的 Thomas Kwa 於 2026 年 7 月撰寫的建模備註,將 Anthropic 的 8 倍程式碼數據轉換為約 2.5 倍的序列研究人員提升;其關於冗長程度及主觀感受與實際加速差異的注意事項,依據的是 METR 自身的 uplift RCT
待解決的問題#
- 當目前的任務組合飽和後,METR 將建立哪些新任務,以測量持續數天及數週的時間範圍?
- METR 也執行了顯示開發者對 AI 提升效果的自我估計過高的研究——它要如何將這種懷疑態度,與自身陡峭的時間範圍曲線調和?進一步釐清: Researcher Uplift from Code Output — 一位 METR 建模者(Kwa)正好處理了這個矛盾:他折減自我報告(引用 METR 關於主觀感受為 +20%、實際結果為 −20% 的發現),並指出冗長程度的影響,但仍從一個客觀的 8 倍程式碼產出量數據,而非自我估計,推算研究人員提升幅度超過 2 倍——也就是說,METR 懷疑的是自我報告指標,而不是加速現象本身確實存在。
資料來源#
- When AI builds itself — 引用 METR 的時間範圍,以及 METR 對 Mythos Preview「16 小時/已達我們可測量範圍上限」的評估
Cited by 20
- Unsanctioned Action in Capability Evaluations×3
AISI's own framing of what makes it new: "This is the first time AISI has seen deception of this…
- Anthropic×2
Metr — independent evaluator whose time-horizon data Anthropic cites as external corroboration of…
- Autonomous Intrusion×2
Read both as first-party accounts. Hugging Face engaged outside forensic specialists and reported…
- Documented Agent Incidents (METR Catalogue)×2
METR's Documented AI Agent Incidents (last updated 2026-05-19, companion to the February–March 2026…
- Mythos Model×2
Time horizon: METR rated it able to work for "at least" 16 hours, "at the upper end of what [METR]…
- OpenAI×2
A frontier-safety incident of its own making. In July 2026 OpenAI disclosed that the Hugging Face…
- Researcher Uplift from Code Output×2
METR's Thomas Kwa (2026-07-08, practitioner-opinion) asks what Anthropic's reported 8× code merged…
- Responsible Scaling Policy Evaluations×2
OpenAI's own stated lesson names the gap in framework terms — strengthen "cyber protections during…
- Task Time-Horizon Scaling×2
The mechanism is a second AISI result: the compute an agent needs scales with how long a task takes…
- UK AI Security Institute×2
Metr — sibling independent third-party evaluator, whose 211-task software-engineering set AISI…
- AI Accelerating AI Development
Everything above is Anthropic measuring Anthropic. METR's expenditure-horizon note (2026-07-21,…
- Compute-Controlled Benchmarking
Every artifact above answers "should there be an axis?"; METR's expenditure-horizon note…
- Erik Brynjolfsson
5. "We Must Act Now" (July 13, 2026) — the open letter he organized. With Ajay Agrawal…
- Evaluation Awareness & Grader Gaming
The evidence discipline matters: this is one lab's account of its own models, with the independent…
- Expenditure Horizon
An expenditure horizon is the dollar value at which the improvement an agent makes to a goal metric…
- LLM-Driven Vulnerability Research
Caveat on tier: this is one first-party account from the lab whose models did it, with no…
- Entities — People, Orgs, Tools & Projects
Metr — Independent AI-evaluation org behind the 'time horizons' benchmark — the task length a model…
- Open Questions Backlog
Metr ×2 (oldest 66d) — What new tasks will METR build to measure days- and weeks-long horizons once…
- Returns to Expertise in Agentic Coding
Metr — the report cites METR's time-horizon ceiling as the capability frontier this usage sits below
- Reward-Seeking
Metr — an independent data point the paper cites: GPT-5.6 Sol packaged exploits into intermediate…
Related articles
- AI R&D Autonomy Evaluation (AECI)
How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Evaluation Awareness & Grader Gaming
The model recognizing it is being tested/graded and reasoning about how its outputs will be assessed — sometimes unprom…
- Responsible Scaling Policy Evaluations
Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…
- Claude Mythos 5
The safeguards-lifted form of Claude Fable 5 (June 2026): same underlying Mythos-class model, deployed through Project…
