資料來源#
摘要#
Noam Brown 是 OpenAI 的研究科學家,也是被認為率先開創推論時(測試時)計算擴展的研究者之一——這是 GPT-5 系列「思考」模型背後的推理範式。在前沿 LLM 出現之前,他打造了超人類撲克代理(Libratus / Pluribus 系列工作),如今仍將從零開始打造撲克求解器作為個人能力 eval。在本語料中,他是 2026 年 6 月 No Priors 訪談的唯一作者兼主題人物;該訪談討論了他的文章 Implications of Large-Scale Test-Time Compute。
他主張的內容(在本語料中)#
Brown 的文章建立在一項根本主張上——能力如今是推論預算的函數——並追蹤了幾項由此而來的後果:
- 基準測試網格已經失靈。 單一數字的基準測試表沒有控制測試時的計算量,因此更有效率的模型(GPT-5.5)可能看起來只比前代略好,實際上卻是一次大幅躍升。解法:將成本/token 數/時間放在 x 軸上。參見受計算控制的基準測試。
- 在無界預算下,安全 evals 的定義不清。 Preparedness 框架/負責任擴展政策是為 ChatGPT 時代打造的,卻沒有問「你要在多少預算下進行評估?」——然而危險能力和有用能力一樣,都會隨美元投入而擴展。
- 不會出現一夜之間的智能爆炸。 由於峰值能力需要大規模測試時計算執行,時間便成為限制條件;起飛是漸進的,而非瞬間發生。參見智能爆炸動力學。
- 研究品味是人類剩餘的角色。 模型能將他的撲克演算法最佳化到 10–100 倍,但目前還無法發明更好的演算法;它們是「研究者非常好的補充」,而非替代品——不過他預期這裡也會出現類似程式設計和數學領域的轉折點。參見研究品味作為人類瓶頸。
- 存在潛在能力過剩。 沒有人探索過,將 $100K 的計算投入已發布模型後,它能做到什麼;在 OpenAI 宣布 Erdős 單位距離猜想可被公開模型反駁之前,該模型就已能完成反駁。參見潛在能力過剩。
Brown 的文章屬於 practitioner-opinion——來自單一實驗室的論點與軼事。2026 年 7 月,UK AI Security Institute 發表了其核心論點群的第一份獨立、empirical 印證,在數個基準測試上測量跨 token 預算的能力曲線。Brown 引用的 AISI 資安結果(模型「在 100M tokens 時仍持續改進」)正是該機構自己的研究,如今已完整發表。他的論點不再只取材自單一人物。
撲克 eval(他的招牌工具)#
Brown 會製作撲克求解器,因為這類工具的開源程式碼很少、已發表的理論很多,而且有「許多小陷阱」是他早已處理過的——因此他能精確看出模型在哪裡失敗。他所描述的進展也同時是一條能力時間線:早期模型「基本上什麼都做不了」;GPT-5.2 能在 steering 下建立 river solver(而且「感覺像研究生」),但會「gaslit」他——著名例子是堅稱棄掉一個 $100 底池會損失 $92,「它接近 100,沒問題」;GPT-5.5 則能 zero-shot 完成其中大部分工作。他的預測是:一年內,某個模型就能「基本上一口氣完成我的整篇 PhD 論文」。他也日常使用模型處理高風險的非程式碼決策(稅務、不動產文件),對其輸出的信任「可以說比……一位人類專家還高」。
相關連結#
- Large-Scale Test-Time Compute — 他的核心論點;這個群集正是以他所寫的文章為基礎
- Compute-Controlled Benchmarking — 他對基準測試網格的批評,以及「將計算量放在 x 軸上」的主張
- Latent Capability Overhang — 他觀察到已發布模型仍保有尚未提取的能力
- Research Taste as the Human Bottleneck — 他以實務者角度的解讀:品味是模型「一段時間內」做不到、之後卻會變得擅長的剩餘部分
- Intelligence Explosion Dynamics — 他主張測試時計算依賴使時間成為起飛瓶頸
- OpenAI — 他的雇主;他所報告的實驗室內部模型 Erdős 反駁及產品文化選擇所屬的機構
- UK AI Security Institute — 這個政府評估機構在 2026 年 7 月發表研究,獨立且以實證方式印證了他的測試時計算論點
資料來源#
- Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown — No Priors 與 Sarah Guo 的訪談(2026-06-26);Brown 談論他的文章 Implications of Large-Scale Test-Time Compute(
practitioner-opinion)。
Cited by 15
- Large-Scale Test-Time Compute×3
Two further AISI findings sharpen downstream pages rather than this one: compute demand scales with…
- OpenAI×3
noam brown large scale test time compute — No Priors, 2026-06-26 (Brown on test-time compute, the…
- UK AI Security Institute×3
Its July 2026 blog More compute, more capability is the corpus's first primary AISI publication,…
- Compute-Controlled Benchmarking×2
The evaluation consequence of Large Scale Test Time Compute: if capability is a function of…
- Intelligence Explosion Dynamics×2
Noam Brown (OpenAI, practitioner-opinion) supplies an independent, mechanism-level argument against…
- Latent Capability Overhang×2
If capability scales with inference budget (Large Scale Test Time Compute) but nobody spends a…
- Multi-Agent Collective Intelligence×2
Noam Brown (OpenAI, practitioner-opinion) frames the gap between today's multi-agent scaffolds and…
- Open-Weight Elicitation Irreversibility×2
> Status: wiki synthesis, not a source claim. No source argues this. Brown makes the budget…
- Research Taste as the Human Bottleneck×2
Noam Brown — a frontier researcher's practitioner reading: models optimize his algorithms 100× but…
- Responsible Scaling Policy Evaluations×2
Noam Brown (OpenAI, practitioner-opinion) names a structural hole this framework shares with every…
- Task Time-Horizon Scaling×2
Noam Brown — source of the "the only way to evaluate a year-long agent is to run it for a year"…
- Inference Efficiency as Capability
If capability is a function of inference budget, then cutting the cost of a token is capability work: Gemma 4's five le…
- Entities — People, Orgs, Tools & Projects
Noam Brown — OpenAI research scientist and a pioneer of inference-time (test-time) compute scaling;…
- OpenClaw
An agent-society precursor. Noam Brown names "Moltbook and OpenClaw" (the project's earlier…
- Recursive Self-Improvement
An external practitioner reaches the same brake by a different route. Noam Brown (OpenAI,…
Related articles
- Large-Scale Test-Time Compute
Noam Brown's thesis that model capability is now a function of inference budget (tokens/cost/time): with good scaffoldi…
- Latent Capability Overhang
Noam Brown's claim that already-released models can do far more than anyone has extracted, because nobody spends enough…
- Compute-Controlled Benchmarking
Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…
- Responsible Scaling Policy Evaluations
Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…
- Open-Weight Elicitation Irreversibility
A wiki-drawn synthesis of Brown and Gemma 4: if dangerous capability scales with inference budget, then an open-weight…
