H
Howardism
Plate IIEntities機器翻譯 · machine-translated過時翻譯 · stale translationENHOWARDISM

Perplexity

PublishedJune 15, 2026FiledEntityDomainEntitiesTagsEntityOrgAI LabDeep ResearchReading3 minSourceAI-synthesised

AI 答案引擎公司;Perplexity Deep Research 的製作者(在其自身的 DRACO 基準測試中為領先系統),也是 DRACO 的發布者;在其編排中將 Claude Opus 4.5/4.6 作為基礎模型——同時是 Anthropic 的客戶與基準測試競爭者

Perplexity 插圖

資料來源#

摘要#

Perplexity 是一家 AI 答案引擎/搜尋公司。在本資料集中,它是 DRACO 基準測試的作者(arXiv:2602.11685,2026 年 2 月,與一位 Harvard 合著者共同完成),也是 Perplexity Deep Research 的製作者;後者是能動式 deep-research 系統,在該基準測試的每個領域與評分規準軸上都名列第一。它是本 wiki 中第一個其貢獻屬於生產環境來源評估的供應商;這項評估取自其自身的部署流量。

它做什麼(在本資料集中)#

  • Perplexity Deep Research — 一個深度研究代理,會拆解查詢、反覆從許多來源擷取資料,並綜合成附有引用的報告。在 DRACO 上,它的標準化分數為 70.5%(Opus 4.6 基礎模型)/72.8% 通過率,領先 Gemini Deep Research、OpenAI Deep Research(o3/o4-mini),以及未經編排、搭配工具的 Claude Opus 4.5/4.6。其特徵是:在深度研究系統中同時擁有最高分數與最低延遲,輸入 token 足跡也最大(約 779k/任務)——偏重擷取、精簡輸出。
  • DRACO — Perplexity 從自家的 Deep Research 查詢中抽樣數千萬筆(2025 年 9 月至 10 月),接著去識別化、擴充、篩選並策展成 100 個由專家評分規準評測的任務(Production-Sourced Evaluation)。已在 Hugging Face 公開發布。

值得注意的結構性事實:客戶也是競爭者#

Perplexity Deep Research 使用 Claude Opus 4.5 / 4.6 作為其基礎模型(依據論文的實驗設定)。因此在 DRACO 上,Perplexity 的編排式產品(以 Opus 為基礎)會與未經編排的 Anthropic Opus 模型進行基準比較——而且以約 10 個百分點擊敗它們。Perplexity 同時是 Anthropic 的 API 客戶,也是證明其編排層能在 Anthropic 模型之上增加大量價值的實體。這是「超越基礎模型的編排」發現最清晰的具體實例,也是對 Harness Shrinkage as Models Improve 的即時反例資料點。

相關連結#

  • DRACO Benchmark — Perplexity 撰寫的基準測試,其產品在此測試中領先
  • Deep Research Agents — Perplexity Deep Research 是此系統類別中最具代表性的領先實例
  • Production-Sourced Evaluation — DRACO 的方法:以 Perplexity 自身生產環境流量建立基準測試
  • Anthropic — Perplexity 將 Claude Opus 4.5/4.6 作為基礎模型,並且與未經編排的 Opus 進行基準比較;同時是客戶與競爭者
  • Google DeepMind — 競爭者(評估了 Gemini Deep Research);Perplexity 也選用其 Gemini-3-Pro 作為 DRACO 的主要裁判模型
  • LLM-as-a-Judge — DRACO 的評分方法;Perplexity 透過人類對齊研究選定裁判

開放問題#

  • 供應商發布自家產品獲勝的基準測試,顯然存在誘因問題——隨著時間推移,DRACO 的可信度如何維持?Perplexity 是否真的會執行可自動化的更新?
  • Perplexity 依賴 Anthropic(以及其他公司)提供基礎模型,同時又在最終產品上與它們競爭——如果基礎模型製造商推出自己的深度研究模式,這項編排優勢能維持多久?

資料來源#

§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 7
  • DRACO Benchmark×2

    Perplexity / Anthropic / Google Deepmind — benchmark author; makers of evaluated systems and the…

  • Anthropic

    Perplexity — Anthropic API customer and deep-research competitor: runs Opus 4.5/4.6 as base models,…

  • Deep Research Agents

    Perplexity — builder of the leading deep-research system and of DRACO

  • Google DeepMind

    Perplexity — deep-research competitor whose DRACO benchmark uses DeepMind's Gemini-3-Pro as…

  • Entities — People, Orgs, Tools & Projects

    Perplexity — AI answer-engine company; maker of Perplexity Deep Research (the leading system on its…

  • Open Questions Backlog

    Perplexity ×2 (oldest 58d) — A vendor publishing a benchmark its own product wins is an obvious…

  • OpenAI

    Perplexity — a deep-research competitor that runs Anthropic (not OpenAI) base models; OpenAI Deep…

Related articles
  • DRACO Benchmark

    Perplexity's benchmark of 100 production-sourced deep-research tasks (10 domains, 40 countries) graded by 26-expert rub…

  • Anthropic

    AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…

  • Deep Research Agents

    Agentic systems that decompose a complex query, iteratively search diverse sources, and synthesize a structured, cited…

  • Google DeepMind

    Google's AI lab; built AlphaProof Nexus; Gemini models, AlphaProof, AlphaEvolve, and the open-weight Gemma line; opens…

  • LLM-as-a-Judge

    Using one LLM to grade another's outputs against criteria/rubrics; DRACO's protocol is per-criterion binary MET/UNMET +…