資料來源#
摘要#
Faros AI 的方法論論點,也是其 2026 年報告中最關鍵衝突的依據:在 AI 快速轉型期間,感知落後於現實,因此以問卷為基礎的工程研究會系統性漏掉下游損害,而系統遙測能近乎即時地捕捉到這些損害。Faros 從工程系統(任務追蹤器、IDE、靜態分析、CI/CD、版本控制、事件管理)取得研究結果,而不是依據開發者的感受,並利用這項區分直接反駁 Google 的 DORA 2025 結論。
vendor-claim來源——Faros 自家的平台就是遙測工具,因此「遙測勝過問卷」同時也是該平台的銷售論點。這項方法論觀點本身仍然成立,但這個結論恰好有利於供應商的產品。完整的證據說明請見 Acceleration Whiplash。
為何感知落後於現實#
Faros 提出的機制是:在個人層級,開發者確實變得更有效率——任務完成量上升、程式碼流動更快、工具感覺更強大——因此問卷捕捉到真實且正面的感受。 問卷無法捕捉的是下游發生的事:「審查佇列悄悄堆積、production 中累積的事件、抵達客戶手中的錯誤。」等到這些後果反映在人們的感受中時,「幾個月已經過去,而訊號早已過時。」遙測取自工作實際發生的系統,因此不會延遲。其主張是:對人力配置、工具與流程做出重大決策的工程領導者「需要盡可能接近即時的資料……而不是事後人們對工作的感受。」
DORA 的矛盾#
這是一項已標記的跨來源矛盾。DORA 的 2025 State of AI-Assisted Software Development 得出結論:AI 會放大既有的優勢與弱點,而且強健的工程基礎能抵禦 AI 的負面影響。 Faros 的遙測則聲稱「不支持這能成為保護因素」:高績效組織與其他組織一樣,都經歷相同的下游惡化(請見 Acceleration Whiplash 中關於成熟度獨立性的發現)。
依據方法與誘因衡量這項衝突:
- DORA 2025——以問卷為基礎;規模大、運行時間長、相對不受供應商影響(Google/DevOps Research)。優點:廣度與連續性。根據 Faros 的說法,缺點:快速轉型期間的感知延遲。
- Faros 2026——以遙測為基礎;公司內部的縱向比較(低採用率與高採用率季度),在 p<0.05 時的 Spearman ρ。優點:測量行為而非感受,且接近即時。缺點:
vendor-claim——Faros 銷售這個平台,而「成熟的實務無法拯救你,你需要可見性與情境引擎」正是能擴大其市場的結論。
兩者都不是乾淨俐落的勝負。誠實的解讀是:Faros 對問卷測量的批評是站得住腳的(感知落後確實存在),但其成熟度完全沒有保護作用的實質主張應在供應商誘因的視角下看待——這是最有利於銷售該測量工具的結論。值得持續與未來的 DORA 版本及任何非供應商遙測研究進行比對。
相關連結#
- Acceleration Whiplash——成熟度獨立性的發現建立在這套「遙測勝過問卷」的方法論上
- Production-Sourced Evaluation——同樣「從真實系統而非代理指標進行測量」的直覺,應用於模型評估;遙測與問卷是它在工程指標上的近親
- Evals as Product Spec——Cat Wu 的 evals 將規格編碼其中;遙測則編碼實際發布的內容——兩者都偏好真實基準訊號,而非自我回報
- Verification as the New Bottleneck——Fiona Fung 提醒,應將 PR 週期時間拆成漏斗區段,而不是閱讀總體數字;這正是「仔細設置測量工具,否則訊號會誤導」的同一種紀律
- Compounding Data Moat——擁有遙測串流本身就是護城河;這份報告展示了資料資產能帶來什麼
- Agentic Coding Work-Composition Shift——
empirical的近親:Anthropic 以 Clio 為基礎的 400K 工作階段遙測,同樣讀取行為而非感受,但它是研究產物(經過驗證的分類器與控制),而不是以vendor-claim為主的潛在客戶開發報告;其工作階段層級的樂觀,對上 Faros 組織層級的悲觀,正是本頁所命名的感受與系統之分 - Conversation-to-Delegation Shift——第三項主要的使用遙測研究(OpenAI/Codex,
empirical),並延伸本頁的論點:當使用轉為委派時,即使互動次數指標(活躍使用者、聊天)也會過時——應改為追蹤複雜度、執行時間、並行、重用與輸出 - Anthropic Economic Index——解決本頁二分法的計畫:它將每人的使用遙測與問卷回覆連結起來(Cadences 報告,約 9,700 名連結受訪者),將遙測與問卷視為互補,而非競爭
- AI Usage Cadences——AEI 的持續每小時遙測,是將本頁「細緻地測量真實系統」原則推向時間解析度的做法
- Review as the Control Point——本頁開放問題所要求的非供應商遙測,也是對其方法論論點的深化:CMU 重新抓取了 250 萬多筆 GitHub PR(行為而非感受),卻發現同一批跡象在合理的分析選擇下支持相反的結論——因此遙測在延遲上勝過問卷,但表層遙測若沒有因果模型,就無法裁定原因(Pearl:「資料極其愚蠢」)。遙測的優勢是真實存在,但有其邊界
- Market-Priced AI Exposure (the AI Premium)——最純粹的已實現遙測(每筆觀察都是付費請求,是行為而非感受),擴展至跨供應商廣度(400 多個模型、約占全球 token 的 2%),再透過股價共同變動轉化為市場定價訊號;這是此脈絡中的金融側入口——但也提醒我們,遙測自身的蒐集機制會造成偏差(OpenRouter 偏向開發者),因此已實現不等於具代表性
- AI Investment Story, Not Efficiency Story——將本頁二分法帶入新創經濟學:以每位員工營收來看,Emergence 的已測量股權表收入(AI 公司低於非 AI 同業)與 AWS 的創辦人自我回報問卷(AI 原生公司高於基準)相衝突——這是問卷與財務遙測工具的分歧,與 Faros 對 DORA 案例相同的感受與系統形狀,也以相同方式解決(衡量實際收入,標記而非平均)
- Firm AI-Spend Intensity and Headcount Growth——同樣將行為而非感受的直覺應用於AI 採用及其勞動影響:Ramp 從實際 AI 供應商的付款紀錄讀取採用情形(而不是「你使用 AI 嗎?」問卷;同一期間的回答範圍為 18%→78%),並將其與 Revelio 的勞動力記錄連結——這是一種論文明確用來對抗「混亂的問卷與暴露量測量」的揭示式採用工具
- AI Product Economics Maturation——將問卷軸推向最柔軟的形式:ICONIQ 的高管問卷在自我回報之上疊加前瞻預測(2026P/2027P 的利潤率與 RPE),因此其樂觀軌跡是對未來的感受,與遙測相隔兩層。當其預測的利潤率擴張(→2027 年達 59%)遇上 Emergence 的已測量成長毛利率壓縮時,本頁的處方不變——衡量實際收入,標記而非平均,並將該預測標為預測級
開放問題#
- 問卷與遙測測量的是不同事物(感受到的生產力 vs. 系統成果);「矛盾」是否部分是分類錯誤——兩者在各自層級都是真的——而不是其中一者錯了?
- 是否存在規模足夠大的非供應商遙測資料集,能獨立於 Faros 的商業框架裁定成熟度保護問題?部分已獲回答:CMU 的 arXiv 2607.07980 正好提供了這項資料——一項非供應商、250 萬多筆 PR 的 GitHub 遙測研究——而且它(a)發現代理無審查率正收斂至人類基準,而非擴大的差距;(b)主張該效果的方向由團隊實務決定,更接近 DORA。問題在於:它自己的標題式結論是遙測具有方向不穩定性,因此它能平衡 Faros,卻無法乾淨地解決成熟度問題——誠實的判決是:「單靠表層遙測,無論是否來自供應商,都無法裁定這件事。」這項判決的具體例子是 The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence:兩項遙測研究在審查問題上的相反標題,當指標、母體、時間軸與作者單位對齊後便消解了——資料集從未互相矛盾,矛盾的只有框架。
資料來源#
- AI Engineering Report 2026: The Acceleration Whiplash——「A direct counterpoint to DORA's 2025 findings」;Research Methodology;Report's Purpose
- DORA, 2025 State of AI-Assisted Software Development(Faros 引用):https://dora.dev/research/2025/dora-report/
Cited by 36
- Acceleration Whiplash×7
Telemetry Vs Survey Measurement — the methodological basis for the maturity-independence claim and…
- Agentic Coding Work-Composition Shift×3
Telemetry Vs Survey Measurement — both studies measure behavior not feeling; this is the empirical…
- AI Product Economics Maturation×3
The margin story is projected, not banked. Notably this runs more optimistic than the measured…
- Faros AI×3
AI Engineering Report 2026: The Acceleration Whiplash — the paradox "sharpened into a crisis." See…
- Firm AI-Spend Intensity and Headcount Growth×3
Telemetry Vs Survey Measurement — a third measurement instrument: revealed AI-vendor spend linked…
- Outsource Your Thinking, Not Your Understanding×3
Telemetry Vs Survey Measurement — the reverse case to that page's thesis: here the instrumented…
- The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence×3
The kicker is that Faros prescribes the very behavior CMU observes emerging organically. Faros's…
- AI and Market Power×2
On the diffusion level, and the instrument aperture. OECD ICT usage statistics put AI adoption at…
- AI Investment Story, Not Efficiency Story×2
Telemetry Vs Survey Measurement — the instrument-split lens for the AWS-vs-Emergence RPE conflict:…
- AI Native Product Cadence×2
Telemetry Vs Survey Measurement — the instrument problem behind "motion vs. progress": the readouts…
- AI Usage Cadences×2
Telemetry Vs Survey Measurement — same "measure the real system, at higher fidelity" instinct,…
- Anthropic Economic Index×2
Telemetry Vs Survey Measurement — the AEI is the case that resolves the dichotomy by linking usage…
- Community Smells Under AI Adoption×2
Telemetry Vs Survey Measurement — a rare instance of the instrument tension appearing inside a…
- Conversation-to-Delegation Shift×2
This is the same "measure what the system actually did, not the proxy" instinct as Telemetry Vs…
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated×2
Where the older six are projections of what AI could do (per raters, patents, or rubrics), Steele &…
- Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping?×2
Telemetry Vs Survey Measurement — is the Faros–DORA "contradiction" partly a category error, both…
- Open Questions Backlog×2
Telemetry Vs Survey Measurement: Is there a non-vendor telemetry dataset large enough to adjudicate…
- The Open-Weight Frontier Gap×2
Telemetry Vs Survey Measurement — why the 5.8% is a floor-and-ceiling problem rather than a number:…
- Organizational Complements to AI×2
Telemetry Vs Survey Measurement — the instrument caveat that travels with the German HR evidence…
- Review as the Control Point×2
Telemetry Vs Survey Measurement — sharpens the methodology debate: this is the non-vendor telemetry…
- Agent-Generated Test Quality
Telemetry Vs Survey Measurement — where the missing human baseline both of these cuts report stops…
- AI as Primary Author
The ordering is the finding. For the overlapping period the self-reported figure is the lowest of…
- Compounding Data Moat
Telemetry Vs Survey Measurement — Faros Ai's cross-org SDLC telemetry is a compounding data asset;…
- When Knowledge Layers Disagree: Context Files vs Memory, and Conflicting Sources at Compile Time
Attach provenance and evidence tier to every claim; weigh by method and incentive, never average.…
- Controlled Variance: AI's Edge as Reduced Dispersion
Telemetry Vs Survey Measurement — the third instrument, and the only one that identifies causation.…
- Efficiency Debt of AI-Generated Code
Telemetry Vs Survey Measurement — a new instrument shape: first-party engineering telemetry with a…
- Evals as Product Spec
Telemetry Vs Survey Measurement — Faros Ai's "measure what actually shipped, not how people feel"…
- Google AI & Economy ATLAS
Telemetry Vs Survey Measurement — ATLAS is a telemetry instrument that repeatedly checks itself…
- Market-Priced AI Exposure (the AI Premium)
Telemetry Vs Survey Measurement — realized paid requests are the purest telemetry (behavior, not…
- AI Coding Practice
Telemetry Vs Survey Measurement — Perception lags reality: survey-based research (DORA) misses…
- Post-Acceptance Edit Behavior
Telemetry Vs Survey Measurement — a third instrument shape for that page's taxonomy: pre-commit…
- Production-Sourced Evaluation
Telemetry Vs Survey Measurement — Faros Ai's telemetry-over-survey stance is the…
- Security Debt of Agent-Generated Code
Telemetry Vs Survey Measurement — why the matched human baseline this page keeps asking for is…
- Standardize the Infrastructure, Not the Tools
What does per-team AI usage analytics get used for once it exists — cost containment, capacity…
- Usage-Telemetry Classifier Validation
Telemetry Vs Survey Measurement — telemetry beats self-report on latency and scale, but this page…
- Verification as the New Bottleneck
Discount appropriately. Both quantities are perception measures by DX's own definitions, from a…
Related articles
- Acceleration Whiplash
Faros 2026: AI floods a human-paced SDLC with output it can't absorb — throughput up (tasks +34%, epics +66%), quality…
- Open Questions Backlog
_456 actionable open questions across 205 pages · 107 predictions · 9 notes · 147 in progress · 69 watching (entities),…
- Organizational Complements to AI
The general-purpose-technology argument: AI productivity gains depend on complementary workflow, skill, and org-design…
- Returns to Expertise in Agentic Coding
Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the act…
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
