Sources#
- 5 takeaways from the State of Software Delivery Q2 Pulse report
- AI and Job Postings: From Destruction to Creation?
- AI Engineering Report 2026: The Acceleration Whiplash
- AI-Augmented Human Resource Management? Insights from German companies
- Helping People Choose Careers in the Age of AI
- Ramp's latest data on China vs. the American AI Labs
- The State of AI Impact in Engineering: Q2 2026
- Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews
Summary#
Faros AI's methodological argument, and the basis for the most consequential conflict in its 2026 report: during rapid AI transformation, perception lags reality, so survey-based engineering research systematically misses the downstream damage that system telemetry catches in near-real-time. Faros draws its findings from engineering systems (task trackers, IDEs, static analysis, CI/CD, version control, incident management) rather than from how developers feel, and uses that distinction to directly contradict Google's DORA 2025 conclusions.
vendor-claimsource — Faros's own platform is the telemetry instrument, so "telemetry beats surveys" is also a sales argument for that platform. The methodological point stands on its own merits, but the conclusion conveniently favors the vendor's product. See Acceleration Whiplash for the full evidence note.
Why perception lags reality#
The mechanism Faros proposes: at the individual level developers genuinely are more productive — task completion is up, code flows faster, the tools feel powerful — so surveys capture real, positive feeling. What surveys cannot capture is what happens downstream: "the review queues quietly backing up, the incidents accumulating in production, the bugs reaching customers." By the time those consequences show up in how people feel, "months have passed and the signal is already stale." Telemetry, drawn from the systems where work actually happens, does not lag. The claim: engineering leaders making consequential decisions about headcount, tooling, and process "need data as close to real time as possible… not how people feel about the work after the fact."
The DORA contradiction#
This is a flagged inter-source contradiction. DORA's 2025 State of AI-Assisted Software Development concluded that AI amplifies existing strengths and weaknesses, and that strong engineering foundations protect against AI's downsides. Faros's telemetry, it claims, "does not support that as a protective factor": high-performing organizations experience the same downstream deterioration as everyone else (see the maturity-independence finding in Acceleration Whiplash).
Weighing the conflict by method and incentive:
- DORA 2025 — survey-based; large, long-running, vendor-neutral-ish (Google/DevOps Research). Strength: breadth and continuity. Weakness, per Faros: perception lag during fast transitions.
- Faros 2026 — telemetry-based; within-company longitudinal comparison (low- vs high-adoption quarters), Spearman ρ at p<0.05. Strength: measures behavior, not feeling, near-real-time. Weakness:
vendor-claim— Faros sells the platform, and "your mature practices won't save you, you need visibility + a context engine" is precisely the conclusion that grows its market.
Neither is a clean win. The honest read: Faros's measurement critique of surveys is sound (lagging perception is real), but its substantive claim that maturity offers zero protection should be held with the vendor incentive in view — it is the conclusion most favorable to selling the instrument. Worth tracking against future DORA editions and any non-vendor telemetry study.
The family effect: instruments agree with their data source, not with their construct#
The sharpest evidence that this page's dichotomy is a real fault line rather than a framing device comes from outside engineering metrics. Steele & Cruz (arXiv 2607.15506) put seven occupational AI-exposure instruments on the same O*NET occupations and correlate them pairwise. All seven claim to measure the same thing. What predicts whether two of them agree is which data source they were built from:
- The two built from 2025 Anthropic Claude usage correlate at ρ = 0.89.
- The two built from theoretical generative-AI capability (GPT-4 task ratings; 2,000 MTurk ability ratings) are the next-strongest pair — despite differing in level of analysis, rater, and question asked.
- Across families, correspondence largely collapses; patent-mining, ML-rubric, and bottleneck-based instruments show "very little correspondence with each other or with later measures."
Two consequences for this page. First, the telemetry/survey split is not just a latency difference (behavior now vs. feeling later) — it is a partition of the answer space. Instruments in different families do not produce the same ranking, and in Steele & Cruz's case they do not even produce the same sign on the exposure-salary gradient. Second, the paper is the first entry in the vault where telemetry is run forward into a projection rather than reported as an observed count: query-volume ventiles rank the tasks, and a hand-set schedule of assumed automation ceilings supplies the levels. That is a hybrid — telemetry's ordering, assumption's magnitudes — and it inherits the credibility of the ranking without inheriting it for the numbers. See Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated for the full head-to-head.
The third instrument: randomization, and the one thing neither telemetry nor survey can do#
This page's dichotomy has a missing third term. Telemetry and surveys are both observational — one records behavior as it happened, the other records how it felt. Neither can say what would have happened otherwise, which is the question every claim on this page is ultimately about. A randomized experiment can, and Jabarian & Henkel (arXiv 2607.28222) is the vault's cleanest instance: 67,056 job applicants randomized between AI voice interviewers and human recruiters, pre-registered, with administrative employment outcomes.
What makes it useful here is that the paper runs all three instruments on the same event and they point in different directions:
| Instrument | What it says about AI-conducted interviews |
|---|---|
| Administrative records (offers, starts, retention) | AI-interviewed applicants receive 12% more offers, start more often, and are retained more (all p<0.001) |
| Applicant survey (post-interview CX survey) | AI interviews are rated significantly less natural (p=0.014), pulling the composite perceived-quality index in the humans' favor (p=0.076) |
| Expert forecast (recruiter survey, pre-disclosure) | 61% of recruiters expected AI-led interviews to be of lower quality; 36% expected lower offer rates; 48% expected lower retention |
The forecast is simply wrong — 131 experienced practitioners, surveyed before results were disclosed, predicted the sign backwards. The applicant survey is not wrong; it measures a real and different thing (naturalness) that does not track the outcome. And the administrative records are the only reason we know which is which.
Two consequences for this page. First, the felt-vs-system split it was built on is a special case of a bigger problem: an observational instrument is a claim about correlation regardless of how good its latency is, which is why CMU's 2.5M-PR telemetry study could support opposite conclusions from the same traces. Randomization is the rung above both, not a third flavor at the same level. Second, the honest cost of that rung: it took a firm partnership, an IRB, a pre-registration, and three months of a live hiring pipeline to answer one question about one task in one labor market. Telemetry and surveys are cheap and broad; the instrument that actually identifies the effect is expensive and narrow, which is exactly why the vault has hundreds of the former and one of the latter.
The survey arm, made concrete: what only asking can see, and what it costs#
Everything above argues against surveys. Kalff & Simbeck's N=410 study of German HR departments (arXiv 2607.13839, empirical) is the survey arm of this contrast in its plainest form — a self-reported adoption census of one business function, whose authors list "reliance on self-reported survey data for quantitative measures" among their own limitations. It is useful here for the opposite reason to the rest of the page: it contains the clearest instance the vault has of something no telemetry can see, and the clearest instance of the price of asking.
The thing only a survey can measure. Of the 410 valid respondents, 183 report using AI tools informally, on personal devices, regardless of employer policy or the availability of company-provided systems — including at companies whose official policy restricts or prohibits generative AI. The paper draws the boundary itself: "AI solutions requiring complex, company-specific data cannot be used informally, as they are accessible only through an official rollout." So corporate telemetry sees exactly the half of adoption that went through IT, and is blind by construction to a channel roughly as large. This is the CircleCI aperture problem in its limiting case — the instrument's coverage is whatever the vendor's product touches, and shadow adoption touches nothing the firm operates. On this question the ordering inverts: the survey is not the lagging instrument, it is the only instrument.
And the price of asking, in an unusually legible form. The same paper documents its own construct collapsing. Group discussions "often devolved into debates over the meaning of AI"; participants' perceptions were "strongly shaped by their exposure to generative AI applications popular in the media," so "some participants did not recognize machine-learning systems that underpin traditional HR analytics — such as those used to measure turnover risk or identify workforce trends — as 'actual AI.'" That is not ordinary measurement noise, because it is directional and it points at the paper's own headline: the conclusion is that predictive analytics plays only a minor role, and the tools respondents are least likely to count as AI are precisely the predictive-ML ones. A license- or spend-based instrument (Ramp's payment traces) has no such problem — it counts what was deployed, not what gets called AI.
Worse, the ambiguity is not merely fuzzy, it is contested by interested parties in both directions: vendors "leverage AI branding to generate interest and support business cases," while "the AI aspects of HR tools may be downplayed to avoid scrutiny from co-determination bodies." When the label is a strategic instrument for the people answering the question, no amount of question wording recovers the construct — a third failure mode, distinct from perception lag (this page's original argument) and from the family effect (instruments agreeing with their data source). Perception lag says respondents report an old truth; the family effect says instruments disagree about a stable construct; this says the construct itself is being moved by the respondents while you measure it.
The payment rail: the platform-records family gets a second member, and the first case where survey and behaviour agree#
The Connections entry below names Indeed's job postings as a fourth instrument family — a platform's by-product record of transactions it brokers, neither telemetry nor survey nor randomization. Ramp's monthly AI Index (empirical) is that family's second member, and putting the two side by side sharpens what the family is: whoever runs the rail can count what crosses it, for free, in near-real time, for exactly the population that transacts there. Indeed brokers vacancies and can therefore count labor demand; Ramp brokers corporate payments and can therefore count who firms pay for AI. Both are published by the operator's own research arm, so both fold the vendor incentive and the instrument into one party the way this page's Faros entry does.
The aperture bites in a checkable way here. Ramp's June-2026 vendor shares put OpenAI at 39.5% and Anthropic at 42.4% of businesses — and Google at 6.4% and Microsoft at 1.7%, with Google flat in the 4.6–6.4% band for the entire 42-month series while overall AI adoption went 7.5% → 55.0%. The most likely reading is not that Microsoft and Google lost the enterprise; it is that enterprise agreements, negotiated invoices and cloud committed-spend drawdowns do not cross a corporate-card rail the way a per-seat or per-API subscription does. That is the CircleCI aperture problem restated for payments: the instrument's universe is the purchases that fit its rail, and a vendor whose sales motion routes around that rail is structurally undercounted regardless of its actual share.
And the part this page did not previously have an example of: a survey and a behavioural instrument agreeing. The family effect above predicts that instruments track their data source rather than their construct, and most of this page is disagreements. Here two instruments from different families, on different populations, converge over the same six months on the same reordering: ICONIQ's Q2-2026 exec survey of ~305 AI-building software companies has Anthropic going 51% → 81% of respondents and passing OpenAI (77% → 71%), while Ramp's payment records have Anthropic passing OpenAI in May 2026 (42.4% vs 39.5% by June) with OpenAI down ~1.9pp from its November-2025 peak. Self-report and receipts, a builder cohort and a whole card base, same direction and roughly the same timing. Convergence across families is weak evidence taken alone and strong evidence taken here, precisely because this page's default expectation is that it does not happen — and it is the cleanest case in the vault of the instrument question being settled by agreement rather than adjudicated.
"The control group is no longer viable" — a claim to split, not to accept (2026-08)#
DX's Q2 2026 AI-impact readout (vendor-claim, 500+ customer organizations) opens with the sharpest methodological assertion any source in this vault has made about its own instrument, and it is aimed squarely at this page's subject:
"When we first began tracking the impact of AI on engineering teams, our primary goal was to measure AI cohorts against historical baselines… With industry-wide AI adoption exceeding 90%, comparing AI users against a non-user control group is no longer a viable measurement strategy."
The claim is true of one control design and false of another, and the two are not usually distinguished. What saturates at >90% is organization- and developer-level adoption — whether a person or a firm uses AI at all. What has not saturated is the artifact-level mix: the share of individual changes, PRs or completions that an AI actually wrote.
- The between-firms (or between-developers) adopter-vs-non-adopter contrast is genuinely dying, and DX is right about it. At >90% adoption the non-adopter arm is a residual of holdouts selected on whatever made them hold out, which is a worse confound every quarter. This is also the design DX's own panel and Faros's adoption-depth cross-section are built on, and the reason neither can separate "AI code is worse" from "orgs that adopt hardest differ in other ways."
- The within-firm, within-codebase provenance contrast is alive and was running while DX declared it dead. Tran et al. (Google) compare AI-authored against human-authored changes inside one monorepo, same window, stratified on change size, over 3.52M submissions. Adoption in that population is effectively total — everyone has the tools — and the control cohort survives anyway, because AI's share of submitted code ran 28.99% to 68.62%, leaving roughly three in ten changes human-written at the end of the window. Universal adoption and a usable control cohort coexist without tension the moment the unit of comparison stops being the person.
So the correct statement is narrower and more useful than DX's: the control group moved down a level rather than disappearing. The population-level arm is gone; the per-change arm is not.
Why the claim is also structurally self-serving, in the way this page's Faros entry already documents. DX sells developer-productivity measurement, and its instrument is a survey-plus-SDLC-telemetry panel across customer organizations — an instrument that can only ever run the between-firms contrast. The design it declares non-viable is the one it cannot run; the remedy it proposes (longitudinal within-panel trends against historical baselines) is the product. That does not make the observation wrong, and the honest reading is that DX has correctly identified a real measurement crisis and misattributed its cause.
The cause is instrumentation, not adoption. A per-change control cohort requires authoring-time provenance — a record, made while the code is being written, of what wrote it. That is a decision taken before the artifact exists, by whoever owns the editor, and it is why Google could build both arms and nobody outside such a company has. The alternative available externally is a corpus defined by agent authorship (AIDev), which yields one arm by construction and no matched human comparison — which is exactly why Dipongkor et al. and Jhanglani et al. and Sakib et al. each measure a level and not a delta, and why DECODE can see only accepted completions. Four studies missing the same arm, for a reason that has nothing to do with adoption rates. Post-hoc AI-code detectors are the obvious substitute and Tran et al. reject them outright ("generalize poorly across models and settings").
The prescription that follows for reading this vault: when a source says it could not build a control group, ask whether it lacked non-adopters (increasingly unavoidable) or lacked provenance (a fixable instrumentation gap that most publishers of AI-impact numbers have not paid for).
Connections#
-
Community Smells Under AI Adoption — the instrument tension appearing inside one instrument: the same survey's free-text layer reports AI displacing teammates while its structural models estimate peer interaction rising. The authors' three-part reconciliation — individual change vs cross-respondent association, concentrated-and-conditional effects, and a non-significant direct path indicating mediation rather than absence — is a reusable frame for reading self-report here
-
The Open-Weight Frontier Gap — the limit of a payment-rail instrument on the question it is most often used for: Ramp's 5.8%-of-AI-spenders model-serving figure is the vault's main quantitative handle on open/Chinese-model adoption, and it can only see firms that pay a serving vendor — downloading open weights and running them on your own GPUs generates no transaction and is invisible by construction, the same blind spot shadow AI creates for corporate telemetry above
-
Organizational Complements to AI — the substantive home of this source, and the reason its instrument matters: adoption below the organizational level (183 of 410 using AI outside any rollout) is a complements story that firm-level telemetry cannot reach, so the complements literature's usage-share and spend-trace instruments are structurally blind to the fastest-diffusing half
-
Controlled Variance: AI's Edge as Reduced Dispersion — the third instrument above: the vault's only randomized causal estimate of AI substituting for a human in an expert conversational task, and a single design in which administrative records, an applicant survey, and an expert forecast disagree about the same intervention
-
Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — where the seven-instrument comparison lives; the family effect above is the measurement-methodology reading of it, and its telemetry-built seventh instrument is the hybrid case this page's dichotomy did not previously have a slot for
-
Usage-Telemetry Classifier Validation — the price of the telemetry side of this argument: system logs beat self-report on latency and scale, but the economic claims drawn from them run through an LLM classifier whose exact accuracy (22.6% at the O*NET task level) was never published until Google ATLAS did it
-
AI Native Product Cadence — the limit of the telemetry side: OpenAI's Akshay Nathan reports the classic engineering proxies (commits, LOC, PRs, tokens) decoupling from goal attainment under agents, so better instrumentation of the same signals measures motion rather than progress — telemetry beats survey on latency, but only if the logged quantity still tracks the outcome
-
Acceleration Whiplash — the maturity-independence finding rests on this telemetry-over-survey methodology
-
Production-Sourced Evaluation — the same "measure from the real system, not a proxy" instinct applied to model evals; telemetry-vs-survey is its engineering-metrics cousin
-
Evals as Product Spec — Cat Wu's evals encode the spec; telemetry encodes what actually shipped — both prefer ground-truth signal over self-report
-
Verification as the New Bottleneck — Fiona Fung's warning to break PR-cycle-time into funnel chunks rather than read the aggregate is the same "instrument carefully or the signal misleads" discipline. That page also carries the instrument-aperture case: CircleCI's Q2 2026 Pulse (
vendor-claim, 20M+ workflows) is behavior-not-feeling telemetry like Faros's, but from a single platform on a single signal — it sees CI workflow runs and nothing else. That aperture fixes what it can claim: it counts validation cycles (its Merge Efficiency Ratio) with a precision no survey could reach, and cannot observe a production incident, a customer-facing bug, or a review queue at all. So its cheerful main-branch success rate (70.8%→76.7%) and Faros's grim incidents-per-PR (+242.7%) are barely in conflict — each vendor reports the half of the SDLC its own product instruments. The corollary for this page's thesis: telemetry beats surveys on latency, but its coverage is whatever the vendor's product happens to touch, and a vendor's headline metric is reliably one its instrument can see and its product can improve -
Compounding Data Moat — owning the telemetry stream is itself a moat; the report is a demonstration of what the data asset enables
-
Agentic Coding Work-Composition Shift — the
empiricalcousin: Anthropic's Clio-based 400K-session telemetry reads behavior-not-feeling the same way, but is a research artifact (validated classifiers, controls) rather than avendor-claimlead-gen report — and its session-layer optimism vs Faros's org-layer pessimism is the felt-vs-system split this page names -
Conversation-to-Delegation Shift — the third major usage-telemetry study (OpenAI/Codex,
empirical), and it extends this page's argument: as usage becomes delegation, even interaction-count metrics (active users, chats) go stale — track complexity, runtime, concurrency, reuse, output instead -
Anthropic Economic Index — the program that resolves this page's dichotomy: it links usage telemetry to survey responses per person (the Cadences report, ~9,700 linked respondents), treating telemetry and survey as complements rather than rivals
-
AI Usage Cadences — the AEI's continuous hourly telemetry is this page's "measure the real system finely" principle pushed to time resolution
-
Review as the Control Point — the non-vendor telemetry this page's open question asked for, and a deepening of its methodology argument: CMU re-scraped 2.5M+ GitHub PRs (behavior, not feeling) yet found the same traces support opposite conclusions under defensible analysis choices — so telemetry beats surveys on latency, but surface telemetry can't adjudicate why without a causal model (Pearl: "data are profoundly dumb"). Telemetry's edge is real and bounded
-
Market-Priced AI Exposure (the AI Premium) — the purest realized telemetry (every observation is a paid request, behavior not feeling), pushed to cross-provider breadth (400+ models, ~2% of global tokens) and then turned into a market-priced signal via equity-price comovement; the finance-side entry in this thread — but a reminder that telemetry's own collection mechanism biases it (OpenRouter is developer-skewed), so realized ≠ representative
-
AI Investment Story, Not Efficiency Story — this page's dichotomy carried into startup economics: on revenue-per-employee, Emergence's measured cap-table receipts (AI companies below non-AI peers) conflict with AWS's self-reported founder survey (AI-natives above the baseline) — a survey-vs-financial-telemetry instrument split, same felt-vs-system shape as the Faros-vs-DORA case, resolved the same way (weight the receipts, flag rather than average)
-
Firm AI-Spend Intensity and Headcount Growth — also the home of a fourth instrument family this page had no slot for: platform administrative records, in Indeed Hiring Lab's job-postings series. Not telemetry (nobody's usage is logged), not survey (nobody is asked), not randomization (nothing is assigned) — it is the by-product record of transactions a platform brokers, which makes it behavior-not-feeling and near-real-time on the demand side, where usage telemetry sees only the supply side. It inherits the aperture problem in a sharper form than CircleCI's: the instrument is one job board, so its universe is whoever posts there, and "US job postings fell 7%" is a statement about Indeed's marketplace before it is a statement about the economy. It also collapses the vendor-incentive and the instrument into one party in the way this page's Faros entry does — Indeed's research arm publishing a rebound story about Indeed's own board — while the substantive limit is different from Faros's: the numbers are administrative and hard to dispute, and it is the causal reading ("Claude Code launched, then postings rebounded") that the design cannot carry. Beyond that, the same behavior-not-feeling instinct applied to AI adoption and its labor effect: Ramp reads adoption off actual AI-vendor payment traces (not "do you use AI?" surveys, whose answers for the same period span 18%→78%) and joins them to Revelio workforce records — a revealed-adoption instrument the paper positions explicitly against "messy surveys and exposure measurements"
-
AI Product Economics Maturation — the survey axis pushed to its softest form: ICONIQ's exec survey layers forward projections on top of self-report (2026P/2027P margin and RPE), so its rosy trajectory is felt-about-the-future, two removes from telemetry. Where its projected margin expansion (→59% by 2027) meets Emergence's measured growth-margin compression, this page's prescription applies unchanged — weight the receipts, flag rather than average, and mark the projection prediction-grade
-
Outsource Your Thinking, Not Your Understanding — the reverse case to this page's thesis: comprehension debt is damage the telemetry misses, because the shipped artifact is fine (Shopify reports AI-assisted reversion rates at pre-AI baselines) and the human is what changed
-
Standardize the Infrastructure, Not the Tools — the instrument that makes org-wide AI telemetry possible at all (one gateway, one meter), and the open question of what per-team usage analytics get used for once they exist
-
Efficiency Debt of AI-Generated Code — first-party telemetry with a control cohort, which is the shape this page's whole Faros argument has been missing. Google instruments its own monorepo at byte-level authoring provenance and compares AI-authored against human-authored code inside the same window and size stratum — so unlike Faros's low-adoption-vs-high-adoption cross-section it can separate "AI code is worse" from "orgs that adopt hardest differ in other ways." Two lessons for this page. The aperture rule holds and cuts deeper than usual: the instrument sees everything inside one company's pipeline and nothing outside it, which is total coverage of a population of one. And the vendor incentive runs the opposite direction from every other entry here — the org measuring the code is the org that built the tools that wrote it, and the result is on balance reassuring (revert rate below parity) with a discussion attributing the residual weaknesses to "historical default system configurations." Two telemetry sources, two directional COIs, opposite signs, and only one of them has a control group
-
AI and Market Power — the fifth instrument family, and the only one with a real claim to representativeness: compulsory national statistical surveys run by statistical offices to common Eurostat guidelines (France's TIC, Portugal's IUTICE), linked to administrative balance sheets. It is a survey, so it inherits the self-report limits this page catalogues — but it is a sampled, weighted, mandatory one rather than a self-selected panel, and it anchors the low end of the vault's adoption spread: 20.2% of OECD enterprises using AI in 2025, next to Census BTOS's 18%, Ramp's ~55% of eligible businesses on the payment rail, and executive surveys at 69–78%. The aperture argument gets its cleanest illustration here — five instruments, one phenomenon, a 4× spread, and the most representative one reads lowest
-
Post-Acceptance Edit Behavior — a sixth instrument family, and the only one that observes an artifact which is then destroyed: pre-commit editor telemetry. DECODE saves the working file every time a developer pauses for one second, so it sees the AI code that gets deleted 23 minutes after acceptance — code that exists in no repository, no PR, no survey response and no monorepo history, and is therefore invisible to every other instrument on this page. That is the aperture argument's mirror image: the layers below the commit are not a smaller view of the same population, they contain events the higher layers structurally cannot record. The weakness is the familiar one in a sharp form — the aperture is one opt-in VS Code extension whose users installed a model-comparison tool, self-selected in a way a mandatory statistical survey or a company's full monorepo history is not
Derived#
- Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? — resolves the category-error question and carries the instrument-weighting rule ("weight instruments by what they can see") across the whole oversight cluster
Open Questions#
-
Is there a non-vendor telemetry dataset large enough to adjudicate the maturity-protection question independently of Faros's commercial framing? Partially answered: CMU's arXiv 2607.07980 supplies exactly this — a non-vendor, 2.5M+-PR GitHub telemetry study — and it (a) finds the agent no-review rate converging toward the human baseline rather than a widening gap, and (b) argues the effect's sign is set by team practice, closer to DORA. The catch: its own headline is that the telemetry is direction-unstable, so it counterweights Faros without cleanly settling the maturity question — the honest verdict is "surface telemetry alone, vendor or not, can't adjudicate this." The worked example of that verdict is The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence: the two telemetry studies' opposite under-review headlines dissolve once metric, population, time axis, and authorship unit are aligned — the datasets never disagreed, only the framings did.
-
Does anchoring an adoption survey's definition of "AI" change the answer, and in the predicted direction? Kalff & Simbeck document respondents excluding predictive ML from the category while counting ChatGPT in it, which would bias reported adoption down for exactly the tool class their headline finding is about. Falsifiable two ways: a split-ballot survey where one arm names the underlying technique ("software that scores turnover risk") instead of the label, or validating firm-level self-reported adoption against vendor-license/spend records for the same firms — the Ramp instrument applied as a criterion rather than a substitute. Related evidence (2026-08-12), not an answer: DX reports a self-reported AI-generated code share of 34% to 52% across Q1-Q2 2026 in 500+ organizations, while Google's authoring-time provenance over the overlapping window reads 28.99% to 68.62%. Different populations (a cross-industry customer panel vs one C++ monorepo), different constructs (code share vs adoption), no matched firms, and DX's question wording sits in a gated PDF the vault does not hold — so this settles nothing. It is nonetheless the corpus's first side-by-side of a self-reported code share against a provenance-measured one, and the direction is the one Kalff & Simbeck's construct-collapse mechanism predicts: self-report reads below the instrument that counts bytes, not above it. The criterion validation this bullet asks for is that same comparison run on the same firms.
Resolved Questions#
- Surveys and telemetry measure different things (felt productivity vs. system outcomes); is the "contradiction" partly a category error — both true at their own layer — rather than one being wrong? Answered: Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping? — yes for the measurement halves: surveys capture felt productivity (genuinely real at the individual layer — task completion did rise) while telemetry captures system outcomes that haven't propagated into feeling yet; during a fast transition the layers legitimately diverge, and the AEI's linked telemetry+survey design already treats them as complements. But not entirely: the maturity-protection claim is one proposition about one layer (system outcomes) and remains substantively contested — Faros's "no protection" aligns with its vendor incentive, CMU's moderator theory argues the sign is team-set (closer to DORA) but measures no maturity effect. The category-error dissolution cleans the framing; the one real disagreement stays open (tracked in the non-vendor-dataset question above).
Sources#
- AI Engineering Report 2026: The Acceleration Whiplash — "A direct counterpoint to DORA's 2025 findings"; Research Methodology; Report's Purpose
- DORA, 2025 State of AI-Assisted Software Development (cited by Faros): https://dora.dev/research/2025/dora-report/
- 5 takeaways from the State of Software Delivery Q2 Pulse report — Jacob Schmitt, CircleCI blog (2026-07-08),
vendor-claim: the single-platform telemetry aperture and the Software Delivery Data Explorer benchmark framing - Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews — Jabarian & Henkel, arXiv 2607.28222 (2026-07-30),
empirical, pre-registered field experiment. §3.1 (administrative outcomes), §3 intro (the recruiter forecast: 36% lower offers / 48% lower retention / 61% lower interview quality), §5.2 (the applicant survey's naturalness result and the composite quality index). Full treatment and parse warnings at Controlled Variance: AI's Edge as Reduced Dispersion. - Helping People Choose Careers in the Age of AI — Steele & Cruz, arXiv 2607.15506 (2026-07-16),
empirical. §4 (the query-based construction) and §4.4 + Fig. 6 (correspondence among the seven instruments). Table 1 and Table 2 verified clean at ingest; Table 5 and appendix Table A1 are damaged in the parse and are not cited here. The paper contains a source-internal contradiction in its own summary statistics — see the Sources note on Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated. - AI and Job Postings: From Destruction to Creation? — Guillermo Gallacher, AI and Job Postings: From Destruction to Creation? (Indeed Hiring Lab, 2026-07-08;
empirical). Cited here for the instrument rather than the substance: a platform's own administrative transaction records (job postings), analyzed and published by that platform, with an uncontrolled before/after design around a product launch. Substance and full evidence note at Firm AI-Spend Intensity and Headcount Growth; chart-only quantities in the source (per-country shares, scatter-plot correlation statistics) are not quoted anywhere in this vault. - Ramp's latest data on China vs. the American AI Labs — Ara Kharazian, Ramp's latest data on China vs. the American AI Labs (Ramp AI Index, 2026-07-08;
empirical). Cited here for the instrument, not the substance: a payments platform's own administrative records, published as a monthly market index by that platform's lead economist. COI and sampling: Ramp measures its own corporate-card/bill-pay customers — a business-spend-active, VC-forward-skewed base, not a random sample of US businesses — and has a commercial interest in owning the authoritative AI-adoption dataset. The Google/Microsoft series used for the aperture argument comes from the raw file's recovered Datawrapper chart datasets (the letter names neither vendor); substance and full evidence note at Firm AI-Spend Intensity and Headcount Growth, the open-weight demand reading at The Open-Weight Frontier Gap - The State of AI Impact in Engineering: Q2 2026 — Justin Reock, The State of AI Impact in Engineering: Q2 2026 (DX, Engineering Enablement newsletter, 2026-07-22). Evidence tier corrected
empirical(as ingested) tovendor-claimat compile — see the Source Notes entry in Sources; in short, a developer-productivity vendor's lead-gen readout over a self-selected panel of its own 500+ customer organizations, mixing self-report with SDLC telemetry, with the methodology in a gated PDF the vault does not hold. Cited here for the instrument argument above (the opening control-group passage) and for the self-reported 34%-to-52% code share. Web article, no docling parse, so no table-collapse or table-shift discipline applies; the on-page charts are images and every figure quoted anywhere in this vault appears in the newsletter's own prose. Substance at Acceleration Whiplash and AI as Primary Author - AI-Augmented Human Resource Management? Insights from German companies — Kalff & Simbeck, arXiv 2607.13839 (2026-07-15 / v2 07-20;
empirical, mixed methods, N=410 German HR managers). Cited here for the instrument, not the substance: §3 (survey design and the self-report limitation), §4.1 (the 183-of-410 informal-use finding and the boundary that company-specific AI "cannot be used informally"), §4.2 (the construct collapse — "did not recognize machine-learning systems… as 'actual AI'" — and the two-directional strategic use of the AI label by vendors and by firms facing co-determination), §5 (stated limitations). This document required a glyph-level repair at ingest — docling emitted every digit and every DOI/URL label as glyph names (/two_os,/D_SC), restored by a deterministic 1:1 substitution and re-verified against the PDF, so every number quoted from it traces through that repair. Table 1 is row-shifted and is cited nowhere; Table 2 verified clean. Full treatment and the complements reading at Organizational Complements to AI.
Cited by 36
- Acceleration Whiplash×7
Telemetry Vs Survey Measurement — the methodological basis for the maturity-independence claim and…
- Agentic Coding Work-Composition Shift×3
Telemetry Vs Survey Measurement — both studies measure behavior not feeling; this is the empirical…
- AI Product Economics Maturation×3
The margin story is projected, not banked. Notably this runs more optimistic than the measured…
- Faros AI×3
AI Engineering Report 2026: The Acceleration Whiplash — the paradox "sharpened into a crisis." See…
- Firm AI-Spend Intensity and Headcount Growth×3
Telemetry Vs Survey Measurement — a third measurement instrument: revealed AI-vendor spend linked…
- Outsource Your Thinking, Not Your Understanding×3
Telemetry Vs Survey Measurement — the reverse case to that page's thesis: here the instrumented…
- The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence×3
The kicker is that Faros prescribes the very behavior CMU observes emerging organically. Faros's…
- AI and Market Power×2
On the diffusion level, and the instrument aperture. OECD ICT usage statistics put AI adoption at…
- AI Investment Story, Not Efficiency Story×2
Telemetry Vs Survey Measurement — the instrument-split lens for the AWS-vs-Emergence RPE conflict:…
- AI Native Product Cadence×2
Telemetry Vs Survey Measurement — the instrument problem behind "motion vs. progress": the readouts…
- AI Usage Cadences×2
Telemetry Vs Survey Measurement — same "measure the real system, at higher fidelity" instinct,…
- Anthropic Economic Index×2
Telemetry Vs Survey Measurement — the AEI is the case that resolves the dichotomy by linking usage…
- Community Smells Under AI Adoption×2
Telemetry Vs Survey Measurement — a rare instance of the instrument tension appearing inside a…
- Conversation-to-Delegation Shift×2
This is the same "measure what the system actually did, not the proxy" instinct as Telemetry Vs…
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated×2
Where the older six are projections of what AI could do (per raters, patents, or rubrics), Steele &…
- Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping?×2
Telemetry Vs Survey Measurement — is the Faros–DORA "contradiction" partly a category error, both…
- Open Questions Backlog×2
Telemetry Vs Survey Measurement: Is there a non-vendor telemetry dataset large enough to adjudicate…
- The Open-Weight Frontier Gap×2
Telemetry Vs Survey Measurement — why the 5.8% is a floor-and-ceiling problem rather than a number:…
- Organizational Complements to AI×2
Telemetry Vs Survey Measurement — the instrument caveat that travels with the German HR evidence…
- Review as the Control Point×2
Telemetry Vs Survey Measurement — sharpens the methodology debate: this is the non-vendor telemetry…
- Agent-Generated Test Quality
Telemetry Vs Survey Measurement — where the missing human baseline both of these cuts report stops…
- AI as Primary Author
The ordering is the finding. For the overlapping period the self-reported figure is the lowest of…
- Compounding Data Moat
Telemetry Vs Survey Measurement — Faros Ai's cross-org SDLC telemetry is a compounding data asset;…
- When Knowledge Layers Disagree: Context Files vs Memory, and Conflicting Sources at Compile Time
Attach provenance and evidence tier to every claim; weigh by method and incentive, never average.…
- Controlled Variance: AI's Edge as Reduced Dispersion
Telemetry Vs Survey Measurement — the third instrument, and the only one that identifies causation.…
- Efficiency Debt of AI-Generated Code
Telemetry Vs Survey Measurement — a new instrument shape: first-party engineering telemetry with a…
- Evals as Product Spec
Telemetry Vs Survey Measurement — Faros Ai's "measure what actually shipped, not how people feel"…
- Google AI & Economy ATLAS
Telemetry Vs Survey Measurement — ATLAS is a telemetry instrument that repeatedly checks itself…
- Market-Priced AI Exposure (the AI Premium)
Telemetry Vs Survey Measurement — realized paid requests are the purest telemetry (behavior, not…
- AI Coding Practice
Telemetry Vs Survey Measurement — Perception lags reality: survey-based research (DORA) misses…
- Post-Acceptance Edit Behavior
Telemetry Vs Survey Measurement — a third instrument shape for that page's taxonomy: pre-commit…
- Production-Sourced Evaluation
Telemetry Vs Survey Measurement — Faros Ai's telemetry-over-survey stance is the…
- Security Debt of Agent-Generated Code
Telemetry Vs Survey Measurement — why the matched human baseline this page keeps asking for is…
- Standardize the Infrastructure, Not the Tools
What does per-team AI usage analytics get used for once it exists — cost containment, capacity…
- Usage-Telemetry Classifier Validation
Telemetry Vs Survey Measurement — telemetry beats self-report on latency and scale, but this page…
- Verification as the New Bottleneck
Discount appropriately. Both quantities are perception measures by DX's own definitions, from a…
Related articles
- Acceleration Whiplash
Faros 2026: AI floods a human-paced SDLC with output it can't absorb — throughput up (tasks +34%, epics +66%), quality…
- Open Questions Backlog
_456 actionable open questions across 205 pages · 107 predictions · 9 notes · 147 in progress · 69 watching (entities),…
- Organizational Complements to AI
The general-purpose-technology argument: AI productivity gains depend on complementary workflow, skill, and org-design…
- Returns to Expertise in Agentic Coding
Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the act…
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
