Howardism · Vol. 03Plate II · No. 02
AI Economics & Labor, in order.
Notes23DomainAI Economics & LaborOpen Qs68Newest11 Aug 2026Oldest8 May 2026
Work, wages, org design, and the economics of AI-driven labor.
Map of Content for the ai-economics-and-labor domain — 23 concepts. AI's measured economic footprint: usage telemetry, labor-market effects, returns to expertise, organizational complements, and framing effects on accountability. Curated entry point; see Home for all domains.
- AI and Market Power — OECD AI Papers No. 62 on French and Portuguese firm microdata plus global patent and start-up databases: non-GenAI adopters hold 7.5×/3.2× the market share of non-users, but the premium is selection (dies once broadband, digitalisation and lagged productivity enter) and adopters gain no market-share rank or markup growth over five years; firm-level GenAI exposure is inverted-U in size and market share while monotone in productivity and tertiary education; global AI-patent concentration fell 32–60% over 2001–21 yet correlates positively with sales concentration within markets; AI patents raise markups only in ICT (+7.95% interaction); and GenAI start-ups take ~130% more VC and are ~21% more likely to be acquired by incumbents
- AI Brain Fry — Kropp et al. 2026/03: mental fatigue from excessive AI oversight increases minor errors +11%, major errors +39%; cognitive cost surface for both tool and employee framings
- AI Employee Framing — Kropp et al. (HBR May 2026, n=1,261): framing AI agents as "employees" vs "tools" cuts personal accountability −9pp, increases escalation +44%, reduces error catching −18%, no adoption gain
- AI Usage Cadences — AEI Cadences report: continuous hourly telemetry reveals AI usage carries the rhythms of daily life — personal use spikes 35%→~50% on weekends, recipes 2.3× at 6pm, sleep advice pre-dawn, tax queries 8× around the Apr-15 deadline; off-hours work skews toward higher-wage occupations
- The Automation–Optimism Link — AEI Cadences survey finding: people who use Claude in more automated ways are MORE optimistic across all six job-quality dimensions (pay, security, job-finding, meaning, autonomy, human interaction), report their skills growing more valuable, and show no learning deficit — inverting the common delegation→deskilling-anxiety narrative
- Context Advantage, Not Taste — Andrew Ng's reframing of the residual human contribution: not 'taste' but an information asymmetry — 'so long as the human knows something the AI does not, human-in-the-loop is needed.' Recasts the wiki's central open question (is taste a ceiling or the next jagged valley?) as a category error, and makes the human role a closable engineering gap rather than a moat
- Controlled Variance: AI's Edge as Reduced Dispersion — Jabarian & Henkel (arXiv 2607.28222): a pre-registered natural field experiment randomizing 70,884 job applicants between AI voice interviewers and human recruiters — offer rate 8.70%→9.73% (+12%), job starts +18%, one-month retention +18%, no productivity decline, with humans making every hiring decision in both arms. The mechanism the authors name is controlled variance: the AI follows the firm's interview protocol more consistently (topic order τ 0.53 vs 0.33, question similarity 0.59 vs 0.43, significantly lower cross-interview variance) while still adapting per applicant and using richer vocabulary — AI wins by being less dispersed, not more capable. The wiki's only randomized causal estimate of AI substituting for a human in an expert conversational task
- Conversation Artifacts — AEI Cadences report: the 'artifact' (the primary output a user takes away) as a new unit of economic analysis — 93% of conversations produce one, artifact type predicts work/personal/coursework use, compute (tokens) scales with the artifact's economic value, and Claude's output sits ~1 education-year above the prompt
- Conversation-to-Delegation Shift — OpenAI's Codex usage study (June 2026): the move from conversational AI ('asking') to agentic AI ('delegated production'), measured by Codex's share of output tokens across three populations — 99.8% OpenAI / 63.3% organizational / 16.5% individual — with adoption spreading beyond developers; standard usage metrics (active users, chats) become less informative as the unit shifts from a conversation to a delegated workflow
- Experimental Learning Impact of Generative AI — Contractor & Reyes (arXiv 2607.08849): a randomized, proctored experiment with 211 undergraduates finds off-the-shelf AI access raises immediate test scores +0.27 SD, ~76% of which persists a week later on unaided tests, and lifts essay quality only after AI is removed — but the durable gains belong almost entirely to 'augmentation' users (AI as tutor/explainer) while 'automation' users' (AI-drafts-the-text) short-run gains vanish once AI is gone; the objective, measured-skill counterpart to the AEI self-report that learning both persists and can be hollow depending on use mode
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — Four distinct ways to measure AI's reach into an occupation — observed exposure (tasks seen done with Claude), theoretical exposure (tasks an LLM could do), reported exposure (what workers say AI can do today), and anticipated exposure (what they expect in 12 months) — plus their orderings (theoretical > reported > observed), the GDP/experience/automation gradients the AEI survey reveals, and Steele & Cruz's seven-instrument head-to-head showing the instruments cluster by data source rather than by construct, nominate eleven distinct occupations across twelve most-exposed slots, and flip even the sign of the exposure-salary relationship by vintage
- Firm AI-Spend Intensity and Headcount Growth — Ramp × Revelio panel of 21,559 US firms: high-intensity AI-vendor spenders grow headcount ~10% (entry-level ~12%) over the 24 months after adoption while low-intensity adopters show no change — an intensity-gated learning-curve effect, read against Indeed's senior-tilted postings rebound and the Ramp AI Index adoption-breadth cut.
- The Household Production Boundary — Google ATLAS's most novel contribution — 86.5% of conversational AI usage happens outside formal work, human time allocation predicts where AI questions go (slope 0.77, ~50% of variance), high-friction bureaucracy over-indexes ~20× with half of those queries outside business hours, and 0.5–5% household time savings values at $15–149B/yr in the US that GDP cannot see by construction
- Human-AI Accountability Redesign — HBR five-pillar prescription: span-of-control redesign, role redesign, performance management reset, decision-rights/escalation/consequences, agentic-unit-not-human-role design
- Market-Priced AI Exposure (the AI Premium) — Borri-Liu-Tsyvinski: market-implied AI exposure built from 380T tokens of realized OpenRouter consumption — an AI Factor, rolling firm-level AI Betas, and a priced 64 bps/week long-short premium concentrated on frontier/paid use; the implied skill map is orthogonal to task-based exposure measures, and tool-call tokens rising to 52% signal an agentic economy.
- Organizational Complements to AI — The general-purpose-technology argument: AI productivity gains depend on complementary workflow, skill, and org-design changes (David's electrification analogy, Brynjolfsson's paradox) — OpenAI's Codex natural experiment (99.8% vs 16.5% usage of the same model) shows the gap is complements; also home to the HAT substitution model and Kalff & Simbeck's institutional complement.
- Owning Your Externalized Cognition — Garry Tan's ownership axis on skill files: once your judgment is written down as executable markdown it is an asset with a holder, and the same file is either portable career capital or an extraction, depending only on whose repo it sits in — the appropriation counterpart to the cognitive-commons erosion argument, asserted from a keynote stage with no measurement behind it
- Post-Scarcity Macroeconomics — Musk's claim that once digital intelligence acquires end effectors the economy goes quasi-infinite, so money 'won't matter' by 2036: the load-bearing argument is a deflation one — create money slower than output grows and prices still fall — which makes universal transfers non-inflationary and taxation moot; the transition path is the part he concedes he cannot describe
- Returns to Expertise in Agentic Coding — Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the actions and 5× the output per prompt, reach verified success ~2× as often, and abandon stuck sessions far less; every occupation lands within 7pp of software engineers; gains are concentrated novice→intermediate, with mastery adding little
- The Solo-Authorship Rebound — Matsui (arXiv 2607.10780): across 300M+ OpenAlex works and 26 fields, the decades-long decline in solo-authored papers halts or reverses at ChatGPT's November 2022 release — positive trend break in 23 of 26 fields, largest in Engineering (+2.5 pp/yr) and Business (+1.9), absent in Chemistry and Physics and negative in Arts and Humanities. It survives conditioning on author history and is strongest among authors who had never published alone; solo papers stay near their authors' coauthored content while narrowing 23% in breadth and tilting toward computational work. A solo paper is proposed as an observable behavioral trace of AI substituting for a human collaborator — but the design is an interrupted time series with no untreated unit, roughly half the pooled break is venue composition, and the disciplinary ordering, not any single number, is the actual argument
- Task Crossover — OpenAI's Work at the Frontier (800K+ US ChatGPT work messages mapped to O*NET, July 2026): 16.8% of work messages and 43.5% of occupation-specific ones concern tasks historically belonging to another occupation — jobs reorganizing before job descriptions change. Borrowing and lending are separate directions (design borrows 35.2% and lends 1.7%; engineering lends 7.4%), financial calculation and software troubleshooting travel to all seven other groups, and crossover falls as workspace size rises (18.9% at 2–5 seats → 16.3% at 101+)
- Task Saturation: Broad but Shallow AI Diffusion — Google ATLAS's marquee work finding — AI reaches 68% of detailed occupations (88.4% of US employment) but only 21% of the tasks in the median occupation, with end-to-end automation the intent of just 6.5% of non-routine-cognitive conversations vs 26.9% for routine-cognitive; the extensive margin is gated by physicality, the intensive margin concentrates in non-routine cognitive work, and usage over-indexes most on the lowest-expertise cognitive tasks
- The Tragedy of the Cognitive Commons — Lovett (HRD Review, July 2026): professional expertise is a profession-level commons whose regeneration mechanism — entry-level work — AI is removing. Distinguishes Internalized Mastery (built through cognitive struggle) from Distributed Mastery (orchestrating AI), and names the Validation Tether: substantive oversight of AI requires the expertise AI adoption erodes. Its sharpest claim is that junior labor's operational necessity was the hidden governance mechanism all along — regeneration was a side effect of business, never a decision
Open questions 68 open
- AI and Market Power3 open
- SourceIs the non-GenAI null an artefact of a binary adoption measure? The paper's own future-work list asks for "measures of AI intensity" instead of the ICT survey's yes/no, and Firm AI-Spend Intensity and Headcount Growth finds intensity is the whole effect. Would a spend- or intensity-graded adoption variable on the same French and Portuguese panels recover a market-power effect the dummy hides?
- SourceDoes the size/market-share inverted-U in GenAI exposure survive contact with observed GenAI adoption? Exposure here is occupational composition, not use. If small, skill-dense, highly exposed firms turn out not to adopt at rates matching their exposure, the "window of contestability" reading collapses into a statement about who employs analysts.
- SourceDiffusion or consolidation? The 21% GenAI acquisition premium is consistent with technology transfer and with killer acquisitions, and the paper cannot separate them without acquirer-type data. Does post-acquisition patenting or product continuation at acquired GenAI start-ups differ by acquirer size and market position?
- AI Usage Cadences3 open
- SourceTime-of-day rests on IP-inferred location; how much noise do VPNs, travel, and datacenter-routed API traffic inject into the "sleep advice pre-dawn" style claims?
- SourceThe weekend personal-use spike is largest in high-income countries — is that a genuine work/life boundary difference, or a composition effect (who uses Claude for what, where)?
- WaitContinuous sampling is new; are these cadences stable, or will they drift as the user base shifts toward lower-wage tasks (the report's own diffusion trend)?
- SourceDoes the asymmetry regenerate faster than it transfers? The whole human role, under this frame, rests on the answer. Nobody in the corpus has posed it.
- SourceNg prefers the frame because it "gives us a clearer path to helping AI systems get better." That is a reason to adopt the frame, not evidence that it's true. What would distinguish a context asymmetry from a capability gap empirically? (Returns to Expertise in Agentic Coding is the closest thing to an instrument.)
- SourceIf the human's contribution is context injection, is the human replaceable by better context plumbing — memory, retrieval, continuous production telemetry — rather than by a better model? That would put the expiry of human-in-the-loop on the infrastructure roadmap, not the scaling curve.
- SourceNg writes from 0-to-1 consumer products. Does the frame survive contact with domains where the missing thing is a concept rather than a fact?
- NowHow much of the +12% is controlled variance in information collection versus the removal of the interviewer's discretion to abort? Human recruiters screen-out mid-interview 25% of the time against the AI's 7%, which mechanically suppresses human-arm offers before the evaluation stage. Falsifiable: re-estimate the treatment effect on the subsample of interviews that reached completion in both arms, or instrument the screen-out decision. The paper reports both numbers and never decomposes them.
- SourceThe AI system is never identified — no model, vendor, or version, only that Google Cloud supplied infrastructure. Is "controlled variance" a property of a 2025-generation voice agent under this firm's prompt, or of AI-conducted interviews generally? Nothing in the paper lets a replication know what it is replicating.
- SourceRetention ≥1 month is both the quality proxy and the metric the recruiting firm is paid on by its clients. Does the AI advantage survive on an outcome the intermediary is not compensated for — client-side performance at 12 months, promotion, or wage growth? The four-month estimate, the longest horizon measured, already fails significance under recruiter clustering.
- Conversation Artifacts3 open
- SourceTokens are a proxy for both compute cost and output value, but verbose models inflate tokens per unit of intent (the same critique Conversation-to-Delegation Shift raises); how much of "compute tracks value" is genuine value vs. models simply emitting more?
- SourceThe reading-level "+1 year" gap may be register (terse prompts, polished replies) rather than substance; can it be separated from genuine elevation of content?
- SourceArtifact classification is first-party and single-model-graded; do the 30+ categories and the work/personal/coursework split survive independent replication?
- SourceThe token-share metric rewards verbose agentic output. How much of the 99.8% / 63.3% / 16.5% spread is a genuine work shift vs. agentic tools simply emitting more tokens per unit of human intent?
- SourceOpenAI-internal is a frontier preview by assumption. Does the external organizational curve actually trace the OpenAI path (the paper's implicit claim), or does it plateau where adoption frictions don't vanish?
- Source"Asking is half of ChatGPT, doing is most of Codex" — but the two tools self-select different work. How much of the asking→doing contrast is the shift itself vs. routing pre-existing "doing" tasks to the tool built for them?
- SourceTime-on-task is held fixed by the lab; the authors flag that real-world learning depends on how students reallocate saved time. Does the augmentation dividend survive once students can spend the hour AI frees on something else entirely?
- SourceGains skew to the able (upper GPA/SAT quartiles). Is the widening-gaps signal a durable property of unrestricted AI, or an artifact of a high-ceiling elite sample where the bottom quartile has little room to move?
- SourceThe augmentation/automation choice is endogenous to incentives (grade inflation and signaling-motivated students push toward automation). Can incentive or interface design shift the mix toward augmentation at scale — and would that reverse the deskilling half?
- SourceDoes the same use-mode split govern workplace skill accumulation (the open question The Automation–Optimism Link and AI Brain Fry leave for workers), or is a proctored one-week academic task too unlike on-the-job learning to transfer?
- SourceBinned midpoint coding biases the exposure slopes toward zero; how much of the "uniform rising tide" is substance vs. coding artifact (the report checks robustness with a ≥60%-of-tasks indicator, but the levels remain self-reported)?
- SourceReported exposure exceeds observed partly because the survey reaches heavy users; what does the reported/observed gap look like in a representative sample? Sharpened: realized-consumption measurement (Borri-Liu-Tsyvinski) adds a market-implied instrument built on 380T tokens of actual paid requests — but it is skewed toward developers/sophisticated users (OpenRouter is ~2% of global tokens), a different non-representativeness than the survey's heavy-user skew. The lesson: no current AI-exposure instrument is representative; each collection mechanism biases in its own direction, so the reported/observed/market-implied gaps are partly artifacts of who each method reaches. A representative census remains the open target. Sharpened again by Steele & Cruz, which makes the population question concrete across seven instruments at once — 2,000 MTurk respondents (Felten), GPT-4 as rater (Eloundou), whoever files AI patents (Webb), Crowdflower workers against a rubric (Brynjolfsson), 70 jobs hand-coded by two researchers (Frey), and Claude/ChatGPT users in 2025 (Massenkoff; Steele & Cruz). Every instrument biases toward whoever it reaches, and the head-to-head shows the consequence is not a level shift that a rescaling would fix: the instruments produce different job rankings, nearly disjoint most-exposed lists, and opposite signs on the exposure-salary gradient.
- SourceDoes averaging across instruments reduce error or merely blend incompatible biases? Steele & Cruz's cross-model average is a diversification argument, not a validated one — no instrument in the set has been scored against realized labor-market outcomes, and the average's apparent stability partly reflects dropping the two instruments that disagreed most (Massenkoff for redundancy, Frey for anomaly). Falsifiable: score all seven, plus the average, against subsequent occupation-level employment and wage changes. Complication (2026-08-04): Indeed's postings data shows the exposure–outcome relationship changing sign between windows on a single instrument (2022–2026 negative, 2025–2026 positive), so any such validation scores the window as much as the instrument, and a scoring period must be pre-specified rather than chosen after the fact.
- WaitThe experience gradient rests on what workers believe AI can't do (judgment, relational work) — a belief that could be either durable comparative advantage or the next capability to fall. Which, and when?
- SourceWhat operational mechanism converts intensive AI spend into hiring? The paper establishes the correlation (adopters, especially intensive ones, grow) but explicitly cannot say why — product acceleration, sales productivity, engineering leverage, support automation, faster analysis, or new business lines are all candidates, and the firms that cracked it have no incentive to share. Related evidence (2026-08-04): the monthly AI Index narrows what the top of the intensity distribution is buying — the $248-PEPM cohort is defined by paying model-serving and inference platforms, i.e. building on APIs rather than buying more seats. That is a characterisation of the spend, not of the mechanism, but it points the candidate list toward engineering/product leverage and away from enterprise chat rollout.
- WaitDoes the effect diffuse beyond Information as adoption cohorts mature? Significant gains are, so far, an Information-sector phenomenon; the authors intend to update with later cohorts and post-24-month windows. Will professional services, finance, and non-technical sectors follow, or is the coding-agent workflow special? Independent corroboration of the sector boundary (2026-08-11): OECD AI Papers No. 62 finds the markup premium from AI patenting is significant only in ICT (
AI × ICT+7.95%, while standaloneAIturns negative with fixed effects) across ~600K firm-years in 21 European countries — a different outcome (markups, not headcount), a different instrument (patents, not spend), a different continent, and the same sector line. Their reading is the sharper version of the question: AI pays where it is the firm's output, not where it is an input. Not an answer to diffusion-over-time, but two instruments now agree on where the effect currently lives. - WaitIs the entry-level growth durable or a lead-indicator that later reverses? Gains compound through month 24 on thinning samples; whether the +12% entry-level result holds (or inverts toward the Brynjolfsson "Canaries" pattern) as high-intensity adopters mature past 24 months is unresolved. Countervailing signal (2026-08-04): Indeed Hiring Lab finds the May 2025 – May 2026 software-postings rebound is 71% senior roles, on a later window than this panel's average and on the demand flow rather than the headcount stock — not an answer (different unit, different population, no control group), but the first vault evidence pointing the other way on composition.
- WaitIs the premium a durable risk price or an early-diffusion artifact? The authors flag the short, fast-moving sample and call it "the current price of AI exposure." Does the transition-risk premium persist, shrink, or invert as AI diffusion matures?
- SourceHow much does the developer skew move the answer? OpenRouter's slice is unrepresentative; would a representative realized-consumption panel (if one existed) price the same firms and skills, or is the frontier/intensive-margin concentration partly a sampling artifact of who uses OpenRouter?
- SourceWhy is the market-implied skill map orthogonal to every task-based measure (<2% variance)? Is market-implied exposure capturing genuinely different information (forward-looking rents, complement/substitute value rather than technical automability), or is it noisier — and which should labor-impact forecasts trust? Partially answered: Steele & Cruz's seven-instrument head-to-head (see Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated) removes the framing that made this look like an indictment — the task-based measures are largely orthogonal to each other too. Only two pairs correlate strongly and both share a data source (ρ=0.89 between the two Anthropic-usage instruments; GPT-4-rated and MTurk-rated theoretical capability next). Webb's patent measure, Brynjolfsson's ML rubric, and Frey's bottleneck model show "very little correspondence with each other or with later measures," and the exposure-salary gradient flips sign between the older and newer instruments. So being uncorrelated with the task-based family is not evidence of noise — there is no coherent family to be uncorrelated with. The second half stays open, and is now harder: no instrument in either camp has been scored against realized labor-market outcomes, so "which should forecasts trust" has no empirical answer yet.
- WaitThe agentic premium is only "early evidence" (imprecise). Does a positive agentic premium survive a longer sample, and does the falling price-per-agentic-token (caching + cheap-model routing) erode the dollar-side signal even as token volume explodes?
- WaitDoes "Science most negative / interaction most positive" hold out of sample? The finding cuts against the intuition that AI automates cognitive work last; is the market right, or pricing a transient narrative?
- SourceThe "digital production diffuses faster than electrification" claim is asserted from one favorable internal case. Do external organizations actually redesign workflows quickly, or does the low cost of tool adoption mask slow, expensive process redesign (the real complement)? Partially answered — and the split is between the two halves of the question. Kalff & Simbeck find both happening at once in the same 410 firms: the low-threshold half diffuses faster than the organization, with 183 of 410 respondents using AI informally on personal devices regardless of employer policy, while the half that needs process redesign stalls exactly where the electrification analogy predicts — advanced analytics "seldom economically or logistically viable" without centralised data and standardised processes, and 20.2% of departments using no AI tool at all. So tool adoption does mask the absence of process redesign, but not by making it look fast: the two run on separate tracks, and the visible one requires no organizational change to happen. Still short of settling it — self-reported, cross-sectional, one function, one country, and no measure of redesign speed where it does occur.
- SourceWhich complement is the true binding constraint — access/permissions, skills, or review capacity? The paper lists all; it doesn't decompose their relative weight. Sharpened, not answered: the list itself is incomplete. Kalff & Simbeck's German evidence adds an institutional complement (works-council co-determination under BetrVG §87(1) no. 6, EU AI Act high-risk classification) that is not internal to the firm at all and that determines which capability is adoptable rather than how well it is used — so any decomposition needs a fifth term whose weight varies by jurisdiction rather than by firm. Their own most-cited internal blocker is data centralisation, which is closest to "access."
- SourceIf complements, not capability, gate value, does model progress have diminishing near-term returns until orgs catch up — and how long is that lag for agentic AI specifically?
- SourceHAT's P2 (middle-management vulnerability) is the vault's most-quoted-but-least-tested substitution claim, and nothing here can settle it: Ramp × Revelio resolves seniority only to entry-level / non-entry / manager-plus, which cannot distinguish "middle layers thinned first" from "manager-plus grew more slowly than the bottom." Falsifiable with a level-resolved employment panel (Revelio or matched employer-employee data cut by reporting depth, not seniority band) tracking layer counts before and after intensive AI adoption. The same instrument would settle P4, since it could compare regulated against unregulated industries on the same measure — though P4 now has its first field evidence from German HR (see the ledger row), which confirms the direction on self-report and leaves the panel test outstanding.
- WaitCorollary 5 claims irreversibility: once an AI agent is strictly cheaper on risk-adjusted grounds, the optimal allocation never reverts, given fixed human costs and negligible switching costs. The vault has no case either way, and this is falsifiable only by a future event — a documented instance of a firm re-staffing with humans a role it had already automated, for reasons other than a regulatory shock or a rise in AI risk (both of which the corollary's own conditions exempt). Watch for it in the same firm-level adoption panels; a reversal with human costs and regulation unchanged would falsify the corollary, and a long clean run of non-reversal would be weak support for it.
- SourceThe ownership variable is asserted as a boolean (your repo vs the company's), but employment contracts, work-for-hire doctrine, trade-secret law, and non-competes already govern externalized judgment. Does any jurisdiction or litigated case actually treat an employee-authored skill file as portable personal property rather than work product — and has any employer yet claimed ownership of one?
- SourceTan's compounding curve (week 4 flywheel, week 12 library-that-answers) is a personal anecdote against telemetry showing skills are copied once and rarely maintained (Agentic Work Systematization). Does any longitudinal measurement of individual skill-file libraries show quality or coverage improving with age, as opposed to accumulating?
- WaitHis first objection bets that better models raise the value of a personal library while Harness Shrinkage as Models Improve predicts scaffolding gets absorbed. These are separable — harness vs library — but no source tests the library half. Does a model generation that absorbs harness complexity also shrink the measured advantage of personal context?
- WaitMusk's deflation prediction is falsifiable and dated: does the price level of manufactured goods and AI-delivered services fall as robot deployment scales, or do input constraints (energy, land, minerals) keep it rising? Trigger: goods-vs-services price divergence through 2030.
- SourceThe end-effector premise is the checkable half of the abundance case — does the physically-gated share of tasks that Task Saturation: Broad but Shallow AI Diffusion measures actually fall as humanoid deployment scales, and at what rate?
- NowIf validation capacity is a commons and the Stockfish threshold is reached unevenly across domains, the commons is destroyed before the threshold arrives in the domains that still need validators. Is there any domain where the ordering has been observed?
- WaitThe forward test the report itself names: do the returns to expertise persist, narrow, or invert as models improve? A decrease would mean models are absorbing the judgment users currently supply.
- SourceOutcomes are transcript-inferred (verified success leans on git activity + explicit affirmation). How much of the management edge — and the whole success gradient — is real outcome vs. who-narrates-success-in-the-transcript?
- SourceThe study excludes headless / SDK / IDE usage (a "substantial share"). Does the returns-to-expertise pattern hold in non-interactive and pipeline use, where there is no human steering mid-session at all?
- WaitIs "intermediate captures most of the benefit" stable, or an artifact of current model capability — i.e., will the concave curve flatten further (everyone converges) or steepen (mastery starts to separate again) as models get better?
- Task Crossover3 open
- SourceCrossover is measured on consumer-surface ChatGPT messages from Business-account users. Does it hold in agentic/API/Codex usage, where work is delegated rather than typed — or does the occupational boundary reassert itself when the unit is a task handed to an agent? The report explicitly excludes Enterprise, so the most structured populations are unobserved.
- SourceThe dataset records what people attempted and nothing about outcome — the authors say so directly. Is borrowed work done as well as the specialist would have done it, and where does crossover stop being role expansion and start being unreviewed amateur output? A crossover measure joined to a quality or review-coverage measure would settle it, and nothing in the corpus currently does.
- SourceIs the workspace-size gradient about specialist availability (OpenAI's substitution story) or about permission and norms (a large firm's marketer may be allowed to touch less)? The two predict opposite things as small firms grow, and 2.5pp across a descriptive seat-count proxy is thin evidence for either.
- WaitATLAS is a two-week snapshot with no time dimension, while the AEI reports automation share rising. Does median task saturation move at all over a year, and in which direction?
- SourceThe expertise inversion (2.6× on lowest-expertise non-routine cognitive tasks) is measured on consumer surfaces. Does it hold on enterprise and agentic-coding traffic, where the task mix is deliberately harder?
- WaitAutor & Thompson predict opposite wage effects depending on whether AI absorbs an occupation's expert or inexpert tasks. ATLAS's snapshot points at inexpert. What signal would show the crossover if it happens?
- SourceSelection vs. treatment: tenure controls attenuate but don't eliminate the enthusiast-selects-into-delegation story. Does a within-person design (sentiment before/after adopting automated workflows) hold the effect?
- Self-reported "no learning loss" cannot detect real atrophy; is there an objective skill measure that agrees, or does measured skill diverge from felt skill (the AI Brain Fry direction)? Partially answered: Contractor & Reyes's randomized experiment supplies exactly the objective, unaided skill measure this survey lacks — and gives a both answer. It agrees that learning can persist under AI (augmentation users hold +0.29 SD test gains a week later, unaided), so felt-and-measured can align. But it also finds the divergence the question feared: automation users' gains vanish once AI is removed — and "automation share," the survey's own axis, pools both types, so a flat self-reported learning curve can hide a real deskilling half. (Different population — elite undergrads in a proctored lab, not workers — so this sharpens rather than closes the workplace-atrophy question.)
- SourceThe sample is heavily computer/math + management and 88% men; how much of the automation–optimism link survives in a representative population? Partially answered in a population about as far from this one as the vault contains: Jabarian & Henkel surveyed 2,764 Filipino entry-level customer-service applicants (60% female, wages ≈$280–435/month) and found 47% expect AI's workplace impact on themselves to be positive against 19% negative, with the belief predicting delegation choice — 77% of optimists, 72% of balanced, and 65% of pessimists chose an AI voice agent over a human recruiter to interview them. So the belief→delegation association survives a low-wage, non-Western, majority-female, non-technical population. Two things it does not settle. The direction is still unidentified (this is choice given belief, not sentiment given usage), and the same paper shows the association inverts by position: the recruiters, whose own task was the one being automated, split 68% "AI will have a significant personal impact" but only 12% "generally positive" — a quarter of the applicants' rate, inside the same firm. Optimism may track being served by AI rather than delegating to it, and this survey cannot separate those.
- SourceThe \$15–149B range rests on an assumed 0.5–5% time saving because no causal estimate exists for AI in the household. What experiment would measure actual household time savings, and does the effect survive contact with one?
- SourceATLAS argues gains skew to women (30% more productive household time) but could reverse given the AI adoption gender gap. Which effect dominates in current data?
- WaitIf AI substitutes household self-service for purchased professional services (tax prep, legal advice, therapy), measured GDP falls while welfare rises. Is that substitution detectable yet in the market-services data Coyle cites?
- SourceRoughly half the pooled break is venue composition (+1.72 → +0.75 pp/yr inside continuously observed venues), and OpenAlex changed its author-disambiguation pipeline in July 2023 — inside the window. Does the break survive in a corpus with stable curation and stable disambiguation (Scopus, Web of Science, or a single large publisher's internal records) over the same period?
- SourceSolo papers narrow 23% in content breadth while showing no movement toward new territory. Is that scope discipline (the author does what they can verify alone) or capacity limit (the LLM covers the execution but not the range a second mind supplied) — and does the quality of solo output diverge from coauthored output on citations, replication or retraction? The paper measures quantity and content and explicitly not quality.
- WaitThe break's attribution rests on a cross-field ordering the author concedes is an ordering, not a test. If the halt is LLM-driven it should track LLM capability, so the ordering should shift as models improve — fields whose execution work is newly automatable (lab protocol design, instrument control) should join late. If it is a one-time re-sorting, the ordering freezes and the halt decays.
- SourceMechanism 2 has never been directly measured in a workplace. Does a cohort that entered an AI-heavy profession after 2023 show lower unaided task accuracy than a matched earlier cohort at the same tenure? The paper specifies the design (no-AI assessment stratified by cohort and AI exposure); nobody has run it.
- SourceThe cohort evidence is a snapshot ending Sept 2025 in the most AI-exposed occupations. Does the 22–25 employment decline persist, reverse, or re-sort as agentic tooling matures — and if entry-level postings recover, does that restore the developmental content of the work or just its headcount? Recovery of positions and recovery of the regeneration mechanism are not the same event, and only the first is currently instrumented. Partially answered (2026-08-04): Indeed Hiring Lab instruments the first half — postings in the most-exposed occupation did rebound (US software development +15% since February 2025 against overall postings −7%) — and the composition answers the sub-question in this page's favour: 71% of the May 2025 – May 2026 increase is senior roles, 37% AI-titled, with the author himself conceding the market "could still be experiencing a seniority-biased technological change." So the recovery is real and is not entry-level, on this instrument. It leaves the harder half untouched: nothing there measures the developmental content of any role, the data are one job board's vacancy flow analyzed by that job board, and Ramp's firm panel finds entry-level headcount growing fastest on a different unit.
- SourceThe framework predicts differential depletion by its five factors. Do software engineering, financial analysis and legal research actually diverge from medicine and engineering on validation-capability measures — or does regulatory intensity turn out to be weaker protection than the model assumes?