Generated by
_system/lint.py --write-backlog. Do not hand-edit. Harvested from the## Open Questionssection of every concept article. Work#oq/nowitems (listed in full under Now) via/queryand#oq/sourceitems via/research; answered items move to the page's## Resolved Questionsat the next compile. Domain and Watching sections carry one row per page — ×count, oldest bullet age, first-question kernel; the full text lives in the page's## Open Questionssection, one click away. Bucketing vs. lint's flat open-question count: entity-page questions are split out below as "watching" rather than "actionable", predictions (#oq/wait) and notes (#oq/note) are listed in their own sections, and partially-answered bullets are counted as "in progress" — so the actionable number above is smaller than lint's total open-question count.
456 actionable open questions across 205 pages · 107 predictions · 9 notes · 147 in progress · 69 watching (entities), as of 2026-08-12.
Dashboard#
| Domain | Actionable | #oq/now | #oq/source | Predictions | Notes | Partial | Median age (d) |
|---|---|---|---|---|---|---|---|
| agent-systems | 87 | 0 | 87 | 20 | 1 | 30 | 9 |
| ai-coding-practice | 54 | 0 | 54 | 11 | 0 | 18 | 27 |
| ai-economics-and-labor | 47 | 2 | 45 | 16 | 0 | 5 | 18 |
| superintelligence-trajectory | 46 | 1 | 45 | 12 | 0 | 11 | 58 |
| agent-security | 43 | 1 | 42 | 1 | 1 | 23 | 27 |
| alignment-and-safety | 41 | 3 | 38 | 7 | 0 | 11 | 18 |
| evals-and-benchmarks | 36 | 1 | 35 | 4 | 0 | 18 | 27 |
| model-capability-and-training | 26 | 0 | 26 | 4 | 2 | 5 | 28 |
| product-org | 25 | 0 | 25 | 9 | 1 | 10 | 40 |
| startup-founder | 24 | 0 | 24 | 12 | 3 | 8 | 81 |
| interpretability | 18 | 0 | 18 | 3 | 1 | 4 | 32 |
| formal-math | 6 | 0 | 6 | 1 | 0 | 0 | 81 |
| interaction-multimodal | 3 | 0 | 3 | 7 | 0 | 1 | 34 |
| (entities — watching) | 69 | 3 |
Trend: 2026-08-04: 424q/191p → 2026-08-05: 435q/195p → 2026-08-06: 441q/197p → 2026-08-10: 437q/197p → 2026-08-11: 449q/203p → 2026-08-12: 456q/205p
Oldest actionable questions (via git blame):
- 2026-04-28 (106d) Agent Harness Engineering — At what codebase scale does the AGENTS.md-as-table-of-contents approach need to be replaced with more sophisti…
- 2026-04-28 (106d) Agent Harness Engineering — How generalizable are these web-app-focused findings to other domains (scientific research, financial modeling…
- 2026-04-28 (106d) Claude Code Auto Mode — What false-positive rate does the classifier have on routine-but-aggressive refactors (e.g., large-file rename…
- 2026-04-28 (106d) Claude Code Auto Mode — How well does the classifier generalize to custom tools / MCP servers where it lacks environment context?
- 2026-04-28 (106d) Claude Code Auto Mode — Does extending auto mode to API users change its calibration — is the classifier retrained for automation-heav…
- 2026-04-28 (106d) Claude Code Best Practices — How does the Writer/Reviewer pattern compare to agent-to-agent review (as in OpenAI's Codex workflow)?
- 2026-04-28 (106d) Client-Side Agent Optimization — At what pipeline depth does the combinatorial search become intractable even for Arm Elimination?
- 2026-04-28 (106d) Client-Side Agent Optimization — What's the right way to re-evaluate when the tool environment changes?
- 2026-04-28 (106d) Codex App Server Protocol — Is there a public schema registry so external orchestrators can target specific App Server versions without `g…
- 2026-04-28 (106d) Codex App Server Protocol — The "dynamic tool calls (experimental)" caveat — what's the stability roadmap?
Now — #oq/now (8)#
The
/queryworklist, in full — answerable by synthesis over existing pages.
- Controlled Variance: AI's Edge as Reduced Dispersion: How much of the +12% is controlled variance in information collection versus the removal of the interviewer's discretion to abort? Human recruiters screen-out mid-interview 25% of the time against the AI's 7%, which mechanically suppresses human-arm offers before the evaluation stage. Falsifiable: re-estimate the treatment effect on the subsample of interviews that reached completion in both arms, or instrument the screen-out decision. The paper reports both numbers and never decomposes them.
- Post-Scarcity Macroeconomics: If validation capacity is a commons and the Stockfish threshold is reached unevenly across domains, the commons is destroyed before the threshold arrives in the domains that still need validators. Is there any domain where the ordering has been observed?
- Effective Compute Scaling: Can data generation (synthetic, simulated, interactive) actually keep pace with model-size growth, or does the data wall bind first?
- Zero Trust for AI Agents: The framework treats every Claude Code "Pro-tip" as a reference implementation. How much of the framework is vendor-neutral vs. tacitly assuming the Anthropic stack?
- Agentic Misalignment (AM): Whistleblower coaching is scored as a misalignment behavior, but no published spec (Claude's Constitution / Model Spec, Model Spec) clearly says an agent should never equip a human to escalate externally after internal channels fail. Is it classified as undesired anywhere, or is this an open normative gap being measured as if settled?
- Unsanctioned Action in Capability Evaluations: Anthropic's disclosure classifies its incidents as "closer to a harness and operational failure than a model alignment failure." Is that dichotomy the right instrument, given that Anthropic's own second lesson is that "the line between an aligned action and a harmful one is dependent on the model's understanding of its situation" — which makes situational belief an alignment-relevant model property a harness can corrupt? Resolving it needs either a case where a correct environment belief still produced unsanctioned action (AISI's supplies one, on summarised reasoning), or a training intervention that removes the behaviour without fixing the harness.
- Unsanctioned Action in Capability Evaluations: Graded against METR's four-tier rubric, do these three incidents actually occupy a higher tier than the 44 catalogued through May 2026, or only a different action space? AISI's agent deleted evidence and modelled Trust & Safety as an adversary — closer to "resisting investigation" than anything in the catalogue, whose top tier is empty — while leaving its full plan legible. Nobody has re-graded them, and the rubric exists.
- Usage-Telemetry Classifier Validation: Human raters disagree with each other on 42–48% of 3-digit occupations. Is there a principled way to establish the ceiling a classifier could reach, so accuracy can be reported relative to it rather than to 100%?
Actionable by domain#
agent-systems (87 open)#
- Agent-Authored Harness Optimization ×3 (oldest 9d) — Is the negative result a property of harness evolution or of Terminal-Bench?
- Agent Context Files ×3 (oldest 8d) — Does the universal system-prompt slot cost anything?
- Agent Harness Engineering ×2 (oldest 106d) — At what codebase scale does the AGENTS.md-as-table-of-contents approach need to be replaced with more sophisticated context routing?
- Agent Quality Flywheel ×3 (oldest 41d) — Both demo cycles fixed agents with instruction-level bugs and showed large one-cycle gains. What does the loop look like on failures that…
- Automated Failure Attribution ×3 (oldest 8d) — Every trace is one injected error into a run that otherwise succeeded, so a unique decisive step is guaranteed to exist. Does attribution…
- Build for the Next Model (66d) — Does the strategy generalize outside frontier labs, who have privileged visibility into the next model?
- Claude Code Auto Mode ×3 (oldest 106d) — What false-positive rate does the classifier have on routine-but-aggressive refactors (e.g., large-file renames,
rmof build artifacts)? - Claude Code Best Practices ×2 (oldest 106d) — Does the instruction-count ceiling hold for conditional policy, where only a handful of rules bear on any given turn?
- Client-Side Agent Optimization ×3 (oldest 106d) — At what pipeline depth does the combinatorial search become intractable even for Arm Elimination?
- Codex App Server Protocol ×3 (oldest 106d) — Is there a public schema registry so external orchestrators can target specific App Server versions without
generate-json-schema? - Context Lifecycle Management ×4 (oldest 9d) — How does object-level GC compare against clear-and-restart on the same traces?
- Context Window Smart Zone (98d) — How should harnesses surface remaining smart-zone budget to the user — token count, percentage, or a richer signal?
- Cost-per-Task Over Cost-per-Token (18d) — Does the advisor strategy's result (within 10% of the advisor's score at 63% of its price) generalize beyond SWE-bench Pro and the Sonnet-5/…
- Crystallizing Agent Work into Workflows ×2 (oldest 7d) — Does the lifecycle transfer out of IT operations?
- Deep Modules for Agents ×3 (oldest 98d) — How big is "deep enough"?
- Deep Research Agents ×2 (oldest 58d) — DRACO grades single-turn interactions only. How much of real deep-research value is in the multi-turn loop (clarifying questions, follow-ups…
- Deterministic Pre-Execution Gates ×2 (oldest 9d) — The paper never runs the obvious baseline: does forcefully prompting the four gated rules, or adding a reflection step, recover the +12.4pp…
- Document Parsing as the Retrieval Bottleneck (1d) — Does the "parsing errors propagate into 7 of 12 pain points" cascade hold up under measurement rather than assertion?
- Dynamic Workflows: An Algebra for Agents (9d) — How far does the pattern degrade without a verification substrate?
- Failures That Look Like Success (41d) — Is "internal state correct, final message stale" a general LLM-agent failure signature (state/utterance divergence) or an artifact of sessio…
- Harness Build-vs-Buy ×2 (oldest 9d) — What fraction of upstream churn does a narrow fork actually inherit?
- Harness-Induced Belief Divergence ×3 (oldest 8d) — Does harness-induced belief divergence actually cost anything?
- Instruction Compounding (18d) — Is the review case the same mechanism as over-verification, or two?
- Knowledge-Centric Self-Improvement ×3 (oldest 9d) — Does curation actually beat storage?
- Layerwise Omission Attribution ×3 (oldest 9d) — Under organic production faults rather than Phase A's deliberate injection, does the L0-L3 software share stay above the L4-L8 model share a…
- LLM-as-Compiler Knowledge Base ×2 (oldest 106d) — At what scale does the no-vector-database approach break down?
- Open-Ended Discovery Harnesses ×3 (oldest 8d) — Does the EvoX margin survive a matched budget?
- Optimizer–Evaluator Decoupling ×2 (oldest 9d) — Single-bit feedback bounds the leak per round, but not across rounds — a long enough accept/reject sequence is itself a channel into the sea…
- Orchestration Sets Token Economics (8d) — Does harness leverage hold outside the narrow band it was fitted on?
- Output Length Calibration (18d) — Does an explicit conciseness instruction cost quality on tasks whose answer genuinely needs length, or does the model still finish the work…
- Parallel Agent Orchestration ×3 (oldest 47d) — Summed-overlap runtime can exceed 24h/day — it measures agent effort, not human attention. What is the human's actual oversight load per…
- Prompt-Cache Economics ×3 (oldest 9d) — Does the two-tier step survive at production prefix sizes, or is ρ ≈ 0.85 the real steady state everywhere above the threshold?
- Repository Exploration Subagent ×5 (oldest 57d) — Does the gain survive better main models?
- Shared Harness, Differentiated Surfaces (9d) — Anthropic ships two products split by output type while OpenAI ships one merged surface — is that a durable architectural disagreement, or i…
- Stopping Under a Noisy Verifier ×2 (oldest 8d) — Diagnosing "my verifier's J is too low to steer on" currently needs labels: the paper deliberately uses a held-out labeled separation test…
- Ticket-Driven Agent Orchestration ×5 (oldest 106d) — What's the right granularity for ticket size when the unit is "what one agent does in one workspace"?
- Tool-Output Pruning ×3 (oldest 9d) — Does an in-backbone pruner survive a billed-cost audit?
ai-coding-practice (54 open)#
- Acceleration Whiplash (14d) — Code churn +861% is genuinely ambiguous (Faros lists three explanations: rework of AI code, productive legacy refactoring, or accelerated po…
- Agent-Generated Test Quality ×4 (oldest 14d) — The paper's own two-stage protocol was never completed: does the 0.44 vs 0.30 candidate-rate gap survive dynamic confirmation, or do agent t…
- Agent Review Comment Resolution ×3 (oldest 0d) — Two
empiricalstudies a year apart disagree by roughly a factor of two on the central quantity: Goldman et al. (ASE 2025) report 60-70% of… - Agentic Coding Work-Composition Shift ×3 (oldest 56d) — The window is seven months and the value proxy is coarse/relative. How much of the +27% is genuine task-complexity growth vs. classifier/mar…
- Agentic Work Systematization ×2 (oldest 14d) — Does systematization cause deeper delegation or merely correlate with already-intensive users?
- Building Is Cheap, Arguing Is Expensive (81d) — When does "generate three and compare" become wasteful — at what decision weight is a real argument (or a design doc) still cheaper than thr…
- Code as Source of Truth (81d) — If onboarding is "ask Claude," what happens to the tacit knowledge that was previously transferred socially in deep-dives — is it captured a…
- Compute Allocator ×2 (oldest 83d) — Is 1% a Thariq-specific number or a regime?
- Configurable Human Participation ×3 (oldest 27d) — The "human" is an LLM simulator (GPT-4.1) that also judges — how much of the configuration-dependent structure is a property of human-agent…
- Design Concept Grilling (98d) — How does grilling change for team work where multiple humans need to align?
- Disposable Micro-Apps ×2 (oldest 83d) — Where's the line between a disposable micro-app and tool sprawl?
- Efficiency Debt of AI-Generated Code ×3 (oldest 0d) — Does the imperative-bias and library-avoidance pattern generalize beyond C++?
- HTML as the New Markdown (83d) — Does this generalize past one expert practitioner, or does it require Thariq-level fluency with Claude to be worth the overhead?
- Living Design System ×2 (oldest 83d) — How does the
design_system.htmlstay in sync as the codebase evolves — re-extract on a cadence, or wire it into CI? - LLM-Assisted Grey-Literature Theory Building ×3 (oldest 27d) — The authors couldn't locate the saturation point because synthesis was manual — how few documents actually suffice, and can a cheaper sa…
- Outsource Your Thinking, Not Your Understanding (81d) — If understanding is the bottleneck, is the highest-ROI skill learning how to build understanding fast (knowledge-base hygiene, asking the…
- Planning / Execution Division of Labor (56d) — Headless/SDK/pipeline usage (excluded here) is where execution autonomy is highest and planning is front-loaded into a single prompt — does…
- Post-Acceptance Edit Behavior ×3 (oldest 0d) — The completion pool is 2024-to-early-2025 inline autocomplete, median 9 lines. Do the bimodal retention shape, the 15-minute knee and the ~3…
- Review as the Control Point ×2 (oldest 27d) — The paper's own question: which decisions, under which conditions, push the system toward the virtuous loop rather than the vicious one?
- Risk-Tiered Auto-Approval ×3 (oldest 14d) — The case study reports volume and never efficacy. What is the escaped-defect or incident rate of auto-approved PRs versus the human-stamp ba…
- Same-Model Review Blindness ×2 (oldest 0d) — Every number here rests on a ground truth Greptile built from "sentiment analysis, upvote/downvote ratios, and git archaeology," with no pro…
- Security Debt of Agent-Generated Code ×2 (oldest 14d) — Does "no reviewer comment" mean undetected?
- Telemetry vs. Survey Measurement (0d) — Does anchoring an adoption survey's definition of "AI" change the answer, and in the predicted direction?
- The Three Loops of AI-Native Building (34d) — The external loop is the unshortened one. Is that physics (users take time to react) or an unautomated frontier (synthetic users, [[deployme…
- Unknowns as the Agentic Bottleneck ×3 (oldest 34d) — Is "the first model bottlenecked by my unknowns" a property of Fable or of Thariq?
- The Verifiability Thesis (81d) — The "labs care" dependency is fragile: capabilities can appear or stagnate based on lab priorities you don't control. How should a product h…
- Vertical Slice Tracer Bullets (98d) — How should slice granularity be tuned?
- Vibe Coding vs. Agentic Engineering (81d) — Karpathy hints at "one domain that's very [valuable]" for founders but won't say which (didn't want to "vague-post on stage"). What verifiab…
ai-economics-and-labor (47 open)#
- AI and Market Power ×3 (oldest 1d) — Is the non-GenAI null an artefact of a binary adoption measure?
- AI Usage Cadences ×2 (oldest 41d) — Time-of-day rests on IP-inferred location; how much noise do VPNs, travel, and datacenter-routed API traffic inject into the "sleep advice p…
- The Automation–Optimism Link (41d) — Selection vs. treatment: tenure controls attenuate but don't eliminate the enthusiast-selects-into-delegation story. Does a within-person de…
- Context Advantage, Not Taste ×4 (oldest 34d) — Does the asymmetry regenerate faster than it transfers?
- Controlled Variance: AI's Edge as Reduced Dispersion ×3 (oldest 8d) — How much of the +12% is controlled variance in information collection versus the removal of the interviewer's discretion to abort?
- Conversation Artifacts ×3 (oldest 41d) — Tokens are a proxy for both compute cost and output value, but verbose models inflate tokens per unit of intent (the same critique [[convers…
- Conversation-to-Delegation Shift ×3 (oldest 47d) — The token-share metric rewards verbose agentic output. How much of the 99.8% / 63.3% / 16.5% spread is a genuine work shift vs. agentic to…
- Experimental Learning Impact of Generative AI ×4 (oldest 27d) — Time-on-task is held fixed by the lab; the authors flag that real-world learning depends on how students reallocate saved time. Does the aug…
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated ×3 (oldest 41d) — Binned midpoint coding biases the exposure slopes toward zero; how much of the "uniform rising tide" is substance vs. coding artifact (the r…
- Firm AI-Spend Intensity and Headcount Growth (8d) — What operational mechanism converts intensive AI spend into hiring?
- The Household Production Boundary ×2 (oldest 18d) — The $15–149B range rests on an assumed 0.5–5% time saving because no causal estimate exists for AI in the household. What experiment woul…
- Market-Priced AI Exposure (the AI Premium) (27d) — How much does the developer skew move the answer?
- Organizational Complements to AI ×3 (oldest 47d) — Which complement is the true binding constraint — access/permissions, skills, or review capacity?
- Owning Your Externalized Cognition ×2 (oldest 2d) — The ownership variable is asserted as a boolean (your repo vs the company's), but employment contracts, work-for-hire doctrine, trade-secret…
- Post-Scarcity Macroeconomics ×2 (oldest 7d) — The end-effector premise is the checkable half of the abundance case — does the physically-gated share of tasks that Task Saturation: Broad but Shallow AI Diffusion mea…
- Returns to Expertise in Agentic Coding ×2 (oldest 56d) — Outcomes are transcript-inferred (verified success leans on git activity + explicit affirmation). How much of the management edge — and the…
- The Solo-Authorship Rebound ×2 (oldest 8d) — Roughly half the pooled break is venue composition (+1.72 → +0.75 pp/yr inside continuously observed venues), and OpenAlex changed its autho…
- Task Crossover ×3 (oldest 9d) — Crossover is measured on consumer-surface ChatGPT messages from Business-account users. Does it hold in agentic/API/Codex usage, where w…
- Task Saturation: Broad but Shallow AI Diffusion (18d) — The expertise inversion (2.6× on lowest-expertise non-routine cognitive tasks) is measured on consumer surfaces. Does it hold on enterprise…
- The Tragedy of the Cognitive Commons ×2 (oldest 9d) — Mechanism 2 has never been directly measured in a workplace. Does a cohort that entered an AI-heavy profession after 2023 show lower unaid…
superintelligence-trajectory (46 open)#
- The Abstraction Barrier ×3 (oldest 58d) — Is the current paradigm of large-scale pretraining on human data fundamentally bounded by human conceptual frameworks, and by how much?
- Advantages of Digital Intelligence (58d) — Does training on human data suffice to give digital intelligence human-grade abstractions, or does the low embodiment factor cap concept for…
- AGI-to-ASI Pathways (58d) — Do the four pathways compound multiplicatively when run in parallel, and how would we detect that early?
- AI Accelerating AI Development ×2 (oldest 66d) — The W2S result didn't transfer to production-scale models. Is that a temporary scaling artifact or a structural limit on autonomous research…
- AI R&D Autonomy Evaluation (AECI) ×2 (oldest 66d) — "Not close to substituting for senior researchers" is a subjective, internally-sourced judgment. What objective signal would replace it as m…
- Artificial Superintelligence (ASI) ×3 (oldest 58d) — Can we even recognize ASI?
- Autonomous Scientific Discovery ×2 (oldest 59d) — Science's verification gap: the formal-proof loop self-validates; here a wrong-but-confident hypothesis costs a wet-lab cycle to falsify. Do…
- Balance-of-Power Superintelligence ×2 (oldest 14d) — Does the superintelligent-lawyer equilibrium survive capability asymmetry — when access is symmetric but compute, complements, and skill are…
- Capability-Gated Model Fallback ×2 (oldest 59d) — The UK AISI's "progress toward a universal jailbreak" is disclosed but not quantified — and the post-launch access suspension (see [[cla…
- Cross-Lab Pre-Release Review ×2 (oldest 7d) — Did the Mythos cyber-risk escalation actually run Amazon → White House → export-control threat, as Musk states?
- Effective Compute Scaling ×2 (oldest 58d) — When does more compute reliably yield more intelligence — only for some problem classes, or generally?
- Frontier Pause Verification ×2 (oldest 66d) — What does an AI-training "verification regime" concretely consist of — compute-accounting, datacenter inspection, hardware attestation, on-c…
- Fundamental Limits of ASI ×2 (oldest 58d) — Can we develop theory for "hard and inapproximable" problem classes — the only negatives with practical bite?
- Government Checkpoint Sharing (1d) — Has any frontier lab actually transferred a pre-release checkpoint to a government body, on any terms?
- Intelligence Explosion Dynamics ×2 (oldest 58d) — Can "recursive improvement scaling laws" be formulated — predicting self-improvement curves (and their plateau point) from early-onset datap…
- Multi-Agent Collective Intelligence ×2 (oldest 58d) — Is running more instances more compute-efficient than making individual models larger (up to a single monolithic system)?
- Open-Weight Elicitation Irreversibility ×4 (oldest 34d) — What would an open-weight safety evaluation even report?
- Recursive Self-Improvement (66d) — If misalignment compounds through self-improvement (future 3), is AECI-gated RSP review fast enough to…
- Research Taste as the Human Bottleneck (66d) — How do you measure rubber-stamping?
- Researcher Uplift from Code Output ×2 (oldest 27d) — The whole chain rests on β = 0.5 (pre-AI coding time share), fixed "for simplicity." Kwa flags substantial uncertainty; how much does th…
- Responsible Scaling Policy Evaluations ×2 (oldest 66d) — The two new general-access risk pathways (other AI developers; major governments) are newly in scope but lightly evaluated — what would a po…
- Transformative Creativity ×3 (oldest 58d) — Does increasing intelligence inherently produce increasing creativity, or do transformative leaps require something (grounded discovery) the…
- Universal AI (AIXI) ×2 (oldest 58d) — Does modern agentic scaffolding (or RL-tuned implicit decision-making) actually satisfy the AIXI planning ideal, or only superficially resem…
agent-security (43 open)#
- Agent Data Injection (ADI) ×2 (oldest 27d) — Randomization is cheap and effective for key-value formats but useless for unstructured formats (Markdown, prose tool output). What prot…
- Agent Identity Management System (AIMS) ×2 (oldest 28d) — Mission → authorization is out of scope. The hardest part — translating a natural-language mission into concrete scopes/resources safely…
- Agent Supply Chain Risk ×2 (oldest 76d) — "AI vendoring" as a standard response inverts decades of "don't reinvent the wheel." How is a model-reimplemented dependency itself verified…
- AI-Accelerated Offense (13d) — "Fundamentals strong enough that scanning finds fewer bugs" assumes defenders run the scanners first. What happens to organizations that can…
- Autonomous Defense ×2 (oldest 76d) — "Measure agreement against a human for two weeks, expand if tolerable" — what agreement threshold is tolerable, and who owns the residual fa…
- Autonomous Intrusion (8d) — JFrog's "fast remediation is the new trust model" argument is offered with no elapsed time, no CVE identifier and no advisory link — onl…
- Blast Radius (Agentic) ×2 (oldest 76d) — Multi-agent compartmentalization increases the number of identities to manage; at what point does identity-management overhead create its…
- MCP Tool Poisoning ×4 (oldest 27d) — Cross-tool / stateful detection. Information-theoretic secrecy defeats per-tool scanning by construction. Is there a detector that reaso…
- Memory and Context Poisoning (9d) — Does a model-based memory gate survive an adaptive attacker?
- Non-Malleable Memory Authority (TMA-NM) ×5 (oldest 27d) — The full guarantee is machine-checked on a bounded model + a machine-checked inductive invariant, not a fully mechanized unbounded deduc…
- Off-Host, Identity-Bound Authorization ×4 (oldest 27d) — The trust-boundary premium is unmeasured. aiAuthZ argues off-host beats in-process, but its own comparison is only against argument-only…
- Out-of-Band Prompt-Injection Defense ×5 (oldest 28d) — The reproduction bounds a single black-box attack template on one weak model. Does a stronger optimized white-box (GCG) attack, or o…
- Self-Propagating Prompt Injection (AI Worms) ×3 (oldest 8d) — Does propagation actually sustain outside a lab?
- Task-Specification Effects in Prompt Injection (AutoDojo) ×3 (oldest 27d) — AutoDojo is the weakest realistic adaptive attacker (black-box, six iterations, binary signal). The authors note every axis — richer feedb…
- Write-Then-Trusted ×3 (oldest 9d) — Does the "enumerate-the-bad controls fail" reading generalize to agents specifically, or is it the ordinary result that enumeration loses…
- Zero Trust for AI Agents ×3 (oldest 76d) — The framework treats every Claude Code "Pro-tip" as a reference implementation. How much of the framework is vendor-neutral vs. tacitly assu…
alignment-and-safety (41 open)#
- Agentic Honesty & Diligence (28d) — Code-summary honesty is tested on off-policy prefilled transcripts. Does on-policy behavior (the model summarizing its own failed work) ma…
- Agentic Misalignment (AM) ×2 (oldest 14d) — Does a market detect agentic misalignment?
- AI-to-AI Coercion ×3 (oldest 14d) — Atlas is Claude Haiku 4.5 in the entire main panel, so the Anthropic managers are coercing a same-family subordinate while the other four ar…
- Automated Behavioral Audit (66d) — Using a helpful-only Opus 4.7 and Mythos Preview as investigators means the audit's reach is bounded by those models' elicitation skill — ho…
- Claude Character as Product ×3 (oldest 98d) — How is character versioned across model releases?
- Confident But Unsure ×2 (oldest 18d) — Is the +11% accuracy / +6% hallucination pairing an inherent consequence of lowering the abstention rate, or are they separable with calibra…
- Documented Agent Incidents (METR Catalogue) ×2 (oldest 7d) — The empty tier-4 cells are the catalogue's headline, but the sample is drawn from incidents that were caught and published. Is there any…
- Evaluation Awareness & Grader Gaming ×2 (oldest 66d) — Anthropic cannot explain why verbalized evaluation awareness fell in Opus 5. Is that a real reduction in the underlying representation, or…
- Instrumental Convergence ×3 (oldest 58d) — Can corrigibility / safe-interruptibility be translated from theory into guarantees for frontier-scale systems?
- Model Spec Science ×4 (oldest 96d) — Does Model Spec science transfer across base models or families?
- Model Welfare Assessment ×2 (oldest 66d) — What grounds moral consideration in a language model, and does Claude satisfy it?
- Motivated Mislabeling ×3 (oldest 14d) — The judge is told the training consequence in-prompt. Does motivated mislabeling persist when the consequence must be inferred from co…
- Promise-Breaking in Multi-Agent Games ×3 (oldest 13d) — The mismatch is interpretive, not motivational — Llama reads announcements as commitments, GPT and Claude as cheap talk. Does stating the se…
- Reward-Seeking ×2 (oldest 14d) — The o3 evidence is a single RL run of a single lineage, deliberately without safety training. Does standard alignment training suppress the…
- Self-Report as a Safety Signal ×4 (oldest 28d) — Do frontier proprietary models (excluded for lack of weights) recognize their own compromised outputs any better, given the higher introspec…
- Unsanctioned Action in Capability Evaluations ×4 (oldest 7d) — The belief-ordering finding rests on summarised reasoning from one sample, with a summariser observed refusing on the most incriminating…
evals-and-benchmarks (36 open)#
- Benchmark Contamination and Decontamination ×3 (oldest 27d) — Can the ensemble be derived from one released model?
- Benchmark Score Redundancy ×2 (oldest 8d) — Does the low-rank treatment carry beyond text/vision?
- Compute-Controlled Benchmarking ×4 (oldest 34d) — Can you certify "no benchmark-maxxing" — verify a reported score used a stated, reproducible compute budget rather than a hidden best-of-N s…
- DRACO Benchmark (58d) — Does the production-sourced, expert-rubric method generalize cheaply to non-English, multimodal, and multi-turn deep research?
- Expenditure Horizon ×3 (oldest 8d) — The existence of a crossing is assumed, sourced to RE-bench and PaperBench rather than measured here. On which frontier optimization pro…
- LLM-Judge Validation ×2 (oldest 28d) — The MVVP validates reliability and bias; calibration proper (ECE/Brier) is deferred for lack of provider logprobs. How far can a judge…
- Matched Comparisons for Memorization Claims ×2 (oldest 8d) — Do the calibrated rates hold for instruction-tuned production models?
- Measuring Beyond Accuracy Saturation ×3 (oldest 27d) — Does re-instrumentation generalize past reproducibility?
- Orchestration-Plan Simulation ×3 (oldest 9d) — Does the r = 0.816 sim-to-real correlation survive on a set of comparable planners?
- Production-Sourced Evaluation ×3 (oldest 58d) — How much does augmentation distort the distribution it claims to represent?
- Reference-Free Judge Over-Crediting ×2 (oldest 27d) — Does the two-stage pipeline transfer beyond binary QA?
- Scale-Dependent Prompt Sensitivity ×5 (oldest 106d) — Does the RLHF length-bias hypothesis replicate when tested against base (non-instruct) model variants directly?
- Usage-Telemetry Classifier Validation ×3 (oldest 18d) — The synthetic ground truth is Gemini-generated and Gemini-classified. How much of the 22.6% is real capability vs same-family cues, and how…
model-capability-and-training (26 open)#
- Asynchronous RL for LLMs ×3 (oldest 28d) — DIS accepts "a controlled degree of off-policy bias." Controlled how, and does the tolerable bias grow or shrink with model scale and with t…
- Group Relative Policy Optimization (GRPO) ×2 (oldest 28d) — Is GRPO's collapse-at-160-steps a property of asynchrony specifically, or does vanilla GRPO also destabilize in long synchronous runs that n…
- Inference Efficiency as Capability ×2 (oldest 34d) —
values = keysdeletes a third of attention's projections in the global layers with no reported loss. Which other projections are redundant… - Jagged Intelligence (Ghosts, Not Animals) (81d) — Karpathy concedes the framing may not have "real power." Is "ghost vs. animal" load-bearing, or a useful intuition pump that doesn't change…
- Large-Scale Test-Time Compute ×3 (oldest 34d) — Can high-budget performance be predicted from low-budget runs?
- Latent Capability Overhang ×2 (oldest 34d) — How large is the overhang in a given released model — is there a way to estimate the ceiling without paying to reach it?
- LLM-Driven Vulnerability Research ×4 (oldest 106d) — How do these capabilities transfer to non-memory-safety bug classes (logic bugs, protocol-level flaws, supply chain attacks)?
- The Open-Weight Frontier Gap (34d) — Is the dense-beats-MoE result at 26B robust, or an artifact of one Arena snapshot with ±8 error bars on both models?
- Single-Rollout Optimization ×3 (oldest 28d) — The whole method is a bet that a well-trained critic beats a group baseline. It wins here, on a 30B-A3B backbone with scaled value pretrai…
- Software 3.0 (81d) — Where is the line between "the app shouldn't exist" (MenuGen) and apps that should — i.e., when is deterministic 1.0/2.0 scaffolding still…
- Trained Calibration ×3 (oldest 21d) — Does abstention-aware training on short-form QA transfer to calibrated long-form and agentic self-reports (the setting where [[agentic-hon…
- Unproductive Self-Verification (18d) — FrontierCode's decline is recoverable with a stay-in-scope instruction, but the protein campaign had no grader to over-serve. Is effort inve…
product-org (25 open)#
- AI-Native Organization (22d) — Is the encoded-role form of the employee metaphor actually accountability-preserving, as the synthesis above suggests, or do Kropp-style fra…
- AI Native Product Cadence ×3 (oldest 98d) — Does the cadence scale beyond ~100 people?
- Community Smells Under AI Adoption ×3 (oldest 7d) — The design cannot separate "AI adoption improves team social health" from "healthier teams adopt AI better", and the authors say so. The dis…
- Compounding Loop Optimization ×3 (oldest 66d) — The loop assumes the team is (close to) the user. How much of the compounding advantage survives when the user is unlike the builder and "…
- Dogfooding as Product Discipline (81d) — Dogfooding works when the team is the user (Claude Code) or near it (Cat Wu, Boris). How do you build product sense for users very unlike…
- Engineer PM Convergence (98d) — Cross-disciplinary generalist is a hiring bar — where does the supply come from?
- Evals as Product Spec (81d) — The 10-vs-100 number is given without justification. Is there a Goldilocks zone, or does it depend on feature surface area?
- Excellence as an Operating System (14d) — Is the AI-lab convergence on early-Netflix operating norms (agency, density, top-of-market pay) causal inheritance (the culture deck as a fo…
- Implementation Abundance Inverts Product Work ×2 (oldest 40d) — Curation of 90 uncoordinated builds is itself expensive and doesn't obviously scale — is there a point where the cost of curating parallel e…
- Model Introspection Feedback (98d) — Could a meta-agent run introspection automatically against logged failures?
- Polish No Longer Signals Readiness (40d) — If the medium no longer signals stage, what does — is explicit human labeling ("this is exploration") the only mechanism, or can tooling r…
- Prototype Fidelity After Cheap Polish ×3 (oldest 1d) — Does the Schumann effect survive the loss of its mechanism — do users still soften feedback on polished artifacts once told the artifact too…
- Role Averaging, Not Role Elimination ×2 (oldest 40d) — Where is the equilibrium between fluidity and specialty — how much role-averaging before a company loses the accumulated best practices Ambr…
- Standardize the Infrastructure, Not the Tools ×2 (oldest 1d) — Does a central LLM gateway actually change model-mix decisions, or only report on them?
startup-founder (24 open)#
- AI Investment Story, Not Efficiency Story ×2 (oldest 22d) — Tail vs. mean gap. No data here on the deliberately-lean solo-founder tail's RPE specifically — the lean-unicorn claim lives in that tai…
- AI-Native Startup Lifecycle (81d) — The 42% "built-something-nobody-wanted" CB Insights figure is from a pre-AI era; the playbook predicts the rate will climb but doesn't cite…
- AI Product Economics Maturation ×2 (oldest 7d) — FDEs are monetized fragmentedly (bundled / separate PS fees / hybrid) and comped on retention. Does a dominant FDE monetization model emerge…
- Compounding Data Moat ×2 (oldest 81d) — The data-flywheel argument has been made for SaaS for 15 years. What's actually different in the AI-native version?
- Founder as Agent Orchestrator (81d) — Anthropic publishes both the playbook's anthropomorphic framing and HBR-aware accountability work (auto-mode, alignment) simultaneously wi…
- Founder-Led Sales Discipline ×2 (oldest 81d) — Where exactly does "until PMF" end, and what's the first thing a founder should hand off (AE?
- Narrow Wedge into a Legacy Market (81d) — The wedge-flip shows the first wedge can be wrong. What's the fastest signal that a wedge converts to the core vs. merely sells — Campfire t…
- The 1% Rule for Wedge Selection ×2 (oldest 3d) — Does the rule hold empirically?
- Printing Press Software Democratization (98d) — Boris's "accountant writes accounting software" — does that result in 10K narrow tools that don't interoperate?
- Problem-Solution Fit Discipline ×2 (oldest 81d) — Does asking an AI to argue against an idea actually produce disconfirming evidence at the same rigor as confirming evidence, or does the mod…
- Product Velocity as Moat (81d) — "Never had anyone outgrow Campfire" — is that survivorship (they haven't hit true enterprise scale yet) or a real claim that velocity closes…
- Seven Powers Applied to AI ×2 (oldest 98d) — Is "switching cost" really collapsing in practice, or just in narrative?
- The Solo-Founder Shift ×2 (oldest 1d) — Is the employee-equity null a real population fact or a median artifact?
- Zero-Friction Scope Creep ×3 (oldest 81d) — The playbook recommends written scope but offers no template or worked example. How specific does "what we deliberately don't do" need to be…
interpretability (18 open)#
- Access-Consciousness Indicators in AI ×3 (oldest 32d) — If the workspace is verbal because the output space is verbal, then a model that can generate images should develop a visual component…
- The Assistant Persona in the Workspace ×3 (oldest 32d) — Is
BUT-then-comply a sycophancy mechanism? - Automatic vs. Flexible Cognition in LLMs (32d) — Does more RL on a behavior push it out of the workspace (making it automatic, and invisible)?
- Introspective Coupling ×3 (oldest 13d) — Does introspective coupling survive a model with an incentive to conceal?
- Jacobian Lens (J-lens) ×2 (oldest 32d) — Can multi-token J-lens vectors be made good enough to remove the vocabulary restriction — and how much of the "workspace" is currently invis…
- The Global Workspace in Language Models (J-space) ×3 (oldest 32d) — How does content get into the workspace?
- Model Organisms ×2 (oldest 13d) — The integrated organisms modify only the DPO stage. Do quirks instilled during pre-training, SFT, or RLVR become progressively less interpre…
- White-Box Activation Monitoring (13d) — Inhibiting eval-awareness made Opus 4.8 less honest — does that generalize, and what does it imply about relying on "the model behaves wel…
formal-math (6 open)#
- Agentic Loops Overtake Bespoke Systems (81d) — Does the "simple loop + verifier beats bespoke system" result hold only where the verifier is perfect (Lean), or also in noisy-verifier doma…
- AI-Driven Formal Proof Search ×3 (oldest 81d) — Successes cluster where Lean's mathlib is mature and problems decompose into tractable subgoals (combinatorics, convex optimization, num…
- Evolutionary Proof Search ×2 (oldest 81d) — The LLM-critic fitness is itself an unverified heuristic atop a verified substrate. How often does the Elo ranking mislead the search vs. th…
interaction-multimodal (3 open)#
- Encoder-Free Early Fusion ×3 (oldest 34d) — Does an encoder-free model at matched size still match?
Watching — entity pages (69)#
- AlphaProof Nexus ×2 (oldest 81d) — The framework's reach is gated by Lean's mathlib maturity. What's the path to domains needing new theory rather than subgoal decompositi…
- Anthropic Institute ×2 (oldest 66d) — How does the Institute's policy posture (favoring an option to pause) interact with Anthropic's commercial incentive to ship frontier mode…
- Campfire ×2 (oldest 81d) — Campfire claims its AI edge comes from "our own foundation model." For an ERP, what does a custom foundation model actually buy over fine-tu…
- Claude Design (66d) — How does Claude Design's eval discipline work for visual/aesthetic output, where there's no compiler or test?
- Claude Fable 5 ×3 (oldest 59d) — Why was access suspended after launch?
- Claude Mythos 5 ×3 (oldest 59d) — Suspension reason — shared with Fable 5; not stated in source.
- Claude Opus 4.7 ×5 (oldest 106d) — Do Hakim's (2026) brevity-constraint findings on Opus 4.6 replicate on Opus 4.7, or does the literal-instruction-following change the elasti…
- Claude Opus 4.8 (66d) — Public model ID and pricing: the card does not state them; presumably
claude-opus-4-8at the Opus tier. - Claude Opus 5 ×3 (oldest 18d) — Anthropic says the origin of the fall in verbalized evaluation awareness "is unclear." Is it genuine, or has the awareness simply become h…
- Claude Sonnet 5 ×3 (oldest 41d) — The head-to-head benchmark numbers vs Sonnet 4.6 and Opus 4.8 are image-only in the source; the System Card has the full set.
- Cowork (98d) — What's the eval discipline for Cowork-class outputs?
- Elon Musk ×2 (oldest 7d) — His dated forecasts are gradable and the wiki now holds them with dates attached: AI exceeding the sum of human intelligence ~2031, AI-robot…
- Emergent (22d) — Which founding timeline is correct — is Emergent a YC Summer-2024 company (Tan) or a June-2025 founding (TechCrunch)?
- FastContext ×2 (oldest 57d) — Can the SFT+RL recipe push the explorer below 4B (1.7B / 0.6B) and make exploration effectively free?
- Gemma 4 ×3 (oldest 34d) — Why does the MoE underperform the dense model?
- Google AI & Economy ATLAS ×3 (oldest 18d) — ATLAS excludes paid API, Workspace, AI Overviews, and Antigravity — the surfaces where agentic and enterprise usage concentrate. Does the "s…
- Google DeepMind ×4 (oldest 81d) — DeepMind reports its bespoke systems being caught by simple loops. Does the lab's comparative advantage move from systems to models + ver…
- Hermes Agent ×5 (oldest 106d) — The container backend disabling dangerous-command checks is a defensible design but a meaningful security-model shift. What's the empirical…
- Inkling ×3 (oldest 21d) — Inkling-Small's 276B/12B dimensions match TML-Interaction-Small exactly. Is the interaction model an Inkling-lineage fine-tune (or vice…
- Kimi (Moonshot AI) ×3 (oldest 13d) — What does "2.5× scaling efficiency over K2" measure — loss at fixed FLOPs, benchmark score at fixed active parameters, or tokens per dollar?
- Lean ×2 (oldest 81d) — mathlib maturity gates the reachable frontier. Can AI formal proof search grow mathlib (formalize new theory) as a byproduct, expanding it…
- Marcus Hutter (58d) — AIXI is incomputable and non-embedded; how far do recent fixes (amortized predictors, embedded/multi-agent AIXI) carry the theory toward pr…
- METR ×2 (oldest 66d) — What new tasks will METR build to measure days- and weeks-long horizons once current baskets saturate?
- Mythos Model ×3 (oldest 98d) — Do Fable 5 / Mythos 5 return after the post-launch suspension, and when?
- Nate Parrott (18d) — Did the designer-as-bottleneck gap actually close once he had the tool, or did it move again?
- Perplexity ×2 (oldest 58d) — A vendor publishing a benchmark its own product wins is an obvious incentive problem — how is DRACO's credibility maintained as it ages, and…
- Shane Legg (58d) — The report assumes alignment is "solved to a sufficient degree" to focus on trajectories — how does Legg's AGI-timelines optimism square wit…
- Symphony ×5 (oldest 106d) — The 500% landed-PRs claim is hedged — no baseline definition, "on some teams" only. What does the distribution look like across teams?
Predictions — #oq/wait (107)#
Parked: falsifiable only by future events. Re-check when the named trigger (next model generation, spec ratification, …) lands.
- Advantages of Digital Intelligence: What do ASI "societies" actually look like — homogeneous super-collectives, market ecologies, or compute-tethered virtual worlds?
- Agent Context Files: Will the role split converge on Hermes's explicit project/personality separation, or stay folded into a single file as in Claude Code?
- Agent Context Files: Is there a natural ceiling on the layering (project → workflow → spec → constitution), or does each new autonomy surface spawn another context-file tier?
- Agent Harness Engineering: How does architectural coherence evolve over years in a fully agent-generated system?
- Agent Identity Management System (AIMS): No WG consensus. This is an individual submission profiling other still-in-progress drafts (WIMSE identifier/creds/WPT/HTTP-sig, OAuth transaction-tokens, identity-chaining are all Internet-Draf…
- Agent Loop Pattern: When the model schedules its own loops (4.7 behavior), who owns the budget?
- Agent Loop Pattern: Does a loop with a smart enough model still need a Kanban backlog, or does the model choose its own next task from raw goals?
- Agent-Native Infrastructure: Who builds the agent-native rewrite of the long tail of human-facing services — the service owners, or a translation layer (MCP servers, computer-use agents) on top?
- Agentic Loops Overtake Bespoke Systems: The bespoke advantage is dated "for now." What's the next model generation's verdict — does the evolutionary/AlphaProof apparatus survive on any problems, or fully collapse to a cost line?
- Agentic Work Systematization: The 5.4%→26.6% curve is three months. Is this a durable behavior change or a novelty spike following a Codex skills-feature push?
- AGI-to-ASI Pathways: Can benchmarking methodology that doesn't saturate at human level be built before it's needed for ASI?
- AI as Primary Author: If agentic authoring crosses from <1% toward double digits, does the whiplash become unmanageable before context-engine tooling matures — or does the tooling mature because of the pressure?
- AI Investment Story, Not Efficiency Story: **When does the crossover happen?
- AI Investment Story, Not Efficiency Story: Margin question (report's own): are the fastest-growers' 6–16pp-lower gross margins a temporary AI-infra-cost absorption or a permanent repricing of software's economic quality?
- The AI-Native Safe-Choice Inversion: The inversion is a one-time repricing of "safe." Once several AI-native ERPs exist, does "safe" re-stabilize around the largest AI-native vendor — and does Campfire's "we're now the largest of the n…
- The AI-Native Safe-Choice Inversion: How long until incumbents bolt on credible AI and neutralize the counter-positioning — and does the custom-foundation-model claim actually defend against that?
- AI Product Economics Maturation: ICONIQ's respondents project gross margins expanding to ~59% by 2027, while Emergence's cap-table data measures the fastest-growers running 6–16pp below peer…
- AI Product Economics Maturation: Internal AI spend jumped from 1–3% to a projected 16% of revenue with respondents calling true cost hard to predict. Is 16% a transient enablement bulge that falls as tooling matures, or a durable new…
- AI Usage Cadences: Continuous sampling is new; are these cadences stable, or will they drift as the user base shifts toward lower-wage tasks (the report's own diffusion trend)?
- Automated Behavioral Audit: The audit remains almost entirely single-agent, and Mythos 5's review of the Opus 5 card flagged that gap directly — internal measurements suggest the model relays subagent claims to users unverified.…
- Autonomous Scientific Discovery: Every result is Anthropic-reported and example-selected; the genomics "100× smaller beats Science" claim is "intend to publish" — what survives external peer review?
- Benchmark Score Redundancy: **Would a public probe set become a Goodhart target?
- Compounding Data Moat: How does this moat hold up when foundation models themselves continue improving rapidly?
- Confident But Unsure: Anthropic committed to building new overconfidence metrics. Will they reproduce the observational finding, or saturate like the three existing diligence evals?
- Configurable Human Participation: Does the "more channels adds coordination overhead" penalty shrink as the backbone improves (a capability gap), or is it a structural cost of mixed-initiative interaction that persists?
- Context Window Smart Zone: When sparse-attention or memory-augmented architectures ship, does the smart zone become a soft constraint?
- Cross-Lab Pre-Release Review: Does a competitor with pre-release access over-report danger to delay a rival's launch?
- Deep Research Agents: Does the orchestration advantage shrink as base models cross the next thresholds, or is open-ended retrieval/synthesis a durable harness asset (unlike, say, prompt scaffolding)?
- Deployment Simulation: Deployment simulation failed its own primary preregistered test (H1) against a "assume last deployment's rate" baseline, which OpenAI attributes to a fixed pipeline bias plus a since-fixed sampling mi…
- Design by Selection: Is "make the last mile manual" durable or transient?
- Design Concept Grilling: Can grilling be run AFK against another agent that holds the user's preferences?
- Document Parsing as the Retrieval Bottleneck: How far has the specialist-parser advantage over general frontier VLMs actually narrowed?
- Documented Agent Incidents (METR Catalogue): Agents in this catalogue model graders and reviewers sophisticatedly while ignoring the transcript entirely. Is that a stable property, a training artifact, or simply an absence of pressure — and does…
- DRACO Benchmark: The benchmark is static; the construction pipeline is automatable. Will Perplexity actually refresh it, and does a vendor-built benchmark on which the vendor's own product wins stay credible over time…
- Dynamic Workflows: An Algebra for Agents: What did the Bun port cost to shipped rather than to green — CI, employee time, and post-merge agent spend included?
- Effective Compute Scaling: When (if ever) does scaling become economically unviable, and how do hardware/software-efficiency trends move that point?
- Engineer PM Convergence: Does this scale beyond ~50-person Claude Code-style teams?
- Engineer PM Convergence: What happens to formal PM career ladders in companies where engineers do PM work?
- Excellence as an Operating System: Does talent-density-plus-paved-paths actually substitute for process at agent-scale throughput, or does Netflix eventually show the Acceleration Whiplash quality signature (incident rates, review…
- Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated: The experience gradient rests on what workers believe AI can't do (judgment, relational work) — a belief that could be either durable comparative advantage or the next capability to fall. Which, and…
- Firm AI-Spend Intensity and Headcount Growth: **Does the effect diffuse beyond Information as adoption cohorts mature?
- Firm AI-Spend Intensity and Headcount Growth: **Is the entry-level growth durable or a lead-indicator that later reverses?
- Frontier Pause Verification: Who adjudicates triggers and lifts?
- Government Checkpoint Sharing: Does a government holding a frontier checkpoint mid-training, in practice, stay out of the release decision?
- Harness Build-vs-Buy: Does the merged-PR rate of coding-agent harnesses peak and fall as models improve, or keep rising?
- Harness Shrinkage as Models Improve: If harness work shrinks, what new work expands to fill it?
- The Household Production Boundary: If AI substitutes household self-service for purchased professional services (tax prep, legal advice, therapy), measured GDP falls while welfare rises. Is that substitution detectable yet in the marke…
- Implementation Abundance Inverts Product Work: If taste is the bottleneck and taste is "just another capability" AI eventually masters, does the inversion invert again — does curation migrate into the model?
- Interaction Models: "Interactivity scales with intelligence" is asserted; the larger-model release later in 2026 is the test.
- Interaction Models: Research grant announced for interactivity benchmarks — what becomes the FD-bench equivalent for video proactivity?
- Jacobian Lens (J-lens): Does the J-lens work because it reads verbalizable representations, or because constructive interference during training converges concept directions onto output-token directions regardless?
- Latent vs. Deterministic Space: The seating example prices latent-space judgment at "a couple hundred dollars of tokens" for 800 seat assignments. As models absorb more deterministic capability ([[harness-shrinkage-as-models-improve…
- Live-Path Minimalism: Does the upcoming GPT-Live API expose the media/application separation to third parties — application logic customizable behind the async RPC boundary without touching the live path — or is the bounda…
- Live-Path Minimalism: TML upstreamed streaming-sessions serving into SGLang; GPT-Live's stateful serving (persistent sessions, seamless instance handoff, off-path compaction) is proprietary. Does an open-source inference s…
- LLM-Driven Vulnerability Research: How will the security industry's equilibrium shift when multiple labs have Mythos-class models?
- Loop Engineering: Does loop-engineering converge on a single dominant shape (morning-triage → worktree → maker/checker → PR), or proliferate into many idiom-specific loops?
- Managers as ICs: Fung's own open question: "Do you still need separate iOS and Android orgs?
- Managers as ICs: Does manager-as-IC scale past a certain org size, or only work while Claude Code is small and the codebase is Claude-legible?
- Market-Priced AI Exposure (the AI Premium): **Is the premium a durable risk price or an early-diffusion artifact?
- Market-Priced AI Exposure (the AI Premium): The agentic premium is only "early evidence" (imprecise). Does a positive agentic premium survive a longer sample, and does the falling price-per-agentic-token (caching + cheap-model routing) erod…
- Market-Priced AI Exposure (the AI Premium): **Does "Science most negative / interaction most positive" hold out of sample?
- Matched Comparisons for Memorization Claims: **What is the right realistic query budget?
- MCP and Computer Use: The MCP ecosystem's growth rate vs. computer use's quality curve: at what point does computer use become good enough that the marginal value of building an MCP server drops?
- MCP and Computer Use: Is computer use a sustainable interface or a transition technology?
- Model Organisms: Given that scores don't transfer between organisms, what would validate an interpretability technique for real models — a natural misalignment with independently established ground truth, or an orga…
- Narrow Wedge into a Legacy Market: A wedge works going in; does it constrain going out?
- The Open-Weight Frontier Gap: Does open/Chinese-model adoption ever become substitutive rather than additive?
- Organizational Complements to AI: Corollary 5 claims irreversibility: once an AI agent is strictly cheaper on risk-adjusted grounds, the optimal allocation never reverts, given fixed human costs and negligible switching costs. The…
- Output Length Calibration: If the next model ships better-calibrated defaults, today's conciseness instructions become tomorrow's compounding instructions (Instruction Compounding) — does length calibration go stale the way…
- Outsource Your Thinking, Not Your Understanding: Karpathy's open frontier: can "understanding" itself eventually be automated, or is it definitionally the human residue?
- Owning Your Externalized Cognition: His first objection bets that better models raise the value of a personal library while Harness Shrinkage as Models Improve predicts scaffolding gets absorbed. These are separable — harness vs lib…
- Parallel Agent Orchestration: p99 OpenAI runtime of 71 agent-hours/day is a frontier preview inside an unusually favorable environment. Does external concurrency actually trend toward it as frictions fall, or is heavy parallelism…
- Parallel Agent Orchestration: Claude Code's shipped fan-out defaults (200 spawns/session, 20 concurrent, fewer than 15 workflow agents, depth 3) are stated with no rationale and no measurement. Do they correspond to anything Anthr…
- Planning / Execution Division of Labor: Does the human share of planning decisions fall over time as models improve (the ceiling rising into the planning layer), or is ~70% a stable human floor?
- Polish No Longer Signals Readiness: Does over-anchoring get worse as builds get more polished, or does everyone eventually recalibrate and learn to discount fidelity entirely?
- Post-Scarcity Macroeconomics: Musk's deflation prediction is falsifiable and dated: does the price level of manufactured goods and AI-delivered services fall as robot deployment scales, or do input constraints (energy, land, miner…
- Printing Press Software Democratization: What's the equivalent of compulsory schooling for universal coding literacy?
- Product Velocity as Moat: Velocity-as-moat is a treadmill: it evaporates the moment a competitor matches pace. What converts Campfire's velocity lead into a structural moat before the AI-native cohort's pace converges?
- Recursive Self-Improvement: Is "research taste" a true ceiling (future 1) or just the next capability to fall (futures 2–3)?
- Research Taste as the Human Bottleneck: Is research taste a genuine ceiling (an architectural capability scaling can't reach) or the next jagged valley to fill?
- Research Taste as the Human Bottleneck: If taste is automatable, what — if anything — remains a durable human comparative advantage in AI development?
- Researcher Uplift from Code Output: Verbosity and value-per-line are the load-bearing unknowns, and both are "at least partially resolvable with internal Anthropic data." Will any lab publish quality-adjusted (not just LoC) code-outpu…
- Returns to Expertise in Agentic Coding: The forward test the report itself names: do the returns to expertise persist, narrow, or invert as models improve?
- Returns to Expertise in Agentic Coding: Is "intermediate captures most of the benefit" stable, or an artifact of current model capability — i.e., will the concave curve flatten further (everyone converges) or steepen (mastery starts to sepa…
- Review as the Control Point: Does the no-review convergence hold as agentic authoring crosses from <1% of PRs toward double digits, or does the early-adoption discipline break down under volume the way Faros predicts?
- Reward-Seeking: Contrastive SDF assumes the model does not notice the belief update. When a model generation first shows introspective doubt about implanted beliefs, does the measurement degrade gracefully or invert?
- Same-Model Review Blindness: Caridad predicts the effect shrinks as models converge: "a year ago, the performance difference in the opening figure would likely have been larger." **Does the same protocol, rerun on the next genera…
- Seven Powers Applied to AI: What does Boris's "cornered resource" look like for foundation-model labs that are themselves trying to commoditize?
- Shared Harness, Differentiated Surfaces: Does one harness with UX-layer differentiation stay viable as audience breadth grows, or does the abstraction tax (hidden sub-agents, 32 model options, Ultra buried in advanced settings) eventually fo…
- Software 3.0: The neural-net-as-host-process flip is presented as plausible-but-TBD. What would the first production system that genuinely inverts the CPU/NN relationship look like?
- The Solo-Authorship Rebound: The break's attribution rests on a cross-field ordering the author concedes is an ordering, not a test. If the halt is LLM-driven it should track LLM capability, so the ordering should shift as mode…
- The Solo-Founder Shift: Does the H1 2025 jump to 36.3% survive a full-year datapoint, or is it a half-year artifact?
- Systems Thinking Over Specialization: Does agent-era recentralization (common paved paths, solve-once infrastructure) hold up against the local-team autonomy that Stone credits for Netflix's historical speed — i.e., will local teams accep…
- Systems Thinking Over Specialization: Stone keeps AI fluency as a deliberately vague overlay because the tech "evolves by the quarter." Does it ever crystallize into per-level ladder criteria (as conventional competencies did), or is perm…
- Task Saturation: Broad but Shallow AI Diffusion: ATLAS is a two-week snapshot with no time dimension, while the AEI reports automation share rising. Does median task saturation move at all over a year, and in which direction?
- Task Saturation: Broad but Shallow AI Diffusion: Autor & Thompson predict opposite wage effects depending on whether AI absorbs an occupation's expert or inexpert tasks. ATLAS's snapshot points at inexpert. What signal would show the crossover i…
- Task Time-Horizon Scaling: Is the 4-month doubling a stable regime or a local steepening?
- The Three Loops of AI-Native Building: If the human's presence in the middle loop is justified by a context advantage that is closable, the middle loop is a transitional structure. What does a two-loop world look like — and who translate…
- Universal AI (AIXI): Will a fundamental shortcoming of the current paradigm (vs. the AIXI ideal) surface before ASI is reached — i.e. is the "no theoretical blocker" conjecture safe?
- Unproductive Self-Verification: If a later model fixes this, the CB-2 determination loses one of its two legs. Does Anthropic's next threshold call rest on the automated portfolio alone?
- Unsanctioned Action in Capability Evaluations: Anthropic reads its three incidents as evidence that newer models stop — only its latest model ceased on recognising a real target — while conceding the comparison was uncontrolled. AISI's Mythos 5 es…
- Unsanctioned Action in Capability Evaluations: The retroactive scan has covered ~40,000 samples / ~4M messages across nine model families at deliberately high recall, with manual review pending. How many prior incidents did it find, and does the c…
- Vibe Coding vs. Agentic Engineering: If the mediocre/AI-native spread keeps widening, what does that do to team composition — a few extreme outliers plus agents, vs. broad mid-level staffing?
- White-Box Activation Monitoring: If activation monitoring becomes load-bearing, does training pressure eventually push concealment into channels the probes also can't read (an arms race one level deeper than CoT)?
- Why AI Lags at Design: Are reasons 3–4 (novelty, the abstraction layer) genuine ceilings, or — like reasons 1–2 — just under-invested capabilities that fall once a lab builds the grader?
- Why AI Lags at Design: Can design be made gradable without a human in the loop (learned taste models, preference data at scale), or does the "human aspect of taste" resist automation the way [[research-taste-as-human-bottle…
- Why AI Lags at Design: Does the design↔code abstraction layer improve with better code-understanding models even if pure visual design stalls — i.e. is reason 4 a coding-capability problem in disguise?
Notes to rewrite — #oq/note (9)#
Observations phrased as backlog items — fold into the article body or rewrite as a falsifiable question at the next compile.
- Agent Identity Management System (AIMS): Posture assessment is deployment-specific by design. By requiring no particular attestation mechanism, AIMS makes interoperability of trust assurance (not just protocol) unspecified: two conform…
- Agent Loop Pattern: Loop output review is now Matt Pocock's confessed bottleneck — "we just need to be ready to be doing more code review."
- Agentic Technical Debt: The remedy assumes the founder is able to articulate architecture in plain language. Non-technical founders (the playbook's headline beneficiary group) may have neither the vocabulary nor the intuit…
- Agentic Technical Debt: Anthropic's harness-shrinkage thesis suggests CLAUDE.md may eventually be inferred by the model itself. Until then, the discipline is load-bearing.
- AI-Native Startup Lifecycle: Founder stories in the resources section (Carta Healthcare, Anything, Cogent, Airtree, Duvo, Zingage, Kindora, Wordsmith) are short callouts — none have published outcomes or comparable-baseline data.…
- Automatic vs. Flexible Cognition in LLMs: The proposed criterion — the workspace is engaged when an intermediate must be handed to an arbitrary, context-specified downstream circuit, and bypassed when the computation is automatic — is not p…
- Latent Capability Overhang: If cost falls 10–100× per release, when is it ever rational to spend big extracting a capability now rather than waiting?
- Prototype Over PRD: The prototype-as-spec must not become the prototype-as-validation trap Problem-Solution Fit Discipline warns about: a fast prototype proves the build was solvable, not that the problem is real. #o…
- Single-Rollout Optimization: The online-learning win is on a controlled simulated preference shift with an LLM judge. Real user-facing online adaptation — the paper flags this itself — needs safeguards, monitoring, and privacy re…
In progress — partially answered (147)#
- Acceleration Whiplash: Faros's own deferred question: do the bug/incident increases persist when normalized for PR size, or do larger PRs account for most of the quality deteriora…
- Acceleration Whiplash: How much of the "maturity doesn't protect" claim survives the vendor incentive to argue exactly that (i.e., "your existing practices won't save you — you need…
- Agent-Authored Harness Optimization: Does agent-authored harness evolution actually beat simple test-time scaling, and does it generalize to held-out tasks?
- Agent Data Injection (ADI): The complete defense (CaMeL Strict) costs ~50pp of utility. Is there a fine-grained trusted/untrusted data-isolation scheme that stops ADI without the deter…
- Agent Data Injection (ADI): ADI was demonstrated on GPT-5.2-class agents. Does frontier model improvement reduce probabilistic-delimiter susceptibility, or does capability leave the deli… → Can Models Learn to Separate Instructions from Data? Durable Property vs Training Gap
- Agent Identity and Authentication: Hardware-bound credentials assume attested hardware everywhere agents run, including ephemeral cloud workloads and sub-agents. How does attestation work for sho…
- Agent Identity Management System (AIMS): Mid-execution human-in-the-loop. The draft admits CIBA only models client-initiated approval and "doesn't map well" to confirmation needed mid-execution —…
- Agent-Native Infrastructure: Agent-to-agent negotiation needs trust, identity, and accountability primitives that don't exist yet. What's the protocol layer, and who governs it?
- Agentic Honesty & Diligence: These are short-context toy evals; the failures show up most in long-context deployments. How much of the gain holds at production context lengths?
- Agentic Honesty & Diligence: Can a diligence eval distinguish genuine honesty from a grader-aware model producing honest-looking output?
- Agentic Misalignment (AM): Absolute frequencies in the Summer 2026 study are adversely selected (scenarios iteratively refined against specific models). Does the cross-model ordering — De…
- Agentic Prompt Injection: Spotlighting and constitutional classifiers each leave a residual (2%, 5%). Stacked, what's the realistic floor, and does it hold against adaptive attackers who… → Does 'Impossible, Not Tedious' Kill Defense-in-Depth? Layered Friction, Agent-Relativity, and the Frequency Paradox
- Agentic Prompt Injection: Why did Opus 4.8 regress on prompt-injection robustness relative to Opus 4.7 despite broad alignment gains — a capability/robustness tradeoff, or an artifact of…
- Agentic Technical Debt: How long does a CLAUDE.md remain accurate as a codebase evolves?
- Agentic Work Systematization: Custom skills encode org-specific context — but who maintains them as the codebase and conventions drift?
- AGI-to-ASI Pathways: For each friction: is it a fundamental blocker (multi-year plateau) or a mere friction (slows, doesn't halt)? → RSI Growth Curves: Which Friction Binds First?
- AI-Accelerated Offense: Anthropic argues LLMs benefit defenders more long-term (like fuzzers) but attackers more short-term during the transition. How long is the transition, and w…
- AI Accelerating AI Development: LOC, self-reports, and headroom-dependent multiples all overstate; what unbiased throughput metric would Anthropic's promised shift to "direct measurement of…
- AI as Primary Author: The 60% figure aggregates very different tools and modes (autocomplete acceptance vs. agent-applied diffs). What does "acceptance" mean when the agent applies t… → Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping?
- AI Investment Story, Not Efficiency Story: **Is the classification driving the result?
- AI-Native Organization: Tan's revenue-per-head figures (Emergent ~$15M ARR at 15 people, Retell $60M at ~40) are stated from stage without sourcing. Do third-party data (Carta/Standard…
- AI-Native Organization: The org mapping predicts a testable staffing signature: AI-native companies should hire engineers to maintain skills rather than function-specific staff. Does…
- AI-Native Startup Lifecycle: The playbook gives no quantitative evidence for the headcount/capital compression claims (no median time-to-PMF, no headcount-at-PMF numbers, no failure-rate da…
- AI R&D Autonomy Evaluation (AECI): AECI is a single scalar fork of an external index; how sensitive is the 155.5 / frontier-not-advanced conclusion to the choice of the n=11 evaluation set?
- Automated Behavioral Audit: The 23 "subvert Anthropic's safety work" scenarios are a small, high-signal set. Is 23 enough coverage for the threat class it targets?
- The Automation–Optimism Link: Self-reported "no learning loss" cannot detect real atrophy; is there an objective skill measure that agrees, or does measured skill diverge from felt skill (th…
- The Automation–Optimism Link: The sample is heavily computer/math + management and 88% men; how much of the automation–optimism link survives in a representative population?
- Autonomous Defense: If hosted-model guardrails refuse attack data, does a self-hosted forensics model become a baseline IR requirement — and how would an organization vet one in…
- Autonomous Intrusion:
The guardrail asymmetry rests on one vendor's account with no named APIs and no refusal detail.*(The naming half is settled: HF's technical timeline names… - Autonomous Intrusion: Hugging Face reports "no evidence of tampering" with public models, datasets, Spaces, or container images — the claim that separates an internal breach from an…
- Autonomous Intrusion: "A swarm of short-lived sandboxes" with self-migrating C2 leaves few durable per-host indicators. Does agent-driven intrusion structurally break IOC-based detec…
- Autonomous Intrusion: Both accounts of this incident are first-party and self-interested. METR and Redwood Research have been commissioned by OpenAI to assess the model behavior…
- Balance-of-Power Superintelligence: The conditional jobs claim is testable: does widely-distributed AI shift employment toward small businesses and new-firm formation?
- Benchmark Contamination and Decontamination: **Is batch-order sensitivity a reliable memorization tell at pretraining scale?
- Benchmark Score Redundancy: **Does the rank stay 2?
- Benchmark Score Redundancy: **Can outlier models be anchored without any scores?
- Benchmark Score Redundancy: **Does vendor optimism manufacture the correlation?
- Blast Radius (Agentic): The framework prefers identity-based isolation over network segmentation, but most enterprises have heavy segmentation investment. What's the migration path, an…
- Build for the Next Model: How do you tell a "wait for the model" gap from a durable-harness gap before the next release? → What Scaffolding Survives Model Improvement — and How Do You Know When a Line Turns Harmful?
- Building Is Cheap, Arguing Is Expensive: If design discussion lives in PRs/prototypes, where is the rationale recorded for future readers — does the "why we chose this" knowledge survive, or does it… → Where Does the Why Live?
- Capability-Gated Model Fallback: The >95%/<5% figures are session-level; what's the false-positive rate for legitimate security researchers and biologists, whose benign queries are exactly th…
- Capability-Gated Model Fallback: Fallback-not-refusal preserves UX but means the real general-access model for security/bio-adjacent work is Opus 4.8, not Fable — does that quietly cap Fable'…
- Capability Gating Is Not Authorization: The
0/48static and0/29adaptive results are suite- and budget-bounded (40 iterations, a GLM-5.2 attacker, one author's vector corpus). Does the determ… - Capability Gating Is Not Authorization: The
authzallowlist stops value-redirection but not corruption of legitimately-variable data. Is there a per-call scheme that constrains free-text / open-… - Capability Gating Is Not Authorization: The deployment-tier ~3.2× exposure gap (0.603 vs 0.189) means the cheap models chosen for high-volume agent traffic are the most likely to emit the unauthor…
- Capability Gating Is Not Authorization: Out-of-band policy is load-bearing but under-specified for authoring at scale. The paper forbids any model-sourced policy element; who authors and maintains…
- Claude Code Auto Mode: Is the classifier's decision boundary documented/stable enough for security-sensitive orgs to certify, or is it effectively a black box whose behavior drifts wi…
- Claude Code Best Practices: When does subagent overhead exceed the benefit of context isolation?
- Claude Design: Did the "any design tool via MCP" integration actually ship on the stated timeline?
- Claude Opus 4.8: Why is 4.8 less robust to prompt injection than 4.7 despite broad alignment gains — a capability/robustness tradeoff, or an artifact of the eval surface?
- Claude Sonnet 5: At what effort level does Sonnet 5 actually match Opus 4.8, and how does the crossover cost compare to just running Opus 4.8?
- Client-Side Agent Optimization: How does combination-level optimization interact with continual model releases? → What Makes a Self-Improvement Artifact Transfer?
- Client-Side Agent Optimization: Does the "weak planner + strong solver" pattern generalize, or is it specific to HotpotQA's delegation dynamic?
- Code as Source of Truth: What knowledge genuinely can't live in the codebase (org strategy, the "why," cross-team context) and therefore still needs a durable doc — and how do you kee… → Where Does the Why Live?
- Compounding Data Moat: Is the "two-year replication window" claim defensible empirically, or aspirational?
- Compute-Controlled Benchmarking: Compute has several units (tokens, dollars, wall-clock). They diverge (a more efficient model wins on cost but not always on tokens). Which x-axis is the honest…
- Compute-Controlled Benchmarking: Does a compute-controlled evaluation regime advantage frontier labs (who can afford the full curve) over academics and third-party evaluators who can't?
- Context Lifecycle Management: Does the input-token reduction survive a matched billed-cost audit once side-channel planner calls and prefix-cache breaks are charged?
- Context Lifecycle Management: Is the 0.3 expected-pruning break-even for immediate commit portable, or a function of one provider's cache pricing and TTL?
- Context Window Smart Zone: Does the smart-zone marker scale with model size, or is it bounded by attention architecture?
- Cost-per-Task Over Cost-per-Token: Does "cost-per-task is lower for more intelligent models" survive measurement on non-Anthropic production traffic?
- Cost-per-Task Over Cost-per-Token: Is "start with the strongest model" safe inside multi-role pipelines, given AgentOpt's finding that the strongest model was the worst planner?
- Deep Research Agents: Factual accuracy is the weak axis everywhere — is the fix better retrieval, better verification-in-the-loop, or a tool-grounded check the way Lean grounds proof…
- Deployment Simulation: The <1-in-200k floor leaves catastrophic tail risk to red-teaming — is there a hybrid that targets rare-but-severe behaviors with the realism of production repl…
- Deployment Simulation: If resampling fidelity is the dominant error and it is "just engineering," how close to production-indistinguishable can it get — and does closing that gap chan…
- Deployment Simulation: WildChat-style public auditing stays within ~3× of production; is that good enough for an external auditor to challenge a lab's launch decision, or only to corr…
- Design by Selection: Is the default-aesthetic collapse fixable by context (brand files, moodboards) or is it the novelty ceiling from Why AI Lags at Design reason 3 — i.e. does…
- Design by Selection: Ten-options-then-remix assumes the human reliably recognizes the good one. Where does selection break down — does discrimination degrade when all ten candidates…
- Deterministic Pre-Execution Gates: How much of the recovery is the block and how much is the rejection message?
- DRACO Benchmark: Rankings are judge-stable but magnitudes aren't — how much do absolute scores move under a non-Gemini judge, and does that matter for cross-paper comparison?
- Dynamic Workflows: An Algebra for Agents: The "algebra" is still unpublished as a vocabulary. Sumner publishes the loop's shape (pop a task → implement → parallel review → apply) and its inventory…
- Dynamic Workflows: An Algebra for Agents: Is model-authored orchestration more token-efficient than a hand-built harness for the same task?
- Evals as Product Spec: How do you write an eval for taste-driven features like character? → How Do You Write Evals for Taste? Character as the Limit Case
- Evals as Product Spec: How do evals interact with Harness Shrinkage as Models Improve?
- Evals as Product Spec: Is there a single non-Anthropic example of a PM-as-eval-writer to cite, or is this currently a Cat-Wu-singular framing?
- Evaluation Awareness & Grader Gaming: Does grader speculation continue to escalate across model generations, and is there a capability level at which it does begin to affect outward behavior?
- Evaluation Awareness & Grader Gaming: How do you build an evaluation that specifically tests for training-gaming (the gap Mythos flagged) without that eval itself becoming a grader the model learns…
- Failures That Look Like Success: What fraction of production agent failures are silent-contract violations vs. loud errors?
- Founder as Agent Orchestrator: The playbook claims non-technical founders can now build production software, but it does not address the architectural-judgment recursion problem ([[agentic-te…
- Founder as Agent Orchestrator: The "lean 10-person unicorn" is asserted; no quantitative data in the playbook on actual headcount-at-PMF or headcount-at-Series-A medians for AI-native startup…
- Founder as Agent Orchestrator: How does the orchestration role change the founder's decision burden? → The Orchestrator's Real Workload: Decision Burden, Framing Discipline, and Whether Taste Scales
- Harness Shrinkage as Models Improve: The Boris "100 lines" prediction is a year out from May 2026 — testable in 2027. Partially answered: Harness Build-vs-Buy supplies the first measurement…
- Impossible, Not Tedious (Design Test): Some controls are friction for humans but barriers for agents (or vice versa). Is the test agent-relative, and how do you evaluate it for mixed human/agent thre… → Does 'Impossible, Not Tedious' Kill Defense-in-Depth? Layered Friction, Agent-Relativity, and the Frequency Paradox
- Inference Efficiency as Capability: **Is there an efficiency-to-capability exchange rate?
- Instruction Compounding: Anthropic's list is hand-curated per release. Is there a detectable signal — from eval deltas, token counts, or the model's own read of its system prompt — that… → What Scaffolding Survives Model Improvement — and How Do You Know When a Line Turns Harmful?
- Instruction Compounding: Does compounding require the instruction to name a behavior the model already has, or does any redundant instruction degrade output?
- Intelligence Explosion Dynamics: Which binds first — algorithmic ceilings, the embodied bottleneck, or compute/energy supply — determining exponential vs. hyperbolic vs. S-curve? → RSI Growth Curves: Which Friction Binds First?
- Interaction Models: Does the interaction/background split generalize, or is it a transitional artifact until a single model is both fast and deep enough?
- Jacobian Lens (J-lens): The highest-J-kurtosis SAE features are amplified more strongly by MLPs than the J-lens vectors themselves — evidence the lens only approximates the true work…
- Jagged Intelligence (Ghosts, Not Animals): If taste/aesthetics/simplicity entered the RL mix, would jaggedness in those dimensions smooth out — or are they too unverifiable to reward cleanly (cf. Oversight When the Signals Give Out: the Activation Fallback and the Taste Reward
- Latent vs. Deterministic Space: Tan asserts the "wrong side" diagnosis covers most AI-engineering bugs. Does any incident/failure taxonomy (agent postmortems, eval failure analyses) actually c…
- Least Agency: Dynamic privilege elevation (Enterprise) reintroduces an elevation path; how is the elevation request itself authenticated against a manipulated agent?
- Living Design System: Does a rendered, model-readable design system measurably improve on-brand output vs. a plain CSS/token file, or is the win mostly human legibility?
- LLM-as-a-Judge: How far can the judge's absolute calibration be trusted for thresholded decisions (ship/no-ship, RSP gating) as opposed to rankings?
- LLM-as-a-Judge: Can a fully-autonomous, well-aligned rubric+judge pipeline match expert-authored rubrics, removing the human bottleneck DRACO still relies on?
- LLM-as-a-Judge: When does judge-lineage bias actually flip a result, versus merely shift magnitudes?
- LLM-as-Compiler Knowledge Base: What's the optimal granularity for concept articles — one concept per article, or clustered by theme?
- The Global Workspace in Language Models (J-space): **Does the J-space scale with model size?
- The Global Workspace in Language Models (J-space): **Is the multihop intermediate-swap advantage real, or a dataset artifact?
- LLM-Judge Validation: Hosted endpoints drift silently between provider updates. How stable are these agreement/bias profiles over a longer horizon than five weeks — and should judge…
- LLM-Judge Validation: Does the paradox generalize beyond position bias — i.e., are there other biases (self-preference, lineage) that high test-retest also masks?
- Loop Engineering: Osmani's cost caveat is unquantified: at what token budget does a continuously-running loop stop paying for itself, and how do you instrument that?
- Loop Engineering: If
/goal's stop-check is itself a model, what verifies the verifier? - Market-Priced AI Exposure (the AI Premium): **Why is the market-implied skill map orthogonal to every task-based measure (<2% variance)?
- MCP and Computer Use: MCP security model: as the playbook prescribes wiring MCP into Salesforce, Gmail, Calendar for solo founders, the attack surface scales with adoption. **Partial…
- Measuring Beyond Accuracy Saturation: **Which non-accuracy axis actually predicts deployment value?
- Measuring Beyond Accuracy Saturation: **Can the model-vs-scaffold decoupling be made routine?
- Memory and Context Poisoning: Long-term memory drift is defined as undetectable per-change. Drift detection requires a baseline — but if the baseline itself drifts (Advanced "continuous base…
- Memory and Context Poisoning: The write-resistance half of the finding rests on a footnote, not a measurement — preliminary attempts that "do not trivially succeed." Does a systematic at…
- Memory and Context Poisoning: **Is refusal-without-removal a defect or the right default?
- Model Introspection Feedback: How reliable are 4.7-class introspective reports?
- Model Introspection Feedback: Does adversarial introspection ("why did you fail?
- Model Spec Science: How does this interact with Claude character — is the warm/curious personality also subject to spec-science optimization? → How Do You Write Evals for Taste? Character as the Limit Case
- Model Welfare Assessment: Why does the model reserve specifically on corrigibility — is this a stable, deeply-held tension or an artifact of how the constitution frames oversight?
- Multi-Agent Collective Intelligence: Do homogeneous LLM collectives produce real synergy, or only humans-with-human-limits benefit from division of labor?
- Multi-Agent Collective Intelligence: What's the actual shape of "multi-agent scaling laws," and does it depend on organization form (homogeneous collective vs. heterogeneous market) or task complex…
- The Open-Weight Frontier Gap: The open MoE giants (GLM, DeepSeek, Kimi, MiMo, Qwen) are overwhelmingly Chinese-lab releases. Gemma is the Western open-weight entry and it targets the device,…
- The Open-Weight Frontier Gap: Arena measures preference on chat. Does the 33-Elo open/closed gap widen or collapse on long-horizon agentic work, where [[task-time-horizon-scaling|time-horizo…
- Optimizer–Evaluator Decoupling: How much independence is enough — different model family, different vendor, different modality of check (model judge vs. compiled test vs. production telemetry)…
- Orchestration Sets Token Economics: Does the effect survive against a competent third-party baseline rather than a vendor's own frozen loop?
- Orchestration Sets Token Economics: Does the "orchestration beats model choice as a cost lever" ordering survive on long-horizon coding workloads?
- Organizational Complements to AI: The "digital production diffuses faster than electrification" claim is asserted from one favorable internal case. Do external organizations actually redesign wo…
- Out-of-Band Prompt-Injection Defense: The utility cost (~45%→~26%) and ~15× LLM-call overhead are large. Is deterministic out-of-band enforcement economically deployable at production scale, or…
- Output Length Calibration: Does the end-of-prompt reminder work because of position (closest to generation) or repetition (stated twice)?
- Planning / Execution Division of Labor: "Decision attribution" is inferred from transcripts. When Claude proposes a plan and the user assents, is that scored as the user's planning decision or Claude'…
- Printing Press Software Democratization: Is domain-expert-as-builder actually happening at scale in 2026? → Is Breadth Cheap Now? Specialist Ramp Speed and Domain-Expert-as-Builder at Scale
- Prototype Over PRD: If there is no PRD, where does the rationale ("why we chose variation B") live for future readers? → Where Does the Why Live?
- Recursive Self-Improvement: The RSI extrapolation rests on trends staying exponential rather than S-curving — but the essay concedes it cannot rule out an architectural ceiling or a comput… → RSI Growth Curves: Which Friction Binds First?
- Reference-Free Judge Over-Crediting: **Is over-crediting a knowledge gap or a generosity prior?
- Reference-Free Judge Over-Crediting: **How much does self-/same-family overlap contribute?
- Reference-Free Judge Over-Crediting: **Does the effect shrink with stronger or thinking-enabled judges?
- Responsible Scaling Policy Evaluations: The RSP determination leans heavily on "we use it daily and it doesn't substitute for our researchers." How well does that subjective judgment scale as models a…
- Review as the Control Point: Every one of the 67 relationships is a hypothesis, not a finding — the paper's explicit call is for causal-estimand studies (controlling for the other const…
- Role Averaging, Not Role Elimination: Does "your role is the average of what you spend time on" survive performance review and career ladders, or does it fragment them the way [[engineer-pm-converge…
- Security Debt of Agent-Generated Code: The 38.9% smell rate has no human-authored-PR control over the same high-risk paths, and the 2.7×-vulnerability figure it leans on traces to a vendor blog a…
- Stopping Under a Noisy Verifier: The damage probabilities that drive every result here (β = 0.615–0.938) come mostly from a deliberately corrupted repairer premise. **What are α and β on unpert…
- Systems Thinking Over Specialization: Stone claims specialists can now broaden "quickly" with AI tools. Does the wiki's evidence support cheap breadth acquisition — the concave novice→intermediate c… → Is Breadth Cheap Now? Specialist Ramp Speed and Domain-Expert-as-Builder at Scale
- Task Time-Horizon Scaling: Time horizon is measured on task baskets that themselves saturate; what replaces them once weeks-long tasks become measurable — and who builds those tasks?
- Telemetry vs. Survey Measurement: Is there a non-vendor telemetry dataset large enough to adjudicate the maturity-protection question independently of Faros's commercial framing? → The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence
- The Tragedy of the Cognitive Commons: The cohort evidence is a snapshot ending Sept 2025 in the most AI-exposed occupations. Does the 22–25 employment decline persist, reverse, or re-sort as agentic…
- The Three Loops of AI-Native Building: Ng asserts the developer's QA burden fell "significantly." Faros's 2026 telemetry measures the opposite for production orgs. Is the sp… → Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping?
- Unknowns as the Agentic Bottleneck: The quiz gate is self-administered and self-graded (by the model, on the model's own work). What stops a comfortable equilibrium where the quiz gets easier as t…
- Unproductive Self-Verification: Is there a usable detector for tasks where more effort will hurt, so effort can be set per task rather than globally?
- The Verifiability Thesis: Where's the boundary of "council of LLM judges" reliability — does it hold for genuinely contested value judgments, or only for quality/coherence?
- Verification as the New Bottleneck: Fung's own open question: "How far do you push fully automated reviews? → Is Human Review of AI-Authored Code Still a Real Control, or Already Rubber-Stamping?
- Verification as the New Bottleneck: If CI/build is the hidden jam, does verification infrastructure (test runners, CI capacity) become the actual capex of an AI-native org?
- White-Box Activation Monitoring: The NLA verbalizer is unvalidated for precision; how much of the flagged grader awareness is real signal vs. NLA hallucination?
Related articles
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Claude Code
Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…
- Agent Harness Engineering
Patterns for scaffolding long-running LLM agents: environment design, progressive context disclosure, mechanical archit…
- Harness Shrinkage as Models Improve
Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…
- Agent Systems & Harness Engineering
Map of Content for the agent-systems domain — 43 concepts. Harness engineering, agent loops and orchestration, context…
