H
Howardism
Howardism · Vol. 03Plate II · No. 02

Startup & Founder, in order.

Notes16DomainStartup & FounderOpen Qs47Newest11 Aug 2026Oldest6 May 2026

Building AI-native companies: speed, moats, and lifecycle.

Map of Content for the startup-founder domain — 16 concepts. Curated entry point; see Home for all domains.

  • Agentic Technical Debt — Debt that compounds (not just accumulates) because each agentic-coding session re-derives architectural decisions without persistent CLAUDE.md; surfaces late as a forced rewrite
  • AI Investment Story, Not Efficiency Story — Emergence Capital's Beyond Benchmarks 2026 counterintuitive finding: across every revenue segment non-AI companies out-earn AI companies on revenue-per-employee (~39% at the top decile — the only percentile the report splits AI vs non-AI), so AI is still an investment/staffing bet rather than a realized efficiency gain — reconciled with the lean-unicorn narrative via investment-phase staffing and the complements-lag (gains trail adoption), with AI-native RPE growing faster and starting to close the gap; AWS's 2026 founder survey is the vault's third RPE reading and points the other way (55% of AI-natives clear $400K/head, 156% growth); ICONIQ's Q2-2026 exec survey is the fourth, adding a forward RPE projection ($272K→$496K at high-growth firms by 2027) — flagged as a survey-self-report vs cap-table instrument split, projections marked prediction-grade, not averaged
  • The AI-Native Safe-Choice Inversion — Buying the legacy incumbent used to be "safe"; post-AI, being the incumbent = not AI-native; boards give buyers air cover; a counter-positioning play
  • AI-Native Startup Lifecycle — Anthropic's May 2026 reframing of Idea/MVP/Launch/Scale assuming AI infrastructure: each stage's headcount/capital/skill gates dissolve; lean unicorn as deliberate target
  • AI Product Economics Maturation — ICONIQ Q2 2026 exec survey (~305 AI-building software companies): AI crosses from experiment to P&L line — AI products 32%→42% of revenue, gross margin 45%→53%→59%, consumption/outcome pricing rising (blending 1.7 models), provider mix reshuffled (Anthropic 51%→81%, now #1), internal AI spend 11%→16% of revenue with hard-to-predict cost overruns, and FDEs monetized as a permanent revenue-driving GTM motion — forward-year figures are self-reported projections, prediction-grade
  • Compounding Data Moat — Anthropic's prescription for Scale-stage defensibility: time-locked behavioral fingerprint + domain-encoded edge cases + workflow lock-in via APIs/integrations beyond what migration agents can port
  • Founder as Agent Orchestrator — Founder role shift: less individual contributor, more orchestrator of specialized AI assistants; non-technical founders unblocked; lean 10-person unicorn structurally enabled
  • Founder-Led Sales Discipline — Stay founder-led until PMF; don't offload sales to an AE or an agent; explicit tension with Founder as Agent Orchestrator
  • Narrow Wedge into a Legacy Market — Disrupt without being feature-complete: be the best for a narrow customer profile (tech cos outgrowing QuickBooks); Google-Sheets MVP; the wedge-flip lesson
  • The 1% Rule for Wedge Selection — Jeff Dean's test for what a startup should build: run the general models on your candidate problem and pick one where they succeed 0-1% of the time, not 20% — partial success means the capability is already arriving and the next release will take the market. The exact inverse of build-for-the-next-model, and the two only reconcile on who owns the surface the release lifts
  • Printing Press Software Democratization — Boris Cherny's analogy: 1400s literacy expansion → AI software-writing expansion; domain knowledge displaces coding skill; 10× more disruption-grade startups predicted
  • Problem-Solution Fit Discipline — Idea-stage thesis: three defenses against premature building (time, resources, belief friction) all eroded; AI as devil's advocate is the antidote to confirmation-bias-with-research-engine
  • Product Velocity as Moat — Shipping speed as differentiator + trust signal ("you'll scale with us"); a treadmill that must convert into durable lock-in
  • Seven Powers Applied to AI — Helmer/Acquired framework re-evaluated for AI: switching costs and process power erode; network effects, scale, cornered resources persist; counter-positioning amplifies
  • The Solo-Founder Shift — Carta cap-table data on tens of thousands of U.S. startups: the solo-founded share of new companies rose 23.7% (2019) to 36.3% (H1 2025), with dilution, round sizes and employee equity grants near-identical to co-founded teams and median founder ownership at exit 75% higher. The firm-scale twin of the solo-authorship rebound — same left-tail instrument, same period, same attributed mechanism, same composition weakness — but the data measure ownership and timing, never revenue, so they characterize the lean tail's structure without touching its efficiency; and the report's own numbers show solo founders hiring their first employee earlier than co-founded teams, making the organization of one a transitional state rather than a destination
  • Zero-Friction Scope Creep — MVP failure mode when agentic coding removes the cost-based forcing function against scope creep; antidote is written scope + evidence-based amendment criteria

Open questions 47 open

    • SourceHow long does a CLAUDE.md remain accurate as a codebase evolves? The playbook gestures at session-by-session updates; no data on rot rate. (Partially answered — not answered — by Khatri 2026: rot rate is still unmeasured, but the question's stakes move. If a Good/Excellent-rated file buys no correctness over having none, then a stale file costs correspondingly little correctness too, and the rot that matters is in the environment-fact content (test cost, deployment invariants) that carried the one measured effect. The measurement still owed is a longitudinal one: does a file's accuracy decay track anything observable in agent behaviour?)
    • NoteThe remedy assumes the founder is able to articulate architecture in plain language. Non-technical founders (the playbook's headline beneficiary group) may have neither the vocabulary nor the intuition to do this well — a recursion failure the playbook doesn't address. (Deflated, not resolved, by Khatri 2026 — if the file doesn't move correctness, the founder's inability to write a good one costs less than this bullet assumes; see Founder as Agent Orchestrator.)
    • NoteAnthropic's harness-shrinkage thesis suggests CLAUDE.md may eventually be inferred by the model itself. Until then, the discipline is load-bearing.
    • SourceIs the classification driving the result? "AI company" is Emergence's label. If AI companies are disproportionately younger (more likely pre-revenue-inflection) than the non-AI cohort at the same revenue band, some of the RPE gap is an age/stage artifact, not an AI effect. The report doesn't publish a stage-matched comparison. (Partially answered on a different outcome, 2026-08-11: OECD AI Papers No. 62 runs exactly this test on market share instead of RPE, with adoption measured by a compulsory national statistical survey rather than a label. The raw gap is enormous — AI users hold 7.5× (France) and 3.2× (Portugal) the average market share of non-users — and it dies under controls: the AI user coefficient on market-share decile goes 0.0657\\\ → 0.0530\\ → 0.0266 ns in France and 0.0681\\\ → 0.0283 ns → 0.0199 ns in Portugal, killed mainly by lagged productivity (0.18–0.19, 7–9× the AI coefficient it displaces). So on the closest available analogue, the answer is yes, selection is doing the work — but note the direction: there the selection inflates the AI cohort's apparent advantage, whereas here the suspicion is that stage/age composition deflates it. The RPE half is untouched; a stage-matched financial comparison is still what would settle it.)
    • WaitWhen does the crossover happen? AI-native RPE is growing faster and already leads on growth; at $100M+ top decile it grew +58% vs −6%. Does the level gap close within a year or two, and does it invert (AI companies more efficient per head) — the point at which "efficiency story" becomes true?
    • SourceTail vs. mean gap. No data here on the deliberately-lean solo-founder tail's RPE specifically — the lean-unicorn claim lives in that tail, which the population medians can't isolate. (Partly informed: Emergent, a celebrated lean-tail exhibit, checks in at ~$600K/head at $120M ARR — below this cohort's $100M+ top-decile AI figure ($960K), suggesting the tail's scaled RPE is less exceptional than the low-headcount snapshots imply. One vendor-claim datapoint, not a cohort.)
    • SourceWhich instrument is right for the frontier AI-native subset? Two empirical-tagged sources disagree in direction — cap-table financials say AI companies earn ~39% less per head, a founder survey says AI-natives clear $400K/head at 55% and grow 156%. The disagreement is confounded by instrument (measured vs self-reported) and reference class (matched-band AI-vs-non-AI vs AI-native-vs-all-startups). Only a matched-segment, financial-data RPE study of the deliberately-lean AI-native frontier specifically — not the broad "AI company" label — would settle whether the survey optimism or the cap-table pessimism describes that tail. (ICONIQ's fourth reading adds a forward trajectory — RPE projected +84% by 2027 — but it too is self-report, and projected, so it deepens the survey-side optimism rather than adjudicating it.)
    • WaitMargin question (report's own): are the fastest-growers' 6–16pp-lower gross margins a temporary AI-infra-cost absorption or a permanent repricing of software's economic quality?
    • WaitICONIQ's respondents project gross margins expanding to ~59% by 2027, while Emergence's cap-table data measures the fastest-growers running 6–16pp below peers today. Does the projected margin expansion materialize, or is it survey optimism that regresses toward the measured growth-margin tradeoff as these companies scale?
    • SourceFDEs are monetized fragmentedly (bundled / separate PS fees / hybrid) and comped on retention. Does a dominant FDE monetization model emerge, and does the "Revenue Driver" self-framing (38%) survive a margin analysis — i.e. are FDEs actually accretive, or a services drag reclassified as growth? (Still open, and pointedly so: the corpus's most prominent July-2026 coverage of the role discusses supply, scarcity, and vendor structure at length and says nothing about pricing, bundling, or margin. The answer will come from a filing or an operator's P&L, not from role coverage.)
    • SourceEnterprises are building internal FDE teams specifically to avoid exposing proprietary business processes to their model vendor (C&T via TechCrunch, vendor-claim, motive documented via one recruiter and one vendor CEO; the behaviour mostly not yet observed — Ode reports no client asking it to build such a team). Does that in-housing actually happen at scale, and if it does, does it cap the FDE-as-revenue-driver motion at exactly the accounts worth the most — i.e. is the labs' delivery-layer integration self-limiting? Falsifiable from job-postings data (internal FDE-titled roles at non-vendor enterprises) against vendor-services revenue disclosures.
    • WaitInternal AI spend jumped from 1–3% to a projected 16% of revenue with respondents calling true cost hard to predict. Is 16% a transient enablement bulge that falls as tooling matures, or a durable new cost line for software companies?
    • The playbook gives no quantitative evidence for the headcount/capital compression claims (no median time-to-PMF, no headcount-at-PMF numbers, no failure-rate data). The "lean 10-person unicorn" is asserted as deliberate target without case-study evidence in the doc itself. (Partially answered: Emergence Capital, June 2026 now supplies headcount-at-round medians — Seed 6.2 (−39% from the 2021 peak of 10.3), Series A 16.8, Series B 48.2 — plus days-to-first-hire 214→284 and capital concentration (44% of venture to AI). Still missing: median time-to-PMF, headcount-at-PMF specifically, and failure-rate data; and the Carta cohort is market-wide, not the lean-AI-native subset. See AI Investment Story, Not Efficiency Story for the efficiency counter-signal in the same data.)
    • NoteFounder stories in the resources section (Carta Healthcare, Anything, Cogent, Airtree, Duvo, Zingage, Kindora, Wordsmith) are short callouts — none have published outcomes or comparable-baseline data.
    • SourceThe 42% "built-something-nobody-wanted" CB Insights figure is from a pre-AI era; the playbook predicts the rate will climb but doesn't cite a 2026 measurement.
    • ResolvedTension with HBR's accountability findings (above) is unresolved. The playbook's orchestration framing reads as the exact framing HBR's experimental conditions tested against. Answered: Orchestration vs Employee Framing: Reconciling the Founder's Playbook with HBR's Accountability Evidence resolved this operationally in May 2026 — orchestration as workflow design (agents, handoffs, review gates, decision rights) survives HBR's critique; orchestration as a coworker mental model (naming, delegation-without-scope) is what produces the −9pp/+44%/−18% effects, and the playbook's lifecycle needs only the former. The Orchestrator's Real Workload: Decision Burden, Framing Discipline, and Whether Taste Scales adds the July 2026 reinforcement: decision-rights gating now has measured backing (control-channel authorization 100% on safety-critical actions vs 51/54% for advisory channels), while the framing effect compounds with brain-fry in the same direction (felt control, decayed review). The residual — why Anthropic's founder marketing ignores its own framing-discipline work — is a question about Anthropic, tracked on Founder as Agent Orchestrator as #oq/source.
    • SourceIs the "two-year replication window" claim defensible empirically, or aspirational? The playbook does not cite measurement. (Partially answered, at the far end only: DroneDeploy is a ~decade accumulation whose value arrived discontinuously when vision models did, priced at $845M — one case, told by its own investor, with no counterfactual for how fast a late entrant could have caught up once the demand was legible. What it suggests is that "two years" is the wrong unit: the binding variable is whether accumulation started before the monetizing capability was foreseeable, not elapsed calendar time. A real answer still needs a matched pair — two vertical products, one with a pre-capability archive and one without, competing after the capability lands.)
    • WaitHow does this moat hold up when foundation models themselves continue improving rapidly? If a generalist model in 2027 has internalized enough vertical context to handle 340B drug claims natively, does the vertical-edge-case moat erode?
    • SourceThe data-flywheel argument has been made for SaaS for 15 years. What's actually different in the AI-native version? Probably: the data improves the model in addition to the product, but the playbook doesn't make this distinction precisely.
    • SourceThe "customers build APIs on top of you" lock-in is structurally similar to platform plays (Salesforce AppExchange, Shopify apps). Is the moat type really new, or just newly accessible to lean startups?
    • SourceThe playbook claims non-technical founders can now build production software, but it does not address the architectural-judgment recursion problem (Agentic Technical Debt): non-technical founders may not have the vocabulary to write effective CLAUDE.md. How does that scale? (Partially answered — by deflation — Khatri 2026, arXiv 2607.27250, empirical: in a 288-run two-agent ablation, having a Good/Excellent-rated AGENTS.md produces no measurable correctness gain over having none (bounded <10–15pp), and the real file never converts a near-miss to a pass in a 36-cell probe. If the file the non-technical founder cannot write buys ~0 correctness, the recursion is a smaller tax than the playbook's own framing implies. Two reasons this doesn't close the question: the ablation measures single-session task correctness, not the cross-session architectural coherence this recursion is actually about, and the deficit it identifies as gating — implementation skill: feature design, pattern selection, exact wiring — is precisely the judgment a non-technical founder also lacks and cannot delegate to a file. The recursion may not run through the context file at all; it runs through review.)
    • The "lean 10-person unicorn" is asserted; no quantitative data in the playbook on actual headcount-at-PMF or headcount-at-Series-A medians for AI-native startups vs. the prior cohort. (Partially answered: Emergence Capital, June 2026 gives Carta round-medians with a 2020–2025 series — Series A 16.8 (down from the 2021 peak of 25.9), Seed 6.2 (from 10.3), Series B 48.2 — the prior-cohort comparison the playbook lacked, plus the AI-vs-all-tech headcount-allocation split (engineering-heavy, lean support). Still open: these are round-medians, not headcount-at-PMF; and the Carta cohort is market-wide, not split AI vs non-AI at Seed/A.)
    • NowHow does the orchestration role change the founder's decision burden? Fewer hands-on tasks but more parallel agent oversight; net cognitive load is unclear and may be higher (see AI Brain Fry). Partially answered: The Orchestrator's Real Workload: Decision Burden, Framing Discipline, and Whether Taste Scales — higher and reshaped: execution load is exchanged for oversight load at an unfavorable rate, because the incoming work is the error-prone kind (+39% major errors under fatigue), the invisible kind (rubber-stamped planning decisions are transcript-indistinguishable from judgment), and non-monotonic in value (HAS-Bench's returns-curve peak — over-intervention breaks tasks). The load is bounded only by deliberate structure: bounded parallelism, sampled review, high-stakes gates. Still unmeasured: founder-side oversight load directly — concurrency telemetry sums agent-hours, not human attention.
    • SourceAnthropic publishes both the playbook's anthropomorphic framing and HBR-aware accountability work (auto-mode, alignment) simultaneously without engaging the framing literature directly. The synthesis in Orchestration vs Employee Framing: Reconciling the Founder's Playbook with HBR's Accountability Evidence reconciles the tension at the operational level — orchestration as workflow design preserves accountability; orchestration as mental model of agents-as-coworkers does not — but the open question of why the playbook's marketing language doesn't reflect Anthropic's own framing-discipline work remains.
    • SourceWhere exactly does "until PMF" end, and what's the first thing a founder should hand off (AE? agent? both)? Glasgow still does it post-Series-B, suggesting the boundary is fuzzy.
    • SourceDoes Glasgow's anti-offload stance generalize, or is it specific to high-trust, mission-critical enterprise sales (ERP) where "they're buying you" — would a PLG/SMB motion delegate to agents far earlier?
    • WaitA wedge works going in; does it constrain going out? Campfire now serves public companies — at what point does "narrow-but-best" require becoming the broad incumbent it displaced, re-incurring NetSuite's complexity?
    • SourceThe wedge-flip shows the first wedge can be wrong. What's the fastest signal that a wedge converts to the core vs. merely sells — Campfire took ~3 months; can it be read sooner?
    • NowIs domain-expert-as-builder actually happening at scale in 2026? Anecdotes (shop owners, microcontroller hobbyists) yes; primary-job software building by non-engineers, less clear. (Partially answered: Anthropic's 400K-session study finds non-software occupations reach verified success in code-producing sessions within ~7pp of software engineers — the strongest evidence yet that the claim holds, at least within Claude Code's user base. Market-scale corroboration: Emergent reports 200K+ non-technical paying customers — trucking companies, factories, and construction businesses building their own ERPs, property managers building CRM tools (TechCrunch, July 2026, vendor-claim) — Boris's "the accountant writes the accounting software" observed as a paying market, not just inside one vendor's telemetry.) Further advanced: Is Breadth Cheap Now? Specialist Ramp Speed and Domain-Expert-as-Builder at Scale sorts all the evidence into three tiers — capability parity (measured: within-7pp), market existence (demonstrated, vendor-claimed: Emergent's 200K+ non-technical builders; AI responsibilities in 28–40% of business job descriptions), and primary-job building as population-level practice (still unshown: every measured population is selection-biased toward adopters, complements gate realized value, and the ATLAS composition shows experts pointing AI at their own inexpert tasks rather than non-experts becoming builders). The gating variable is now complements + retained understanding, not capability.
    • WaitWhat's the equivalent of compulsory schooling for universal coding literacy? Or does that not happen and we get a long tail of self-taught builders?
    • SourceBoris's "accountant writes accounting software" — does that result in 10K narrow tools that don't interoperate? What's the integration story?
    • SourceDoes asking an AI to argue against an idea actually produce disconfirming evidence at the same rigor as confirming evidence, or does the model still bias toward the framing the founder presents? Worth measuring.
    • SourceHas anyone measured 2026 startup failure rates with AI-built products? The "42% will climb" claim is asserted without measurement.
    • ResolvedThe playbook recommends "ask Claude to make the most compelling argument for why a competitor would succeed while you do not." How does this interact with Anthropic's published character training (sycophancy resistance, devil's-advocate willingness)? Answered: Playbook Boundary Conditions: the Devil's-Advocate Substrate and the Prototype's Edge — complementary, not conflicting: the prompted moves are framing-compliance tasks that work on any instruction-follower (none requires disagreeing with the founder), while character training supplies the unprompted pushback the prompts can't manufacture — portable technique, vendor-specific safety net, and the residual gap (framing bias within the assigned adversarial task) is the sibling #oq/source above.
    • WaitVelocity-as-moat is a treadmill: it evaporates the moment a competitor matches pace. What converts Campfire's velocity lead into a structural moat before the AI-native cohort's pace converges?
    • Source"Never had anyone outgrow Campfire" — is that survivorship (they haven't hit true enterprise scale yet) or a real claim that velocity closes the breadth gap faster than customers grow into it?
    • SourceIs "switching cost" really collapsing in practice, or just in narrative? Anthropic's own retention numbers, Salesforce churn, etc. would test this.
    • WaitWhat does Boris's "cornered resource" look like for foundation-model labs that are themselves trying to commoditize? Internal contradiction or transient phase?
    • SourceCounter-positioning — explicitly the "incumbent can't follow" power — should amplify under AI. Is anyone running this play deliberately?
    • SourceDoes the rule hold empirically? Nothing here tests whether markets where models scored ~20% in 2025 were absorbed faster than markets where they scored ~0%. The data to check it (benchmark-era capability snapshots against startup outcomes) exists in principle.
    • SourceWhat is the 2026 cost of a defensible niche model? Dean asserts "maybe it doesn't take that much compute"; a founder needs the number, and the corpus doesn't have it.
    • WaitThe inversion is a one-time repricing of "safe." Once several AI-native ERPs exist, does "safe" re-stabilize around the largest AI-native vendor — and does Campfire's "we're now the largest of the new cohort" claim reflect a land-grab for that position?
    • WaitHow long until incumbents bolt on credible AI and neutralize the counter-positioning — and does the custom-foundation-model claim actually defend against that?
    • WaitDoes the H1 2025 jump to 36.3% survive a full-year datapoint, or is it a half-year artifact? Every prior step is 0.6–2.7pp and this one is 5.8pp. Trigger: Carta's 2025 full-year or 2026 update to this series.
    • SourceIs the employee-equity null a real population fact or a median artifact? Carta reports near-identical medians; the publisher claims 2–5× among founders in his own program. A distributional cut — variance or upper decile of first-five grants, split by founding-team size — would settle it, and neither party publishes one.
    • SourceDoes the solo-founded tail differ from co-founded companies on revenue per head? This dataset cannot say — it holds cap tables, not revenue — and it is the missing half of AI Investment Story, Not Efficiency Story's tail question.
    • SourceThe playbook recommends written scope but offers no template or worked example. How specific does "what we deliberately don't do" need to be to actually block requests?
    • SourceIs there a measurable threshold where scope creep crosses into outright pivot territory? The playbook gestures at "losing direction" without a metric.
    • SourceHow does this interact with Cat Wu's 1-day shipping cadence? Anthropic's internal practice ships fast but with strong product judgment; how does that judgment translate for a first-time founder?