H
Howardism
Howardism · Vol. 03Plate II · No. 02

Product & Org, in order.

Notes17DomainProduct & OrgOpen Qs45Newest11 Aug 2026Oldest6 May 2026

Product cadence, org design, and the AI-native team.

Map of Content for the product-org domain — 17 concepts. Curated entry point; see Home for all domains.

  • AI-Native Organization — Garry Tan's org-design mapping: skill files = employees, resolver tables = org charts, filing rules = process, trigger evals = performance reviews — a company whose operations are encoded as markdown that agents execute, with engineers hired to maintain the skills; claimed record revenue-per-head (Emergent ~$15M ARR at 15 people, Retell $60M at ~40)
  • AI Native Product Cadence (hub) — Cat Wu's 6mo→1mo→1day cadence at Anthropic: research-preview branding, mission-as-tiebreaker, evergreen launch room, lighter PRDs, weekly metrics readouts
  • Community Smells Under AI Adoption — PLS-SEM on 152 software professionals: AI adoption is associated with fewer socio-technical anti-patterns, by two different mechanisms — indirectly in specialization work (AI → more peer consultation → less knowledge fragmentation) and directly in coordination work (AI → better communication quality, with interaction frequency unchanged) — while a vocal minority of the same respondents report in free text that AI replaced their teammates
  • Compounding Loop Optimization — Dan Carey's discipline of instrumenting and automating every recurring step of the build loop — because when internal tooling is an-afternoon-cheap, each optimization pays back ×(50–100 iterations per project)
  • Dogfooding as Product Discipline — Product sense is built by relentless first-hand use ("ant food"); Mr. Peanut catch; cross-source (Cat Wu vibe-checks, Glasgow founder-led sales)
  • Engineer PM Convergence — Generalists across disciplines; product taste as bottleneck skill; Anthropic Claude Code team as case study; "just do things" cultural substrate
  • Evals as Product Spec — Cat Wu's framing of evals as the emerging core PM skill: ten great evals beats a hundred mediocre; encode what done looks like for ambiguous AI features; companion to introspection (hypothesis) and vibe-check (direction)
  • Excellence as an Operating System — Elizabeth Stone's account of Netflix culture: talent density, agency, and accountability are not values but a mechanism for excellence — resist process even when things go wrong (blameless retros + individual responsibility instead), run the keeper test in both directions; Lenny's observation that top AI labs now converge on the early Netflix culture deck
  • Implementation Abundance Inverts Product Work — Andrew Ambrosino's inversion thesis: when talking to a frontier model can stand up any feature from scratch, implementation stops being the expensive step you derisk up front — so the process runs backwards and the costly work becomes curating the 90 uncoordinated builds people already produced; taste is the new bottleneck
  • Managers as ICs — Every Claude Code manager starts as an IC; flat org; agentic coding collapsed the onboarding cost that pushed managers out of the codebase
  • Model Introspection Feedback — Cat Wu's underrated technique: ask the model why it failed; treat answer as harness-debugging signal not model criticism; caveats around model self-report fidelity
  • Polish No Longer Signals Readiness — Andrew Ambrosino's observation that the medium used to encode process-stage — a production-looking artifact meant late-stage, derisked, design-and-business-approved — but cheap implementation divorces polish from maturity: a 90-person exploration can look ready-to-ship while being early design work, and over-anchoring on it ('can we release this now?') is the trap
  • Prototype Fidelity After Cheap Polish — Hundhausen's argument that GenAI decoupled polish from effort, invalidating the empirical basis of the low-fidelity-first playbook: the classic finding was that polish suppresses feedback because it signals sunk effort, and that signal is now false while the psychological barrier likely persists — plus the revival of Boehm's evolutionary prototyping and three unanswered research questions
  • Prototype Over PRD — Dan Carey's prototype-replaces-PRD method: record a why-not-what conversation, transcribe it, hand the transcript to Claude, ask for a few prototype variations; the prototype is the spec, not a downstream artifact
  • Role Averaging, Not Role Elimination — Andrew Ambrosino's nuanced OpenAI-side take on role collapse: your role is 'the average of what you spend your time on' and tool-gatekeeping is eroding — but eliminating roles dangerously eliminates specialties with knowable best practices ('getting rid of the product role is a terrible idea'), and 'zone defense' coverage plus managers remain necessary because not everyone can work on everything in both breadth and depth
  • Standardize the Infrastructure, Not the Tools — Shopify's inversion of the one-tool-per-job norm for AI: route every coding agent through a central LLM proxy so leadership gets cost control, per-team usage analytics, and model portability, while engineers keep free tool choice — buying optionality under uncertainty about which model or workflow wins, with MCP servers extending the same governs-access-not-engineers principle to internal systems
  • Systems Thinking Over Specialization — Elizabeth Stone's Netflix hiring thesis: in an agent-heavy org the scarce profile is the systems thinker who abstracts across business domains into paved paths, design systems, and source-of-truth data — narrow specialists shrink to a few irreplaceable niches; AI fluency becomes a cross-level career-ladder overlay, and the trainable move is 'step out one click'

Open questions 45 open

    • SourceDoes the cadence scale beyond ~100 people? Anthropic itself is bigger (~30-40 PMs alone), but the Claude Code team that visibly drives cadence is small.
    • SourceWhat's the equivalent of research-preview branding for B2B enterprise launches where customers expect stability? Cat doesn't address.
    • SourceHow much of the cadence is structural (process choices) vs cultural (talent density)? Probably both, ratio unclear.
    • SourceTan's revenue-per-head figures (Emergent ~$15M ARR at 15 people, Retell $60M at ~40) are stated from stage without sourcing. Do third-party data (Carta/Standard Metrics cohorts, press-verified ARR) corroborate record revenue-per-head at AI-native YC companies, or do these examples regress toward the AI Investment Story, Not Efficiency Story mean on inspection? (Partially answered: TechCrunch, 2026-07-15 corroborates Emergent as a real, fast-growing $1.5B unicorn — $120M company-reported ARR, 200K+ paying customers — so the direction holds. But the record per-head claim compresses: at scale it is ~$600K/head (200 employees), below the $100M+ top-decile AI-company RPE of $960K AI Investment Story, Not Efficiency Story and about half the ~$1M/head of Tan's own 15-people/$15M snapshot — the per-head extreme is a low-headcount-phase artifact that regresses as the company staffs up. Caveats keeping this open: TechCrunch's figures are themselves company-reported vendor-claim, not Carta-audited, and the Retell half ($60M at ~40 ≈ $1.5M/head) remains unverified. AWS's June-2026 founder survey adds a population reading (55% of AI-natives self-report $400K+/head) but it is self-report, not the cap-table/press verification this question asks for. See Emergent.)
    • SourceThe org mapping predicts a testable staffing signature: AI-native companies should hire engineers to maintain skills rather than function-specific staff. Does job-posting data show a "skill maintainer / agent ops" role emerging as a distinct hiring category? (Partially answered: ICONIQ, Q2 2026 confirms the composition shift the signature predicts — 45% of ~305 AI-builders plan a "different mix of roles (fewer operational, more AI-fluent talent)," function-level headcount reallocates toward R&D/Product/Sales and away from Customer Support/G&A, and G&A operators are "removing finance-ops and order-management roles… redirecting budget to strategic and AI-specific functions." It also names the concrete new engineering hiring categories: forward-deployed engineers (~50% scaling as a permanent motion) and AI safety / trust & reliability engineers. What it does not supply is the specific "skill maintainer / agent ops" title from job-posting data — ICONIQ measures function-level headcount intent and two named roles, not an occupational taxonomy. The direct test (a "skill maintainer / agent ops" posting category) still needs job-posting/occupational-emergence data, e.g. the un-ingested arXiv 2606.22769 "Agent Systems Engineer" signal from the 2026-07-21 research pass. See the restructuring section above and AI Product Economics Maturation for the FDE detail.)
    • SourceIs the encoded-role form of the employee metaphor actually accountability-preserving, as the synthesis above suggests, or do Kropp-style framing effects attach to skill-files-as-employees too once teams talk about them that way? No study has tested framing effects on artifact-level anthropomorphism.
    • SourceThe design cannot separate "AI adoption improves team social health" from "healthier teams adopt AI better", and the authors say so. The discriminating study is the one they name: longitudinal or quasi-experimental tracking of AI adoption and team dynamics over time, controlling for communication culture, org maturity, leadership practice and seniority. Until then every coefficient here is an association.
    • SourceThe Information Sharing path — AI adoption directly worsening documentation and information governance — is the only harm signal in five models and lands at p = .069, below the β ≈ .20 the sample can reliably detect. Does it survive at n ≈ 400, and does it strengthen in teams without a documentation discipline? This is the falsifiable half of the paper's "governance-dependent" conclusion.
    • SourceThe study measures peer-interaction frequency, not what the interaction contains or what expertise is retained. If AI raises the count of specialization-oriented exchanges while lowering their depth (the comprehension register), frequency would rise exactly as measured while the underlying transactive memory thins. Nothing here distinguishes the two.
    • SourceThe loop assumes the team is (close to) the user. How much of the compounding advantage survives when the user is unlike the builder and "talk to users" can't be same-room?
    • SourceWhere is the line between worthwhile internal tooling and yak-shaving? Carey's "afternoon" bar is the heuristic, but Cat Wu warns that over-customizing setups "becomes distraction."
    • SourceDoes Claude-as-first-pass-on-all-feedback ever filter out the rare signal that doesn't cluster? Automating triage optimizes the common case; the tail is where surprising bets come from.
    • SourceDogfooding works when the team is the user (Claude Code) or near it (Cat Wu, Boris). How do you build product sense for users very unlike you — does "talk to customers" fully substitute, as Glasgow/Fung's small-business work suggests?
    • ResolvedCan dogfooding scale, or does it implicitly cap how large an AI-native product org can stay taste-driven before it reverts to dashboards? Answered: The Orchestrator's Real Workload: Decision Burden, Framing Discipline, and Whether Taste Scales — false binary: dogfooding itself never scales (first-hand use is per-person and breaks when the team stops being the user), but the taste it produces scales through three named mechanisms — encoding into runnable artifacts (Evals as Product Spec), concentrating the rare-trusted-evaluator role plus vibe-check rituals rather than diffusing taste with headcount (Claude Character as Product), and AI-extended contact surface (Carey's Claude-first-pass on every user conversation). The cap variable is not org size but team-user distance plus encoding discipline: an org reverts to dashboards when it stops converting felt use into evals and rituals, at any size.
    • WaitDoes this scale beyond ~50-person Claude Code-style teams? Boris hedges: "I think this is going to be a question for years."
    • WaitWhat happens to formal PM career ladders in companies where engineers do PM work? Open at Anthropic per Cat.
    • SourceCross-disciplinary generalist is a hiring bar — where does the supply come from? Career changers, or new-grad bias toward AI-native education?
    • How do you write an eval for taste-driven features like character? Amanda's role is canonical for being eval-resistant; Cat names her as someone who is good at evals here, but doesn't describe the technique. Partially answered: How Do You Write Evals for Taste? Character as the Limit Case — the technique is a pipeline (conviction → dogfood-sourced failure modes → MSM-style variant A/B measurement → ~10 interpretable evals); proven on the safety/values core but still tacit on the warm/witty aesthetic surface.
    • SourceThe 10-vs-100 number is given without justification. Is there a Goldilocks zone, or does it depend on feature surface area? Client-Side Agent Optimization's framing of combos suggests evals also have a combinatorial explosion problem.
    • SourceHow do evals interact with Harness Shrinkage as Models Improve? When a harness asset shrinks because the model now handles it natively, the evals built around the old harness may become artifacts rather than guardrails. Does Anthropic retire evals or repurpose them? Partially answered: Boris Cherny (YC interview, 2026-07-27, practitioner-opinion) — retire: evals live "one, two, three model generations," then saturate and get thrown away and rebuilt from observed struggle; what persists is the authoring practice, not the artifact. Still open: whether any eval class (safety, character) is exempt from the saturation cycle.
    • Is there a single non-Anthropic example of a PM-as-eval-writer to cite, or is this currently a Cat-Wu-singular framing? The Matt Pocock workshop reaches the same place from a different vocabulary, but no third source has been ingested yet. Partially answered (with a twist): Google's Agent Quality Flywheel is a third-party arrival at eval-as-the-quality-surface — but its answer is to have the coding agent author the eval, compressing the human role to stating the worry and approving the plan.
    • SourceIs the AI-lab convergence on early-Netflix operating norms (agency, density, top-of-market pay) causal inheritance (the culture deck as a founding document for lab founders) or convergent evolution under the same constraint (scarce elite talent)? A history of lab founding cultures could settle it.
    • WaitDoes talent-density-plus-paved-paths actually substitute for process at agent-scale throughput, or does Netflix eventually show the Acceleration Whiplash quality signature (incident rates, review latency) like Faros's high-maturity cohort? Trigger: future Netflix engineering telemetry or tech-blog disclosures.
    • SourceCuration of 90 uncoordinated builds is itself expensive and doesn't obviously scale — is there a point where the cost of curating parallel exploration exceeds the cost it replaced? ("zone defense" is Ambrosino's partial answer.)
    • WaitIf taste is the bottleneck and taste is "just another capability" AI eventually masters, does the inversion invert again — does curation migrate into the model?
    • SourceThe 90-uncoordinated-builds picture assumes abundant tokens and an agentic culture; how much of the inversion survives outside a frontier lab that gives everyone "unlimited tokens"?
    • WaitFung's own open question: "Do you still need separate iOS and Android orgs?" — if engineers flex across platforms via Claude, the traditional platform-split org may dissolve too. How far does flattening go?
    • WaitDoes manager-as-IC scale past a certain org size, or only work while Claude Code is small and the codebase is Claude-legible?
    • How reliable are 4.7-class introspective reports? Anthropic's interpretability research suggests partial fidelity but not full. Empirically, Cat reports it's good enough to drive harness fixes — but unclear at what model scale this technique becomes load-bearing. Partially answered: Self-Report as a Safety Signal — reliability is context-dependent. In a benign debugging setting the report is good enough to drive harness fixes; in an adversarial safety setting, open-weight models (3B–70B) fail to recognize their own compromised outputs 27.3% of the time, and the recognition that exists is the refusal circuit firing late rather than genuine own-output introspection. So the channel Cat relies on is a weak safety signal even where it is a useful debugging one. Also partially answered: Introspective Coupling puts a floor under the scale question — an untrained 8B model manages only 14–18% exact match at predicting its own counterfactual behavior, so at that scale the channel carries essentially nothing without explanation training.
    • Does adversarial introspection ("why did you fail?") yield different signal than neutral ("walk me through your reasoning")? Worth probing. Partially answered: Self-Report as a Safety Signal finds self-attribution is heavily framing-dependent — an "intention" probe and a "tampering" probe elicit qualitatively different answers on the same models (some families deny tampering ~100% of the time regardless), so the phrasing of the introspective question materially changes the signal.
    • SourceCould a meta-agent run introspection automatically against logged failures? Sounds tractable but no public implementation.
    • SourceIf the medium no longer signals stage, what does — is explicit human labeling ("this is exploration") the only mechanism, or can tooling re-attach the signal (e.g. a visible "exploration / preview / prod" marker on every build)?
    • WaitDoes over-anchoring get worse as builds get more polished, or does everyone eventually recalibrate and learn to discount fidelity entirely?
    • SourceDoes the Schumann effect survive the loss of its mechanism — do users still soften feedback on polished artifacts once told the artifact took an hour? The article's Question 1, and the field's key unknown.
    • SourceWhich feedback is lost to polish: strategic (workflow, information architecture) or tactical (visual polish)? The distinction determines whether the low-fi-first playbook mattered for the reasons its advocates claimed.
    • SourceDoes GenAI-generated prototype code actually evolve into production, or rebuild? Hundhausen poses it as open; this corpus's debt evidence suggests rebuild, but no source measures prototype-to-production survival directly.
    • NowIf there is no PRD, where does the rationale ("why we chose variation B") live for future readers? Same rationale-capture gap flagged in Building Is Cheap, Arguing Is Expensive. Partially answered: Where Does the Why Live? — the why is well-homed at authoring time (it is the recorded why-not-what conversation) but orphaned at read time, since the artifact that carried it is deleted. Still open: whether a durable read-time home exists that doesn't reintroduce the PRD.
    • NoteThe prototype-as-spec must not become the prototype-as-validation trap Problem-Solution Fit Discipline warns about: a fast prototype proves the build was solvable, not that the problem is real.
    • ResolvedWhere does prototype-over-PRD break down? Carey's domain is a visual design tool where a prototype is the product surface; for backend/infra/data work the prototype may not capture the spec (cf. AI Native Product Cadence's "full PRD for heavy-infra features"). Answered: Playbook Boundary Conditions: the Devil's-Advocate Substrate and the Prototype's Edge — the boundary is observable-surface-vs-invariant, not backend-vs-frontend: every domain has its own tracer artifact (three PRs, vertical slice, ten evals, design_system.html), so what breaks at the backend is the clickable prototype, not artifact-over-document; the PRD survives where no artifact's surface covers the risk — cross-cutting invariants and cross-team coordination.
    • SourceWhere is the equilibrium between fluidity and specialty — how much role-averaging before a company loses the accumulated best practices Ambrosino warns about?
    • SourceZone defense assumes enough high-taste people to cover the whole company; does it degrade in orgs without OpenAI's talent density, collapsing back to top-down planning?
    • SourceDoes "your role is the average of what you spend time on" survive performance review and career ladders, or does it fragment them the way Cat Wu flags ("we're sacrificing product consistency")? Partially answered: Netflix's move (Systems Thinking Over Specialization) is to leave per-level criteria untouched and add a cross-level AI-fluency overlay, explicitly because the tech shifts too fast to encode per level — one large-org existence proof that ladders survive by absorbing fluidity as an overlay rather than rewriting levels.
    • SourceDoes a central LLM gateway actually change model-mix decisions, or only report on them? The claimed benefit is portability; no source in the corpus records an org exercising it.
    • SourceWhat does per-team AI usage analytics get used for once it exists — cost containment, capacity planning, or performance evaluation of engineers? The third would collide with everything Telemetry vs. Survey Measurement establishes about what instrumented output data can and cannot support.
    • WaitDoes agent-era recentralization (common paved paths, solve-once infrastructure) hold up against the local-team autonomy that Stone credits for Netflix's historical speed — i.e., will local teams accept the paved path when their problem doesn't fit it, or does shadow infrastructure reappear?
    • WaitStone keeps AI fluency as a deliberately vague overlay because the tech "evolves by the quarter." Does it ever crystallize into per-level ladder criteria (as conventional competencies did), or is permanent-overlay the stable state? Trigger: Netflix's next ladder revision.
    • NowStone claims specialists can now broaden "quickly" with AI tools. Does the wiki's evidence support cheap breadth acquisition — the concave novice→intermediate curve in Returns to Expertise in Agentic Coding suggests yes for working grasp, but is there evidence on speed of cross-domain ramp for experienced specialists? Partially answered: Is Breadth Cheap Now? Specialist Ramp Speed and Domain-Expert-as-Builder at Scale — split the claim: tool-in-hand performance breadth is measurably cheap (the concave curve makes the needed increment small; the expertise meta-skills — framing precision, verify-specification, who-corrects-whom — transfer across domains, per the management edge, so an experienced specialist enters above the novice floor; AI-assisted onboarding compresses ramp further). Retained-capability breadth is unproven and the only randomized evidence cuts against it: automation-mode gains vanish when the tool is removed while self-report hides the deficit — the augmentation/automation usage split decides which good you get. Ramp speed itself is measured nowhere; the mechanism argument stands in for it.