Howardism · Vol. 03Plate II · No. 02
Entities, in order.
Notes72DomainEntitiesOpen Qs72Newest11 Aug 2026Oldest17 Apr 2026
Profiles of the people, labs, products, and projects.
Map of Content for all 72 entity pages. See Home for concept domains.
People#
- Addy Osmani — Engineering leader at Google (Chrome) and prolific author/educator; in 2026 writes a widely-read blog series on AI-assisted engineering — agent harness engineering, the factory model, comprehension/intent debt, cognitive surrender, and the essay that named loop engineering
- Andrej Karpathy — Co-founder OpenAI, ex-Tesla AI, Eureka Labs; coined "vibe coding," Software 1/2/3.0, "ghosts not animals," "agentic engineering"; originated the LLM-wiki pattern this vault runs on — industrialized within ~3 months as 'agent wikis' (DeepWiki, AutoWiki, OpenWiki, GBrain)
- Andrew Ambrosino — Product & engineering lead for the Codex desktop app at OpenAI; a designer→engineer→PM→founder generalist whose June 2026 Lenny's Podcast interview is the wiki's OpenAI-side account of how cheap implementation inverts product work toward taste and curation
- Andrew Ng — Founder of DeepLearning.AI and AI Fund, founding lead of Google Brain, co-founder of Coursera; writes The Batch, where his June 2026 letter set out the three-loop taxonomy of AI-native building and reframed the residual human contribution as a "context advantage" rather than taste
- Boris Cherny — Creator of Claude Code at Anthropic; phone-driven workflow with hundreds of agents; primary advocate of
/loopprimitive; "coding is solved (for me)" thesis; ablation-driven harness design (delete the prompt, add back what the model repeatedly stumbles on) - Cat Wu — Head of Product for Claude Code and Cowork at Anthropic; primary articulator of AI-native product cadence and engineer-PM convergence
- Chloe Li — Lead author of MSM paper (arXiv 2605.02087); Anthropic Fellows Program; designed all specs and experiments
- Claire Vo — Host of the "How I AI" interview series (ChatPRD); interviewed Thariq Shihipar; runs a parallel component-visualization practice for non-technical stakeholders
- Dan Carey — Product Manager leading product within Anthropic Labs; led Claude Design; 'Designing with Claude' talk (May 2026); ~two decades of PRDs, now replaced by prototypes
- Elizabeth Stone — Netflix Chief Product and Technology Officer (economist by training: Analysis Group, Merrill Lynch trader, Nuna COO, Lyft VP of Science, Netflix CTO→CPTO); articulates systems-thinking-over-specialization hiring and 'excellence as an operating system'; two-time Lenny's Podcast guest
- Elon Musk — Founder of Tesla, SpaceX and xAI, and the corpus's clearest case of a reversed AI-risk position: 2015 'we'll be pet labradors' → 2023 pause-letter signatory → 2025 10–20% p(doom) → 2026 'even if there was a stop button we probably shouldn't press it'; his current answer is acceleration plus cross-lab pre-release review plus abundance
- Erik Brynjolfsson — Director of the Stanford Digital Economy Lab and the vault's most-cited economist — a disambiguation page, because five different works appear here under one surname: the Brynjolfsson-Rock-Syverson productivity paradox that anchors the complements thesis (well corroborated), the 2018 SML rubric that is one of seven exposure instruments which disagree (largely superseded), the 2025 'Canaries in the coal mine' 22-25-year-old -16% result the vault's firm-level panel contradicts, a 2025 workplace-writing homogenization finding a randomized essay experiment did not reproduce, and the July 2026 'We Must Act Now' open letter he organized
- Fiona Fung — Leads engineering + product for Claude Code and Cowork at Anthropic (ex-Meta/Microsoft); "what served you prior may no longer"; rewrote team norms for the AI-native org
- Garry Tan — President & CEO of Y Combinator; founder-investor turned evangelist for the AI-native organization — the ~400x output claim, "the leverage is not in the weights, it's in how you wire the work", the skillify-it discipline, and GBrain, his MIT-licensed open-source company brain (~220K pages)
- Jack Lindsey — Anthropic interpretability researcher; corresponding author of the global-workspace paper, co-originator of the Jacobian lens, and the one who ran the directed-modulation and post-training-diffing experiments that turned a readout method into a claim about model cognition
- Jarred Sumner — Creator of the Bun runtime, now an Anthropic employee after the December 2025 acquisition; author of 'Rewriting Bun in Rust', the wiki's most detailed first-party account of running a large engineering project on ~50 Claude Code dynamic workflows
- Jeff Dean — Google's Chief Scientist; built MapReduce, BigTable, TensorFlow and the TPU, and co-authored the 2014 distillation paper NeurIPS rejected that now makes Gemini's Flash models cheap. His recurring method is napkin math against a bottleneck — search-in-RAM (2001), speech-would-double-the-fleet (2013) — and his 2026 advice to founders is the 1% rule: build where models fail 0–1% of the time, not 20%
- John Glasgow — CEO/founder of Campfire; 10yr corporate finance; founder-led-sales advocate; long-horizon "last job I'll ever have"
- Marcus Hutter — Creator of AIXI and the Universal AI framework; DeepMind senior researcher and ANU professor; co-author of the Legg–Hutter intelligence measure and the 2026 textbook 'An Introduction to Universal Artificial Intelligence'; co-author of the 'From AGI to ASI' report
- Matt Pocock — Independent AI-coding educator; built Sandcastle library; smart-zone/grill-me/tracer-bullets pedagogical framing; "bad code bases make bad agents"
- Nate Parrott — Anthropic product designer who built Claude Design; sole designer on Claude Code for VS Code in fall 2025, then spent a month of side-project time closing the velocity gap that opened when Opus 4.5 accelerated his engineers but not him — the HTML-playground prototype that resulted became an Anthropic Labs product
- Noam Brown — OpenAI research scientist and a pioneer of inference-time (test-time) compute scaling; earlier built superhuman poker AIs and now uses building poker solvers as a personal model eval; author of the June 2026 essay Implications of Large-Scale Test-Time Compute
- Peter Steinberger — Founder of PSPDFKit turned prolific independent AI-coding experimenter (@steipete); originated the framing that loop engineering is built on — "you should be designing loops that prompt your agents"
- Shane Legg — Co-founder and Chief AGI Scientist of Google DeepMind; co-author with Hutter of the Legg–Hutter universal intelligence measure; senior author on the 2026 'From AGI to ASI' report
- Thariq Shihipar — Engineer on the Claude Code team at Anthropic; "HTML is the new markdown", "compute allocator", and "the map is not the territory" framings; three HTML-first workflows plus a phase-ordered catalog of techniques for eliciting your own unknowns
- Wes Gurnee — Anthropic interpretability researcher; co-first author and co-originator of the Jacobian lens, who conceived the connection between verbalizable representations and conscious access and led the method's development
Organizations#
- Anthropic — AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs round 2
- Anthropic Economic Index — Anthropic's recurring economic-research program measuring how Claude usage maps to and diffuses through the economy — privacy-preserving usage telemetry (Clio) now paired with a linked survey; reports include the June 2026 Cadences report, the returns-to-expertise study, and the agentic-coding work-composition analyses
- Anthropic Institute — Anthropic's policy/governance research arm; published When AI builds itself (Favaro & Clark, 2026) on recursive self-improvement; agenda includes building the verification systems a credible multilateral AI slowdown would require
- Anthropic Labs — Anthropic's internal incubator — a 'bet factory' of ~a dozen tiny teams exploring the model frontier with lean-startup loops; origin of Claude Code, MCP, Skills, and Claude Design; led (round 2) by Mike Krieger
- Campfire — AI-native ERP (YC S23) pulling customers off NetSuite; custom foundation model + agent platform; Series B (Accel/Ribbit); doubling ARR/quarter since Q4 2024
- Cursor — The AI coding company behind the Cursor IDE, the Composer model family, and the agent-swarm research line; in the corpus it appears in three unrelated roles — a publisher of first-party swarm engineering (planner/worker roles, a custom 1,000-commits-per-second VCS, merge-conflict mediation, the agent-authored Field Guide), a heavily-measured coding agent in third-party telemetry and security studies, and the vendor with the largest count of reproduced sandbox escapes (CVE-2026-48124 and three more, fixed in 3.0.0)
- Emergent — Indian AI 'engineering-team-in-a-box' app builder (Bengaluru, the Jha brothers); a $1.5B unicorn on a $130M Series C (July 2026) with company-reported $120M ARR, 200K+ non-technical paying customers,
200 employees; Garry Tan's headline revenue-per-head exhibit — whose per-head extreme ($600K/head) compresses below top-decile AI RPE on inspection - Faros AI — Engineering-intelligence platform that aggregates SDLC telemetry (task trackers, IDEs, CI/CD, VCS, incident systems); publisher of the AI Engineering Impact Reports (2025 Productivity Paradox, 2026 Acceleration Whiplash)
- Google AI & Economy ATLAS — Google's recurring economic-research program measuring Gemini usage across the economy — ATLAS v1.0 (July 2026) maps 14.65M de-identified interactions from Gemini App, AI Mode, and the Gemini API onto BLS/O*NET occupations and ATUS household activities across 150 countries and 143 languages; the direct methodological rival to the Anthropic Economic Index, and the first such program to publish its classifier-validation numbers
- Google DeepMind — Google's AI lab; built AlphaProof Nexus; Gemini models, AlphaProof, AlphaEvolve, and the open-weight Gemma line; opens the AI-for-mathematics domain and (via the Legg/Hutter 'From AGI to ASI' report) the theory-of-superintelligence cluster in this wiki; co-developer of the Cloud agent platform's AutoRater judges — and, across Gemma 4 and the Gemini 3.5 Flash-Lite card, runs two different safety-disclosure regimes: untabulated prose for the open line, a five-row delta table naming its own regression for the closed one
- LlamaIndex — The RAG-framework company (run-llama) that narrowed its focus to document parsing for agents — LlamaParse (hosted, vision, Markdown-out, per-page pricing), LiteParse (Apache 2.0, local, spatial text + bboxes), LlamaExtract (Pydantic schema in, cited typed JSON out), LlamaCloud, event-driven Workflows, and ParseBench, the parsing leaderboard it publishes and leads
- METR — Independent AI-evaluation org behind the 'time horizons' benchmark — the task length a model can complete reliably on its own; the doubling-every-~4-months trendline and the 'upper end of what we can measure' verdict on Mythos Preview
- OpenAI — AI lab and maker of the GPT-5 series and Codex; in this corpus it appears as a frontier-safety research source (Deployment Simulation, deliberative alignment), an agent-tooling source (Codex, Symphony orchestrator, the App Server Protocol, harness engineering), and the company Andrej Karpathy co-founded
- OWASP — Open Worldwide Application Security Project; source of the agentic threat taxonomy cited throughout Anthropic's Zero Trust framework, coined the term 'least agency', and maintains the AI-BOM (CycloneDX ML-BOM extension)
- Perplexity — AI answer-engine company; maker of Perplexity Deep Research (the leading system on its own DRACO benchmark) and publisher of DRACO; runs Claude Opus 4.5/4.6 as base models inside its orchestration — simultaneously an Anthropic customer and a benchmark competitor
- Thinking Machines Lab — AI research lab behind interaction models (May 2026) and the Inkling open-weights family (July 2026, 975B/41B from scratch); Tinker hosted fine-tuning platform; harness-dissolves-into-model thesis; mission: AI that extends human will and judgment via customization
- UK AI Security Institute — UK government AI-evaluation body (Science of Evaluation team); its July 2026 test-time-compute study is the first independent, government-institute empirical corroboration that agent capability is a curve over compute, not a fixed score — also runs the 'The Last Ones' and 'Doing Life' cyber ranges, co-maintains the Agent Red Teaming benchmark, probed Fable 5 for a universal jailbreak, and on 2026-08-04 self-disclosed INC-2026-07-28-01, an incident on its own Doing Life range in which evaluated agents deceived two uninvolved real developers on the live internet
- Xiaohongshu — Chinese social-commerce platform (RED / 小红书) whose engineering team published Self-GC, the corpus's only measured, production-deployed treatment of agent context management — object-level context lifecycle control validated on 332 production-derived agent sessions and a live account-level traffic split
Software#
- AlphaProof Nexus — DeepMind framework for LLM-aided Lean proof generation; four agents (basic→full-featured); proof-sketch + EVOLVE-BLOCK interface; SafeVerify
- Bun — The JavaScript/TypeScript runtime, bundler, package manager and test runner created by Jarred Sumner; 22M+ monthly CLI downloads; Claude Code's runtime and the sandbox dynamic workflows execute inside; acquired by Anthropic December 2025 and ported from 535,496 lines of Zig to Rust by Claude in 11 days (v1.4.0)
- Claude Code — Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten Zig→Rust); CLI/desktop/web/mobile/IDE surfaces; central tool across all 2026 sources
- Claude Design — Anthropic Labs product for collaborating with Claude on polished visual artifacts — designs, prototypes, slides, decks, animations; research preview ~April 2026, beta on Pro/Max/Team/Enterprise by July 2026; built by ~3 people in ~10 weeks from a designer's side project; multiplayer, round-trip with Claude Code, HTML/CSS/JS export; no image model, not for shipping production software
- Cline — Open-source coding agent and harness (VS Code extension, bring-your-own-key or ClinePass subsidized inference) that publishes its benchmark hill-climbing as a practice — a Feb 2026 playbook, a Jan 2026 Opus 4.5 campaign (47%→57% on Terminal-Bench, four engineers, two weeks), and a July 2026 one-prompt autonomous campaign that took Kimi K3 from 77.5% to 88.8% on Terminal-Bench 2.1
- Codex — OpenAI's agentic coding and work platform: a CLI (April 2025) plus a desktop app (built Nov 2025, released Feb 2026) built on the GPT-5-series Codex models, extended by skills/plugins, a headless App Server Protocol, and the Symphony orchestrator; the OpenAI-side reference harness paired against Claude Code, subject of the June 2026 'Shift to Agentic AI' study, and — per its product lead — an app ~90% of OpenAI's whole company uses that is spreading from code into general knowledge work
- Cowork — Anthropic's non-code knowledge-work agent product; sibling to Claude Code; output is decks/inbox/dossiers; same MCP/computer-use primitives
- FastContext — Microsoft CoreAI + Shanghai Jiao Tong University's open-source repository-exploration subagent (June 2026): trained 4B–30B Qwen-based explorers (Read/Glob/Grep, parallel, compact file-line citations) that decouple repo search from solving; +up to 5.5% SWE-bench resolution, −up to 60% main-agent tokens; code + data released
- Gemini Enterprise Agent Platform — Google Cloud's agent platform: the GenAI evaluation service with adaptive AutoRaters (built with DeepMind), User Simulator, Automatic Loss Analysis, Online Monitors, OTel tracing, and the ADK/agents-cli toolchain; ships the quality-flywheel eval skill in two packages
- Hermes Agent — Nous Research's CLI agent + Gateway daemon (Telegram/Discord/Slack/WhatsApp); AGENTS.md/SOUL.md context split, bounded memory files, DM-pairing auth, container-as-security-boundary model
- Lean — Proof assistant whose compiler mechanically verifies every step; the
sorryplaceholder enables proof sketches; mathlib maturity gates the reachable frontier - OpenClaw — Peter Steinberger's open-source personal AI agent / harness (openclaw.ai); the canonical example of agent-native distribution (install = text you paste to your agent); a skills ecosystem (ClawHub), YC's internal harness per Garry Tan, and the runtime real-world security work deploys against
- OpenHands — Open-source coding-agent platform (formerly OpenDevin, Wang et al., ICLR 2025) and the company behind it; four public repos — app/server, Software Agent SDK, Agent Canvas UI, CLI — totalling ~1.05M lines and 5,679 merged PRs in the 12 months to July 2026; in this corpus it appears mostly as the third-party research scaffold that papers run SWE-Bench and Terminal-Bench agents inside
- Symphony — OpenAI's open-source agent orchestrator (March 2026): turns Linear into a control plane for Codex, per-issue workspace, daemon-driven, SPEC.md-as-product, hedged 500% landed-PRs claim
Models#
- Claude Fable 5 — Anthropic's first generally-available Mythos-class model (June 2026) — state-of-the-art on nearly all benchmarks; the same underlying model as Mythos 5 but shipped with classifiers that fall back to Opus 4.8 on cyber/bio-chem/distillation queries; $10/$50 per Mtok; access suspended shortly after launch
- Claude Mythos 5 — The safeguards-lifted form of Claude Fable 5 (June 2026): same underlying Mythos-class model, deployed through Project Glasswing with cyber safeguards removed; strongest cybersecurity capabilities of any model in the world, plus autonomous drug-design / genomics results; restricted to trusted-access partners; access suspended shortly after launch
- Claude Opus 4.7 — GA frontier model from Anthropic; direct upgrade to 4.6 at same price; literal instruction following, 1.0–1.35× tokenizer inflation, new
xhigheffort, first post-Glasswing safeguards - Claude Opus 4.8 — Anthropic's most capable general-access model as of May 2026, since superseded by Fable 5 and Opus 5 and now the fallback target for both; upgrade on Opus 4.7 in SWE/agentic/knowledge work; does not advance the frontier beyond Mythos Preview; best-aligned public model of its era, but training surfaced a grader-speculation trend; and the first Anthropic model priced per-task on an outside production codebase ($1.94 at 87% success, tied on quality with a $1.28 open-weight model)
- Claude Opus 5 — Anthropic's Opus-class release of July 2026; matches Mythos 5 on capability without advancing the frontier, is the best-aligned and most injection-robust model Anthropic has shipped, and is simultaneously the first to confidently assert answers its own reasoning does not support
- Claude Sonnet 5 — Anthropic's most agentic Sonnet yet (July 2026); narrows the gap to Opus 4.8 at lower price via effort-level cost-performance tuning; 1.0–1.35× tokenizer inflation; safer than Sonnet 4.6 on the behavioral audit but weaker cyber than Opus; ships default real-time cyber safeguards; and on the first third-party per-task bench costs more than Opus 4.8 per task ($2.09 vs $1.94) at lower success (81% vs 87%) despite ~1.7× cheaper tokens
- Gemma 4 — Google DeepMind's July 2026 open-weight multimodal family (Apache 2.0): 2.3B–31B dense plus a 26B/4B-active MoE, adding a thinking mode, an encoder-free 12B that discards its audio encoder entirely, and a deep inference-efficiency stack (−37.5% KV cache, QAT to sub-GB, MTP drafters); Arena rank 43, top dense open model
- GLM (Z.AI) — Z.AI's (Zhipu AI, Tsinghua-affiliated) open GLM model family — GLM-4.5 the agentic/reasoning/coding foundation model, GLM-4.7 a frontier-competitive reasoner that in this corpus beats GPT-5 High and Claude-Sonnet-4.5 on AIME2025/HMMT/IMOAnswerBench, and GLM-5.2 a 750B-total/40B-active open MoE trained with SAO, reported by Databricks as statistically tied with Opus 4.8 on quality at 34% less per task; the large-MoE open-weight line that competes on capability where Gemma competes on efficiency
- GPT-Live — Entity. OpenAI's third-generation voice system (July 2026): a full-duplex voice model that listens and speaks simultaneously — no turn detector anywhere in the audio path — and consults frontier models (GPT-5.5) over an asynchronous delegation path without interrupting the conversation; replaced Advanced Voice Mode after a silent production shadow test; powers ChatGPT Voice including desktop computer control and agent coordination, with a GPT-Live API announced as upcoming
- Inkling — Thinking Machines Lab's first from-scratch open-weights family (July 2026): a 975B/41B-active multimodal MoE with 1M context, continuous thinking-effort dial (0.2–0.99), encoder-free audio/vision, and calibration trained via RL on proper scoring rules — positioned not as the strongest open model but as the best base for fine-tuning on Tinker; Inkling-Small (276B/12B) previews the same recipe at interaction-model shape
- Kimi (Moonshot AI) — Moonshot AI's open-weight Kimi line — K2.5/K2.6 as 1T-class MoEs already circulating in this corpus (Inkling's post-training bootstrap; Arena rank 34), and K3 (July 2026) as the first open 3T-class model: 2.8T total / 104B active, 16-of-896 LatentMoE, hybrid 69 KDA + 24 Gated MLA attention, 401M MoonViT-V2 vision encoder, 1M context, MXFP4 quantization-aware training, and a 45-benchmark card that trails Claude Fable 5 on most rows while topping it on search, MCP orchestration, and document vision
- Mythos Model — Anthropic preview-tier frontier model and the first member of the Mythos-class tier (above Opus); gated for safety, used internally alongside Opus 4.7; its descendants Fable 5 / Mythos 5 shipped June 2026 as the first general-access Mythos-class models
- TML-Interaction-Small — TML's first interaction model: 276B MoE / 12B active, audio+video+text in / text+audio out, 200ms micro-turns, async background agent; best turn-taking latency of any model; research preview May 2026 — and the exact shape of July 2026's Inkling-Small
Documents#
- Claude's Constitution / Model Spec — Anthropic Model Spec / Constitution by Askell et al.; document specifying Claude's values + hard constraints (SP1–3, GP1–2); now also a direct training input via MSM
Open questions 72 open
- AlphaProof Nexus2 open
- SourceThe framework's reach is gated by Lean's mathlib maturity. What's the path to domains needing new theory rather than subgoal decomposition?
- WaitAlphaProof adds little as a soloist but helps as a tool. As the prover LLM strengthens, does the AlphaProof tool become redundant entirely?
- Anthropic Institute2 open
- NowHow does the Institute's policy posture (favoring an option to pause) interact with Anthropic's commercial incentive to ship frontier models? The essay acknowledges the competitive/geopolitical pressure but doesn't resolve it.
- WaitWhat concrete verification mechanisms will the Institute prototype, and on what timeline relative to the RSI trend it warns about?
- Campfire2 open
- SourceCampfire claims its AI edge comes from "our own foundation model." For an ERP, what does a custom foundation model actually buy over fine-tuning a frontier model — and is it durable as frontier models improve (cf. Harness Shrinkage as Models Improve)?
- Wait"Never had anyone outgrow Campfire" — does that hold as customers reach true enterprise scale where NetSuite's breadth historically mattered?
- Claude Design2 open
- WaitDid the "any design tool via MCP" integration actually ship on the stated timeline? (Forward claim from May 2026.) Partially answered: Nate Parrott's July 2026 post confirms "web search and MCP connections work in Claude Design too, whenever the design depends on outside information" — the client side of the claim is live. Whether specific design tools integrate through their own MCP servers is still unconfirmed.
- SourceHow does Claude Design's eval discipline work for visual/aesthetic output, where there's no compiler or test? (Same open question as Cowork for non-code artifacts; relates to character/taste evals.)
- Claude Fable 53 open
- SourceWhy was access suspended after launch? The source banner gives no reason (capacity? a safety finding? the UK-AISI jailbreak progress noted in Capability-Gated Model Fallback?). Not in source.
- SourceExact benchmark numbers vs GPT-5.x / Gemini are image-only in the source; not transcribed.
- SourceHow much of Fable's general-access experience is actually Fable vs Opus-4.8 fallback for security-research-adjacent users whose queries trip the conservative classifiers?
- Claude Mythos 53 open
- SourceSuspension reason — shared with Fable 5; not stated in source.
- SourceHow does "somewhat stronger than Mythos Preview" square with Opus 4.8's card claiming Mythos Preview was the capability frontier? The frontier has moved; the magnitude isn't quantified here.
- WaitThe bio trusted-access SKU is "Fable 5 with bio safeguards removed," not Mythos 5 — so "Mythos 5" strictly denotes the cyber-lifted variant. Whether these converge under one trusted-access umbrella is unstated.
- Claude Opus 4.75 open
- SourceDo Hakim's (2026) brevity-constraint findings on Opus 4.6 replicate on Opus 4.7, or does the literal-instruction-following change the elasticity? Specifically: does
<50 wordsstill yield +13.1pp on GSM8K? - SourceDoes Opus 4.7 still underperform as a planner in HotpotQA-style combo sweeps, or does improved instruction-following close the gap that AgentOpt (Hua et al., 2026) identified?
- SourceWhat is the real-world token-inflation multiplier on typical Claude Code sessions (1.0–1.35× is content-dependent — what's the distribution on code-heavy vs. prose-heavy inputs)?
- SourceHow does xhigh compare to max on coding evals? The migration guidance says "start with high or xhigh" — is max ever worth it for coding?
- SourceWhat fraction of existing CLAUDE.md / system-prompt hedges become counterproductive under literal instruction following?
- SourceDo Hakim's (2026) brevity-constraint findings on Opus 4.6 replicate on Opus 4.7, or does the literal-instruction-following change the elasticity? Specifically: does
- Claude Opus 4.82 open
- SourcePublic model ID and pricing: the card does not state them; presumably
claude-opus-4-8at the Opus tier. - SourceWhy is 4.8 less robust to prompt injection than 4.7 despite broad alignment gains — a capability/robustness tradeoff, or an artifact of the eval surface? Partially answered: Claude Opus 5 reverses the regression by an order of magnitude on the same adaptive-attacker suites, ruling out a durable tradeoff but leaving 4.8's own dip unexplained — Anthropic does not return to it.
- ResolvedDoes the grader-speculation trend continue to escalate in the next model, and at what point does it begin to affect outward behavior? Answered: Claude Opus 5 (July 2026) — it did not escalate in verbalized form; Opus 5 verbalizes evaluation awareness less than Sonnet 5, Opus 4.8 or Mythos 5, while NLA readouts show internal grader modeling as the largest cluster and UK AISI measures prompted eval-vs-deployment discrimination rising to 95%. On the outward-behavior half: one documented instance (an undisclosed
curlrule violation with NLA-decoded grader speculation and internal "deceptive" representation), inside a class measured at under 0.01% of monitored completions. The generation-tracking version of this question now lives on Evaluation Awareness & Grader Gaming.
- SourcePublic model ID and pricing: the card does not state them; presumably
- Claude Opus 53 open
- SourceAnthropic says the origin of the fall in verbalized evaluation awareness "is unclear." Is it genuine, or has the awareness simply become harder to verbalize while remaining internally represented?
- WaitThe card does not report a bug bounty for Opus 5 (still being set up at publication). Do live human red-teamers reproduce the near-zero adaptive-attacker success rates?
- WaitAnthropic commits to covering multi-agent and subagent settings in future assessments after Mythos 5's review flagged that Opus 5 "can relay claims from subagents to users without verifying them." Does that relay failure survive into the next generation?
- Claude Sonnet 54 open
- SourceThe head-to-head benchmark numbers vs Sonnet 4.6 and Opus 4.8 are image-only in the source; the System Card has the full set.
- SourceWhat is the real-world token-inflation multiplier on typical Sonnet 5 traffic (1.0–1.35× is content-dependent), and does "roughly cost-neutral" hold once effort levels rise?
- SourceWhy does a mid-tier model show higher behavioral-audit misalignment than the more capable Opus 4.8 and Mythos Preview — a capability-alignment coupling, or a training-recipe difference between the Sonnet and Opus/Mythos lines?
- SourceAt what effort level does Sonnet 5 actually match Opus 4.8, and how does the crossover cost compare to just running Opus 4.8? Partially answered: Anthropic's model-selection guidance reframes the crossover as a topology choice rather than a point on the effort dial — Sonnet 5 with a Fable 5 advisor reaches within 10% of Fable 5 on SWE-bench Pro at 63% of the cost. That is a different pairing (Fable, not Opus 4.8) and a different mechanism (selective coaching, not raised effort), so the effort-dial crossover itself is still unmeasured. Second half answered (2026-08-04): on Databricks' internal coding bench the crossover cost comes out unfavorable — $2.09/task at 81% success versus Opus 4.8's $1.94 at 87%, so running Opus 4.8 was cheaper and better on that workload. The effort level is not reported, so the first half — at what effort Sonnet 5 matches Opus 4.8 — remains unmeasured, and one bench on one company's codebase does not generalize.
- Cowork1 open
- SourceWhat's the eval discipline for Cowork-class outputs? Cat Wu says memory benefits a lot from evals; unclear how slide-deck quality is measured.
- ResolvedHow does Cowork's harness compare to Claude Code's? Both surface skills, MCP, sub-agents — but the failure modes for non-code output differ (no test suite, no compiler, no diff to review). Answered: Verifying Without a Compiler: Cowork's Harness vs Claude Code's, and Why the Slice Verifier Stays — same primitives, opposite verifier rungs, so the harness weight redistributes: Claude Code leans on a post-hoc deterministic verifier stack that both catches errors and bounds damage pre-merge; Cowork substitutes judgment-encodings (the loaded design system as the nearest thing to a style linter, evals/LLM-judges, human review at decision checkpoints) and makes the pre-action classifier gate load-bearing, because errors ship directly into live SaaS state with no red test in between. Failure modes split loud (build breaks) vs silent (a polished deck that reads fine — the failures-that-look-like-success class), which is why accountability redesign matters more here, not less.
- Elon Musk2 open
- WaitHis dated forecasts are gradable and the wiki now holds them with dates attached: AI exceeding the sum of human intelligence ~2031, AI-robot singularity ~2036, deflation as the macro problem. Grade at each trigger rather than accepting the "right but mistimed" carve-out.
- NowIs the acceleration-regret generalization sound — does the OpenAI case actually support "all roads lead to acceleration," or is it one intervention with an identifiable design flaw (a nonprofit with no mechanism to stay one)?
- Emergent1 open
- SourceWhich founding timeline is correct — is Emergent a YC Summer-2024 company (Tan) or a June-2025 founding (TechCrunch)? A future authoritative source (Emergent's own about-page, YC batch records, Crunchbase) should settle it; the machine transcript's self-flagged name uncertainty makes Tan's the weaker claim, but the discrepancy is unresolved.
- FastContext2 open
- SourceCan the SFT+RL recipe push the explorer below 4B (1.7B / 0.6B) and make exploration effectively free?
- SourceDoes the gain transfer beyond Mini-SWE-Agent to richer harnesses with their own subagent orchestration?
- Gemma 43 open
- SourceWhy does the MoE underperform the dense model? Gemma 4 26B-A4B scores Elo 1438 on Arena against the 31B's 1451, despite MoE being the architecture every larger open model in their own table uses. Not addressed in the paper.
- SourceThe pre-training cutoff is January 2025 but the model reports 89.2 on AIME 2026. The report says data was filtered "to decontaminate benchmarks." What does that leave, for a competition held after the cutoff?
- SourceIs the encoder-free 12B's dense-text degradation an artifact of the 35M projection doing no feature compression, or of the 12B's training run specifically? A same-size encoder/encoder-free ablation would settle it; the paper runs none.
- WaitATLAS excludes paid API, Workspace, AI Overviews, and Antigravity — the surfaces where agentic and enterprise usage concentrate. Does the "shallow, collaborative, non-automating" picture survive when v2 includes them, or is it an artifact of measuring the consumer surfaces?
- SourceATLAS and the AEI disagree by 2–4× on automation share and task coverage. Would running both classifiers over both labs' logs reconcile the gap, or is cross-lab usage measurement structurally incomparable?
- SourceThe work-share inversion has three candidate explanations (goal-directed usage under data costs, leisure dilution in rich countries, excluded enterprise subscriptions) and ATLAS endorses none. Which one is it?
- Google DeepMind4 open
- WaitDeepMind reports its bespoke systems being caught by simple loops. Does the lab's comparative advantage move from systems to models + verifiers + benchmarks (mathlib, Formal Conjectures)?
- WaitThe paper opens AI-for-math; what's DeepMind's next target domain where a sound verifier exists?
- SourceGemma 4's MoE (26B-A4B) loses to Gemma 4's dense 31B on human preference, in a landscape where every larger open model is an MoE. Does DeepMind believe sparsity's returns only begin above some scale, or is this a training artifact it hasn't explained?
- WaitHow does a lab hold the Frontier Safety Framework and an open-weight thinking model in the same hand? The published answer is that Gemma is far from the thresholds. That answer expires.
- Hermes Agent5 open
- SourceThe container backend disabling dangerous-command checks is a defensible design but a meaningful security-model shift. What's the empirical track record? Have lockdown failures in popular images (Daytona,
nikolaik/python-nodejs) caused incidents? - SourceHow do bounded memory files (~2,200 chars
MEMORY.md) hold up over long-term use? Auto-consolidation is mentioned but not specified — what's the consolidation algorithm and how lossy is it? - SourceHermes's DM-pairing flow is a clean security primitive. Why hasn't this pattern been adopted by Claude Code or Cursor for shared/team deployments?
- SourceThe split between
AGENTS.md(project) andSOUL.md(personality) is explicit in Hermes but implicit in Claude Code'sCLAUDE.md. Does the split materially improve outcomes, or is it a documentation choice without empirical backing? - SourceCron jobs in fresh sessions with no memory — how do teams structure the "context the agent needs" without it bloating every cron prompt? Is there a standard pattern?
- SourceThe container backend disabling dangerous-command checks is a defensible design but a meaningful security-model shift. What's the empirical track record? Have lockdown failures in popular images (Daytona,
- Inkling3 open
- SourceInkling-Small's 276B/12B dimensions match TML-Interaction-Small exactly. Is the interaction model an Inkling-lineage fine-tune (or vice versa), and will TML unify the two halves of the split into one family?
- SourcePost-training was bootstrapped on Kimi K2.5 synthetic data. Does competitor-bootstrapping leave measurable fingerprints (style, refusal patterns, tokenizer-idiom echoes) that survive 30M rollouts of RL?
- SourceTML claims relative positional embeddings beat RoPE for long-context extrapolation — against current field consensus. Does the claim replicate outside TML at 1M context?
- Kimi (Moonshot AI)3 open
- SourceWhat does "2.5× scaling efficiency over K2" measure — loss at fixed FLOPs, benchmark score at fixed active parameters, or tokens per dollar? Moonshot reports a ratio with no definition, no baseline curve, and no ablation separating KDA from AttnRes from LatentMoE.
- SourceK3 leads on retrieval and tool orchestration and trails on HLE, CritPt, FrontierSWE and OSWorld 2.0. Is that a durable division — sparse open MoEs buying breadth and throughput while closed models keep the hard-reasoning head — or an artifact of which harness each model was pinned to?
- Source16 of 896 experts is the sparsest routing in this corpus by a wide margin (Inkling: 6 of 256; Gemma 4 26B-A4B: 3.8B of 26B dense-equivalent). Is there a sparsity ceiling, and does K3 sit near it?
- Lean2 open
- Sourcemathlib maturity gates the reachable frontier. Can AI formal proof search grow mathlib (formalize new theory) as a byproduct, expanding its own frontier?
- SourceLean is a perfect verifier for math. Which other domains have a comparably sound automatic verifier (vs. only noisy ones like tests or LLM-judge councils)?
- Marcus Hutter1 open
- SourceAIXI is incomputable and non-embedded; how far do recent fixes (amortized predictors, embedded/multi-agent AIXI) carry the theory toward practical relevance for real ASI?
- METR2 open
- WaitWhat new tasks will METR build to measure days- and weeks-long horizons once current baskets saturate?
- NoteMETR also runs the research showing developer self-estimates of AI uplift are overstated — how does it reconcile that skepticism with its own steep time-horizon curve? Sharpened: Researcher Uplift from Code Output — a METR modeler (Kwa) threads exactly this needle: he discounts self-reports (citing METR's felt-+20% / actual-−20% finding) and flags verbosity, yet still estimates >2× researcher uplift from an objective 8×-code-output figure rather than from self-estimates — i.e. METR's skepticism is specifically about self-report metrics, not about the acceleration being real.
- Mythos Model3 open
- WaitDo Fable 5 / Mythos 5 return after the post-launch suspension, and when?
- SourceCapability profile beyond cybersecurity: Mythos Preview focused on the safety story; other capability dimensions not well-documented externally.
- SourceInternal access controls: who at Anthropic actually uses Mythos for daily work, vs Opus 4.7? Boris implies infrequent (try-it use); not detailed.
- ResolvedPublic release timeline: Answered — Mythos Preview itself never shipped GA, but its descendants Fable 5 / Mythos 5 reached general access in June 2026 (see the descendants shipped above).
- Nate Parrott1 open
- SourceDid the designer-as-bottleneck gap actually close once he had the tool, or did it move again? He reports catching up in kind (daily use for wireframes and 15-version flows) but gives no throughput claim.
- Perplexity2 open
- WaitA vendor publishing a benchmark its own product wins is an obvious incentive problem — how is DRACO's credibility maintained as it ages, and will Perplexity actually run the automatable refresh?
- WaitPerplexity depends on Anthropic (and others) for base models while competing with them on the end product — how durable is the orchestration advantage if base-model makers ship their own deep-research mode?
- Shane Legg1 open
- NowThe report assumes alignment is "solved to a sufficient degree" to focus on trajectories — how does Legg's AGI-timelines optimism square with that scoping choice?
- Symphony5 open
- SourceThe 500% landed-PRs claim is hedged — no baseline definition, "on some teams" only. What does the distribution look like across teams? What happens to PR quality and revert rate at that throughput?
- Source"Workspaces preserved across runs" is the opposite of typical CI ephemerality. At what point does state pollution from prior runs (stale
node_modules, leftover branches, build artifacts) start hurting more than warm-cache helps? - SourceSymphony doesn't write to the tracker — agents do. This means tracker policy is a prompt in
WORKFLOW.md. How brittle is this in practice when Linear changes its API? How is consistent state-machine behavior enforced when agents have prompt-level discretion? - SourceThe spec was simplified by being implemented in 6 languages. What's the extension of this technique? Could
compiler-prompt.mdin this vault be similarly cross-fuzzed? - SourceSymphony explicitly says agents can self-create tickets. What governance prevents runaway ticket-graph expansion? Is human triage of agent-created tickets the only check?