H
Howardism
Plate IIEntities中文HOWARDISM

Google DeepMind

PublishedMay 23, 2026FiledEntityDomainEntitiesTagsEntityOrganizationAI LabReading13 minSourceAI-synthesised

Google's AI lab; built AlphaProof Nexus; Gemini models, AlphaProof, AlphaEvolve, and the open-weight Gemma line; opens the AI-for-mathematics domain and (via the Legg/Hutter 'From AGI to ASI' report) the theory-of-superintelligence cluster in this wiki; co-developer of the Cloud agent platform's AutoRater judges — and, across Gemma 4 and the Gemini 3.5 Flash-Lite card, runs two different safety-disclosure regimes: untabulated prose for the open line, a five-row delta table naming its own regression for the closed one

Illustration for Google DeepMind

Sources#

Summary#

Google's AI research lab. In this corpus it appears as the lab behind AI-Driven Formal Proof Search — the team (George Tsoukalas, Anton Kovsharov, Sergey Shirobokov, Swarat Chaudhuri, Pushmeet Kohli et al.) that built AlphaProof Nexus and ran the first large-scale evaluation of LLM-aided formal proof search on open research mathematics (arXiv 2605.22763). It is also the maker of the Gemini model family used throughout (Gemini 3.1 Pro as prover, Gemini 3.0 Flash as rater), the prior AlphaProof olympiad theorem-prover, and AlphaEvolve, whose evolutionary design inspired Evolutionary Proof Search.

Role in the corpus#

DeepMind is the third frontier-lab "voice" in the wiki alongside Anthropic and OpenAI (Symphony / Agent Harness Engineering), and the one that opens the AI-for-mathematics domain. Its contribution is methodological as much as mathematical: the paper's finding that simple agentic loops increasingly rival DeepMind's own bespoke trained systems (Agentic Loops Overtake Bespoke Systems) is a candid, self-undercutting result — a lab that built specialized RL provers reporting that a plain LLM loop is catching up.

It is also the source of the wiki's theory-of-superintelligence cluster. The June 2026 report From AGI to ASI — senior-authored by co-founder Shane Legg with Marcus Hutter (creator of AIXI) and twelve others — maps the four pathways from AGI to ASI, grounds them in the Universal AI upper bound, and frames the frictions (the The Abstraction Barrier, the data wall, deliberate slowdown) as open research questions. Where Anthropic's When AI builds itself argues RSI from internal measurement, DeepMind's report is the theory-first sibling — same question, formal framing.

The open-weight line (Gemma)#

DeepMind runs two model lines with different theories. Gemini is the closed frontier line. Gemma is the open-weight line — Apache 2.0, aimed at "varied hardware environments" and edge deployment rather than at the leaderboard.

Gemma 4 (July 2026) is the corpus's entry point into that line, and it is the wiki's only substantial source on the deployment side of the stack: KV-cache reduction, quantization-aware training, speculative decoding, encoder removal (Inference Efficiency as Capability). It places DeepMind in a third posture beyond the two above — not the frontier lab, not the theorist, but the one shipping capability that anyone can download and nobody can recall.

That posture generates a tension the wiki records rather than resolves. DeepMind authors the Frontier Safety Framework (2024) and publishes an open-weight model with a thinking mode whose safety evaluations are reported as prose with no tables and no compute budget. The report is careful about capability and casual about safety, in the same PDF. See Open-Weight Elicitation Irreversibility — a structural argument, not an alarm about Gemma 4 specifically, which sits at Arena rank 43.

Gemma 4's 12B is also the second independent instance of Encoder-Free Early Fusion, arrived at for memory reasons where Thinking Machines arrived at it for latency. Two labs, orthogonal objectives, same architectural verdict.

Systems and models referenced#

  • Gemini 3.1 Pro / 3.0 Flash / 3.1 Flash-Lite — the LLM backbone; Pro for proving, Flash for rating; the smaller variants solved no problems (capability is sharply scale-gated — Scale-Dependent Prompt Sensitivity).
  • Gemini 3.5 Flash-Lite (2026-07-21) — the efficiency tier of the Gemini 3 family; see below. 1M-token input across text/image/audio/video, 64K output, knowledge cutoff March 2026 (with the card conceding some domains are stuck at January 2025). Shipped to the Gemini App, AI Studio, the Gemini API and Gemini Enterprise.
  • AlphaProof — DeepMind's RL-trained olympiad-level Lean prover; used inside Nexus as a focused subgoal tool (and the system behind earlier IMO results).
  • AlphaEvolve — the evolutionary-coding system whose population/diversity approach Evolutionary Proof Search adapts; also helped formulate the bipartite graph-reconstruction variants in the paper.
  • Formal Conjectures repo — DeepMind's open-source Lean formalizations of Erdős problems, the benchmark for the Erdős runs.
  • AutoRaters — the adaptive LLM-as-a-Judge graders at the core of Google Cloud's Gemini Enterprise Agent Platform evaluation service, developed in close partnership with DeepMind and (per Google) the same ones used to evaluate its own models and first-party agents; the grading engine of the Agent Quality Flywheel.
  • Gemma 4 — the open-weight family (2.3B–31B dense + a 26B/4B-active MoE), Apache 2.0, July 2026. Thinking mode, encoder-free 12B, and the efficiency stack.
  • TPU v5p / v6e + Slice-Granularity Elasticity — the training substrate (4,096–12,288 chips per Gemma 4 model); elasticity reduces the stall from a localized chip failure "from many minutes to a few seconds."
  • Frontier Safety Framework (2024) — the safety commitments Gemma 4's §5 invokes, without reporting numbers against them.

The economics posture (ATLAS, July 2026)#

A fourth posture beyond frontier lab, theorist, and open-weight shipper: economic measurement of its own deployment. ATLAS v1.0 (July 23, 2026) is joint Google / Google DeepMind, and DeepMind's contribution is the instrument — OCTO (Observation Clustering and Taxonomy Organisation), the bespoke clustering and hierarchical-taxonomy tool that groups 14.65M de-identified Gemini conversations before mapping them onto BLS occupations and ATUS activities. Gemini 3.1 Flash-Lite does the classification throughout, including generating the synthetic ground-truth set used to validate itself.

This places DeepMind opposite Anthropic's Economic Index on the wiki's usage-measurement axis, and the report is unusually candid for a first-party artifact: it publishes classifier accuracy numbers no competing program has (Usage-Telemetry Classifier Validation), states seven limitations including the exclusion of Workspace, AI Overviews, and Antigravity, and puts named external economists (Diane Coyle, David Autor) inside the review. The contrast with Gemma 4's untabulated safety prose is worth noting: the same organization is rigorous about measurement when the subject is economics and casual when it is safety. (Refined 2026-07-30 by the Gemini 3.5 Flash-Lite card — the casualness tracks the open line, not the lab: the closed line's card tabulates five safety deltas including one that goes the wrong way. See below.)

The closed line's efficiency tier (Gemini 3.5 Flash-Lite, July 2026)#

The wiki's first primary source on the Gemini side of the two-line strategy. Gemini 3.5 Flash-Lite (2026-07-21, vendor-claim) is an iteration on 3.1 Flash-Lite that inherits its training data, hardware and software sections wholesale — the card documents deltas, not a system. Three things it establishes:

  • The efficiency tier bought a capability tier and raised its price. SWE-Bench Pro 38.3 → 54.2, Terminal-bench 2.1 31.0 → 54.0, OSWorld-Verified 54.3 → 74.0, MLE-Bench 22.0 → 39.2, GDPVal-AA Elo 642 → 1140 — against output pricing that moved $1.50 → $2.50/1M (+67%). DeepMind prints both in one table; the per-dollar consequences are worked through in Inference Efficiency as Capability, and the table's cost rows are analysed as a partial defection from the benchmark grid in Compute-Controlled Benchmarking.
  • A safety regression reported against its own interest. Automated internal evaluations versus 3.1 Flash-Lite: text-to-text safety −8.14pp and multilingual −0.92pp (both improvements), image-to-text unchanged, tone +3.04pp (improvement), and unjustified refusals +5.32pp — a regression, on the axis measuring whether the model can answer borderline prompts rather than refuse them. Human red-teaming is reported as similar-or-improved on child safety and general content policy. That the wrong-way number is in the table at all is the notable part.
  • Frontier Safety cleared by reference to a larger sibling. The assessment concludes no meaningful new capabilities or material increases "relative to frontier safety domains compared to Gemini 3.1 Pro" and no Critical Capability Level thresholds reached. Note the comparator: an efficiency-tier model is cleared against the previous generation's flagship, not against its own predecessor or an absolute threshold — a framing that stays valid exactly as long as the family's ceiling does not move, and is the same relative-determination move the wiki tracks elsewhere.

The disclosure asymmetry is the finding for this page. In the same month, the same lab published Gemma 4's §5 — zero tables, the prose assertion that Gemma 4 keeps "unjustified refusals low", no benchmark named — and Flash-Lite's five-row delta table naming the metric, the direction and the magnitude, including the one that regressed. So DeepMind can measure and publish exactly the thing it declined to quantify in the open-weight report. Whatever explains the gap, it is not that the lab lacks the instrument. The tension recorded above (Open-Weight Elicitation Irreversibility) sharpens rather than resolves: the release that is irreversible is the one whose safety section is prose.

Connections#

  • AI-Driven Formal Proof Search — the paradigm DeepMind demonstrated at research scale
  • Google AI & Economy ATLAS — the joint Google/DeepMind economics program; DeepMind built OCTO, the clustering engine underneath it
  • Usage-Telemetry Classifier Validation — the validation numbers ATLAS published and no rival program has
  • AlphaProof Nexus — its framework
  • Lean — the proof assistant it drives with Gemini
  • Evolutionary Proof Search — adapts DeepMind's AlphaEvolve
  • Agentic Loops Overtake Bespoke Systems — DeepMind's self-undercutting finding about its own bespoke systems
  • Anthropic — peer frontier lab; the two anchor different domains in the corpus (alignment/coding vs. mathematics; and the two RSI framings — empirical vs. theoretical)
  • Scale-Dependent Prompt Sensitivity — Gemini-model scale gating mirrors the broader model-capability-threshold theme
  • Shane Legg — co-founder and Chief AGI Scientist; senior author of From AGI to ASI
  • Marcus Hutter — senior researcher; creator of the AIXI / Universal AI framework the report rests on
  • AGI-to-ASI Pathways — the report's four-pathway map of AI progress beyond AGI
  • Universal AI (AIXI) — the theoretical upper bound DeepMind uses to bound ASI from above
  • DRACO Benchmark — Gemini plays both roles in Perplexity's deep-research benchmark: Gemini Deep Research is an evaluated system, and Gemini-3-Pro is the primary judge model
  • Perplexity — deep-research competitor whose DRACO benchmark uses DeepMind's Gemini-3-Pro as judge-of-record
  • Gemini Enterprise Agent Platform — the Cloud product surface where DeepMind-built AutoRaters ship to customers
  • Agent Quality Flywheel — the eval-fix methodology those AutoRaters power
  • Gemma 4 — the open-weight line; the lab's third posture in this corpus
  • Inference Efficiency as Capability — the deployment-side stack Gemma 4 contributes, absent from the wiki before it; Gemini 3.5 Flash-Lite adds the product-tier version, where efficiency gets more expensive
  • Compute-Controlled Benchmarking — the lab is now on both sides of it: Gemma 4's headline table is the corpus's worked failure, while the Gemini 3.5 Flash-Lite card is the only one to put prices for every compared model in the grid itself
  • Encoder-Free Early Fusion — DeepMind independently corroborates Thinking Machines' design, for memory rather than latency
  • The Open-Weight Frontier Gap — the lab publishes the Arena table that places it 43rd
  • Open-Weight Elicitation Irreversibility — the tension between authoring the Frontier Safety Framework and shipping an unrecallable thinking model
  • Jeff Dean — Google's Chief Scientist, and the corpus's only source on the hardware layer this lab's models run on: the TPU's napkin-math origin, the energy/data-movement ratio underneath every efficiency lever, and the 2014 distillation paper that produces the Gemini Flash line from Pro

Open Questions#

  • DeepMind reports its bespoke systems being caught by simple loops. Does the lab's comparative advantage move from systems to models + verifiers + benchmarks (mathlib, Formal Conjectures)?
  • The paper opens AI-for-math; what's DeepMind's next target domain where a sound verifier exists?
  • Gemma 4's MoE (26B-A4B) loses to Gemma 4's dense 31B on human preference, in a landscape where every larger open model is an MoE. Does DeepMind believe sparsity's returns only begin above some scale, or is this a training artifact it hasn't explained?
  • How does a lab hold the Frontier Safety Framework and an open-weight thinking model in the same hand? The published answer is that Gemma is far from the thresholds. That answer expires.

Sources#

§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 26
  • Gemini Enterprise Agent Platform×3

    Model distribution — the platform is one of the four named launch surfaces for DeepMind's Gemini…

  • Gemma 4×3

    The instrument exists — it was pointed at the closed line instead. Nineteen days later DeepMind…

  • Google AI & Economy ATLAS×3

    Google Deepmind — co-author of the report and builder of OCTO, the clustering tool underneath it

  • Agent Quality Flywheel×2

    Google Cloud's methodology for engineering agent quality instead of vibe-checking it, shipped (June…

  • AlphaProof Nexus×2

    Google Deepmind's framework for LLM-aided formal proof generation in Lean (arXiv 2605.22763).…

  • Inference Efficiency as Capability×2

    Gemma 4 and K3 are efficiency at the level of the architecture — levers inside the model that make…

  • Marcus Hutter×2

    Marcus Hutter is the originator of AIXI and the Universal AI framework — the formal, mathematically…

  • Shane Legg×2

    Shane Legg is a co-founder of DeepMind and a long-standing theorist of machine intelligence. With…

  • Agent Data Injection (ADI)

    Codex / Google Deepmind — Codex and Gemini CLI are equally vulnerable to the origin- and…

  • Anthropic

    Google Deepmind — peer frontier lab; anchors the AI-for-mathematics domain (Ai Driven Formal Proof…

  • Claude Code

    The bash/merge confirmation dialog did not prevent these: because the agent's own displayed…

  • Compute-Controlled Benchmarking

    DeepMind's Gemini 3.5 Flash-Lite card (2026-07-21, vendor-claim) does something none of the others…

  • Cost-per-Task Over Cost-per-Token

    DeepMind's Gemini 3.5 Flash-Lite card (2026-07-21, vendor-claim) is what this page's argument looks…

  • Cross-Lab Pre-Release Review

    He also reports having discussed a related proposal with Demis Hassabis (Google Deepmind) for "a…

  • Deep Research Agents

    Anthropic / Google Deepmind — makers of evaluated systems (Claude Opus; Gemini Deep Research, and…

  • DRACO Benchmark

    Perplexity / Anthropic / Google Deepmind — benchmark author; makers of evaluated systems and the…

  • Encoder-Free Early Fusion

    How much this should move you: not far, and the reason is worth stating rather than resolving. This…

  • Jeff Dean

    Google Deepmind — the lab whose Gemini, AlphaFold, AlphaEvolve and AlphaChip work he cites as the…

  • Lean

    Google Deepmind — the lab building Lean agents at research scale

  • LLM-as-a-Judge

    Google's Gemini Enterprise Agent Platform AutoRaters (developed with Google Deepmind; the grading…

  • Entities — People, Orgs, Tools & Projects

    Google Deepmind — Google's AI lab; built AlphaProof Nexus; Gemini models, AlphaProof, AlphaEvolve,…

  • Open Questions Backlog

    Google Deepmind ×4 (oldest 81d) — DeepMind reports its bespoke systems being caught by simple…

  • Open-Weight Elicitation Irreversibility

    Google Deepmind — publisher of both the Frontier Safety Framework and an open-weight thinking model

  • The Open-Weight Frontier Gap

    Google Deepmind — publishes the table, and places itself 43rd on it

  • Perplexity

    Google Deepmind — competitor (Gemini Deep Research is evaluated) whose Gemini-3-Pro Perplexity also…

  • Write-Then-Trusted

    Claude Code / Codex / Google Deepmind — the affected agent products; the .claude hook-configuration…

Related articles
  • Open Questions Backlog

    _456 actionable open questions across 205 pages · 107 predictions · 9 notes · 147 in progress · 69 watching (entities),…

  • Kimi (Moonshot AI)

    Moonshot AI's open-weight Kimi line — K2.5/K2.6 as 1T-class MoEs already circulating in this corpus (Inkling's post-tra…

  • Inference Efficiency as Capability

    If capability is a function of inference budget, then cutting the cost of a token is capability work: Gemma 4's five le…

  • Large-Scale Test-Time Compute

    Noam Brown's thesis that model capability is now a function of inference budget (tokens/cost/time): with good scaffoldi…

  • AI-Driven Formal Proof Search

    LLM generates Lean, compiler verifies every step → eliminates hallucination; DeepMind resolves 9/353 Erdős + 44/492 OEI…