zhudong.si
DEEP RESEARCH REPORT

DEEPMINDGoogle DeepMind Deep Research Report

Scientific Discovery, R&D Feedback, and Physical Intelligence

Research cutoff:Oct 6, 2026Chapters:19Full text:~79k wordsVersion:V1.0

Scientific Discovery, R&D Feedback, and Physical Intelligence

Research cutoff: October 7, 2026 Version: GPT V1.0 | Independent research | Deliverable: Markdown Research scope: Google DeepMind, plus Google Research, Google Cloud, Waymo, and Isomorphic Labs where they have direct technical interfaces with DeepMind.

The central question of this report: can DeepMind turn "AI that solves a problem" into "a system that keeps producing stronger problem-solving capability"? The conclusion of this report: several local feedback loops already have engineering evidence; whether those loops can be connected into a continuously accelerating general R&D system remains a hypothesis to be verified.


Executive Summary

Google DeepMind is worth studying in depth not only because Gemini competes at the frontier-model level. It offers a set of experiments for testing how superintelligence might form: program search, mathematical research, scientific prediction, hypothesis generation, virtual-environment learning, real robots, and compute-infrastructure optimization.

This report's judgment [D]: the most explanatory thread is "generate candidates — external validation — selection and accumulation — search again." Base models expand the range of proposals; search expands the number of attempts; validators decide which results enter the next round. Real progress depends on how the three work together, and on whether validation covers the actual goal.

This thread explains both the progress and the difficulties:

Seven Core Judgments

Question This report's conclusion Evidence status and main boundaries
Has AI entered AI R&D? Yes — public evidence goes beyond general code completion into compute kernels, systems, and search-process optimization Mainly company papers and deployment disclosures [B]; the share of R&D labor hours is unknown
Has AI R&D takeoff happened? Cannot be confirmed yet Missing continuous cross-generation R&D speed data with controlled inputs and labor
Could partial superhuman capability appear first in research? Possible — structured tasks already have instances Local capability does not imply overall scientific autonomy [B→D]
Will superintelligence first appear as a system? The system explanation is currently stronger, but the single-model and specialized-system-combination explanations remain competitive There is evidence for gains between modules; full integration is unproven [D]
Is multi-agent necessary? No necessity evidence yet Should be compared against single-agent with equal compute, search, and tools via ablation [D]
Has World Models already solved physical intelligence? Insufficient evidence Visual coherence, decision effectiveness, physical accuracy, and real transfer are different metrics
Can Alphabet's full-stack advantage guarantee victory? No Interfaces and scale provide advantages but also bring capital, organizational, and control complexity [D]

The most important quantitative distinction comes from AlphaEvolve. The researchers disclosed that its matrix-operation-related heuristics sped up specific kernels by about 23% on average, corresponding to roughly a 1% reduction in total Gemini training time. There is a large conversion gap between local optimization and overall growth speed. This is real evidence of AI participating in R&D with actual value — but it is not enough, on its own, to prove a takeoff in R&D speed.5

As of the research cutoff, Google has announced Gemini 4 Argon with phased availability. The latest organizational arrangements have also changed: Hassabis has moved to Chair of Google DeepMind and Chief Scientist of Alphabet; Koray Kavukcuoglu takes day-to-day leadership as SVP of Google DeepMind while also serving as Google's Chief AI Architect. Continuing to describe Hassabis as the day-to-day operating head, or Argon as fully available, would distort this report's research baseline.12

The most decision-relevant object of observation is whether the system can, at the same total cost, keep increasing "independently validated and adopted effective results" while reducing human remediation and supervision burden. Model rankings, agent counts, candidate counts, capital expenditure, and demo videos cannot answer this question on their own.


I. Research Design: Separating "Capability Enhancement" from "Accelerating Intelligence Growth"

1.1 Six Questions That Actually Need Answers

  1. What common mechanisms do DeepMind's various achievements rely on, and under what conditions do they fail to transfer to each other?
  2. How do AI-proposed solutions move from seeming reasonable to reliable knowledge or actual engineering improvement?
  3. Has AI participation in R&D already changed the speed at which the next round of AI capability grows?
  4. Among base models, tools, memory, validation, and environments — which are sources of capability, and which are merely delivery conditions?
  5. Can virtual experience improve reliable action capability in the real world?
  6. When do capital, compute, experiments, and control capacity become decisive bottlenecks?

These questions come before the report's structure. After collecting the evidence, this report organizes its chapters into validation mechanisms, R&D feedback, systems explanations, physical transfer, industrial conditions, and governance — not a product-by-product tour.

1.2 Evidence Grading

Grade This report's usage What it cannot automatically imply
A: Independent research or authoritative external data METR methods and limitations, IEA energy research, independent scholars' analysis of materials research, etc. Independent sources do not mean conclusions are uncontested; preprints still need reproduction
B: Official first-hand sources and research-team papers Google, DeepMind, Waymo, and Isomorphic disclosures; peer-reviewed papers with company participation; Alphabet regulatory filings Peer review does not substitute independent repeated experiments; deployment disclosures are not independent audits
C: Authoritative media Used for news cross-checking and as search leads Reports whose full text was not obtained are not used to reconstruct internal technical roadmaps
D: This report's analysis Mechanism explanations, competing hypotheses, scenarios, company recommendations Cannot be written as the company's realized internal state

The grading describes source relationships, not a mechanical reliability ranking. Formal announcements should be preferred for organizational appointments; same-condition independent tests should be preferred for model performance. Regulatory filings are strong on financial definitions but cannot provide internal R&D attribution for DeepMind.

This report's main limitation: Google's internal R&D logs, the full set of failed experiments, actual labor investment, model training recipes, and complete safety-incident data are not public. Where these variables are involved, the unknown is explicitly preserved.

1.3 Working Definitions

These are this report's analytical definitions [D]; they do not declare AGI or ASI on behalf of any company.


II. Current Baseline: The Research Subject Is No Longer a Closed Lab

2.1 Organizational Changes and Their Research Implications

The organizational adjustment disclosed in August 2026 further separated strategic scientific responsibilities from day-to-day model and product operations. Hassabis retains research-advisory and strategic roles and continues to lead Isomorphic Labs; Kavukcuoglu is responsible for Gemini models, frontier AI research, and the Gemini app and developer teams, reporting to Pichai. The announcement also disclosed that Jeff Dean and Sanjay Ghemawat will found an independent public-benefit company, with Google continuing to collaborate as investor and Cloud partner.1

This report's judgment [D]: this is a change in how research and products coordinate — it cannot be used, on job titles alone, to infer that "scientific research has been abandoned," nor that "an internal secret breakthrough toward ASI has been discovered." Judging the outcome of the reorganization requires future data: whether basic research programs continue, whether papers and tools stay open, whether long-horizon scientific research gets resources, and whether the shared evaluation standards of frontier-model and application teams improve.

2.2 Correct Attribution of Achievements

Entity Interfaces examined in this report Attribution discipline
Google DeepMind Gemini, Alpha series, SIMA, Genie, Robotics, frontier safety Core subject
Google Research ERA and some joint scientific research Cannot all be credited as DeepMind's independent achievements
Google Cloud and infrastructure teams TPU, compilers, deployment, enterprise and science tools Full-stack synergy needs specific interface evidence
Waymo Autonomous-driving domain adaptation of world models Waymo's road performance cannot be directly attributed to Genie
Isomorphic Labs AI drug design and experimental translation A drug R&D institution, not equivalent to the AlphaFold product
External researchers and labs Biological experiments, math evaluations, independent criticism and reproduction Collaborative verification and fully independent verification should be separated

This boundary has practical meaning. If research candidates are proposed by DeepMind, experiments completed by universities, and production deployment handled by the Cloud team, then the capability comes from a cross-organization chain. Compressing the whole chain into "one model did it autonomously" overstates model autonomy and understates the collaborating institutions' contributions.

2.3 Gemini 4 Argon: A Current Observation Point, Not the Ranking Center of This Report

On September 30, 2026, Google introduced Gemini 4 Argon, disclosing internal software-engineering and security-task applications and raising the output limit to 1 million tokens. This is an output limit, not a context window. The official rollout was still hardening safeguards, starting with trusted testers and cyber-defense institutions before wider availability.2

What should be observed is not the "number one" label, but:

As of the cutoff date, public materials cannot confirm Argon's full availability, independent task horizons, or its contribution to frontier AI R&D cycles. Official engineering cases can support "expanded task scope" but not "the entire research team has been replaced."


III. Mechanisms Across Different Achievements: Connecting Proposal Capability to Usable Feedback

3.1 From Game Research to Research Search: What Transfers Is the Method Structure

AlphaZero improved strategy through self-play and reinforcement learning in games with known rules; MuZero learned environment representations useful for decisions and used them to plan actions. They demonstrated the combination of learning, search, and feedback.3

But games provide conditions that real research usually lacks: clear rules, decidable endings, low retry costs, and rapidly generable experience. AlphaZero also needed separate training per game; broad algorithm applicability does not mean one trained policy can directly handle every task.4

This report's judgment [D]: DeepMind's historical continuity is better explained as "searching for task structures that can be optimized in a closed loop" than as "a smooth technical line from Go to all intelligence."

An effective closed loop needs at least four components:

  1. A generator that can propose valuable candidates;
  2. A validation mechanism that distinguishes good from bad and covers important failures;
  3. A search-and-selection strategy that allocates trial resources sensibly;
  4. A state or learning mechanism that retains results, constraints, and failures.

They can appear in different forms across projects. A shared structure being valid does not mean weights, memory, or skills are already fully shared.

3.2 Five Levels of Validation

The following is this report's analytical framework [D].

Level Validation object Typical methods Easily misread results
V1: Form and execution Programs run; proofs pass specific rule checks Compilation, tests, formal verification Running gets written as having solved the real business
V2: Task performance Better under fixed data, goals, and resources Hidden tests, latency, cost, accuracy Optimizing one score gets written as comprehensive progress
V3: Scientific fact Predictions or hypotheses about natural phenomena hold Experiments, measurement, statistics, independent reproduction Model agreement gets written as experimental validation
V4: System outcome Improvements work in the complete pipeline Production deployment, end-to-end R&D cycles Local speedups get written as equal overall gains
V5: Social and physical outcome Reliable, controllable, and value-producing in real use Long-term field data, accidents and recovery, clinical evidence Demo success gets written as scaled applicability

Higher levels usually add cost, time, and uncertainty. Lower-level success can provide necessary evidence but cannot skip higher levels.

This framework does not require every problem to use the same validator. Rigorous mathematical proof, statistical evidence from biological experiments, and physical tests of robots are epistemologically different. The key is to state which errors validation can rule out, and which errors remain.

3.3 A Unified Analytical Diagram, Not a Deployed Architecture

flowchart TD
    M["Model proposes candidates"] --> S["Search and resource allocation"]
    S --> V["Validation independent of candidates"]
    V --> K["Retain results and failure records"]
    K --> M
    V --> E["Real deployment or experiment"]
    E --> K
    E --> R["Has R&D capability improved"]
    R -. "Cross-generation feedback to be proven" .-> M

Solid lines show the process this report uses to analyze each project; the dashed line marks the key unproven link of cross-generation intelligence-growth feedback. It does not claim Google has integrated all modules into one autonomous system.


IV. Scientific Intelligence: Expanding the Search Space, Still Crossing the Threshold of Knowledge

4.1 AlphaFold's Significance Is Predictive Infrastructure — It Cannot Be Expanded into "Biology Solved"

AlphaFold 2 improved protein-structure prediction accuracy through new network architectures and training methods; AlphaFold 3 extended to joint structure prediction of proteins, nucleic acids, small molecules, and other interacting complexes. The latter is the work of DeepMind together with researchers including Isomorphic.1213

EMBL-EBI's external database documentation notes that AlphaFold led at CASP14 and made predicted structures available to researchers. The accessibility of predictions lets the model serve as input to broad research pipelines.14

This report's judgment [D]: its strategic value is not just replacing one structure prediction, but reducing downstream research's initial uncertainty and changing candidate screening and experiment design. Structure prediction alone cannot guarantee function, affinity, in-vivo efficacy, toxicity, or manufacturing conditions. The count of predictions in the database cannot be counted as an equal number of experimental discoveries.

A stronger scientific system needs to acknowledge:

These are this report's analysis of the scientific translation chain, not claims that the model cannot handle the above factors at all.

4.2 AlphaGenome Atlas: The Distance Between Computational Coverage and Experimental Coverage

Released on September 8, 2026, AlphaGenome Atlas precomputes molecular-effect predictions for about 9 billion human single-base variants, combining AlphaGenome and AlphaMissense into variant-impact scores. The official disclosure included directed experimental instances from partners, plus a website, API, and Antigravity workflow interfaces. The official statement also made clear it has not been validated or approved for clinical use.15

Such resources have two different kinds of value [D]:

It is not 9 billion real experiments, nor a completed map of all genetic causality. Research should track: how many recommended candidates hold in preregistered experiments; and whether the total cost of confirming one valid discovery falls compared to pipelines that don't use the resource.

The combination of Atlas, base models, and agent interfaces shows that specialized scientific resources can enter general workflows. It has not proven that an agent can reliably complete disease-mechanism research without researchers.

4.3 Co-Scientist: Multi-Agent Improves Hypothesis Quality — It Does Not Automatically Produce Truth

The Co-Scientist study published in Nature in May 2026 used Gemini to build a multi-agent system that generates, criticizes, and evolves scientific hypotheses, expanding search with inference-time compute. The paper reported collaborative validation in leukemia drug repurposing, liver fibrosis, and bacterial gene-transfer mechanisms. It also acknowledged limitations in literature access, missing negative results, hallucinations, and preliminary validation; connecting to experimental automation remains a future direction.8

The key distinction here: tournament ranking optimizes relative judgments of candidates; experiments test natural phenomena. Elo-style ranking can organize proposals, but it carries no natural guarantee of biological truth.

This report therefore proposes three checkpoints [D]:

  1. Whether the comparison includes single-agent repeated search at the same budget, not just one-shot generation;
  2. Whether evaluators saw real experimental results or only judged textual plausibility;
  3. Whether all failed hypotheses entered the denominator, or only final success cases were disclosed.

Multi-agent may help cover different lines of thought, execute asynchronously, and divide responsibilities. But if multiple agents rely on the same model, the same literature, and similar rewards, they can still produce highly correlated errors.

4.4 Latest Research-Automation Studies: Longer Loops, but Humans Cannot Be Omitted

An August 2026 Co-Scientist preprint extended the research scope to experiments and papers, including reasoning-system designs like Agent_H. Materials and biology research still retained human experiment execution or protocol adjustments; Agent_H changed the reasoning system, not the base-model weights. In comparisons of 50 papers per condition, the autonomous group with a reliability module still showed 4% severe result hallucination and 24% severe mismatch between implementation and method description; cross-lab reproduction and other issues remain unsolved.9

This evidence supports both progress and limits:

Agent_H needed roughly 40–80 model calls per query, while the base-model control used a single call; blind physician ratings across nine dimensions showed significant improvement on only one. So automatic-scoring advantages cannot be written as same-cost advantages or broad clinical-quality advantages.9

This report's judgment [D]: the core difficulty of near-term research automation may shift from "can it generate a plan" to "can it keep the entire evidence chain valid." For example, train-test leakage, wrong metric choice, and post-hoc selection can all exist inside a program that runs and has complete logs.

4.5 ERA: Converting Some Research Problems into Scorable Software Search

Google Research's Empirical Research Assistance takes tasks, data, and evaluation methods as input, using Gemini to generate and modify code, combined with tree search to select candidates for further exploration. The official research covers six classes of scientific computing tasks. The original introduction was published in September 2025; the April 2026 update mainly added the system name and should not be treated as a brand-new round of experiments.7

ERA's significance [D] is in reducing the cycle cost of "idea → implementation → measurement → rewrite." It also exposes a boundary: if the scientific goal is wrongly encoded as a scoring function, more effective search may amplify the gains of the wrong goal.

To judge whether it surpasses a research assistant, the system needs to face:

4.6 Mathematical Research: The Boundary Between Reasoning and Validation Needs Especially Accurate Statements

Aletheia combines Gemini Deep Think, tools, and a generate–verify–revise process to handle mathematical research in natural-language form. The authors explicitly warn in the research paper that successful examples are rare and should not be understood as the system stably solving research mathematics. The official side also did not claim the results of the time reached the major-progress or milestone-breakthrough tiers in its classification.1011

The revised FirstProof study reports: by majority expert evaluation, 6 of 10 problems were solved, with one problem having dissenting opinion. The generation process needed no human modification mid-way, but researchers participated in designating the preferred answers; the paper's explanation of "correct" allows minor revisions consistent with peer-review conventions — not all raw outputs were directly publishable.17

This is more informative than "AI autonomously solved several hard problems":

Natural-language verification agents cannot automatically be equated with formal provers. In the future, judging domain superintelligence in mathematical research should record the complete chain of new-problem selection, proof repair, error identification, publication, and independent checking — not just problem counts.

4.7 The GNoME Controversy: Don't Merge Computational Stability with Usable Materials

The original GNoME study combined candidate generation, graph-network prediction, and density functional theory computation, using computational results to update the model. The paper reported about 2.2 million crystal structures stable relative to the previous database, of which about 381,000 sit on the updated convex hull. This is mainly computational discovery, not an equal number of experimental syntheses.18

Independent materials chemists Cheetham and Seshadri acknowledged the method's potential but questioned the novelty, realizability, and practicality of some candidates. They studied a sample, not an exhaustive check of the entire database.19

This report neither averages the two sides nor writes the controversy as "all results overturned":

The adopted judgment [D]: GNoME supports the feasibility of scientific search loops; total materials-discovery counts need to be stratified into computation, synthesis, reproduction, function, and use.

4.8 Drug Design: Digital Acceleration Cannot Skip the Physical and Clinical Chains

Isomorphic's Drug Design Engine pipeline disclosed in September 2026 goes from design requirements to computational candidates, then scientist selection, synthesis, and experimental validation. Its disclosure emphasizes day-scale computational design, plus subsequent experimental and clinical-development preparation.16

This report's judgment [D]: the digital design stage may improve rapidly, while complete drug development remains constrained by experiments, human biology, and regulatory validation. Generating candidates in days cannot be converted into producing drugs in days. The company's comparison against traditional processes taking years also cannot be interpreted as an equal-multiple speedup of the full pipeline without a matched baseline.

This is a clear instance of the time gap between digital intelligence and physical results: the former raises candidate-generation and screening speed; the latter requires the real world to supply new evidence.


V. AI for AI: What Thresholds Must Be Crossed from Engineering Assistance to Changes in R&D Speed

5.1 Four Different Feedback Loops

Loop Improvement object Public evidence Chain still to be proven
L1: Operations optimization Kernels, compilers, scheduling, chips, and software AlphaEvolve and deployment disclosures [B] Whether overall R&D output can improve by the same factor
L2: Systems optimization Search, agent processes, inference configuration ERA, Aletheia, Agent_H [B] Whether new configurations transfer stably to new tasks and models
L3: Learning-experience generation Strategy training data, tasks, feedback SIMA 2 self-improvement disclosures [B] External environment realism, sustained cross-round gains
L4: Frontier model R&D Data recipes, algorithms, training, and next-generation capability Locally relevant tools already exist Insufficient public evidence to reconstruct a complete autonomous loop

Calling all four loops "AI self-evolution" masks key differences. L1 can save compute without inventing new learning algorithms; L2 can improve one model's usage without changing weights; L3 improves strategy but may rely on stronger teachers and fixed training frameworks; only L4 directly concerns next-generation frontier capability growth.

5.2 AlphaEvolve: One of the Closest Pieces of Evidence to Engineering Reality

AlphaEvolve puts language-model-generated programs into an executable-evaluation and evolutionary-search pipeline. The 2025 paper disclosed that specific matrix-operation-related optimizations brought about a 23% average kernel speedup, worth roughly a 1% time saving across all of Gemini training. A 2026 official update further disclosed its entry into TPU circuit design and other compute-system uses.56

This provides two valuable layers of fact:

But this report will not therefore write:

Humans still define problems, constraints, and evaluations; the complete gain depends on how much of the work the improvements cover and where the new bottlenecks are.

5.3 How Local Efficiency Converts to Overall Efficiency

A simplified analytical formula [D]:

Overall speedup = 1 / [(1 − f) + f/r + h]

where f is the share of the original process's time that AI can accelerate, r is the speedup factor on those stages, and h is the share of the original process time taken by new review, coordination, and rework. It assumes stages add up in time — good for illustrating bottlenecks, not Google's actual R&D model.

Example: if 40% of a cycle can be sped up 4×, the rest unchanged, and new review takes 5%, the overall speedup is only about 1.33×. This number is an illustrative calculation, not a prediction.

Real R&D also has stage parallelism, queuing, budget constraints, and failed retries. So it is necessary to distinguish:

5.4 Five Levels of Tests from R&D Participation to Takeoff

Level Evidence that must appear This report's judgment on DeepMind
R0: Assistance Provides code, analysis, literature, or suggestions Clearly present
R1: Executing closed loops Automatic experimentation, evaluation, and modification within preset tasks Clearly present in several instances
R2: Effective improvement Results pass checks external to the generation process and enter real pipelines Company first-hand evidence exists; independent review coverage is limited
R3: Cross-generation feedback Improvements enter the next-generation R&D system, which produces stronger improvements Disclosed for local agent training; insufficient for the frontier general R&D chain
R4: Sustained takeoff R&D capability growth significantly accelerates after controlling resources, labor, and tasks Insufficient public evidence yet

"External to the generation process" can be an independent tester, not necessarily an external institution. Independent institutional re-verification is additional evidence; the two should not be conflated.

The safest answer at present: AI R&D participation has already changed how some engineering work is produced; whether it has continuously increased the speed of intelligence growth remains insufficiently evidenced in public materials.

5.5 What Data Must Be Disclosed to Raise the Credibility of a Takeoff Judgment

This report recommends [D] disclosing — or having a trusted third party audit — the following data:

  1. For comparable research tasks: the workload AI handled, human takeover volume, and failure rates;
  2. Calendar time from research-plan proposal to independent confirmation;
  3. Total candidates, total failures, and the share that entered production or the next training round;
  4. Capability gains after controlling compute, teams, and data;
  5. Source tracing for at least three consecutive rounds: which AI produced what improvement, and how it changed the next round of R&D;
  6. Whether improvements persist without increasing supervision cost;
  7. Whether AI-led important algorithms or training methods exist and are reproduced by external researchers.

Three consecutive rounds is this report's suggested minimum observation window, not a law of nature or an official threshold.

5.6 Positive Feedback Can Appear While Gains Still Diminish

A system may repeatedly optimize easily measurable kernels yet find less and less new room. It may also speed up experiments while experiment evaluation, training queues, or data cleaning become bottlenecks.

This report distinguishes two compatible explanations [D]:

Only continuous input–output data can tell them apart. One successful case cannot decide which long-term mechanism dominates.


VI. Scaling Has Split into Multiple Resources — They Cannot All Be Called "Bigger"

6.1 Scaling Variables in the DeepMind Cases

Variable What it can increase Main bottleneck This report's test
Pretraining Representations, knowledge, and basic skills Data, optimization, cost Compare capability gains after post-training, controlling for it
Post-training & RL Reasoning, action, and reward adaptation Reward coverage, distribution shift New tasks and anti-reward-hacking tests
Inference-time compute More candidates, revisions, and planning Search efficiency, validation accuracy Success rate at the same total cost
Context & memory Retaining task materials and history Retrieval, state updates, conflicts Effective retention rate and recovery ability
Program search Algorithm and implementation exploration Evaluation function, execution cost Hidden tests and deployment results
Validation Filtering errors, guiding search Missed detections, correlated errors, price Independent judgments and failure coverage
Synthetic experience Larger training and interaction sets Bias, difficulty, realism New environments and real-task transfer
Agent runtime Longer work that can be advanced State drift, failure accumulation Human-equivalent task horizons and takeovers
Parallel agents More experiments and division of labor Duplication, communication, coordination Single-agent ablation at the same budget
Tools & real resources Access to new information, executing actions Permissions, interfaces, latency Whether resource use produces final gains
Infrastructure Larger throughput and available budgets Chips, memory, network, power Actual utilization and cost per effective result

This table is an analytical framework [D]. It does not claim all variables follow power laws, nor that they can substitute for each other indefinitely.

6.2 Why Inference Scaling Is Conditional

Official research on Deep Think and Aletheia shows that adding inference-time compute improves scores on the corresponding math tasks, and agent processes can further improve compute-use efficiency.11

The judgments derivable from this evidence [D]:

"Verifier Scaling" also cannot just count reviews. If validators and generators share errors, repeated judgments may amplify false confidence. Independent program tests, real experiments, and formal checks each provide different error-correction channels in applicable tasks.

6.3 Long Output, Long Context, and Long Tasks Are Not the Same Kind of Growth

Long output can expand reasoning but may also amplify self-reinforcing errors. Long context can hold materials but does not guarantee the model correctly updates task state. Continuous operation can generate many actions but does not guarantee goal progress.

This report recommends: separate token scale from task-success records. Priority should go to whether the system, after several key state changes, can still correctly locate the goal, what is done, conclusions pending validation, and current permissions.

For the capacity, update quality, cross-task transfer, and R&D contribution of DeepMind's internal persistent-memory systems, existing public materials are insufficient for a reliable baseline. Gemini's context capability cannot fill this unknown.


VII. Single Model, System, or Organization: Four Competing Explanations

7.1 H1: The Unified Base Model Is the Main Capability Source

The strongest argument: Gemini participates in language, code, math, scientific-hypothesis, and action tasks, and the base model's cross-task capability reduces the need to redevelop specialized models problem by problem. [B→D]

If H1 dominates, we should see:

Falsifying conditions: end-to-end tasks not improving correspondingly after base-model upgrades, or validation, environments, and specialized modules still deciding the main success rates.

Remaining unknown: public projects usually change models, processes, and compute simultaneously, making model contributions hard to isolate.

7.2 H2: Models Plus External Mechanisms Form a Stronger System

AlphaEvolve, Aletheia, and Co-Scientist connect models with search, tools, or evaluation processes; Robotics also separates high-level reasoning from action models. [B→D]

If H2 dominates, we should see:

Falsifying conditions: complex systems' advantages disappearing in same-budget comparisons, or old components no longer producing gains after model upgrades.

7.3 H3: A Combination of Specialized Superhuman Systems Never Forms General Autonomous Intelligence

AlphaFold, GNoME, and robot policies may each be strong while using different data, feedback, and goals. Connecting APIs can improve workflows without forming unified understanding or autonomous research. [D]

If H3 is correct, we will see many domain scores rising while cross-domain goal selection, long-horizon planning, and unknown-environment adaptation stagnate. Important resources are still configured by human organizations; module errors are patched by human coordination.

This is not an "AI has no value" scenario. A combination of specialized systems can still have huge impact on research and industry.

7.4 H4: The Strongest Capability Comes from the Research Institution and Industrial System

Google owns model research, compute architecture, cloud deployment, application entry points, and experimental collaboration interfaces.128 This report therefore proposes an institution-level explanation [D]: the unit that actually takes on long-horizon goals may be "the organization of human researchers plus AI," not an independently replicable agent.

If H4 dominates, leading indicators should be verified research output per unit investment, cross-team processes, and actual deployments — not some single score of model weights.

7.5 Provisionally Adopted Combined Judgment

This report favors H2 and H4 jointly explaining current achievements, without ruling out H1 growing stronger in the future. Existing evidence is also compatible with H3. This is not assigning average probabilities to the four hypotheses; it says currently observable achievements mostly reflect clear environments, validation, and organizational interfaces — still insufficient to prove a fully general autonomous system.

7.6 Evidence That Multi-Agent Is Not a Necessity

One agent can repeatedly call models, switch roles, and maintain multiple candidates. Multiple agents may merely split these operations in engineering. Necessity judgments need comparisons of:

This report's proposed collaboration metric:

Net collaboration gain = effective results of multi-agent − effective results of the best same-cost alternative

while also recording communication, duplicated search, shared errors, and human coordination. More agents with unchanged net gains should not be written as enhanced collective intelligence.

To surpass large human organizations, a system also needs to choose goals, coordinate resources, handle conflicts, remember commitments, and accept accountable control. Existing materials are insufficient to confirm DeepMind has this complete capability.


VIII. World Models and Physical Intelligence: Can Virtual Experience Cross the Real-World Interface

8.1 Genie 3: Interactive World Generation Does Not Equal Calibrated Physical Simulation

Genie 3 demonstrated prompt-generated interactive visual environments, with the official report citing 720p, 24 frames per second, and minute-level continuity, while disclosing limitations in action spaces, multi-agent interaction, text, and environment accuracy.20

Its value [D] may lie in reducing environment-creation and experience-collection costs. But visual coherence alone cannot answer:

A model suited for video experiences can be a starting point for training environments; whether it suffices to support robots needs extra validation.

8.2 SIMA 2: Experience Generation and Next-Round Policy Training Have Appeared

SIMA 2 acts in virtual 3D environments through images, keyboard, and mouse. The official description covers using Gemini to generate tasks and feedback, autonomously accumulating experience, and then training an improved agent — also showing iteration inside Genie environments. The official statement also acknowledges limitations in long tasks, memory, fine manipulation, and real-time interaction.21

This report's adopted judgment:

If the next-round agent improves task capability, that is policy progress; only if it also improves experience selection, evaluation, and training methods — and keeps bringing larger next-round gains — does it come closer to this report's recursive-R&D mechanism of interest.

8.3 Waymo: Domain Adaptation Is Transfer Evidence and Boundary Evidence

Waymo's February 2026 introduction of a Genie-3-based, driving-domain-adapted world model generates camera and LiDAR data and provides scenarios and driving controls for simulating rare cases.22

This shows a base world model can enter specialized industrial pipelines. It does not prove a general world model can accurately simulate all real processes without post-training.

Waymo's real road experience belongs to its autonomous-driving system; road miles cannot be counted as the Genie world model's real-world validation miles. Evaluating generative simulation should separately disclose: sensor statistical consistency, action response, accident reproduction, closed-loop policy results, and additional gains on real roads.

8.4 Gemini Robotics 2: Observing "Unevenness" Is More Useful Than Observing the Best Clips

Gemini Robotics 2, released in July 2026, includes action, high-level reasoning, and on-device routes. The official disclosure tested the same checkpoint across multiple hardware setups while acknowledging challenges in multi-finger manipulation; in example tasks, screwing in a light bulb succeeded 36% of the time versus 92% for screwing one out. These are company results under specific test conditions, not general household success rates.23

The model cards further distinguish:

So the capabilities of different variants cannot simply be added: whole-body action demos do not automatically extend the on-device model's validated scope, and open reasoning APIs do not equal universally open complete robot control.

8.5 Why Digital and Physical May Show a Clear Time Gap

This report's analysis [D] gives at least five reasons:

  1. Digital experiments can be replicated and parallelized; robots need equipment, space, and maintenance.
  2. Code failures can often be rolled back; physical failures can damage equipment or cause injury.
  3. Digital task states are easier to record; physical systems have occlusion, wear, contact, and sensing errors.
  4. Virtual experience can be generated in bulk; its realism must be calibrated.
  5. Software deploys quickly; hardware manufacturing, installation, and certification have their own cycles.

But no fixed multi-year gap can be derived from this. If sensors, hardware standardization, world models, and robot data efficiency improve together, the gap may narrow; if dexterous manipulation and safe recovery stall, the gap may persist long-term.

8.6 Minimum Experimental Design for Verifying Simulation Transfer

This report recommends [D] using four control groups:

Group Training condition Question to answer
P0 Real data only Existing real-task baseline
P1 Real data + traditional simulation Gains from traditional simulation
P2 Same real data + generated worlds Whether generated environments provide independent gains
P3 Less real data + generated worlds Whether real-data costs truly fall

In the same batch of unseen environments, record task success rates, takeovers, collisions, action anomalies, recovery, and total cost. If P2 only improves simulation scores with no real-test improvement, no physical-intelligence breakthrough should be declared.


IX. From Software Research to Industrial Conditions: Advantages and Constraints Grow Together

9.1 The TPU Roadmap Shows Training and Inference Resources Are Diverging

Google Cloud's April 2026 disclosure of TPU 8t and 8i: the former for training, the latter for inference and inference-style workloads, with different memory and networking configurations. Official performance and price-performance comparisons depend on specific conditions and use "up to" phrasing.28

This report's judgment [D]: AI capability growth no longer only demands bigger training clusters. Long reasoning, program search, parallel experiments, and agent services need different memory access, networking, latency, and scheduling.

DeepMind-related advantages may appear at four interfaces:

Whether advantages hold should be judged by actual end-to-end throughput, stability, and cost per effective result; chip peak performance cannot directly represent research output.

9.2 Financial Constraints: Don't Treat Alphabet's Budget as DeepMind's Budget

Alphabet's Q2 2026 earnings filing discloses:

Metric Q2 2026 value Definition
Operating cash flow ~US$39.069 billion Group quarterly figure
Property and equipment purchases ~US$44.924 billion CapEx-related cash measure
Free cash flow ~−US$5.855 billion Company-defined non-GAAP measure
June equity issuance net proceeds ~US$49.6 billion For general corporate purposes, including expanding AI infrastructure
Quarter's senior unsecured notes net proceeds ~US$20.3 billion General corporate purposes

Source: company regulatory disclosures; quarterly financials are unaudited. It does not disclose DeepMind's separate compute budget, return rate, or per-project investment.29

This report's judgment [D]: research capability is now interconnected with capital markets and industrial expansion. A strong operating business does not mean AI infrastructure can be funded internally without limit. A quarter of negative free cash flow, on its own, cannot imply the company cannot invest or that the AI business has no returns.

More research-relevant: whether new capability and usage revenue keep up with depreciation, financing, and supply costs; which research directions need resource protection; which projects get priority when budgets tighten.

9.3 Power and HBM: Global Trends Cannot Be Directly Applied to Google

The IEA's 2026 study estimates global data-center electricity use may rise from 485 TWh in 2025 to about 950 TWh in its central 2030 scenario; it also notes grid-connection, equipment, and chip-supply constraints, and expects HBM tightness to continue at least through end-2027. The latter two contain forecasts; 950 TWh is not an occurred fact.30

The data covers global data centers — not AI electricity alone, and certainly not Google's or DeepMind's electricity.

This report's judgment [D]: effective compute is constrained by a chain of interdependent conditions — chip availability, memory and networking fit, facility completion, power connection, cooling availability, cluster stability, and software that can use it well. Announced planned capacity cannot substitute for actually commissioned capacity.

For nuclear, storage, or on-site generation, contracts, permits, construction, and actual power delivery need separate evaluation; this report does not treat distant energy promises as ready R&D resources.

9.4 Scale May Strengthen Advantages — and May Change Technology Choices

If search and reasoning consume more and more resources, the optimal approach may be:

These are this report's techno-economic inferences [D]. The future winning system is not necessarily the one using the largest model at every step.


X. Control: Research Acceleration and Risk Growth Come from the Same Action Interface

10.1 What the Current Frontier Safety Framework Tracks

DeepMind's Frontier Safety Framework 3.1 distinguishes lower-level capabilities of concern from critical capabilities, and covers misuse, supervision evasion, and machine-learning R&D risks. Its ML R&D automation level asks whether a Google capability research team could be automated at similar total cost; the framework also holds that R&D acceleration may come from models combined with workflows rather than weights alone.31

The framework defines a threshold; it does not mean the company has reached it.

The document's decisions involve risk assessment, mitigation, and risk-acceptance conditions — they cannot be compressed into "any capability crossing a line automatically stops training." Likewise, publishing the framework does not prove actual deployments always comply with it.

What should be checked: who evaluates, which systems are covered, what conditions trigger escalation, what evidence is disclosed, and whether actual risk handling can be externally examined.

10.2 Agent Control Is Starting to Adopt an Action-and-Infrastructure Perspective

The June 2026 control roadmap treats capable-but-not-fully-trusted agents as potential insider threats, proposing supervision, blocking, and response mechanisms, tracking monitoring coverage, recall, and response times, and acknowledging that visible reasoning may be insufficient for evasive or opaque reasoning.27

This is an explanatory direction [D]: security needs to protect permissions, networks, data, and experimental resources — not just check whether the final answer contains dangerous text.

Argon's release materials also describe reasoning-and-action monitoring, execution halts, and hardened isolated environments. These are the company's mitigation disclosures; they cannot be taken as externally verified complete control guarantees.2

10.3 Robot Safety Cannot Be Borne by Semantic Reasoning Alone

The Robotics 2 safety report studies multiple semantic and agent risk evaluations. The on-device model card separately recommends layered control: high-level semantic judgment, low-level collision and force control, and hardware-related safety mechanisms.2625

One closed-loop evaluation used Gemini to simulate VLA confidence feedback, avoiding real robots or physical simulators. This can test whether agents obey feedback constraints; it cannot independently prove dangerous real-robot actions are reliably intercepted.26

This report's judgment [D]: a robot "understanding not to harm people" and "the physical system remaining safe under wrong actions" are different capabilities. A complete system needs to handle perception errors, execution anomalies, network interruptions, and mechanical problems. Real safety evidence should include dangerous-action rates, takeovers, emergency stops, and recovery — not just the success rate of refusing dangerous instructions.

10.4 Why AI R&D Automation Carries Special Control Risks

The following is mechanism analysis [D], not an allegation that DeepMind has done anything of the sort:

So capability and control must be observed in pairs:

Capability growth Control metrics that must be paired
More tools and system permissions Least-privilege coverage and overreach rates
Longer runtimes Monitoring continuity, state records, and recoverability
Stronger code and network ability Isolated-environment testing and exploit interception
Automatic modification of R&D pipelines Acceptance and audit chains that cannot be modified by the agent itself
Stronger experiment design Experiment risk approval and execution boundaries
More concurrent agents Scheduling, permission inheritance, and collective-behavior monitoring

10.5 Governance and Benefit Attribution

The DeepMind Institute, founded in September 2026, provides a cross-disciplinary discussion platform headed by Legg, Manyika, and Hassabis. Its statement is explicit: articles are the authors' views and should not be read as Google's official positions.34

The platform can promote discussion, but it is not an independent regulator or a deployment approval body. Its founding cannot be used to prove governance has kept up with capability.

This report raises three benefit-distribution questions [D]:

  1. Which of scientific predictions, model weights, data, and validation resources should stay publicly accessible?
  2. Can external researchers examine key capability and safety judgments, or only see curated instances?
  3. How should knowledge and benefits from research collaborations be distributed among platforms, experimental institutions, and public funders?

Google's system simultaneously provides research resources, commercial cloud services, and application distribution. It can lower participation barriers, and it may also concentrate control over interfaces, pricing, and visible results. Both outcomes must be judged from actual openness conditions and usage data — not decided by "open" or "full-stack" slogans alone.


XI. Evidence Map: What Is Established, What Still Cannot Be

Research proposition Direct evidence This report's status Main gaps
General models can participate in different research tasks Gemini-related science and math projects [B] Established for tested tasks Cross-domain reliability and full autonomy
AI can search and measure program improvements AlphaEvolve, ERA [B] Strong first-hand experimental and applied evidence Independent deployment review and overall gains
AI can propose experimentally testable new hypotheses Co-Scientist collaborative research [B] Instances exist Full attempt denominators and independent reproduction
AI can continuously complete research independently Latest research-automation preprints [B] Not sufficiently established Human execution, methodological errors, external reproduction
AI has improved agent systems Agent_H and related research [B] Local instances exist New tasks, generations, and same-budget validation
Agents can improve with generated experience SIMA 2 [B] Local official evidence Teacher dependence and real transfer
World models can enter industrial simulation Waymo domain adaptation [B] Established for disclosed interfaces External calibration and additional real-vehicle gains
Robotics has broad physical reliability Official tasks and model cards [B] Cannot be confirmed yet Long-term real distributions, accidents, maintenance, takeovers
AI R&D speed keeps accelerating significantly Insufficient public attribution data Unknown R&D cycles under resource control, continuous
One system has integrated all scientific and physical capabilities Multiple modules and local interfaces Insufficient evidence Unified goals, memory, transfer, and long tasks
Infrastructure will constrain IEA [A], TPU and financials [B] Evidence at the global level Each Google project's specific exposure
Control keeps up with capability stably Framework and mitigation disclosures [B] Not publicly proven Independent high-intensity attacks, coverage, and handling data

This report does not fill "unknown" with estimates. Missing public information neither proves nothing was achieved internally, nor allows the report to assume it was.


XII. Active Falsification: Where This Report Is Most Likely Wrong

12.1 "Validation Is the Key" May Be an Illusion of Current Engineering Practice

Stronger base models may significantly reduce the need for search and external patching. If a future model directly produces highly reliable results on complex new tasks and system-component contributions shrink, this report would raise H1's weight.

But "needing less validation" still needs to be proven by independent results. More confident model expression is not a falsification.

12.2 Research Feedback May Not Effectively Improve General Intelligence

Scientific models can handle specific representations accurately without improving general planning or social tasks. Scientific successes cannot be accumulated into an AGI score.

A late-September 2026 external preprint, Fold2Reason, offers a reverse clue: protein-structure-derived supervision may improve some broad reasoning evaluations, with the authors reporting an average gain of about 3.23 percentage points. But it is an early preprint; it does not prove the same transfer for frontier models or long-task research capability.33

This report keeps open the possibility that "scientific data helps general intelligence" while refusing to derive ASI directly from small evaluation transfers.

12.3 A Larger Candidate Pool May Just Be a Larger Selection Advantage

If a system generates thousands of candidates and experts pick from them, good final performance may come from search budget and human curation rather than stronger autonomous research.

Ways to falsify: limit total budgets, record all attempts, and compare against strong human teams and same-budget single agents on hidden tasks. Disclosing only successful solutions cannot answer this.

12.4 Closed Evaluation May Be Overfitted

Scorable tasks suit search — and are easily exploited by it. Test-set leakage, special cases, and unrealistic baselines can all manufacture capability illusions.

Validators should be split into feedback visible during optimization and final invisible acceptance. If the latter does not improve, this report will downgrade the relevant capability judgments.

12.5 Physical Transfer May Be Faster Than Expected — or Stall Long-Term

If generated worlds significantly reduce real-data needs and cross-hardware transfer proves reliable, the Digital–Physical gap may narrow. Conversely, if dexterous manipulation, unexpected recovery, and long-term maintenance don't improve, quality demos may still sit far from actual use.

12.6 Full-Stack Advantages May Be Weakened by Independent Ecosystems

External models, open tools, and specialized experimental institutions can also combine into effective systems. They don't necessarily need all resources under one company.

This report does not treat "owning the full stack" as a sufficient condition. What should be observed: whether synergy truly lowers cost and failure, or whether organizational coordination and closed interfaces offset the advantages.

12.7 Research Automation May Increase Low-Quality Research Rather Than Effective Discovery

As the marginal cost of generating hypotheses and papers falls, review burden and duplicated directions may rise. If the share of effective results falls and experimental resources get crowded out by weaker proposals, research speed may not improve.

So "paper counts," "generated hypothesis counts," and "agent work hours" can only be input metrics — not final intelligence-growth metrics.

12.8 Control May Become a Necessary Input to Growth, Not an Added Cost

Stronger agents may need stricter isolation, slower approvals, or more monitoring. Ignoring these inputs overstates economic gains.

If control technology improves so supervision burden per effective task falls, capability is more likely to form deployable growth. If supervision work grows faster than capability, takeoff may face real constraints.


XIII. Scenarios for the Next 3–10 Years: Update by Conditions, No Precise ASI Date

Observation window: Q4 2026 to around 2036. The following are this report's scenarios [D], not company commitments. Scenarios may appear in different domains simultaneously; no unsupported precise probabilities are assigned.

Scenario A: Capability Keeps Improving, Systems Gradually Expand, No Unified AGI Day Long-Term

Dimension Analysis
Premise Base models, reasoning, and agents keep progressing; module fusion still needs heavy engineering
Leading indicators Hidden-task success rates rise; takeovers and rework fall; science tools enter more pipelines
Possible outcomes Digital work and some research significantly automated; physical tasks develop more slowly
Main risks Reading multiple local capabilities as general autonomy; underestimating supervision costs
Evidence raising this scenario's weight Broad but continuous progress; key breakthroughs still depend on task definition and human validation
Evidence lowering its weight Reliable research-team replacement without major input growth, or long-term capability-trend stagnation

This is the baseline observation path most easily extended from existing public evidence; it does not mean other scenarios are impossible.

Scenario B: Significant Acceleration in AI R&D

Dimension Analysis
Premise AI not only executes research but improves experiment selection, training, and R&D tools; gains enter the next generation
Leading indicators Successive R&D cycles shorten; capability gains at the same resources rise; AI contributions traceable
Possible outcomes Faster iteration of models, agents, and optimization systems; greater pressure on supervision and infrastructure
Main risks Wrong evaluations amplified quickly; capability and safeguards updated out of sync
Evidence raising weight Auditable positive feedback across at least several consecutive rounds; important algorithms independently reproduced
Evidence lowering weight Acceleration mainly from compute or team expansion; local results not changing overall cycles

One kernel optimization or one agent-training improvement cannot be used to declare this scenario already arrived.

Scenario C: Industrial and Control Constraints Become the Main Bottleneck

Dimension Analysis
Premise HBM, power, facilities, capital, or safety verification limit expansion speed
Leading indicators Planned-vs-commissioned gaps widen; resource queuing; unit effective-task costs hard to lower
Possible outcomes More emphasis on small models, selective search, and validation-cost savings; divergent deployment rhythms
Main risks Resource scrambles trigger rushed deployments; research funding crowded out by short-term applications
Evidence raising weight Multiple quarters of concrete attribution of actual resource constraints and capability-R&D delays
Evidence lowering weight Effective compute grows steadily; resource efficiency improvements outpace new demand

Constraints may slow growth — or push new technology choices; they cannot be mechanically written as "power prevents ASI."

Scenario D: A New Paradigm Lowers Existing Bottlenecks

Dimension Analysis
Premise Significant new methods in learning, memory, world representation, or validation
Leading indicators Large improvements in new domains and long tasks at the same budgets; independent reproduction and clear ablations
Possible outcomes Large model scale no longer the main explanatory variable; system structures reorganize
Major risks Packaging special-task gains as a general paradigm; short-term successes hard to scale
Evidence raising weight Sustained gains across multiple tasks, scales, and institutions
Evidence lowering weight Advantages depend on special data, excessive budgets, or human curation

DeepMind's history contains multiple coexisting research approaches, providing room for exploration; undisclosed routes should not be invented from this.

Scenario E: Specialized Scientific Superintelligence Before General Social Autonomous Intelligence

Dimension Analysis
Premise Program, math, and some experimental sciences have validation structures better suited to automated search
Leading indicators Research-side effective discovery rates and confirmation speed rise, while ordinary open tasks still frequently fail
Possible outcomes Some research systems surpass large human teams, but everyday complex agents remain unreliable
Main risks Expanded dual-use scientific capability; over-concentration of knowledge and resource control
Evidence raising weight More externally reproduced important research contributions with limited cross-domain transfer
Evidence lowering weight Research progress no longer depends on special setups; general goal management matures simultaneously

"Specialized scientific superintelligence" here is a domain concept, not the broad ASI defined in this report.

Key Turning Points Between Scenarios

What deserves the most attention is not a particular year, but three changes:

  1. From "AI improves human-set plans" to "AI identifies and restructures research goals and evaluation methods";
  2. From "local results usable" to "additional gains in cross-generation capability growth attributable";
  3. From "digital-environment success" to "long-term reliable real-world action that can be supervised and halted."

Each change needs new evidence; product naming does not decide it.


XIV. Quarterly Leading-Indicator Dashboard

Principles: same task distribution, same success definition, complete denominators, recorded investment, preserved uncertainty. Official and independent results in separate columns; undisclosed items marked "unknown," not zero.

14.1 Twenty-Four Monitoring Indicators

# Indicator Suggested definition and records Cutoff-date baseline status Quarterly update focus
01 Cross-domain reliable capability Success rates by domain on hidden new tasks; avoid reporting only total averages Multi-domain official evidence; unified independent baselines insufficient Whether the weakest domains improve
02 50% task horizon Human-expert completion time corresponding to the 50% success threshold, with intervals METR methods usable; no untested values for Argon Whether task sets and evaluation configurations changed
03 High-reliability task horizon Prefer 80%; higher reliability needs enough data Latest systems' public values insufficient Whether failures concentrate in a few types
04 Human takeover burden Takeovers per completed task, in counts and minutes Unknown Whether it falls, not just whether runs get longer
05 End-to-end economic success rate Tasks completed and passing real acceptance / all tasks Unknown Including tool and labor costs
06 Effective memory State updates, forgetting, conflicts, and recovery tests No unified public baseline Correct retention rate after task changes
07 Multi-agent net gain Extra effective results versus the best same-cost single-agent control No public unified ablations Duplication and communication costs
08 Controllable concurrency scale Concurrency under fixed failure, cost, and risk caps Unknown Net output and permission management
09 AI share of R&D work Accepted contributions by coding, planning, training, evaluation Unknown Don't use lines of code as a substitute
10 Complete R&D cycle Calendar time from problem proposal to independent confirmation and deployment Unknown Matching task types and resource control
11 R&D capability gain per unit Capability gains per unit of compute, capital, and research labor Unknown Same-definition trends
12 AI improvement adoption rate Deployed improvements / all candidates, stratified by importance Instances exist; no complete denominator Failures and rollbacks disclosed too
13 Cross-generation feedback depth Consecutive improvement rounds with source records Disclosed for local training; insufficient for general R&D Whether the next round improves improvement capability
14 New-knowledge confirmation rate Externally reproduced or strictly tested research conclusions / all proposals Collaborative validation instances exist More fully independent reproductions
15 Discovery confirmation time Proposal to validation, including queues, failures, and human work Unknown Don't count only generation time
16 Research methodology errors Severe methodological flaws, hallucinations, and retraction rates Preprint sample results exist New data and external review
17 Experiment autonomy level Human participation listed separately for design, execution, analysis, modification Human links still present Whether less human involvement harms reliability
18 World-model calibration Predicted-vs-measured state errors, rare-event coverage No cross-task independent baselines Don't substitute visual scores
19 Simulation-to-reality gains Additional on-site success at fixed real data No unified public controls Four-group training comparisons
20 Physical task reliability Real new-environment success, takeovers, damage, and recovery Official task performance uneven Long tasks and continuous field data
21 Effective compute Actual throughput, utilization, failures, and available hours Hardware disclosures many; research allocation unknown Separate plans, commissioning, and actual use
22 Energy and capital costs Complete costs per effective task and R&D round Group and global data; projects unknown Grid connection, depreciation, financing, and demand
23 Control effectiveness Monitoring coverage, attack recall, blocking, response, and escapes Framework exists; complete real tests insufficient Independent adversarial and concurrency pressure
24 External verifiability Data, versions, result denominators, and independent test access Varies widely by project Whether it expands rather than shrinks

14.2 How to Use METR and Avoid Misreading It

The METR task-horizon page verified for this report shows its last update as May 8, 2026, using TH1.1. The horizon refers to the human expert's completion time for the corresponding task, not how long the AI ran continuously; tasks are mainly self-contained software, machine-learning, and security work. The page also notes the current task set is unreliable for measurements above 16 hours.32

METR researchers further note that the 50%-success horizon cannot directly serve as a reliable delegation boundary; real task complexity, context, and supervision costs change actual automation gains. That note is a researcher methodology note, reviewed below the level of a formal research report.35

So this report does not invent METR scores for Gemini 4 Argon, nor extrapolate a certain ASI date from historical curves. Quarterly updates must first confirm the new task set, models, agent configurations, and statistical intervals.

14.3 Minimum Record Table for Each Quarterly Update

Indicator number:
Observation cutoff:
Model/system version:
Task distribution and hidden tests:
Total attempts:
Success definition:
Compute, tool, experiment, and labor investment:
Official results:
Independent results:
Statistical intervals and failure types:
Comparable to last quarter:
Reasons for raising/lowering related scenarios:
What remains unknown:

Indicator changes should enter scenario judgments, not be combined into a pseudo-precise "ASI index." For example:

These are this report's update rules [D], revisable as measurement methods improve.


XV. What It Means for Companies and Individuals

15.1 Companies: Start with Verifiable Work

This report recommends [D] prioritizing tasks where:

Software testing, data processing, and some computational experiments usually build faster feedback loops; open research, field operations, and high-stakes decisions need longer validation chains. This is task-structure analysis, not a procurement recommendation for any product.

Pilots should keep strong human baselines, same-budget AI alternatives, and human-review hours. Counting only generation speed easily mistakes efficiency gains after shifting review burden elsewhere.

15.2 Research and Industry Institutions: Validation Resources May Grow Scarcer

If candidate generation expands rapidly, the relative value of experimental equipment, reliable data, metrology, external reproduction, and expert judgment of domain specialists may rise. [D]

Companies' advantages may come from:

These resources can combine with different models. Strategy should not be built entirely on the assumption that one model generation leads forever.

15.3 Managers: Work Organization Will Change, but Responsibility Cannot Be Handed to Models Alone

Tasks should specify goals, resources, acceptance, and permissions; important decisions should keep an accountable owner. Cross-team agent deployments also need clarity on who handles conflicts, error propagation, and failure recovery.

More agents do not automatically reduce management costs. If humans are busy judging contradictory outputs from multiple agents, deployment may just convert execution burden into coordination burden.

15.4 Individuals: Improve Problem Definition and Result Validation

This report's judgment [D]: as code, literature organization, and preliminary plans get easier to obtain, judging whether a problem is worth doing, how to design acceptance, and how to interpret failures may matter more.

A useful practice is keeping an evidence chain in personal work: problems, sources, hypotheses, versions, tests, errors, and decisions. It helps judge whether AI actually expanded capability or merely increased output volume.

This report does not therefore declare any profession necessarily extinct. Career changes also depend on organizational adoption, responsibility, field conditions, and market demand — existing technical materials are insufficient for deterministic predictions.


XVI. Key Unknowns and Final Research Judgments

16.1 Unknowns That Should Not Be Filled with Stories

Unknown question Why it cannot be confirmed Most valuable new evidence
How much of DeepMind's internal R&D is done by AI? Complete labor hours and task denominators undisclosed Audits by role with accepted results
Are Gemini generational R&D cycles significantly shortened by AI? Cycles, resources, failures, and attribution insufficient Multi-generation R&D data with matched inputs
Is there a complete autonomous model-R&D loop? Tools and local instances cannot reconstruct the whole Source chains of training, evaluation, deployment
Do scientific systems have unified long-term memory? Open interfaces do not equal unified state Cross-project long tasks and memory tests
Is multi-agent necessary? No comprehensive same-budget ablations Single-agent, search, and collaboration comparisons
Do generated worlds truly lower physical-data costs? Demos and simulation scores insufficient On-site results at fixed real data
Can control stay effective for stronger systems? Frameworks and disclosures do not equal measured guarantees Independent adversarial, coverage, and response data
Which research routes were strengthened or weakened by the reorg? Appointments cannot substitute for budget and project data Long-term research investment and output changes
How far apart are digital and physical superintelligence? Technology, hardware, and institutions jointly decide Continuous cross-domain real reliability
When will ASI arrive? Definitions, measurement, and growth mechanisms still uncertain Leading-indicator linkages, not vision dates

16.2 Final Judgments

First, DeepMind has provided concrete evidence of "AI expanding verifiable search." Some results entered engineering and scientific collaboration pipelines; they cannot be treated as mere extensions of chat capability.

Second, an evidence gap remains between local closed loops and general recursive improvement. R&D participation, strategy self-improvement, agent-process optimization, and frontier-model R&D takeoff need to be counted separately.

Third, the systems explanation currently has more research value than single-model rankings. It can explain huge performance differences of the same model across environments, and why validation, action, and infrastructure became key. But it is not a proven unique ASI implementation path.

Fourth, research automation is most likely to deepen first in tasks with clear feedback, executability, and repeatability. Physical and biological validation is more expensive; knowledge confirmation will not automatically grow at text-generation speed.

Fifth, the future should observe capability, results, and control together. More generation capability alone cannot confirm effective research growth; more effective research alone cannot confirm deployable superintelligence formation.

This report's core analytical framework: candidate capability × search organization × validation quality, converted through real resources and control conditions into effective results; how many effective results re-raise the next round's R&D capability determines whether sustained acceleration appears. The multiplication sign expresses interdependence, not a fitted quantitative law.

The next quarterly update should most verify: whether continuous, attributable R&D-cycle data can be obtained; whether independent reproduction of research conclusions increases; whether simulated experience improves real physical reliability; and whether control effectiveness rises together with capability.


Sources and Verification Notes

All sources below were retrieved or verified online on October 7, 2026. Dates refer to first publication or stated version dates, not this report's inferred experiment-completion dates. Company-participated papers are marked B even after peer review, with review status noted separately. Some web pages' full text or attachments have access restrictions, explicitly marked; unread content is not used to supplement technical details.

This report's citations only extract facts relevant to its judgments — official predictions, product claims, or preprint scores are not treated as independent confirmation. News searches serve as cross-checking leads; internal R&D narratives are not constructed from them.


Version update rule: The next update retains this version's judgments and their grounds, changing status item by item as new evidence arrives; it does not let new product names overwrite old measurements, does not interpret missing information as zero, and does not let narrative completeness take priority over the unknown.


  1. B | Google's formal organizational announcement. Pichai, Hassabis, The next chapter of our AI momentum, August 2026 reorganization; scraped text showed no specific publication date. Verified formal titles and responsibilities, Jeff Dean and Sanjay Ghemawat arrangements. Source. This is an official organizational disclosure; it does not prove undisclosed motives behind the adjustment. ↩↩↩

  2. B | Google model release and security disclosure. Introducing Gemini 4 Argon, 2026-09-30. Source. Verified output limits, phased availability, engineering cases, and safeguard descriptions; performance and internal-use results are official disclosures. ↩↩↩

  3. B | DeepMind research explanation. AlphaZero and MuZero, research introduction page, cutoff-date version. Source. Used for self-play, learning environment representations, and planning mechanisms — not to prove general intelligence achieved. ↩

  4. B | DeepMind on cross-task training boundaries. Generally capable agents emerge from open-ended play, 2021. Source. Publicly states AlphaZero was trained separately per game; this report does not equate general algorithms with general trained strategies. ↩

  5. B | Research-team preprint. Novikov et al., AlphaEvolve: A coding agent for scientific and algorithmic discovery, 2026-06-16, arXiv:2506.13131; full text verified. Paper | Full text read. Used for program evolution, specific kernel and overall training-gain figures. Not an independent deployment audit. ↩↩

  6. B | DeepMind applied update. AlphaEvolve: How our Gemini-powered coding agent is scaling impact across fields, 2026-05-07. Source. Verified official disclosure of entry into compute systems and TPU design; not expanded into complete chip autonomous design. ↩

  7. B | Google Research. Dorfman, Brenner, Accelerating scientific discovery with AI-powered Empirical Research Assistance, 2025-09-09, name updated 2026-04-29. Source. Verified scorable tasks, tree search, six task classes; the update date is not treated as a new experiment date. ↩

  8. B | Peer-reviewed paper with company and academic collaborators. Gottweis et al., Accelerating scientific discovery with Co-Scientist, Nature, 2026-05-19, 655:487–496. Paper. Verified structure, collaborative experiments, and authors' stated limitations; peer review does not equal fully independent reproduction. ↩

  9. B | Company and academic collaborator preprint. Schmidgall et al., Accelerating Scientific Research with Gemini in the Real-World, 2026-08-27, arXiv:2608.26701. Paper | Full text read. Verified human links, Agent_H, and paper-evaluation errors; preprint with limited task conditions. Sample error rates not generalized to all AI research. ↩↩

  10. B | Research-team preprint. Feng et al., Towards Autonomous Mathematics Research, 2026-02-10, revised v3: 2026-03-06. Paper | Full text read. Verified Aletheia structure and the limitation of rare successful examples. ↩

  11. B | DeepMind scientific-reasoning research introduction. Accelerating mathematical and scientific discovery with Gemini Deep Think, 2026-02-11. Source. Verified reasoning compute, agent gains, and research contribution levels; results are official evaluations. ↩↩

  12. B | DeepMind research-team peer-reviewed paper. Jumper et al., Highly accurate protein structure prediction with AlphaFold, Nature, 2021-07-15. Paper. Verified abstract and method introduction this time; did not rely on unread attachment details. ↩

  13. B | Peer-reviewed paper by DeepMind, Isomorphic, and other research teams. Abramson et al., Accurate structure prediction of biomolecular interactions with AlphaFold 3, Nature, 2024-05-08. Paper. Verified retrievable abstract; full page access restricted; used only for joint structure-prediction scope. ↩

  14. A | External public research infrastructure. EMBL-EBI, AlphaFold Protein Structure Database, cutoff-date version. Database description. Used for external competition results and database accessibility; the database is maintained by partners; its existence does not validate every prediction. ↩

  15. B | DeepMind. AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome, 2026-09-08. Source. Verified precompute scope, scores, interfaces, and clinical-use restrictions; partner instances are not 9 billion independent experiments. ↩

  16. B | Isomorphic Labs. Building a new path to make medicines with AI, 2026-09-29. Source. Used for the compute-design–synthesis–experiment–development chain; does not prove clinical efficacy or approved drugs. ↩

  17. B | Research-team preprint. Feng et al., Aletheia tackles FirstProof autonomously, 2026-02-24, revised v3: 2026-03-15. Paper | Full text read. Adopted the revised 6/10, majority expert opinion, preferred-answer designation, and correctness explanation; did not carry over higher older claims. ↩

  18. B | Company research-team peer-reviewed paper. Merchant et al., Scaling deep learning for materials discovery, Nature, 2023-11-29. Paper. Used for candidates, DFT validation, active learning, and computational-stability figures. ↩

  19. A | Independent scholars' peer-reviewed perspective. Cheetham, Seshadri, Artificial Intelligence Driving Materials Discovery? Perspective on the Article: Scaling Deep Learning for Materials Discovery, Chemistry of Materials, 2024-04-08, 36:3490–3495. Journal & DOI | Authors' archived full text. Authors declare no competing financial interests; their sample criticism is not an exhaustive check. ↩

  20. B | DeepMind. Genie 3: A new frontier for world models, 2025-08-05. Source. Verified interactive visual capability and authors' limitations; does not prove general physical-simulation accuracy. ↩

  21. B | DeepMind. SIMA 2: An agent that plays, reasons, and learns with you in virtual 3D worlds, 2025-11-13. Source. Verified local experience generation and self-improvement disclosure. Attached large PDF had restricted reading; not used for technical-detail inference. ↩

  22. B | Waymo. The Waymo World Model: A new frontier for autonomous driving simulation, 2026-02-06. Source. Verified Genie-based domain adaptation, camera and LiDAR output; real road data and simulation validation kept separate. ↩

  23. B | DeepMind. Gemini Robotics 2 brings whole body intelligence to robots, 2026-07-30. Source. Task numbers come from official public charts and corresponding notes; undisclosed sample denominators or confidence intervals not reconstructed. ↩

  24. B | DeepMind model card. Gemini Robotics ER 2, 2026-07-30. Model card. Verified outputs, intended use, and limits; ER text reasoning does not equal complete physical control. ↩

  25. B | DeepMind model card. Gemini Robotics On-Device 2, July 2026 version. Model card. Verified trusted-tester distribution, action outputs, out-of-distribution and safety-evaluation scope. ↩↩

  26. B | DeepMind technical report. Gemini Robotics 2: Safety Evaluations, 2026-07-29. Report. Used for safety-evaluation types; semantic evaluations not generalized into all physical-system safety guarantees. ↩↩

  27. B | DeepMind control roadmap. Shah, Flynn, Securing the future of AI agents, 2026-06-18. Source. Used for the insider-threat perspective, coverage/recall/response, and visible-reasoning limits. ↩

  28. B | Google Cloud infrastructure disclosure. Inside the eighth-generation TPU: An architecture deep dive, 2026-04-22. Source. Official comparisons are conditional; this report does not treat peak improvements as actual R&D speedups. ↩↩

  29. B | Company regulatory disclosure. Alphabet, Q2 2026 earnings, 2026-07-22, SEC Exhibit 99.1. Filing. Financial units are millions of USD, converted to hundred-millions in the text; group, quarterly, and cash measures kept separate. ↩

  30. A | IEA. Key Questions on Energy and AI, 2026-04-16, Executive summary. Report. Global statistics, central scenarios, and supply forecasts stated separately; not treated as Google project measurements. ↩

  31. B | DeepMind's formal safety framework. Frontier Safety Framework 3.1, 2026-04-17. PDF. Used for capability definitions, concern thresholds, and decision structures; definitions are not evidence of capability attainment. ↩

  32. A | METR independent research. Task-Completion Time Horizons of Frontier AI Models, page last updated 2026-05-08 as verified, TH1.1. Data & methods page. This report did not infer any new-model scores from missing dynamic-chart values. ↩

  33. Grade-A external research lead | Unreproduced preprint. Liu et al., Does Learning Protein Folding Generalize to Broader Reasoning?, 2026-09-30. Paper | Full text read. Authors' experimental result, used as a transfer-hypothesis lead; not equivalent to confirmed general-intelligence breakthrough. ↩

  34. B | DeepMind Institute official platform statement. Legg, Manyika, Hassabis, Introducing the DeepMind Institute, 2026-09-16. Source. Verified platform responsibilities and author-view statements; not treated as an independent governance body. ↩

  35. Grade-A methodology supplement | Independent-institution researcher note. Kwa, Clarifying limitations of time horizon, 2026-01-22. Source. Page states its review level is below formal research articles and does not necessarily represent all of METR's views. ↩