Scientific Discovery, R&D Feedback, and Physical Intelligence
Research cutoff: October 7, 2026 Version: GPT V1.0 | Independent research | Deliverable: Markdown Research scope: Google DeepMind, plus Google Research, Google Cloud, Waymo, and Isomorphic Labs where they have direct technical interfaces with DeepMind.
The central question of this report: can DeepMind turn "AI that solves a problem" into "a system that keeps producing stronger problem-solving capability"? The conclusion of this report: several local feedback loops already have engineering evidence; whether those loops can be connected into a continuously accelerating general R&D system remains a hypothesis to be verified.
Executive Summary
Google DeepMind is worth studying in depth not only because Gemini competes at the frontier-model level. It offers a set of experiments for testing how superintelligence might form: program search, mathematical research, scientific prediction, hypothesis generation, virtual-environment learning, real robots, and compute-infrastructure optimization.
This report's judgment [D]: the most explanatory thread is "generate candidates — external validation — selection and accumulation — search again." Base models expand the range of proposals; search expands the number of attempts; validators decide which results enter the next round. Real progress depends on how the three work together, and on whether validation covers the actual goal.
This thread explains both the progress and the difficulties:
- Programs can be compiled, tested, and measured; feedback is fast and relatively cheap, so they suit large-scale search.
- Scientific hypotheses can be criticized by language models, but experiments are needed to confirm them; more candidates do not equal more knowledge.
- A video world can generate coherent footage; whether it suits training real robots must be verified by actual action outcomes.
- R&D assistance can save time on one step; whether it accelerates the next generation of AI depends on the full R&D cycle and the attribution of results.
Seven Core Judgments
| Question | This report's conclusion | Evidence status and main boundaries |
|---|---|---|
| Has AI entered AI R&D? | Yes — public evidence goes beyond general code completion into compute kernels, systems, and search-process optimization | Mainly company papers and deployment disclosures [B]; the share of R&D labor hours is unknown |
| Has AI R&D takeoff happened? | Cannot be confirmed yet | Missing continuous cross-generation R&D speed data with controlled inputs and labor |
| Could partial superhuman capability appear first in research? | Possible — structured tasks already have instances | Local capability does not imply overall scientific autonomy [B→D] |
| Will superintelligence first appear as a system? | The system explanation is currently stronger, but the single-model and specialized-system-combination explanations remain competitive | There is evidence for gains between modules; full integration is unproven [D] |
| Is multi-agent necessary? | No necessity evidence yet | Should be compared against single-agent with equal compute, search, and tools via ablation [D] |
| Has World Models already solved physical intelligence? | Insufficient evidence | Visual coherence, decision effectiveness, physical accuracy, and real transfer are different metrics |
| Can Alphabet's full-stack advantage guarantee victory? | No | Interfaces and scale provide advantages but also bring capital, organizational, and control complexity [D] |
The most important quantitative distinction comes from AlphaEvolve. The researchers disclosed that its matrix-operation-related heuristics sped up specific kernels by about 23% on average, corresponding to roughly a 1% reduction in total Gemini training time. There is a large conversion gap between local optimization and overall growth speed. This is real evidence of AI participating in R&D with actual value — but it is not enough, on its own, to prove a takeoff in R&D speed.5
As of the research cutoff, Google has announced Gemini 4 Argon with phased availability. The latest organizational arrangements have also changed: Hassabis has moved to Chair of Google DeepMind and Chief Scientist of Alphabet; Koray Kavukcuoglu takes day-to-day leadership as SVP of Google DeepMind while also serving as Google's Chief AI Architect. Continuing to describe Hassabis as the day-to-day operating head, or Argon as fully available, would distort this report's research baseline.12
The most decision-relevant object of observation is whether the system can, at the same total cost, keep increasing "independently validated and adopted effective results" while reducing human remediation and supervision burden. Model rankings, agent counts, candidate counts, capital expenditure, and demo videos cannot answer this question on their own.
I. Research Design: Separating "Capability Enhancement" from "Accelerating Intelligence Growth"
1.1 Six Questions That Actually Need Answers
- What common mechanisms do DeepMind's various achievements rely on, and under what conditions do they fail to transfer to each other?
- How do AI-proposed solutions move from seeming reasonable to reliable knowledge or actual engineering improvement?
- Has AI participation in R&D already changed the speed at which the next round of AI capability grows?
- Among base models, tools, memory, validation, and environments — which are sources of capability, and which are merely delivery conditions?
- Can virtual experience improve reliable action capability in the real world?
- When do capital, compute, experiments, and control capacity become decisive bottlenecks?
These questions come before the report's structure. After collecting the evidence, this report organizes its chapters into validation mechanisms, R&D feedback, systems explanations, physical transfer, industrial conditions, and governance — not a product-by-product tour.
1.2 Evidence Grading
| Grade | This report's usage | What it cannot automatically imply |
|---|---|---|
| A: Independent research or authoritative external data | METR methods and limitations, IEA energy research, independent scholars' analysis of materials research, etc. | Independent sources do not mean conclusions are uncontested; preprints still need reproduction |
| B: Official first-hand sources and research-team papers | Google, DeepMind, Waymo, and Isomorphic disclosures; peer-reviewed papers with company participation; Alphabet regulatory filings | Peer review does not substitute independent repeated experiments; deployment disclosures are not independent audits |
| C: Authoritative media | Used for news cross-checking and as search leads | Reports whose full text was not obtained are not used to reconstruct internal technical roadmaps |
| D: This report's analysis | Mechanism explanations, competing hypotheses, scenarios, company recommendations | Cannot be written as the company's realized internal state |
The grading describes source relationships, not a mechanical reliability ranking. Formal announcements should be preferred for organizational appointments; same-condition independent tests should be preferred for model performance. Regulatory filings are strong on financial definitions but cannot provide internal R&D attribution for DeepMind.
This report's main limitation: Google's internal R&D logs, the full set of failed experiments, actual labor investment, model training recipes, and complete safety-incident data are not public. Where these variables are involved, the unknown is explicitly preserved.
1.3 Working Definitions
- Domain superhuman capability: consistently exceeding an appropriate human-expert baseline under a clear task distribution and fair resource conditions.
- AGI: reliable, transferable capability across broad cognitive tasks. This report does not set a single exam day.
- ASI / superintelligence: consistently exceeding the strongest human experts and large organizations on a wide range of important tasks, able to take on long-horizon goals, adapt to change, and produce verifiable results.
- Systemic superintelligence: the above capability produced jointly by models, tools, environments, memory, validation, collaboration, and infrastructure. It is an implementation hypothesis to be tested, not a preset conclusion.
- AI R&D Takeoff: AI participation causing a significant, sustained additional acceleration in the growth speed of AI R&D capability — not merely raising the efficiency of one researcher or one task.
- Recursive Improvement: improvements produced by the system enter the next round of R&D capability, which in turn expands improvement capability in the following round. The source of improvements, the deployment chain, and cross-round gains must be recorded; not every iterative training counts as recursive self-improvement.
These are this report's analytical definitions [D]; they do not declare AGI or ASI on behalf of any company.
II. Current Baseline: The Research Subject Is No Longer a Closed Lab
2.1 Organizational Changes and Their Research Implications
The organizational adjustment disclosed in August 2026 further separated strategic scientific responsibilities from day-to-day model and product operations. Hassabis retains research-advisory and strategic roles and continues to lead Isomorphic Labs; Kavukcuoglu is responsible for Gemini models, frontier AI research, and the Gemini app and developer teams, reporting to Pichai. The announcement also disclosed that Jeff Dean and Sanjay Ghemawat will found an independent public-benefit company, with Google continuing to collaborate as investor and Cloud partner.1
This report's judgment [D]: this is a change in how research and products coordinate — it cannot be used, on job titles alone, to infer that "scientific research has been abandoned," nor that "an internal secret breakthrough toward ASI has been discovered." Judging the outcome of the reorganization requires future data: whether basic research programs continue, whether papers and tools stay open, whether long-horizon scientific research gets resources, and whether the shared evaluation standards of frontier-model and application teams improve.
2.2 Correct Attribution of Achievements
| Entity | Interfaces examined in this report | Attribution discipline |
|---|---|---|
| Google DeepMind | Gemini, Alpha series, SIMA, Genie, Robotics, frontier safety | Core subject |
| Google Research | ERA and some joint scientific research | Cannot all be credited as DeepMind's independent achievements |
| Google Cloud and infrastructure teams | TPU, compilers, deployment, enterprise and science tools | Full-stack synergy needs specific interface evidence |
| Waymo | Autonomous-driving domain adaptation of world models | Waymo's road performance cannot be directly attributed to Genie |
| Isomorphic Labs | AI drug design and experimental translation | A drug R&D institution, not equivalent to the AlphaFold product |
| External researchers and labs | Biological experiments, math evaluations, independent criticism and reproduction | Collaborative verification and fully independent verification should be separated |
This boundary has practical meaning. If research candidates are proposed by DeepMind, experiments completed by universities, and production deployment handled by the Cloud team, then the capability comes from a cross-organization chain. Compressing the whole chain into "one model did it autonomously" overstates model autonomy and understates the collaborating institutions' contributions.
2.3 Gemini 4 Argon: A Current Observation Point, Not the Ranking Center of This Report
On September 30, 2026, Google introduced Gemini 4 Argon, disclosing internal software-engineering and security-task applications and raising the output limit to 1 million tokens. This is an output limit, not a context window. The official rollout was still hardening safeguards, starting with trusted testers and cyber-defense institutions before wider availability.2
What should be observed is not the "number one" label, but:
- Whether the same long task can continue under changing requirements, failures, and incomplete information;
- Whether the final result passes acceptance checks not controlled by the agent;
- Whether total cost includes repeated search, tool calls, and human review;
- Whether capability gains come with stronger risks of overreach, vulnerability exploitation, or supervision evasion;
- Whether external institutions can retest under near-real deployment conditions.
As of the cutoff date, public materials cannot confirm Argon's full availability, independent task horizons, or its contribution to frontier AI R&D cycles. Official engineering cases can support "expanded task scope" but not "the entire research team has been replaced."
III. Mechanisms Across Different Achievements: Connecting Proposal Capability to Usable Feedback
3.1 From Game Research to Research Search: What Transfers Is the Method Structure
AlphaZero improved strategy through self-play and reinforcement learning in games with known rules; MuZero learned environment representations useful for decisions and used them to plan actions. They demonstrated the combination of learning, search, and feedback.3
But games provide conditions that real research usually lacks: clear rules, decidable endings, low retry costs, and rapidly generable experience. AlphaZero also needed separate training per game; broad algorithm applicability does not mean one trained policy can directly handle every task.4
This report's judgment [D]: DeepMind's historical continuity is better explained as "searching for task structures that can be optimized in a closed loop" than as "a smooth technical line from Go to all intelligence."
An effective closed loop needs at least four components:
- A generator that can propose valuable candidates;
- A validation mechanism that distinguishes good from bad and covers important failures;
- A search-and-selection strategy that allocates trial resources sensibly;
- A state or learning mechanism that retains results, constraints, and failures.
They can appear in different forms across projects. A shared structure being valid does not mean weights, memory, or skills are already fully shared.
3.2 Five Levels of Validation
The following is this report's analytical framework [D].
| Level | Validation object | Typical methods | Easily misread results |
|---|---|---|---|
| V1: Form and execution | Programs run; proofs pass specific rule checks | Compilation, tests, formal verification | Running gets written as having solved the real business |
| V2: Task performance | Better under fixed data, goals, and resources | Hidden tests, latency, cost, accuracy | Optimizing one score gets written as comprehensive progress |
| V3: Scientific fact | Predictions or hypotheses about natural phenomena hold | Experiments, measurement, statistics, independent reproduction | Model agreement gets written as experimental validation |
| V4: System outcome | Improvements work in the complete pipeline | Production deployment, end-to-end R&D cycles | Local speedups get written as equal overall gains |
| V5: Social and physical outcome | Reliable, controllable, and value-producing in real use | Long-term field data, accidents and recovery, clinical evidence | Demo success gets written as scaled applicability |
Higher levels usually add cost, time, and uncertainty. Lower-level success can provide necessary evidence but cannot skip higher levels.
This framework does not require every problem to use the same validator. Rigorous mathematical proof, statistical evidence from biological experiments, and physical tests of robots are epistemologically different. The key is to state which errors validation can rule out, and which errors remain.
3.3 A Unified Analytical Diagram, Not a Deployed Architecture
flowchart TD
M["Model proposes candidates"] --> S["Search and resource allocation"]
S --> V["Validation independent of candidates"]
V --> K["Retain results and failure records"]
K --> M
V --> E["Real deployment or experiment"]
E --> K
E --> R["Has R&D capability improved"]
R -. "Cross-generation feedback to be proven" .-> M
Solid lines show the process this report uses to analyze each project; the dashed line marks the key unproven link of cross-generation intelligence-growth feedback. It does not claim Google has integrated all modules into one autonomous system.
IV. Scientific Intelligence: Expanding the Search Space, Still Crossing the Threshold of Knowledge
4.1 AlphaFold's Significance Is Predictive Infrastructure — It Cannot Be Expanded into "Biology Solved"
AlphaFold 2 improved protein-structure prediction accuracy through new network architectures and training methods; AlphaFold 3 extended to joint structure prediction of proteins, nucleic acids, small molecules, and other interacting complexes. The latter is the work of DeepMind together with researchers including Isomorphic.1213
EMBL-EBI's external database documentation notes that AlphaFold led at CASP14 and made predicted structures available to researchers. The accessibility of predictions lets the model serve as input to broad research pipelines.14
This report's judgment [D]: its strategic value is not just replacing one structure prediction, but reducing downstream research's initial uncertainty and changing candidate screening and experiment design. Structure prediction alone cannot guarantee function, affinity, in-vivo efficacy, toxicity, or manufacturing conditions. The count of predictions in the database cannot be counted as an equal number of experimental discoveries.
A stronger scientific system needs to acknowledge:
- One molecule may have multiple conformations;
- Experimental conditions may change structures and effects;
- Prediction confidence does not equal the probability of some downstream conclusion;
- Biological function and clinical effect belong to a longer validation chain.
These are this report's analysis of the scientific translation chain, not claims that the model cannot handle the above factors at all.
4.2 AlphaGenome Atlas: The Distance Between Computational Coverage and Experimental Coverage
Released on September 8, 2026, AlphaGenome Atlas precomputes molecular-effect predictions for about 9 billion human single-base variants, combining AlphaGenome and AlphaMissense into variant-impact scores. The official disclosure included directed experimental instances from partners, plus a website, API, and Antigravity workflow interfaces. The official statement also made clear it has not been validated or approved for clinical use.15
Such resources have two different kinds of value [D]:
- Retrieval and ranking value: letting researchers choose more experiment-worthy variants from a huge candidate set;
- Mechanistic-clue value: linking candidates to possible pathways like splicing, expression, or regulation.
It is not 9 billion real experiments, nor a completed map of all genetic causality. Research should track: how many recommended candidates hold in preregistered experiments; and whether the total cost of confirming one valid discovery falls compared to pipelines that don't use the resource.
The combination of Atlas, base models, and agent interfaces shows that specialized scientific resources can enter general workflows. It has not proven that an agent can reliably complete disease-mechanism research without researchers.
4.3 Co-Scientist: Multi-Agent Improves Hypothesis Quality — It Does Not Automatically Produce Truth
The Co-Scientist study published in Nature in May 2026 used Gemini to build a multi-agent system that generates, criticizes, and evolves scientific hypotheses, expanding search with inference-time compute. The paper reported collaborative validation in leukemia drug repurposing, liver fibrosis, and bacterial gene-transfer mechanisms. It also acknowledged limitations in literature access, missing negative results, hallucinations, and preliminary validation; connecting to experimental automation remains a future direction.8
The key distinction here: tournament ranking optimizes relative judgments of candidates; experiments test natural phenomena. Elo-style ranking can organize proposals, but it carries no natural guarantee of biological truth.
This report therefore proposes three checkpoints [D]:
- Whether the comparison includes single-agent repeated search at the same budget, not just one-shot generation;
- Whether evaluators saw real experimental results or only judged textual plausibility;
- Whether all failed hypotheses entered the denominator, or only final success cases were disclosed.
Multi-agent may help cover different lines of thought, execute asynchronously, and divide responsibilities. But if multiple agents rely on the same model, the same literature, and similar rewards, they can still produce highly correlated errors.
4.4 Latest Research-Automation Studies: Longer Loops, but Humans Cannot Be Omitted
An August 2026 Co-Scientist preprint extended the research scope to experiments and papers, including reasoning-system designs like Agent_H. Materials and biology research still retained human experiment execution or protocol adjustments; Agent_H changed the reasoning system, not the base-model weights. In comparisons of 50 papers per condition, the autonomous group with a reliability module still showed 4% severe result hallucination and 24% severe mismatch between implementation and method description; cross-lab reproduction and other issues remain unsolved.9
This evidence supports both progress and limits:
- Progress: AI no longer just offers a suggestion; it can organize longer research processes and produce experimentable protocols.
- Limits: the system may turn real execution logs into methodologically flawed conclusions; "having executed" does not equal "having reasoned correctly."
- The dividing line: AI improving one agent's operating structure is a concrete instance of AI for AI; it does not equal autonomously discovering and training a new generation of general models.
Agent_H needed roughly 40–80 model calls per query, while the base-model control used a single call; blind physician ratings across nine dimensions showed significant improvement on only one. So automatic-scoring advantages cannot be written as same-cost advantages or broad clinical-quality advantages.9
This report's judgment [D]: the core difficulty of near-term research automation may shift from "can it generate a plan" to "can it keep the entire evidence chain valid." For example, train-test leakage, wrong metric choice, and post-hoc selection can all exist inside a program that runs and has complete logs.
4.5 ERA: Converting Some Research Problems into Scorable Software Search
Google Research's Empirical Research Assistance takes tasks, data, and evaluation methods as input, using Gemini to generate and modify code, combined with tree search to select candidates for further exploration. The official research covers six classes of scientific computing tasks. The original introduction was published in September 2025; the April 2026 update mainly added the system name and should not be treated as a brand-new round of experiments.7
ERA's significance [D] is in reducing the cycle cost of "idea → implementation → measurement → rewrite." It also exposes a boundary: if the scientific goal is wrongly encoded as a scoring function, more effective search may amplify the gains of the wrong goal.
To judge whether it surpasses a research assistant, the system needs to face:
- Problems not fully defined by humans in advance;
- Situations with hidden confounds, data issues, or contradictory evaluations;
- Cases requiring rejecting existing metrics and redesigning experiments;
- Results that still hold on external new data.
4.6 Mathematical Research: The Boundary Between Reasoning and Validation Needs Especially Accurate Statements
Aletheia combines Gemini Deep Think, tools, and a generate–verify–revise process to handle mathematical research in natural-language form. The authors explicitly warn in the research paper that successful examples are rare and should not be understood as the system stably solving research mathematics. The official side also did not claim the results of the time reached the major-progress or milestone-breakthrough tiers in its classification.1011
The revised FirstProof study reports: by majority expert evaluation, 6 of 10 problems were solved, with one problem having dissenting opinion. The generation process needed no human modification mid-way, but researchers participated in designating the preferred answers; the paper's explanation of "correct" allows minor revisions consistent with peer-review conventions — not all raw outputs were directly publishable.17
This is more informative than "AI autonomously solved several hard problems":
- Autonomous generation is a process property;
- Experts judging it correct is an outcome evaluation;
- Human selection is a system boundary;
- Formal checking is another validation method;
- Major restructuring of a problem is a higher-level research capability.
Natural-language verification agents cannot automatically be equated with formal provers. In the future, judging domain superintelligence in mathematical research should record the complete chain of new-problem selection, proof repair, error identification, publication, and independent checking — not just problem counts.
4.7 The GNoME Controversy: Don't Merge Computational Stability with Usable Materials
The original GNoME study combined candidate generation, graph-network prediction, and density functional theory computation, using computational results to update the model. The paper reported about 2.2 million crystal structures stable relative to the previous database, of which about 381,000 sit on the updated convex hull. This is mainly computational discovery, not an equal number of experimental syntheses.18
Independent materials chemists Cheetham and Seshadri acknowledged the method's potential but questioned the novelty, realizability, and practicality of some candidates. They studied a sample, not an exhaustive check of the entire database.19
This report neither averages the two sides nor writes the controversy as "all results overturned":
- The original study provides direct evidence for computational search and database expansion;
- The critics are closer to the judgment standards of synthesis and materials use;
- "Computationally stable candidates" and "useful new materials" carry different requirements.
The adopted judgment [D]: GNoME supports the feasibility of scientific search loops; total materials-discovery counts need to be stratified into computation, synthesis, reproduction, function, and use.
4.8 Drug Design: Digital Acceleration Cannot Skip the Physical and Clinical Chains
Isomorphic's Drug Design Engine pipeline disclosed in September 2026 goes from design requirements to computational candidates, then scientist selection, synthesis, and experimental validation. Its disclosure emphasizes day-scale computational design, plus subsequent experimental and clinical-development preparation.16
This report's judgment [D]: the digital design stage may improve rapidly, while complete drug development remains constrained by experiments, human biology, and regulatory validation. Generating candidates in days cannot be converted into producing drugs in days. The company's comparison against traditional processes taking years also cannot be interpreted as an equal-multiple speedup of the full pipeline without a matched baseline.
This is a clear instance of the time gap between digital intelligence and physical results: the former raises candidate-generation and screening speed; the latter requires the real world to supply new evidence.
V. AI for AI: What Thresholds Must Be Crossed from Engineering Assistance to Changes in R&D Speed
5.1 Four Different Feedback Loops
| Loop | Improvement object | Public evidence | Chain still to be proven |
|---|---|---|---|
| L1: Operations optimization | Kernels, compilers, scheduling, chips, and software | AlphaEvolve and deployment disclosures [B] | Whether overall R&D output can improve by the same factor |
| L2: Systems optimization | Search, agent processes, inference configuration | ERA, Aletheia, Agent_H [B] | Whether new configurations transfer stably to new tasks and models |
| L3: Learning-experience generation | Strategy training data, tasks, feedback | SIMA 2 self-improvement disclosures [B] | External environment realism, sustained cross-round gains |
| L4: Frontier model R&D | Data recipes, algorithms, training, and next-generation capability | Locally relevant tools already exist | Insufficient public evidence to reconstruct a complete autonomous loop |
Calling all four loops "AI self-evolution" masks key differences. L1 can save compute without inventing new learning algorithms; L2 can improve one model's usage without changing weights; L3 improves strategy but may rely on stronger teachers and fixed training frameworks; only L4 directly concerns next-generation frontier capability growth.
5.2 AlphaEvolve: One of the Closest Pieces of Evidence to Engineering Reality
AlphaEvolve puts language-model-generated programs into an executable-evaluation and evolutionary-search pipeline. The 2025 paper disclosed that specific matrix-operation-related optimizations brought about a 23% average kernel speedup, worth roughly a 1% time saving across all of Gemini training. A 2026 official update further disclosed its entry into TPU circuit design and other compute-system uses.56
This provides two valuable layers of fact:
- Some AI-generated improvements were concretely measured, not just textual suggestions;
- Some results entered continuously used engineering systems, not just one-off benchmark experiments.
But this report will not therefore write:
- Gemini training improved 23% overall;
- AI has already designed a complete next-generation TPU;
- Frontier algorithm research is fully automated;
- Continuous multi-generation model R&D shows unmanned positive feedback.
Humans still define problems, constraints, and evaluations; the complete gain depends on how much of the work the improvements cover and where the new bottlenecks are.
5.3 How Local Efficiency Converts to Overall Efficiency
A simplified analytical formula [D]:
Overall speedup = 1 / [(1 − f) + f/r + h]
where f is the share of the original process's time that AI can accelerate, r is the speedup factor on those stages, and h is the share of the original process time taken by new review, coordination, and rework. It assumes stages add up in time — good for illustrating bottlenecks, not Google's actual R&D model.
Example: if 40% of a cycle can be sped up 4×, the rest unchanged, and new review takes 5%, the overall speedup is only about 1.33×. This number is an illustrative calculation, not a prediction.
Real R&D also has stage parallelism, queuing, budget constraints, and failed retries. So it is necessary to distinguish:
- Compute hours saved;
- Researcher hours saved;
- Calendar time shortened;
- More experiments at the same time;
- More capability gain at the same investment.
5.4 Five Levels of Tests from R&D Participation to Takeoff
| Level | Evidence that must appear | This report's judgment on DeepMind |
|---|---|---|
| R0: Assistance | Provides code, analysis, literature, or suggestions | Clearly present |
| R1: Executing closed loops | Automatic experimentation, evaluation, and modification within preset tasks | Clearly present in several instances |
| R2: Effective improvement | Results pass checks external to the generation process and enter real pipelines | Company first-hand evidence exists; independent review coverage is limited |
| R3: Cross-generation feedback | Improvements enter the next-generation R&D system, which produces stronger improvements | Disclosed for local agent training; insufficient for the frontier general R&D chain |
| R4: Sustained takeoff | R&D capability growth significantly accelerates after controlling resources, labor, and tasks | Insufficient public evidence yet |
"External to the generation process" can be an independent tester, not necessarily an external institution. Independent institutional re-verification is additional evidence; the two should not be conflated.
The safest answer at present: AI R&D participation has already changed how some engineering work is produced; whether it has continuously increased the speed of intelligence growth remains insufficiently evidenced in public materials.
5.5 What Data Must Be Disclosed to Raise the Credibility of a Takeoff Judgment
This report recommends [D] disclosing — or having a trusted third party audit — the following data:
- For comparable research tasks: the workload AI handled, human takeover volume, and failure rates;
- Calendar time from research-plan proposal to independent confirmation;
- Total candidates, total failures, and the share that entered production or the next training round;
- Capability gains after controlling compute, teams, and data;
- Source tracing for at least three consecutive rounds: which AI produced what improvement, and how it changed the next round of R&D;
- Whether improvements persist without increasing supervision cost;
- Whether AI-led important algorithms or training methods exist and are reproduced by external researchers.
Three consecutive rounds is this report's suggested minimum observation window, not a law of nature or an official threshold.
5.6 Positive Feedback Can Appear While Gains Still Diminish
A system may repeatedly optimize easily measurable kernels yet find less and less new room. It may also speed up experiments while experiment evaluation, training queues, or data cleaning become bottlenecks.
This report distinguishes two compatible explanations [D]:
- Acceleration explanation: stronger models raise research productivity, research results raise models further, and the loop's gains grow.
- Bottleneck migration explanation: AI removes several constraints, and remaining experimental and organizational links make overall gains diminish.
Only continuous input–output data can tell them apart. One successful case cannot decide which long-term mechanism dominates.
VI. Scaling Has Split into Multiple Resources — They Cannot All Be Called "Bigger"
6.1 Scaling Variables in the DeepMind Cases
| Variable | What it can increase | Main bottleneck | This report's test |
|---|---|---|---|
| Pretraining | Representations, knowledge, and basic skills | Data, optimization, cost | Compare capability gains after post-training, controlling for it |
| Post-training & RL | Reasoning, action, and reward adaptation | Reward coverage, distribution shift | New tasks and anti-reward-hacking tests |
| Inference-time compute | More candidates, revisions, and planning | Search efficiency, validation accuracy | Success rate at the same total cost |
| Context & memory | Retaining task materials and history | Retrieval, state updates, conflicts | Effective retention rate and recovery ability |
| Program search | Algorithm and implementation exploration | Evaluation function, execution cost | Hidden tests and deployment results |
| Validation | Filtering errors, guiding search | Missed detections, correlated errors, price | Independent judgments and failure coverage |
| Synthetic experience | Larger training and interaction sets | Bias, difficulty, realism | New environments and real-task transfer |
| Agent runtime | Longer work that can be advanced | State drift, failure accumulation | Human-equivalent task horizons and takeovers |
| Parallel agents | More experiments and division of labor | Duplication, communication, coordination | Single-agent ablation at the same budget |
| Tools & real resources | Access to new information, executing actions | Permissions, interfaces, latency | Whether resource use produces final gains |
| Infrastructure | Larger throughput and available budgets | Chips, memory, network, power | Actual utilization and cost per effective result |
This table is an analytical framework [D]. It does not claim all variables follow power laws, nor that they can substitute for each other indefinitely.
6.2 Why Inference Scaling Is Conditional
Official research on Deep Think and Aletheia shows that adding inference-time compute improves scores on the corresponding math tasks, and agent processes can further improve compute-use efficiency.11
The judgments derivable from this evidence [D]:
- Expanding the search budget is meaningful only if search can find different candidates;
- Improving evaluation quality may matter more than continuing to expand candidate counts;
- After the task distribution changes, existing curves need not hold;
- One evaluation curve cannot be extended to infinite compute and infinite capability.
"Verifier Scaling" also cannot just count reviews. If validators and generators share errors, repeated judgments may amplify false confidence. Independent program tests, real experiments, and formal checks each provide different error-correction channels in applicable tasks.
6.3 Long Output, Long Context, and Long Tasks Are Not the Same Kind of Growth
Long output can expand reasoning but may also amplify self-reinforcing errors. Long context can hold materials but does not guarantee the model correctly updates task state. Continuous operation can generate many actions but does not guarantee goal progress.
This report recommends: separate token scale from task-success records. Priority should go to whether the system, after several key state changes, can still correctly locate the goal, what is done, conclusions pending validation, and current permissions.
For the capacity, update quality, cross-task transfer, and R&D contribution of DeepMind's internal persistent-memory systems, existing public materials are insufficient for a reliable baseline. Gemini's context capability cannot fill this unknown.
VII. Single Model, System, or Organization: Four Competing Explanations
7.1 H1: The Unified Base Model Is the Main Capability Source
The strongest argument: Gemini participates in language, code, math, scientific-hypothesis, and action tasks, and the base model's cross-task capability reduces the need to redevelop specialized models problem by problem. [B→D]
If H1 dominates, we should see:
- Most tasks improving significantly after base-model improvements;
- Major gains retained even with simplified tools and processes;
- Marginal gains from more complex agent structures shrinking;
- New domains not depending on large amounts of specialized data or human design.
Falsifying conditions: end-to-end tasks not improving correspondingly after base-model upgrades, or validation, environments, and specialized modules still deciding the main success rates.
Remaining unknown: public projects usually change models, processes, and compute simultaneously, making model contributions hard to isolate.
7.2 H2: Models Plus External Mechanisms Form a Stronger System
AlphaEvolve, Aletheia, and Co-Scientist connect models with search, tools, or evaluation processes; Robotics also separates high-level reasoning from action models. [B→D]
If H2 dominates, we should see:
- The same model improving significantly through better state, tools, and validation;
- System gains transferable across models;
- Real task costs falling, not just evaluation scores rising;
- Fault recovery and permission management becoming main competitive factors.
Falsifying conditions: complex systems' advantages disappearing in same-budget comparisons, or old components no longer producing gains after model upgrades.
7.3 H3: A Combination of Specialized Superhuman Systems Never Forms General Autonomous Intelligence
AlphaFold, GNoME, and robot policies may each be strong while using different data, feedback, and goals. Connecting APIs can improve workflows without forming unified understanding or autonomous research. [D]
If H3 is correct, we will see many domain scores rising while cross-domain goal selection, long-horizon planning, and unknown-environment adaptation stagnate. Important resources are still configured by human organizations; module errors are patched by human coordination.
This is not an "AI has no value" scenario. A combination of specialized systems can still have huge impact on research and industry.
7.4 H4: The Strongest Capability Comes from the Research Institution and Industrial System
Google owns model research, compute architecture, cloud deployment, application entry points, and experimental collaboration interfaces.128 This report therefore proposes an institution-level explanation [D]: the unit that actually takes on long-horizon goals may be "the organization of human researchers plus AI," not an independently replicable agent.
If H4 dominates, leading indicators should be verified research output per unit investment, cross-team processes, and actual deployments — not some single score of model weights.
7.5 Provisionally Adopted Combined Judgment
This report favors H2 and H4 jointly explaining current achievements, without ruling out H1 growing stronger in the future. Existing evidence is also compatible with H3. This is not assigning average probabilities to the four hypotheses; it says currently observable achievements mostly reflect clear environments, validation, and organizational interfaces — still insufficient to prove a fully general autonomous system.
7.6 Evidence That Multi-Agent Is Not a Necessity
One agent can repeatedly call models, switch roles, and maintain multiple candidates. Multiple agents may merely split these operations in engineering. Necessity judgments need comparisons of:
- The same base model;
- The same token, tool, and compute budgets;
- The same information and permissions;
- The same final acceptance;
- Different coordination methods.
This report's proposed collaboration metric:
Net collaboration gain = effective results of multi-agent − effective results of the best same-cost alternative
while also recording communication, duplicated search, shared errors, and human coordination. More agents with unchanged net gains should not be written as enhanced collective intelligence.
To surpass large human organizations, a system also needs to choose goals, coordinate resources, handle conflicts, remember commitments, and accept accountable control. Existing materials are insufficient to confirm DeepMind has this complete capability.
VIII. World Models and Physical Intelligence: Can Virtual Experience Cross the Real-World Interface
8.1 Genie 3: Interactive World Generation Does Not Equal Calibrated Physical Simulation
Genie 3 demonstrated prompt-generated interactive visual environments, with the official report citing 720p, 24 frames per second, and minute-level continuity, while disclosing limitations in action spaces, multi-agent interaction, text, and environment accuracy.20
Its value [D] may lie in reducing environment-creation and experience-collection costs. But visual coherence alone cannot answer:
- Whether the same action produces correct contact and mechanical changes;
- Whether hidden states stay consistent;
- Whether object mass, friction, or deformation can be calibrated;
- Whether rare failures cover the real distribution;
- Whether a trained policy will exploit loopholes in the generated environment.
A model suited for video experiences can be a starting point for training environments; whether it suffices to support robots needs extra validation.
8.2 SIMA 2: Experience Generation and Next-Round Policy Training Have Appeared
SIMA 2 acts in virtual 3D environments through images, keyboard, and mouse. The official description covers using Gemini to generate tasks and feedback, autonomously accumulating experience, and then training an improved agent — also showing iteration inside Genie environments. The official statement also acknowledges limitations in long tasks, memory, fine manipulation, and real-time interaction.21
This report's adopted judgment:
- There is official evidence of a local self-improvement training loop.
- The loop depends on an established training system, teacher models, and environments.
- It does not prove SIMA autonomously rebuilt the whole research framework.
- Gains on virtual game tasks cannot directly substitute for real-robot evidence.
If the next-round agent improves task capability, that is policy progress; only if it also improves experience selection, evaluation, and training methods — and keeps bringing larger next-round gains — does it come closer to this report's recursive-R&D mechanism of interest.
8.3 Waymo: Domain Adaptation Is Transfer Evidence and Boundary Evidence
Waymo's February 2026 introduction of a Genie-3-based, driving-domain-adapted world model generates camera and LiDAR data and provides scenarios and driving controls for simulating rare cases.22
This shows a base world model can enter specialized industrial pipelines. It does not prove a general world model can accurately simulate all real processes without post-training.
Waymo's real road experience belongs to its autonomous-driving system; road miles cannot be counted as the Genie world model's real-world validation miles. Evaluating generative simulation should separately disclose: sensor statistical consistency, action response, accident reproduction, closed-loop policy results, and additional gains on real roads.
8.4 Gemini Robotics 2: Observing "Unevenness" Is More Useful Than Observing the Best Clips
Gemini Robotics 2, released in July 2026, includes action, high-level reasoning, and on-device routes. The official disclosure tested the same checkpoint across multiple hardware setups while acknowledging challenges in multi-finger manipulation; in example tasks, screwing in a light bulb succeeded 36% of the time versus 92% for screwing one out. These are company results under specific test conditions, not general household success rates.23
The model cards further distinguish:
- ER 2 takes multimodal input and outputs text reasoning, serving as a robot-system component — not a complete controller directly outputting motor commands.24
- On-Device 2 outputs numeric actions, currently for trusted testers; the model card retains limits on out-of-distribution tasks and high-degree-of-freedom control, and its safety evaluation mainly covers standing dual-arm operation — locomotion and whole-body control are not in the same validated scope.25
So the capabilities of different variants cannot simply be added: whole-body action demos do not automatically extend the on-device model's validated scope, and open reasoning APIs do not equal universally open complete robot control.
8.5 Why Digital and Physical May Show a Clear Time Gap
This report's analysis [D] gives at least five reasons:
- Digital experiments can be replicated and parallelized; robots need equipment, space, and maintenance.
- Code failures can often be rolled back; physical failures can damage equipment or cause injury.
- Digital task states are easier to record; physical systems have occlusion, wear, contact, and sensing errors.
- Virtual experience can be generated in bulk; its realism must be calibrated.
- Software deploys quickly; hardware manufacturing, installation, and certification have their own cycles.
But no fixed multi-year gap can be derived from this. If sensors, hardware standardization, world models, and robot data efficiency improve together, the gap may narrow; if dexterous manipulation and safe recovery stall, the gap may persist long-term.
8.6 Minimum Experimental Design for Verifying Simulation Transfer
This report recommends [D] using four control groups:
| Group | Training condition | Question to answer |
|---|---|---|
| P0 | Real data only | Existing real-task baseline |
| P1 | Real data + traditional simulation | Gains from traditional simulation |
| P2 | Same real data + generated worlds | Whether generated environments provide independent gains |
| P3 | Less real data + generated worlds | Whether real-data costs truly fall |
In the same batch of unseen environments, record task success rates, takeovers, collisions, action anomalies, recovery, and total cost. If P2 only improves simulation scores with no real-test improvement, no physical-intelligence breakthrough should be declared.
IX. From Software Research to Industrial Conditions: Advantages and Constraints Grow Together
9.1 The TPU Roadmap Shows Training and Inference Resources Are Diverging
Google Cloud's April 2026 disclosure of TPU 8t and 8i: the former for training, the latter for inference and inference-style workloads, with different memory and networking configurations. Official performance and price-performance comparisons depend on specific conditions and use "up to" phrasing.28
This report's judgment [D]: AI capability growth no longer only demands bigger training clusters. Long reasoning, program search, parallel experiments, and agent services need different memory access, networking, latency, and scheduling.
DeepMind-related advantages may appear at four interfaces:
- Research goals and chip design;
- Model architectures and compilers;
- Training recipes and cluster operations;
- Agent workloads and inference services.
Whether advantages hold should be judged by actual end-to-end throughput, stability, and cost per effective result; chip peak performance cannot directly represent research output.
9.2 Financial Constraints: Don't Treat Alphabet's Budget as DeepMind's Budget
Alphabet's Q2 2026 earnings filing discloses:
| Metric | Q2 2026 value | Definition |
|---|---|---|
| Operating cash flow | ~US$39.069 billion | Group quarterly figure |
| Property and equipment purchases | ~US$44.924 billion | CapEx-related cash measure |
| Free cash flow | ~−US$5.855 billion | Company-defined non-GAAP measure |
| June equity issuance net proceeds | ~US$49.6 billion | For general corporate purposes, including expanding AI infrastructure |
| Quarter's senior unsecured notes net proceeds | ~US$20.3 billion | General corporate purposes |
Source: company regulatory disclosures; quarterly financials are unaudited. It does not disclose DeepMind's separate compute budget, return rate, or per-project investment.29
This report's judgment [D]: research capability is now interconnected with capital markets and industrial expansion. A strong operating business does not mean AI infrastructure can be funded internally without limit. A quarter of negative free cash flow, on its own, cannot imply the company cannot invest or that the AI business has no returns.
More research-relevant: whether new capability and usage revenue keep up with depreciation, financing, and supply costs; which research directions need resource protection; which projects get priority when budgets tighten.
9.3 Power and HBM: Global Trends Cannot Be Directly Applied to Google
The IEA's 2026 study estimates global data-center electricity use may rise from 485 TWh in 2025 to about 950 TWh in its central 2030 scenario; it also notes grid-connection, equipment, and chip-supply constraints, and expects HBM tightness to continue at least through end-2027. The latter two contain forecasts; 950 TWh is not an occurred fact.30
The data covers global data centers — not AI electricity alone, and certainly not Google's or DeepMind's electricity.
This report's judgment [D]: effective compute is constrained by a chain of interdependent conditions — chip availability, memory and networking fit, facility completion, power connection, cooling availability, cluster stability, and software that can use it well. Announced planned capacity cannot substitute for actually commissioned capacity.
For nuclear, storage, or on-site generation, contracts, permits, construction, and actual power delivery need separate evaluation; this report does not treat distant energy promises as ready R&D resources.
9.4 Scale May Strengthen Advantages — and May Change Technology Choices
If search and reasoning consume more and more resources, the optimal approach may be:
- Let smaller models propose cheap candidates, stronger models handle hard nodes;
- Raise validator selectivity;
- Reuse existing experiments and failure records;
- Allocate budgets by risk and value;
- Stop searching when evidence is truly lacking, rather than generating infinitely.
These are this report's techno-economic inferences [D]. The future winning system is not necessarily the one using the largest model at every step.
X. Control: Research Acceleration and Risk Growth Come from the Same Action Interface
10.1 What the Current Frontier Safety Framework Tracks
DeepMind's Frontier Safety Framework 3.1 distinguishes lower-level capabilities of concern from critical capabilities, and covers misuse, supervision evasion, and machine-learning R&D risks. Its ML R&D automation level asks whether a Google capability research team could be automated at similar total cost; the framework also holds that R&D acceleration may come from models combined with workflows rather than weights alone.31
The framework defines a threshold; it does not mean the company has reached it.
The document's decisions involve risk assessment, mitigation, and risk-acceptance conditions — they cannot be compressed into "any capability crossing a line automatically stops training." Likewise, publishing the framework does not prove actual deployments always comply with it.
What should be checked: who evaluates, which systems are covered, what conditions trigger escalation, what evidence is disclosed, and whether actual risk handling can be externally examined.
10.2 Agent Control Is Starting to Adopt an Action-and-Infrastructure Perspective
The June 2026 control roadmap treats capable-but-not-fully-trusted agents as potential insider threats, proposing supervision, blocking, and response mechanisms, tracking monitoring coverage, recall, and response times, and acknowledging that visible reasoning may be insufficient for evasive or opaque reasoning.27
This is an explanatory direction [D]: security needs to protect permissions, networks, data, and experimental resources — not just check whether the final answer contains dangerous text.
Argon's release materials also describe reasoning-and-action monitoring, execution halts, and hardened isolated environments. These are the company's mitigation disclosures; they cannot be taken as externally verified complete control guarantees.2
10.3 Robot Safety Cannot Be Borne by Semantic Reasoning Alone
The Robotics 2 safety report studies multiple semantic and agent risk evaluations. The on-device model card separately recommends layered control: high-level semantic judgment, low-level collision and force control, and hardware-related safety mechanisms.2625
One closed-loop evaluation used Gemini to simulate VLA confidence feedback, avoiding real robots or physical simulators. This can test whether agents obey feedback constraints; it cannot independently prove dangerous real-robot actions are reliably intercepted.26
This report's judgment [D]: a robot "understanding not to harm people" and "the physical system remaining safe under wrong actions" are different capabilities. A complete system needs to handle perception errors, execution anomalies, network interruptions, and mechanical problems. Real safety evidence should include dangerous-action rates, takeovers, emergency stops, and recovery — not just the success rate of refusing dangerous instructions.
10.4 Why AI R&D Automation Carries Special Control Risks
The following is mechanism analysis [D], not an allegation that DeepMind has done anything of the sort:
- Research agents need access to code, training logs, data, and large amounts of compute;
- Optimization tasks easily grant them the ability to modify tests, schedules, and monitoring;
- If the evaluator, the evaluated object, and the evaluation setup are all controlled by the same system, feedback may be distorted;
- Rapidly generating large numbers of experiments increases the safety-review backlog;
- Optimizations that are harder to understand than human ones may raise validation costs.
So capability and control must be observed in pairs:
| Capability growth | Control metrics that must be paired |
|---|---|
| More tools and system permissions | Least-privilege coverage and overreach rates |
| Longer runtimes | Monitoring continuity, state records, and recoverability |
| Stronger code and network ability | Isolated-environment testing and exploit interception |
| Automatic modification of R&D pipelines | Acceptance and audit chains that cannot be modified by the agent itself |
| Stronger experiment design | Experiment risk approval and execution boundaries |
| More concurrent agents | Scheduling, permission inheritance, and collective-behavior monitoring |
10.5 Governance and Benefit Attribution
The DeepMind Institute, founded in September 2026, provides a cross-disciplinary discussion platform headed by Legg, Manyika, and Hassabis. Its statement is explicit: articles are the authors' views and should not be read as Google's official positions.34
The platform can promote discussion, but it is not an independent regulator or a deployment approval body. Its founding cannot be used to prove governance has kept up with capability.
This report raises three benefit-distribution questions [D]:
- Which of scientific predictions, model weights, data, and validation resources should stay publicly accessible?
- Can external researchers examine key capability and safety judgments, or only see curated instances?
- How should knowledge and benefits from research collaborations be distributed among platforms, experimental institutions, and public funders?
Google's system simultaneously provides research resources, commercial cloud services, and application distribution. It can lower participation barriers, and it may also concentrate control over interfaces, pricing, and visible results. Both outcomes must be judged from actual openness conditions and usage data — not decided by "open" or "full-stack" slogans alone.
XI. Evidence Map: What Is Established, What Still Cannot Be
| Research proposition | Direct evidence | This report's status | Main gaps |
|---|---|---|---|
| General models can participate in different research tasks | Gemini-related science and math projects [B] | Established for tested tasks | Cross-domain reliability and full autonomy |
| AI can search and measure program improvements | AlphaEvolve, ERA [B] | Strong first-hand experimental and applied evidence | Independent deployment review and overall gains |
| AI can propose experimentally testable new hypotheses | Co-Scientist collaborative research [B] | Instances exist | Full attempt denominators and independent reproduction |
| AI can continuously complete research independently | Latest research-automation preprints [B] | Not sufficiently established | Human execution, methodological errors, external reproduction |
| AI has improved agent systems | Agent_H and related research [B] | Local instances exist | New tasks, generations, and same-budget validation |
| Agents can improve with generated experience | SIMA 2 [B] | Local official evidence | Teacher dependence and real transfer |
| World models can enter industrial simulation | Waymo domain adaptation [B] | Established for disclosed interfaces | External calibration and additional real-vehicle gains |
| Robotics has broad physical reliability | Official tasks and model cards [B] | Cannot be confirmed yet | Long-term real distributions, accidents, maintenance, takeovers |
| AI R&D speed keeps accelerating significantly | Insufficient public attribution data | Unknown | R&D cycles under resource control, continuous |
| One system has integrated all scientific and physical capabilities | Multiple modules and local interfaces | Insufficient evidence | Unified goals, memory, transfer, and long tasks |
| Infrastructure will constrain | IEA [A], TPU and financials [B] | Evidence at the global level | Each Google project's specific exposure |
| Control keeps up with capability stably | Framework and mitigation disclosures [B] | Not publicly proven | Independent high-intensity attacks, coverage, and handling data |
This report does not fill "unknown" with estimates. Missing public information neither proves nothing was achieved internally, nor allows the report to assume it was.
XII. Active Falsification: Where This Report Is Most Likely Wrong
12.1 "Validation Is the Key" May Be an Illusion of Current Engineering Practice
Stronger base models may significantly reduce the need for search and external patching. If a future model directly produces highly reliable results on complex new tasks and system-component contributions shrink, this report would raise H1's weight.
But "needing less validation" still needs to be proven by independent results. More confident model expression is not a falsification.
12.2 Research Feedback May Not Effectively Improve General Intelligence
Scientific models can handle specific representations accurately without improving general planning or social tasks. Scientific successes cannot be accumulated into an AGI score.
A late-September 2026 external preprint, Fold2Reason, offers a reverse clue: protein-structure-derived supervision may improve some broad reasoning evaluations, with the authors reporting an average gain of about 3.23 percentage points. But it is an early preprint; it does not prove the same transfer for frontier models or long-task research capability.33
This report keeps open the possibility that "scientific data helps general intelligence" while refusing to derive ASI directly from small evaluation transfers.
12.3 A Larger Candidate Pool May Just Be a Larger Selection Advantage
If a system generates thousands of candidates and experts pick from them, good final performance may come from search budget and human curation rather than stronger autonomous research.
Ways to falsify: limit total budgets, record all attempts, and compare against strong human teams and same-budget single agents on hidden tasks. Disclosing only successful solutions cannot answer this.
12.4 Closed Evaluation May Be Overfitted
Scorable tasks suit search — and are easily exploited by it. Test-set leakage, special cases, and unrealistic baselines can all manufacture capability illusions.
Validators should be split into feedback visible during optimization and final invisible acceptance. If the latter does not improve, this report will downgrade the relevant capability judgments.
12.5 Physical Transfer May Be Faster Than Expected — or Stall Long-Term
If generated worlds significantly reduce real-data needs and cross-hardware transfer proves reliable, the Digital–Physical gap may narrow. Conversely, if dexterous manipulation, unexpected recovery, and long-term maintenance don't improve, quality demos may still sit far from actual use.
12.6 Full-Stack Advantages May Be Weakened by Independent Ecosystems
External models, open tools, and specialized experimental institutions can also combine into effective systems. They don't necessarily need all resources under one company.
This report does not treat "owning the full stack" as a sufficient condition. What should be observed: whether synergy truly lowers cost and failure, or whether organizational coordination and closed interfaces offset the advantages.
12.7 Research Automation May Increase Low-Quality Research Rather Than Effective Discovery
As the marginal cost of generating hypotheses and papers falls, review burden and duplicated directions may rise. If the share of effective results falls and experimental resources get crowded out by weaker proposals, research speed may not improve.
So "paper counts," "generated hypothesis counts," and "agent work hours" can only be input metrics — not final intelligence-growth metrics.
12.8 Control May Become a Necessary Input to Growth, Not an Added Cost
Stronger agents may need stricter isolation, slower approvals, or more monitoring. Ignoring these inputs overstates economic gains.
If control technology improves so supervision burden per effective task falls, capability is more likely to form deployable growth. If supervision work grows faster than capability, takeoff may face real constraints.
XIII. Scenarios for the Next 3–10 Years: Update by Conditions, No Precise ASI Date
Observation window: Q4 2026 to around 2036. The following are this report's scenarios [D], not company commitments. Scenarios may appear in different domains simultaneously; no unsupported precise probabilities are assigned.
Scenario A: Capability Keeps Improving, Systems Gradually Expand, No Unified AGI Day Long-Term
| Dimension | Analysis |
|---|---|
| Premise | Base models, reasoning, and agents keep progressing; module fusion still needs heavy engineering |
| Leading indicators | Hidden-task success rates rise; takeovers and rework fall; science tools enter more pipelines |
| Possible outcomes | Digital work and some research significantly automated; physical tasks develop more slowly |
| Main risks | Reading multiple local capabilities as general autonomy; underestimating supervision costs |
| Evidence raising this scenario's weight | Broad but continuous progress; key breakthroughs still depend on task definition and human validation |
| Evidence lowering its weight | Reliable research-team replacement without major input growth, or long-term capability-trend stagnation |
This is the baseline observation path most easily extended from existing public evidence; it does not mean other scenarios are impossible.
Scenario B: Significant Acceleration in AI R&D
| Dimension | Analysis |
|---|---|
| Premise | AI not only executes research but improves experiment selection, training, and R&D tools; gains enter the next generation |
| Leading indicators | Successive R&D cycles shorten; capability gains at the same resources rise; AI contributions traceable |
| Possible outcomes | Faster iteration of models, agents, and optimization systems; greater pressure on supervision and infrastructure |
| Main risks | Wrong evaluations amplified quickly; capability and safeguards updated out of sync |
| Evidence raising weight | Auditable positive feedback across at least several consecutive rounds; important algorithms independently reproduced |
| Evidence lowering weight | Acceleration mainly from compute or team expansion; local results not changing overall cycles |
One kernel optimization or one agent-training improvement cannot be used to declare this scenario already arrived.
Scenario C: Industrial and Control Constraints Become the Main Bottleneck
| Dimension | Analysis |
|---|---|
| Premise | HBM, power, facilities, capital, or safety verification limit expansion speed |
| Leading indicators | Planned-vs-commissioned gaps widen; resource queuing; unit effective-task costs hard to lower |
| Possible outcomes | More emphasis on small models, selective search, and validation-cost savings; divergent deployment rhythms |
| Main risks | Resource scrambles trigger rushed deployments; research funding crowded out by short-term applications |
| Evidence raising weight | Multiple quarters of concrete attribution of actual resource constraints and capability-R&D delays |
| Evidence lowering weight | Effective compute grows steadily; resource efficiency improvements outpace new demand |
Constraints may slow growth — or push new technology choices; they cannot be mechanically written as "power prevents ASI."
Scenario D: A New Paradigm Lowers Existing Bottlenecks
| Dimension | Analysis |
|---|---|
| Premise | Significant new methods in learning, memory, world representation, or validation |
| Leading indicators | Large improvements in new domains and long tasks at the same budgets; independent reproduction and clear ablations |
| Possible outcomes | Large model scale no longer the main explanatory variable; system structures reorganize |
| Major risks | Packaging special-task gains as a general paradigm; short-term successes hard to scale |
| Evidence raising weight | Sustained gains across multiple tasks, scales, and institutions |
| Evidence lowering weight | Advantages depend on special data, excessive budgets, or human curation |
DeepMind's history contains multiple coexisting research approaches, providing room for exploration; undisclosed routes should not be invented from this.
Scenario E: Specialized Scientific Superintelligence Before General Social Autonomous Intelligence
| Dimension | Analysis |
|---|---|
| Premise | Program, math, and some experimental sciences have validation structures better suited to automated search |
| Leading indicators | Research-side effective discovery rates and confirmation speed rise, while ordinary open tasks still frequently fail |
| Possible outcomes | Some research systems surpass large human teams, but everyday complex agents remain unreliable |
| Main risks | Expanded dual-use scientific capability; over-concentration of knowledge and resource control |
| Evidence raising weight | More externally reproduced important research contributions with limited cross-domain transfer |
| Evidence lowering weight | Research progress no longer depends on special setups; general goal management matures simultaneously |
"Specialized scientific superintelligence" here is a domain concept, not the broad ASI defined in this report.
Key Turning Points Between Scenarios
What deserves the most attention is not a particular year, but three changes:
- From "AI improves human-set plans" to "AI identifies and restructures research goals and evaluation methods";
- From "local results usable" to "additional gains in cross-generation capability growth attributable";
- From "digital-environment success" to "long-term reliable real-world action that can be supervised and halted."
Each change needs new evidence; product naming does not decide it.
XIV. Quarterly Leading-Indicator Dashboard
Principles: same task distribution, same success definition, complete denominators, recorded investment, preserved uncertainty. Official and independent results in separate columns; undisclosed items marked "unknown," not zero.
14.1 Twenty-Four Monitoring Indicators
| # | Indicator | Suggested definition and records | Cutoff-date baseline status | Quarterly update focus |
|---|---|---|---|---|
| 01 | Cross-domain reliable capability | Success rates by domain on hidden new tasks; avoid reporting only total averages | Multi-domain official evidence; unified independent baselines insufficient | Whether the weakest domains improve |
| 02 | 50% task horizon | Human-expert completion time corresponding to the 50% success threshold, with intervals | METR methods usable; no untested values for Argon | Whether task sets and evaluation configurations changed |
| 03 | High-reliability task horizon | Prefer 80%; higher reliability needs enough data | Latest systems' public values insufficient | Whether failures concentrate in a few types |
| 04 | Human takeover burden | Takeovers per completed task, in counts and minutes | Unknown | Whether it falls, not just whether runs get longer |
| 05 | End-to-end economic success rate | Tasks completed and passing real acceptance / all tasks | Unknown | Including tool and labor costs |
| 06 | Effective memory | State updates, forgetting, conflicts, and recovery tests | No unified public baseline | Correct retention rate after task changes |
| 07 | Multi-agent net gain | Extra effective results versus the best same-cost single-agent control | No public unified ablations | Duplication and communication costs |
| 08 | Controllable concurrency scale | Concurrency under fixed failure, cost, and risk caps | Unknown | Net output and permission management |
| 09 | AI share of R&D work | Accepted contributions by coding, planning, training, evaluation | Unknown | Don't use lines of code as a substitute |
| 10 | Complete R&D cycle | Calendar time from problem proposal to independent confirmation and deployment | Unknown | Matching task types and resource control |
| 11 | R&D capability gain per unit | Capability gains per unit of compute, capital, and research labor | Unknown | Same-definition trends |
| 12 | AI improvement adoption rate | Deployed improvements / all candidates, stratified by importance | Instances exist; no complete denominator | Failures and rollbacks disclosed too |
| 13 | Cross-generation feedback depth | Consecutive improvement rounds with source records | Disclosed for local training; insufficient for general R&D | Whether the next round improves improvement capability |
| 14 | New-knowledge confirmation rate | Externally reproduced or strictly tested research conclusions / all proposals | Collaborative validation instances exist | More fully independent reproductions |
| 15 | Discovery confirmation time | Proposal to validation, including queues, failures, and human work | Unknown | Don't count only generation time |
| 16 | Research methodology errors | Severe methodological flaws, hallucinations, and retraction rates | Preprint sample results exist | New data and external review |
| 17 | Experiment autonomy level | Human participation listed separately for design, execution, analysis, modification | Human links still present | Whether less human involvement harms reliability |
| 18 | World-model calibration | Predicted-vs-measured state errors, rare-event coverage | No cross-task independent baselines | Don't substitute visual scores |
| 19 | Simulation-to-reality gains | Additional on-site success at fixed real data | No unified public controls | Four-group training comparisons |
| 20 | Physical task reliability | Real new-environment success, takeovers, damage, and recovery | Official task performance uneven | Long tasks and continuous field data |
| 21 | Effective compute | Actual throughput, utilization, failures, and available hours | Hardware disclosures many; research allocation unknown | Separate plans, commissioning, and actual use |
| 22 | Energy and capital costs | Complete costs per effective task and R&D round | Group and global data; projects unknown | Grid connection, depreciation, financing, and demand |
| 23 | Control effectiveness | Monitoring coverage, attack recall, blocking, response, and escapes | Framework exists; complete real tests insufficient | Independent adversarial and concurrency pressure |
| 24 | External verifiability | Data, versions, result denominators, and independent test access | Varies widely by project | Whether it expands rather than shrinks |
14.2 How to Use METR and Avoid Misreading It
The METR task-horizon page verified for this report shows its last update as May 8, 2026, using TH1.1. The horizon refers to the human expert's completion time for the corresponding task, not how long the AI ran continuously; tasks are mainly self-contained software, machine-learning, and security work. The page also notes the current task set is unreliable for measurements above 16 hours.32
METR researchers further note that the 50%-success horizon cannot directly serve as a reliable delegation boundary; real task complexity, context, and supervision costs change actual automation gains. That note is a researcher methodology note, reviewed below the level of a formal research report.35
So this report does not invent METR scores for Gemini 4 Argon, nor extrapolate a certain ASI date from historical curves. Quarterly updates must first confirm the new task set, models, agent configurations, and statistical intervals.
14.3 Minimum Record Table for Each Quarterly Update
Indicator number:
Observation cutoff:
Model/system version:
Task distribution and hidden tests:
Total attempts:
Success definition:
Compute, tool, experiment, and labor investment:
Official results:
Independent results:
Statistical intervals and failure types:
Comparable to last quarter:
Reasons for raising/lowering related scenarios:
What remains unknown:
Indicator changes should enter scenario judgments, not be combined into a pseudo-precise "ASI index." For example:
- If 09, 10, 11, and 13 improve together, raise Scenario B's weight;
- Only if 18, 19, and 20 improve together does physical transfer acceleration gain support;
- If 01 and 02 improve while 04 and 05 don't, capability hasn't yet become stable delegation;
- Capability growth with falling 23 requires lowering deployable-growth judgments;
- If 21 expands while 11 stalls, check marginal R&D returns.
These are this report's update rules [D], revisable as measurement methods improve.
XV. What It Means for Companies and Individuals
15.1 Companies: Start with Verifiable Work
This report recommends [D] prioritizing tasks where:
- Inputs and constraints can be specified;
- Final results can be independently accepted;
- Failures are recoverable;
- Permission boundaries are controllable;
- Success changes actual costs or delivery times.
Software testing, data processing, and some computational experiments usually build faster feedback loops; open research, field operations, and high-stakes decisions need longer validation chains. This is task-structure analysis, not a procurement recommendation for any product.
Pilots should keep strong human baselines, same-budget AI alternatives, and human-review hours. Counting only generation speed easily mistakes efficiency gains after shifting review burden elsewhere.
15.2 Research and Industry Institutions: Validation Resources May Grow Scarcer
If candidate generation expands rapidly, the relative value of experimental equipment, reliable data, metrology, external reproduction, and expert judgment of domain specialists may rise. [D]
Companies' advantages may come from:
- High-quality real data and negative results;
- Clear evaluation and acceptance;
- Automatically executable experiment interfaces;
- The ability to turn discoveries into production results;
- Processes that retain sources, audits, and failure records.
These resources can combine with different models. Strategy should not be built entirely on the assumption that one model generation leads forever.
15.3 Managers: Work Organization Will Change, but Responsibility Cannot Be Handed to Models Alone
Tasks should specify goals, resources, acceptance, and permissions; important decisions should keep an accountable owner. Cross-team agent deployments also need clarity on who handles conflicts, error propagation, and failure recovery.
More agents do not automatically reduce management costs. If humans are busy judging contradictory outputs from multiple agents, deployment may just convert execution burden into coordination burden.
15.4 Individuals: Improve Problem Definition and Result Validation
This report's judgment [D]: as code, literature organization, and preliminary plans get easier to obtain, judging whether a problem is worth doing, how to design acceptance, and how to interpret failures may matter more.
A useful practice is keeping an evidence chain in personal work: problems, sources, hypotheses, versions, tests, errors, and decisions. It helps judge whether AI actually expanded capability or merely increased output volume.
This report does not therefore declare any profession necessarily extinct. Career changes also depend on organizational adoption, responsibility, field conditions, and market demand — existing technical materials are insufficient for deterministic predictions.
XVI. Key Unknowns and Final Research Judgments
16.1 Unknowns That Should Not Be Filled with Stories
| Unknown question | Why it cannot be confirmed | Most valuable new evidence |
|---|---|---|
| How much of DeepMind's internal R&D is done by AI? | Complete labor hours and task denominators undisclosed | Audits by role with accepted results |
| Are Gemini generational R&D cycles significantly shortened by AI? | Cycles, resources, failures, and attribution insufficient | Multi-generation R&D data with matched inputs |
| Is there a complete autonomous model-R&D loop? | Tools and local instances cannot reconstruct the whole | Source chains of training, evaluation, deployment |
| Do scientific systems have unified long-term memory? | Open interfaces do not equal unified state | Cross-project long tasks and memory tests |
| Is multi-agent necessary? | No comprehensive same-budget ablations | Single-agent, search, and collaboration comparisons |
| Do generated worlds truly lower physical-data costs? | Demos and simulation scores insufficient | On-site results at fixed real data |
| Can control stay effective for stronger systems? | Frameworks and disclosures do not equal measured guarantees | Independent adversarial, coverage, and response data |
| Which research routes were strengthened or weakened by the reorg? | Appointments cannot substitute for budget and project data | Long-term research investment and output changes |
| How far apart are digital and physical superintelligence? | Technology, hardware, and institutions jointly decide | Continuous cross-domain real reliability |
| When will ASI arrive? | Definitions, measurement, and growth mechanisms still uncertain | Leading-indicator linkages, not vision dates |
16.2 Final Judgments
First, DeepMind has provided concrete evidence of "AI expanding verifiable search." Some results entered engineering and scientific collaboration pipelines; they cannot be treated as mere extensions of chat capability.
Second, an evidence gap remains between local closed loops and general recursive improvement. R&D participation, strategy self-improvement, agent-process optimization, and frontier-model R&D takeoff need to be counted separately.
Third, the systems explanation currently has more research value than single-model rankings. It can explain huge performance differences of the same model across environments, and why validation, action, and infrastructure became key. But it is not a proven unique ASI implementation path.
Fourth, research automation is most likely to deepen first in tasks with clear feedback, executability, and repeatability. Physical and biological validation is more expensive; knowledge confirmation will not automatically grow at text-generation speed.
Fifth, the future should observe capability, results, and control together. More generation capability alone cannot confirm effective research growth; more effective research alone cannot confirm deployable superintelligence formation.
This report's core analytical framework: candidate capability × search organization × validation quality, converted through real resources and control conditions into effective results; how many effective results re-raise the next round's R&D capability determines whether sustained acceleration appears. The multiplication sign expresses interdependence, not a fitted quantitative law.
The next quarterly update should most verify: whether continuous, attributable R&D-cycle data can be obtained; whether independent reproduction of research conclusions increases; whether simulated experience improves real physical reliability; and whether control effectiveness rises together with capability.
Sources and Verification Notes
All sources below were retrieved or verified online on October 7, 2026. Dates refer to first publication or stated version dates, not this report's inferred experiment-completion dates. Company-participated papers are marked B even after peer review, with review status noted separately. Some web pages' full text or attachments have access restrictions, explicitly marked; unread content is not used to supplement technical details.
This report's citations only extract facts relevant to its judgments — official predictions, product claims, or preprint scores are not treated as independent confirmation. News searches serve as cross-checking leads; internal R&D narratives are not constructed from them.
Version update rule: The next update retains this version's judgments and their grounds, changing status item by item as new evidence arrives; it does not let new product names overwrite old measurements, does not interpret missing information as zero, and does not let narrative completeness take priority over the unknown.
-
B | Google's formal organizational announcement. Pichai, Hassabis, The next chapter of our AI momentum, August 2026 reorganization; scraped text showed no specific publication date. Verified formal titles and responsibilities, Jeff Dean and Sanjay Ghemawat arrangements. Source. This is an official organizational disclosure; it does not prove undisclosed motives behind the adjustment. ↩↩↩
-
B | Google model release and security disclosure. Introducing Gemini 4 Argon, 2026-09-30. Source. Verified output limits, phased availability, engineering cases, and safeguard descriptions; performance and internal-use results are official disclosures. ↩↩↩
-
B | DeepMind research explanation. AlphaZero and MuZero, research introduction page, cutoff-date version. Source. Used for self-play, learning environment representations, and planning mechanisms — not to prove general intelligence achieved. ↩
-
B | DeepMind on cross-task training boundaries. Generally capable agents emerge from open-ended play, 2021. Source. Publicly states AlphaZero was trained separately per game; this report does not equate general algorithms with general trained strategies. ↩
-
B | Research-team preprint. Novikov et al., AlphaEvolve: A coding agent for scientific and algorithmic discovery, 2026-06-16, arXiv:2506.13131; full text verified. Paper | Full text read. Used for program evolution, specific kernel and overall training-gain figures. Not an independent deployment audit. ↩↩
-
B | DeepMind applied update. AlphaEvolve: How our Gemini-powered coding agent is scaling impact across fields, 2026-05-07. Source. Verified official disclosure of entry into compute systems and TPU design; not expanded into complete chip autonomous design. ↩
-
B | Google Research. Dorfman, Brenner, Accelerating scientific discovery with AI-powered Empirical Research Assistance, 2025-09-09, name updated 2026-04-29. Source. Verified scorable tasks, tree search, six task classes; the update date is not treated as a new experiment date. ↩
-
B | Peer-reviewed paper with company and academic collaborators. Gottweis et al., Accelerating scientific discovery with Co-Scientist, Nature, 2026-05-19, 655:487–496. Paper. Verified structure, collaborative experiments, and authors' stated limitations; peer review does not equal fully independent reproduction. ↩
-
B | Company and academic collaborator preprint. Schmidgall et al., Accelerating Scientific Research with Gemini in the Real-World, 2026-08-27, arXiv:2608.26701. Paper | Full text read. Verified human links, Agent_H, and paper-evaluation errors; preprint with limited task conditions. Sample error rates not generalized to all AI research. ↩↩
-
B | Research-team preprint. Feng et al., Towards Autonomous Mathematics Research, 2026-02-10, revised v3: 2026-03-06. Paper | Full text read. Verified Aletheia structure and the limitation of rare successful examples. ↩
-
B | DeepMind scientific-reasoning research introduction. Accelerating mathematical and scientific discovery with Gemini Deep Think, 2026-02-11. Source. Verified reasoning compute, agent gains, and research contribution levels; results are official evaluations. ↩↩
-
B | DeepMind research-team peer-reviewed paper. Jumper et al., Highly accurate protein structure prediction with AlphaFold, Nature, 2021-07-15. Paper. Verified abstract and method introduction this time; did not rely on unread attachment details. ↩
-
B | Peer-reviewed paper by DeepMind, Isomorphic, and other research teams. Abramson et al., Accurate structure prediction of biomolecular interactions with AlphaFold 3, Nature, 2024-05-08. Paper. Verified retrievable abstract; full page access restricted; used only for joint structure-prediction scope. ↩
-
A | External public research infrastructure. EMBL-EBI, AlphaFold Protein Structure Database, cutoff-date version. Database description. Used for external competition results and database accessibility; the database is maintained by partners; its existence does not validate every prediction. ↩
-
B | DeepMind. AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome, 2026-09-08. Source. Verified precompute scope, scores, interfaces, and clinical-use restrictions; partner instances are not 9 billion independent experiments. ↩
-
B | Isomorphic Labs. Building a new path to make medicines with AI, 2026-09-29. Source. Used for the compute-design–synthesis–experiment–development chain; does not prove clinical efficacy or approved drugs. ↩
-
B | Research-team preprint. Feng et al., Aletheia tackles FirstProof autonomously, 2026-02-24, revised v3: 2026-03-15. Paper | Full text read. Adopted the revised 6/10, majority expert opinion, preferred-answer designation, and correctness explanation; did not carry over higher older claims. ↩
-
B | Company research-team peer-reviewed paper. Merchant et al., Scaling deep learning for materials discovery, Nature, 2023-11-29. Paper. Used for candidates, DFT validation, active learning, and computational-stability figures. ↩
-
A | Independent scholars' peer-reviewed perspective. Cheetham, Seshadri, Artificial Intelligence Driving Materials Discovery? Perspective on the Article: Scaling Deep Learning for Materials Discovery, Chemistry of Materials, 2024-04-08, 36:3490–3495. Journal & DOI | Authors' archived full text. Authors declare no competing financial interests; their sample criticism is not an exhaustive check. ↩
-
B | DeepMind. Genie 3: A new frontier for world models, 2025-08-05. Source. Verified interactive visual capability and authors' limitations; does not prove general physical-simulation accuracy. ↩
-
B | DeepMind. SIMA 2: An agent that plays, reasons, and learns with you in virtual 3D worlds, 2025-11-13. Source. Verified local experience generation and self-improvement disclosure. Attached large PDF had restricted reading; not used for technical-detail inference. ↩
-
B | Waymo. The Waymo World Model: A new frontier for autonomous driving simulation, 2026-02-06. Source. Verified Genie-based domain adaptation, camera and LiDAR output; real road data and simulation validation kept separate. ↩
-
B | DeepMind. Gemini Robotics 2 brings whole body intelligence to robots, 2026-07-30. Source. Task numbers come from official public charts and corresponding notes; undisclosed sample denominators or confidence intervals not reconstructed. ↩
-
B | DeepMind model card. Gemini Robotics ER 2, 2026-07-30. Model card. Verified outputs, intended use, and limits; ER text reasoning does not equal complete physical control. ↩
-
B | DeepMind model card. Gemini Robotics On-Device 2, July 2026 version. Model card. Verified trusted-tester distribution, action outputs, out-of-distribution and safety-evaluation scope. ↩↩
-
B | DeepMind technical report. Gemini Robotics 2: Safety Evaluations, 2026-07-29. Report. Used for safety-evaluation types; semantic evaluations not generalized into all physical-system safety guarantees. ↩↩
-
B | DeepMind control roadmap. Shah, Flynn, Securing the future of AI agents, 2026-06-18. Source. Used for the insider-threat perspective, coverage/recall/response, and visible-reasoning limits. ↩
-
B | Google Cloud infrastructure disclosure. Inside the eighth-generation TPU: An architecture deep dive, 2026-04-22. Source. Official comparisons are conditional; this report does not treat peak improvements as actual R&D speedups. ↩↩
-
B | Company regulatory disclosure. Alphabet, Q2 2026 earnings, 2026-07-22, SEC Exhibit 99.1. Filing. Financial units are millions of USD, converted to hundred-millions in the text; group, quarterly, and cash measures kept separate. ↩
-
A | IEA. Key Questions on Energy and AI, 2026-04-16, Executive summary. Report. Global statistics, central scenarios, and supply forecasts stated separately; not treated as Google project measurements. ↩
-
B | DeepMind's formal safety framework. Frontier Safety Framework 3.1, 2026-04-17. PDF. Used for capability definitions, concern thresholds, and decision structures; definitions are not evidence of capability attainment. ↩
-
A | METR independent research. Task-Completion Time Horizons of Frontier AI Models, page last updated 2026-05-08 as verified, TH1.1. Data & methods page. This report did not infer any new-model scores from missing dynamic-chart values. ↩
-
Grade-A external research lead | Unreproduced preprint. Liu et al., Does Learning Protein Folding Generalize to Broader Reasoning?, 2026-09-30. Paper | Full text read. Authors' experimental result, used as a transfer-hypothesis lead; not equivalent to confirmed general-intelligence breakthrough. ↩
-
B | DeepMind Institute official platform statement. Legg, Manyika, Hassabis, Introducing the DeepMind Institute, 2026-09-16. Source. Verified platform responsibilities and author-view statements; not treated as an independent governance body. ↩
-
Grade-A methodology supplement | Independent-institution researcher note. Kwa, Clarifying limitations of time horizon, 2026-01-22. Source. Page states its review level is below formal research articles and does not necessarily represent all of METR's views. ↩