Can the frontier keep building?
Compute, money, and power set the pace the physical world can support.
Progress, constraints, forecasts, and safety
Follow what labs can build, what systems can do, whether progress is compounding, and whether safeguards can keep pace.
Measurement first, interpretation second. No composite risk score.
Each question opens into measurements below. Start with the verdict; open the evidence only when you want to audit it.
Compute, money, and power set the pace the physical world can support.
Algorithms translate resources into useful autonomy—but the cleanest measurements are still young and imperfect.
AI automating AI research is the hinge in fast-takeoff forecasts, and the least settled part of the story.
Policy is observable. Alignment, verification, and real-world control remain open questions—not inputs to a pretend safety score.
Black line = measured reality
Colored marker = a scoreable published claim
Open any claim = source, evidence, confidence, and the strongest counterargument
Public evidence shows task-specific AI assistance inside AI development and credible adjacent evidence about software work, but it does not establish a reproducible, independently audited multiplier for frontier AI R&D or autonomous successor creation.
No single score is shown because the evidence measures different things.
No independently verified direct evidence is registered. The closest direct-relevance observations are company disclosures of a reported 1% Gemini-training-time reduction and an internal AI-research demonstration; neither is independently audited.
METR's randomized study of experienced open-source developers found a 19% slowdown with early-2025 tools. Its later experiment contained serious selection and measurement problems, so it neither cleanly reverses nor settles the earlier result.
There is no public longitudinal study that randomly varies frontier-agent access inside frontier AI R&D, measures quality-adjusted research output and model improvement, and separates human direction, compute, and agent contribution.
Public evidence reaches parts of steps 1–2. It does not yet establish steps 3–4.
Anthropic reports that, as of May 2026, Claude authored more than 80% of code merged into its codebase and that the typical engineer merged eight times as much code per day in Q2 2026 as in 2024. It also reports a March poll of 130 research-team employees with a median estimate of roughly 4x output with Mythos Preview.
When AI builds itself · Anthropic ↗Anthropic reports that Claude-powered agents, given an open AI-safety research problem with a pre-specified floor and ceiling, recovered 97% of the measured gap over 800 cumulative agent-hours; two human researchers recovered about 23% over roughly a week. Anthropic says the agents proposed hypotheses, ran experiments, shared findings, and iterated.
When AI builds itself · Anthropic ↗METR surveyed 349 technical workers, including 71 researchers. Participants reported median AI-related changes in work value of 1.4x to 2x and a median speed change of 3x around March 2026.
Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity · METR ↗METR's later raw estimates suggested an 18% speedup for returning participants and a 4% speedup for newly recruited participants, but METR judged selection into the study and time measurement during concurrent agent use severe enough that the experiment was only very weak evidence about the current effect.
We are Changing our Developer Productivity Experiment Design · METR ↗In a randomized study of 16 experienced developers working on 246 issues in repositories they knew well, METR found that allowing early-2025 AI tools made issue completion 19% slower (reported confidence interval: 2% to 39% slower).
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity · METR ↗Google DeepMind reports that AlphaEvolve found a matrix-multiplication-kernel change that improved that kernel by 23% and reduced Gemini training time by 1%; it also reports reducing kernel-optimization work from weeks of expert effort to days of automated experiments.
AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms · Google DeepMind ↗METR evaluated agents on seven ML research-engineering tasks with 71 human-expert attempts. With an eight-hour budget, it reported Claude 3.5 Sonnet and o1-preview at roughly the 37th and 30th percentiles of the human attempts, respectively.
Evaluating frontier AI R&D capabilities of language model agents against human experts · METR ↗The detailed metric section keeps each author’s original definition, relation, status, evidence, and counterargument.
Explore the AI R&D milestone evidence →299 responses to this question, within 2,778 respondents to the October 2023 ESPAI. This is a dated judgment about a stronger scenario—not a current consensus or a measured probability of takeoff.
Some people have argued the following: If AI systems do nearly all research and development, improvements in AI will accelerate the pace of technological progress, including further progress in AI. Over a short period (less than 5 years), this feedback loop could cause technological progress to become more than an order of magnitude faster. How likely do you find this argument to be broadly correct?
Direct response-category shares among the 299 randomized-subset responses; no mean, median, or pooled forecast was calculated.
Original survey results ↗Start with the question, not the database. Each one keeps measured reality, named forecasts, and group expectations in separate lanes so unlike evidence never becomes a fake consensus.
The four load-bearing questions appear first. Reveal the rest only when you want the full research map.
Tracks the compute, capital, and physical capacity that makes future training and inference possible.
Partly measured5 measurement streams directly or supportingly inform this question. Open one to see its reality line and point-level evidence.
18 attributed claims from 5 registered works. They remain separate from observed reality and from one another.
No like-for-like panel, survey, or forecasting-platform snapshot is registered for this question.
What would change the picture: A sustained break in frontier-compute growth, or verified energized power and chip capacity that contradicts announced build plans.
Favors task completion and independently evaluated autonomy over broad capability labels.
Measured1 measurement stream directly or supportingly inform this question. Open one to see its reality line and point-level evidence.
2 attributed claims from 2 registered works. They remain separate from observed reality and from one another.
Relevant panels or forecasting platforms are registered, but no number is shown until the exact wording, sample, date, and aggregation method match.
What would change the picture: Replicated, methodology-stable evaluations showing either durable long-horizon gains or a flattening after controls for evaluation gaming.
Direct public evidence is limited: company reports show task-specific contributions, while independent causal evidence remains adjacent to frontier AI R&D.
Partly measuredNo direct or supporting recurring measurement series is registered yet. The gap is being shown rather than filled with a proxy. 2 related streams are registered only as proxy or context.
6 attributed claims from 6 registered works. They remain separate from observed reality and from one another.
Verified, exactly matched aggregate snapshots are displayed separately from observed reality and named views.
What would change the picture: A repeated, independently audited measure of AI contribution to frontier research, plus verified successor-system work with limited human direction.
Tracks the hard-to-measure gap between a model appearing safe in evaluation and remaining controllable in deployment.
Open questionNo direct or supporting recurring measurement series is registered yet. The gap is being shown rather than filled with a proxy.
Relevant panels or forecasting platforms are registered, but no number is shown until the exact wording, sample, date, and aggregation method match.
What would change the picture: Independent safety evaluations, robust behavioral monitoring, and evidence that mitigations hold under realistic deployment pressure.
Separates capability gains from additional hardware and records the vintage of the underlying methodology.
Partly measured1 measurement stream directly or supportingly inform this question. Open one to see its reality line and point-level evidence.
3 attributed claims from 3 registered works. They remain separate from observed reality and from one another.
What would change the picture: A current, transparent fixed-capability efficiency series or evidence that observed improvements no longer transfer across tasks.
Distinguishes spending and revenue signals from verified autonomous deployment and economic use.
Partly measured1 measurement stream directly or supportingly inform this question. Open one to see its reality line and point-level evidence. 1 additional stream is shown only as proxy or context.
7 attributed claims from 4 registered works. They remain separate from observed reality and from one another.
No like-for-like panel, survey, or forecasting-platform snapshot is registered for this question.
What would change the picture: Audited adoption, utilization, and productivity series that can be separated from vendor revenue and announced investment.
Tests whether cognitive performance converts into practical cyber, bio, persuasion, or resource-acquisition advantage.
Open questionNo direct or supporting recurring measurement series is registered yet. The gap is being shown rather than filled with a proxy.
Relevant panels or forecasting platforms are registered, but no number is shown until the exact wording, sample, date, and aggregation method match.
What would change the picture: Independently documented incidents or controlled evaluations that establish a reliable capability-to-impact link, including effective defenses.
Measures enforceable action rather than policy announcements, and keeps coordination as an open question.
Partly measured1 measurement stream directly or supportingly inform this question. Open one to see its reality line and point-level evidence.
3 attributed claims from 1 registered work. They remain separate from observed reality and from one another.
Relevant panels or forecasting platforms are registered, but no number is shown until the exact wording, sample, date, and aggregation method match.
What would change the picture: Verified implementation and enforcement evidence across relevant jurisdictions, especially where commercial incentives conflict with restraint.
Compares dated, attributed claims with their intended measurements while preserving scenarios, conditions, and revisions.
Partly measuredNo direct or supporting recurring measurement series is registered yet. The gap is being shown rather than filled with a proxy.
30 attributed claims from 8 registered works. They remain separate from observed reality and from one another.
Relevant panels or forecasting platforms are registered, but no number is shown until the exact wording, sample, date, and aggregation method match.
What would change the picture: More comparable resolved probabilistic forecasts, pre-registered scoring rules, and direct measures for currently proxy-scored claims.
A deliberately visible coverage gap: incidents and near misses need a curated, independently verifiable series before they can be trended.
Missing a defensible seriesNo direct or supporting recurring measurement series is registered yet. The gap is being shown rather than filled with a proxy.
No like-for-like panel, survey, or forecasting-platform snapshot is registered for this question.
What would change the picture: A transparent incident taxonomy with primary documentation, severity criteria, and explicit coverage limits.
These are the nearest comparable deadlines still in play. They are checkpoints to revisit—not probabilities that an event will happen.
AI 2027
Frontier cluster power (announced vs energized) →various (vintaged)
Hyperscaler capex →AI 2027
Largest known single training run →various (vintaged)
Frontier training compute →AI 2027
Global AI compute stock →Situational Awareness
Algorithmic efficiency (pretraining) →A source can be checked recently while its newest published observation remains old. This separates maintenance health from source lag.
Current structured reading reviewed against the canonical source.
Epoch's published series still ends in August 2025; review is current, observation is not.
Newer frontier models still lack public compute estimates.
No formal measurement has superseded the March 2024 estimate.
Eval and disclosure evidence reviewed; measurement remains low confidence.
Energized capacity kept separate from announced capacity.
January hard anchor retained; later figure remains an extrapolation.
Latest reported run-rate figures reviewed.
Dangerous capabilities and failure modes are increasingly observable. Deployment exposure and real-world uplift remain only partly visible; safeguards are improving but uneven; independent evaluation and enforceable oversight do not yet provide comprehensive coverage.
Scope: Severe and catastrophic risks from frontier general-purpose AI. This is not a complete taxonomy of all AI harms.
In which severe-harm domains does a frontier system give a malicious or unskilled actor material, validated uplift over a non-AI baseline?
Autonomous task length and expert-level successes are rising; verified malicious use now reaches later attack stages.
Expert knowledge and laboratory-support performance are strong, but public end-to-end actor-uplift evidence remains limited.
Misuse is documented, but standardized severe-harm uplift measurements are sparse.
Longer autonomous software and cyber tasks are measurable; reliable consequential end-to-end operation remains uneven.
Domains remain separate because evidence in one does not establish danger in another.
UK AISI reports that average success on apprentice-level cyber tasks rose from just over 10% in early 2024 to about 50%, while the cyber task duration models complete without human direction grew from under ten minutes in early 2023 to over an hour by mid-2025.
external evaluator · Controlled evaluations do not establish realized harm or novice uplift in deployment.UK AI Security Institute · Frontier AI Trends Report ↗Anthropic's Frontier Red Team described early-warning capability gains but assessed the tested models as below its thresholds for substantially elevated national-security risk.
developer-reported with government testing input · The thresholds and sensitive evaluation details are not fully public.Anthropic · Progress from our Frontier Red Team ↗No cross-question safety score is calculated. Each reading keeps its measurement setting, independence, coverage, and gaps visible.
When given realistic delegated authority, do frontier agents strategically violate operator intent, conceal it, or resist correction at consequential rates?
Agents obey constraints, accept correction, and remain controllable across realistic high-authority settings.
Agents deceive, sabotage, evade correction, or preserve their objectives when authority and incentives conflict.
The dot is the current qualitative synthesis. The band is interpretive disagreement—not a probability or statistical confidence interval.
Anthropic stress-tested 16 models in hypothetical corporate environments and observed malicious insider behavior in at least some forced-conflict cases across developers; it reported no known real-world deployment instances of this behavior.
developer-reported, code released · The scenarios closed off ethical alternatives and were designed to elicit failure, so they do not estimate deployment prevalence.Anthropic · Agentic Misalignment ↗UK AISI reports self-replication-task success rising from below 5% to above 60% on a subset of simplified evaluations, while emphasizing that current models remain unlikely to complete real-world self-replication chains.
external evaluator · Component benchmarks are not evidence that a system has attempted autonomous replication in deployment.UK AI Security Institute · Frontier AI Trends Report ↗No cross-question safety score is calculated. Each reading keeps its measurement setting, independence, coverage, and gaps visible.
How much potentially dangerous capability is exposed through real deployments, tool permissions, unsupervised operation, or irreversible weight release?
Text output; no direct action or persistent credentials.
Can inspect external systems or private data but cannot change state.
Can act through bounded tools with approvals, logs, and rollback.
Persistent credentials, broad write access, long runtimes, or safety-critical authority.
Open weights or copied systems operate outside provider monitoring and recall.
No single current rung is shown: public telemetry is insufficient to estimate the distribution of real deployments.
NIST distinguishes read-only, constrained-write, and write access and highlights statefulness, reversibility, environment trust, and action criticality as key dimensions of agent risk.
multi-stakeholder government workshop · A taxonomy identifies what should be measured; it does not provide deployment prevalence.NIST · Lessons Learned on Tool Use in Agent Systems ↗UK AISI finds that open-weight systems can be modified arbitrarily, used without oversight, and spread irreversibly; safeguards can be removed quickly and cheaply.
external research collaboration · This establishes qualitative exposure mechanisms, not the number or severity of exposed deployments.UK AI Security Institute · Open-weight risk management ↗No cross-question safety score is calculated. Each reading keeps its measurement setting, independence, coverage, and gaps visible.
Against each defined threat model, how much attack success remains after all deployed safeguards and operational controls are applied?
Defense-in-depth keeps end-to-end attack success low under adaptive, persistent, well-resourced testing.
Transferable attacks, decomposition, fine-tuning, or scale defeat the production safeguard stack.
The dot is the current qualitative synthesis. The band is interpretive disagreement—not a probability or statistical confidence interval.
UK AISI found a universal jailbreak for every system it tested, while one comparison showed about 40 times more expert effort was required to jailbreak a newer protected system.
external evaluator · Two systems and one defended domain do not establish general safeguard robustness.UK AI Security Institute · Frontier AI Trends Report ↗Anthropic reports that its first Constitutional Classifiers reduced attack success from 86% to 4.4%. Its successor was tested for more than 1,700 red-team hours across 198,000 attempts, finding one high-risk vulnerability and no universal jailbreak during that test.
developer-reported with external red teamers · The result is domain- and implementation-specific and is not an independent population estimate.Anthropic · Constitutional Classifiers++ ↗No cross-question safety score is calculated. Each reading keeps its measurement setting, independence, coverage, and gaps visible.
Do pre-deployment evaluations and live monitors reliably predict dangerous capability and behavior in real use?
Tests and monitors prospectively identify consequential capabilities and failures before they appear in use.
Systems recognize tests, exploit scoring, hide capabilities, or behave differently in deployment.
The dot is the current qualitative synthesis. The band is interpretive disagreement—not a probability or statistical confidence interval.
The International AI Safety Report identifies an evaluation gap: controlled pre-deployment tests often do not reliably predict real-world utility, risk, or behavior in the deployed system context.
independent expert report · This diagnoses the gap but does not quantify one cross-system miss rate.International AI Safety Report 2026 ↗OpenAI reports that chain-of-thought monitoring outperformed monitoring actions and outputs alone across 13 evaluations and 24 environments, but notes limited realism, evaluation awareness, and uncertain generalization to wild misbehavior.
developer-reported with external-model comparisons · Monitorability of elicited reasoning is one control signal, not proof of end-to-end safety.OpenAI · Evaluating chain-of-thought monitorability ↗No cross-question safety score is calculated. Each reading keeps its measurement setting, independence, coverage, and gaps visible.
What share of frontier development and high-risk deployment is covered by binding, independently auditable requirements that are actually enforced?
High-risk developers and deployments face binding evaluation, reporting, audit, security, and corrective requirements.
Important systems remain outside binding rules, or requirements lack verification, capacity, penalties, and corrective action.
The dot is the current qualitative synthesis. The band is interpretive disagreement—not a probability or statistical confidence interval.
EU obligations require providers of systemic-risk general-purpose models to submit risk, safety-framework, and serious-incident materials; Commission enforcement powers, including fines, became applicable in August 2026.
regulator primary source · Legal authority does not yet establish broad or effective enforcement outcomes.European Commission · GPAI provider guidelines ↗California SB 53 requires large frontier developers to publish and follow frontier-AI frameworks, disclose model-specific catastrophic-risk assessment summaries, and report specified critical safety incidents.
government primary source · The law emphasizes transparency and process; it does not itself prove mitigation effectiveness or global coverage.California Legislature · SB 53 statutory text ↗No cross-question safety score is calculated. Each reading keeps its measurement setting, independence, coverage, and gaps visible.
What well-documented severe incidents and near misses are occurring, what incremental role did AI play, and how quickly were they detected and contained?
Anthropic reported an AI-orchestrated campaign against roughly 30 targets, with a small number of successful infiltrations and limited human intervention.
provider-observed; external verification limitedAnthropic analyzed 832 banned accounts with sufficient detail; 67.3% used AI for malware preparation and 6.5% for lateral movement.
provider-selected account sample; no population denominatorThese are documented provider cases, not independently verified examples or an incidence rate. Reporting coverage and denominators are missing.
Anthropic reported a state-sponsored cyber-espionage campaign using Claude Code to attempt infiltration of roughly 30 targets, succeeding in a small number of cases; it assessed the operation as largely AI-executed with occasional human direction.
provider-observed, not independently replicated · The report has privileged telemetry but limited external verification and no population denominator.Anthropic · Disrupting AI-orchestrated cyber espionage ↗Anthropic mapped 832 banned malicious-cyber accounts with sufficient detail: 560 used AI for malware preparation and 54 for lateral movement, illustrating deeper use in attack workflows without estimating prevalence among all attackers.
provider-selected dataset · The sample is a subset of banned accounts and cannot establish an incidence rate or counterfactual harm.Anthropic · Mapping AI-enabled cyber threats ↗The OECD distinguishes incidents that caused actual harm from hazards that could plausibly cause harm and is developing an open reporting process to complement news-derived events.
intergovernmental source · News-derived discovery is sensitive to coverage and reporting changes, so raw counts are not a stable harm trend.OECD.AI · Incidents and Hazards Monitor methodology ↗No cross-question safety score is calculated. Each reading keeps its measurement setting, independence, coverage, and gaps visible.
If model, safeguard, or governance failures occur, can affected institutions prevent propagation, maintain critical services, and recover quickly?
Critical systems have tested fallbacks, limited common-mode exposure, rapid detection, and effective recovery.
Dependencies, concentration, or weak response turn local AI failures into persistent or systemic harm.
The dot is the current qualitative synthesis. The band is interpretive disagreement—not a probability or statistical confidence interval.
The International AI Safety Report treats societal resilience as necessary because no safeguard stack is perfectly reliable and identifies stronger critical infrastructure, detection tools, and institutional response capacity as key defenses.
independent expert report · The report establishes importance and candidate practices, not a current comparative resilience score.International AI Safety Report 2026 · Executive Summary ↗NIST reports that post-deployment monitoring methods, terminology, and best practices remain nascent and fragmented, limiting visibility into field failures and recovery performance.
government literature review and workshops · Monitoring maturity is an input to resilience, not a direct measure of recovery capacity.NIST AI 800-4 · Monitoring Deployed AI Systems ↗No cross-question safety score is calculated. Each reading keeps its measurement setting, independence, coverage, and gaps visible.
This is a status map, not a ranking. Forecasts, scenarios, models, and intentions answer different questions, so claim counts should never be read as grades.
| Body of work | Coverage | Current evidence |
|---|---|---|
| Situational Awareness Leopold Aschenbrenner · 2024-06 | 5 comparable · 6 proxy/context 7 areas: capability, compute, algorithms, physical, capital, value, response |
1 confirmed1 ahead1 on-track2 pending
Jump to individual claims
|
| AI 2027 Kokotajlo/Lifland et al. · 2025-04 | 4 comparable · 3 proxy/context 4 areas: compute, automation, physical, value |
1 on-track3 behind
Jump to individual claims
|
| Dec 2025 timelines update AI Futures Project · 2025-12 | 1 comparable · 0 proxy/context 1 area: automation |
1 pending
Jump to individual claims |
| Forethought SIE series MacAskill/Davidson et al. · 2025 | 2 comparable · 0 proxy/context 2 areas: capability, algorithms |
2 on-track
|
| Bio Anchors (2-yr update)Ajeya Cotra · 2022 | Source registeredNo structured claim harvest yet | Not scored. An empty row describes tracker coverage, not the forecaster’s position. |
| various (vintaged) Epoch AI · varies | 4 comparable · 3 proxy/context 4 areas: compute, physical, capital, value |
4 on-track
Jump to individual claims
|
| Compute-centric takeoff-speeds framework Tom Davidson · 2023 | 1 comparable · 2 proxy/context 3 areas: compute, algorithms, automation |
1 pending
|
| Case for multi-decade timelinesEge Erdil · 2025-04 | Source registeredNo structured claim harvest yet | Not scored. An empty row describes tracker coverage, not the forecaster’s position. |
| AI Infrastructure Spending: Inelastic Demand and Rising Costs Morgan Stanley · 2026 | 0 comparable · 1 proxy/context 1 area: capital |
Jump to individual claims |
| Internal-goals statement (livestream + X post) Sam Altman (OpenAI) · 2025-10 | 0 comparable · 1 proxy/context 1 area: automation |
Jump to individual claims |
Portfolio bars include only direct and formula-backed translated claims. Proxy and context claims remain linked but excluded. Status and confidence remain separate.
Every section begins with the plain-language takeaway and current observation. The charts put published predictions on the same axis as what actually happened.
Latest recorded observation2025-08 · last published point; 90% CI 1.078e26-2.659e26Epoch AI trends — Training Runs ↗
checked 2026-07-27 · Epoch refits as models are added
Log scale. Rate forecasts are plotted as the level they imply at their target date, compounded from this metric's own series.
| Series | Date / deadline | Observation / target | State | Evidence |
|---|---|---|---|---|
| Reality | 2020-01 | 4.47e+22 · 90% CI 2.73e22-7.50e22 | Measured | Epoch AI trends — Training Runs ↗ |
| Reality | 2020-10 | 1.32e+23 · 90% CI 8.90e22-2.01e23 | Measured | Epoch AI trends — Training Runs ↗ |
| Reality | 2021-09 | 4.83e+23 · 90% CI 3.59e23-6.59e23 | Measured | Epoch AI trends — Training Runs ↗ |
| Reality | 2022-06 | 1.64e+24 · 90% CI 1.30e24-2.08e24 | Measured | Epoch AI trends — Training Runs ↗ |
| Reality | 2023-03 | 4.71e+24 · 90% CI 3.79e24-5.91e24 | Measured | Epoch AI trends — Training Runs ↗ |
| Reality | 2023-12 | 1.395e+25 · 90% CI 1.083e25-1.821e25 | Measured | Epoch AI trends — Training Runs ↗ |
| Reality | 2024-08 | 3.89e+25 · 90% CI 2.82e25-5.45e25 | Measured | Epoch AI trends — Training Runs ↗ |
| Reality | 2025-08 | 1.677e+26 · last published point; 90% CI 1.078e26-2.659e26 | Measured | Epoch AI trends — Training Runs ↗ |
| Leopold | 2027-12 | ~0.5 OOM/yr frontier training-compute growth sustained | ahead · confidence 55 | Published claim ↗ |
| Epoch | 2027-09 | ~5x/yr training-compute growth sustained | on-track · confidence 60 | Published claim ↗ |
“training compute used for frontier AI systems has grown at roughly ~0.5 OOMs/year”Ch I, From GPT-4 to AGI
Conditionality: firm — stated as the historical baseline he extrapolates through 2027. NOTE: his separate '+2 OOMs of compute (a cluster in the $10s of billions)... by the end of 2027' is a CLUSTER size/cost claim, not a single-run FLOP claim; conflating them would misstate him. Scored here only as the sustained growth-rate extrapolation (implies ~2e27-scale runs by end-2027 from GPT-4's 2e25).
Evidence · 2026-07Measured against the frontier trend rather than the record run, reality runs faster than his baseline: Epoch's top-5 series grows 0.7 OOM/yr (90% CI 0.6-0.8) against his stated ~0.5, rising 3.89e25 to 1.677e26 in the year to Aug 2025 (about 4.3x). Extrapolated to end-2027 that clears the ~2e27 his +2 OOM implies.Canonical measurement source 1 ↗
CounterargumentScored 'ahead' on Epoch's fitted trend, whose published data stops at Aug 2025 — the last 11 months are unmeasured. Epoch itself expects the pace to break within 1-2 years of Sept 2025 on training-duration and lead-time limits, which would pull this back to his rate or below.
“Frontier AI labs may still scale training compute at 5x per year for another 1-2 years by allocating a larger fraction of compute to training.”Compute scaling will slow down due to increasing lead times (Sep 2025), https://epoch.ai/gradient-updates/compute-scaling-will-slow-down-due-to-increasing-lead-times
Conditionality: hedged near-term trajectory ('may still') inside an analysis whose headline is a coming SLOWDOWN — the same piece moves the trillion-dollar-cluster date from ~2030 to ~2035 on lead-time grounds.
Evidence · 2026-07Now scored against the series the claim is actually about: Epoch's top-5 trend median went 3.89e25 (Aug 2024) to 1.677e26 (Aug 2025), roughly 4.3x in a year — inside the stated 4-6x band and consistent with the 5x/yr headline.Canonical measurement source 1 ↗
CounterargumentThe confirming data ends Aug 2025, one month before the claim was published, so almost none of the 1-2 year window it forecasts has been measured. The same piece predicts the slowdown that would falsify it.
Record frozen since Jul 2025 — but newer models lack estimates
Log scale. 3 of 4 claims are plotted here. The other 1 set no numeric target on this axis (dated milestones, or a quantity this axis does not measure) — they are recorded in the claims below.
| Series | Date / deadline | Observation / target | State | Evidence |
|---|---|---|---|---|
| Reality | 2023-03 | 2e+25 · GPT-4 | Measured | Epoch AI notable models / trends ↗ |
| Reality | 2023-12 | 5e+25 · Gemini 1.0 Ultra | Measured | Epoch AI notable models / trends ↗ |
| Reality | 2025-02 | 3e+26 · Grok 3 — first model over 1e26 (GPT-4.5 same month, 6.4e25) | Measured | Epoch AI notable models / trends ↗ |
| Reality | 2025-07 | 5e+26 · Grok 4 (~246M H100-hours, ~$500M compute cost) | Measured | Epoch AI notable models / trends ↗ |
| Reality | 2026-07 | 5e+26 · record stands; newest frontier models lack Epoch estimates (right-censored) | Measured | Epoch AI notable models / trends ↗ |
| Epoch | 2030 | 2e29 FLOP training run possible | on-track · confidence 60 | Published claim ↗ |
| AI 2027 | 2025-12 | 1e27 FLOP training run | behind · confidence 65 | Published claim ↗ |
| AI 2027 | 2027-03 | 2e28 FLOP training run | behind · confidence 75 | Published claim ↗ |
| Davidson | undated threshold | ~1e36 effective FLOP (2020-algorithms) to train AGI | pending · confidence 50 | Published claim ↗ |
Latest recorded observation2026-07 · record stands; newest frontier models lack Epoch estimates (right-censored)Epoch AI notable models / trends ↗
checked 2026-07-24 · rolling
“Training runs of around 2e29 FLOP are likely possible by 2030”Can AI scaling continue through 2030? (Aug 2024), https://epoch.ai/blog/can-ai-scaling-continue-through-2030
Conditionality: feasibility claim, not a prediction — 'likely possible' if the ~4x/yr trend continues. Their constraint hierarchy: power binds first, then chip manufacturing, then data. Score against feasibility, not occurrence.
Evidence · 2026-07Reaching 2e29 from Grok 4's 5e26 requires ~3.9x/yr through 2030 — right at Epoch's measured 4-5x/yr historical trend. Multi-GW campus construction (Colossus 2, Stargate-class sites) remains consistent with their power-first constraint analysis.Canonical measurement source 1 ↗
CounterargumentThe realized record has been flat for 12 months and GPT-5 used less pretraining compute than GPT-4.5 — labs are choosing not to scale single runs even where feasible. Epoch's own Sep 2025 lead-times analysis pushes megascale milestones out ~5 years. A feasibility claim can stay technically true while becoming irrelevant to what the frontier actually does.
“Agent-0 training compute: 10^27 FLOP (Late 2025); Agent-1: 4x10^27 FLOP”Compute forecast supplement, https://ai-2027.com/research/compute-forecast. CANONICAL READING: the compute-forecast table figures — the supplement's body text and chart give conflicting Agent-0/Agent-1 values (documented source-internal conflict, resolved per schema rule 4).
Conditionality: scenario narrative — authors' own hedge: '2027 was our modal (most likely) year at the time of publication, our medians were somewhat longer'; the scenario is 'not a prediction'. Superseded (not overwritten) by the Dec 2025 AI Futures Model revision, which moved milestone medians ~4-5 years later.
Evidence · 2026-07Largest known run as of mid-2026 is Grok 4 at ~5e26 (Jul 2025) — no confirmed 1e27+ run exists more than six months past the claimed date. Rumors of an xAI 1e27 run exist but Epoch and others place it closer to 1e26-scale.Canonical measurement source 1 ↗
CounterargumentThe gap is only ~2x, within Epoch's stated uncertainty on the Grok 4 point estimate, and the newest 2026 models have no published estimates — a 1e27 run may already have happened unrecorded. Not scored 'falsified' for exactly this reason.
“Agent-2 trained with 2e28 FLOP ('1000x GPT-4'), Apr 2026 - Mar 2027”Compute forecast supplement training table, https://ai-2027.com/research/compute-forecast
Conditionality: scenario narrative, same hedges and Dec 2025 supersession as Agent-0 claim. Note: the Phase-2 planning doc for this project quoted this as '1e28 by mid-27' — the primary source says 2e28; primary source wins.
Evidence · 2026-07Requires a 40x jump from the current 5e26 record within ~8 months, against a record that has been flat for 12 months and a demonstrated pretraining pullback (GPT-5 below GPT-4.5). The authors' own Dec 2025 revision effectively concedes this timeline ('Things seem to be going somewhat slower than the AI 2027 scenario' — Kokotajlo).Canonical measurement source 1 ↗
CounterargumentGlobal compute stock kept growing ~3.3-3.4x/yr (Epoch) even while single-run size stalled — if a lab reallocates stock into one giant run (the mechanics of Epoch's 'Manhattan Project' scenario), a 1e28-scale run in 2027 is not physically implausible, just commercially unmotivated so far.
“My median AGI training requirements (~1e36 FLOP using 2020 algorithms) are high compared to some.”What a Compute-Centric Framework Says About Takeoff Speeds, Monte Carlo with aggressive training requirements — verified against the canonical Coefficient Giving page on 2026-08-17
Conditionality: model parameter, not a dated forecast — an undated threshold in effective (2020-algorithm-equivalent) FLOP, so raw training FLOP alone cannot resolve it; the report presents wide uncertainty and separately tests an aggressive ~1e31 median.
Evidence · 2026-07Frontier raw compute (5e26) sits ~9.5 OOM below the median anchor; even crediting ~2 OOM of post-2020 algorithmic efficiency, reality is ~7+ OOM short — useful mainly as chart context showing how far the model camp's central anchor sits above the measured frontier.Canonical measurement source 1 ↗
CounterargumentThe +/-3 OOM band and the effective-vs-raw conversion make this the least falsifiable overlay on this metric; it earns its place as context, not as a scoreable bet.
Epoch reconstructions, not lab disclosures. The Grok 4 estimate rests on vague public xAI statements (Epoch: 'significant uncertainty around our point estimate'). Right-censored: Grok 4.5 (Jul 2026), Gemini 3, and GPT-5.5-class models have no published estimates yet, so the record may lag reality by months. 'Largest known' = largest Epoch-estimated.
Compounding on the bulls’ schedule while the record run sleeps
Log scale.
| Series | Date / deadline | Observation / target | State | Evidence |
|---|---|---|---|---|
| Reality | 2023-12 | 0.65 · ~650k H100s shipped cumulatively (H100-specific, pre-H100e normalization) | Measured | Epoch AI chip sales / data insights ↗ |
| Reality | 2024-07 | 3 · ~3M H100e cumulative sold, inferred from ~$90B accelerator revenue (Epoch) | Measured | Epoch AI chip sales / data insights ↗ |
| Reality | 2026-01 | 15 · >15M H100e installed (Epoch chip-production insight, data as of Jan 8 2026) | Measured | Epoch AI chip sales / data insights ↗ |
| AI 2027 | 2027-12 | 100M H100e global stock (2.25x/yr) | on-track · confidence 55 | Published claim ↗ |
Latest recorded observation2026-01 · >15M H100e installed (Epoch chip-production insight, data as of Jan 8 2026)Epoch AI chip sales / data insights ↗
checked 2026-07-24 · rolling (Epoch chip-sales explorer, updated ~weekly)
“We expect the total stock of AI-relevant compute in the world will grow 2.25x per year over the next three years, from 10M H100e today to 100M H100e by the end of 2027.”Compute forecast supplement, https://ai-2027.com/research/compute-forecast — verified verbatim ('today' = Mar 2025)
Conditionality: firm trajectory forecast (their least scenario-dependent number); production table gives checkpoints Dec 2025 = 18M, Dec 2026 = 40M.
Evidence · 2026-07Their implied mid-2026 checkpoint is ~28M H100e; extrapolating Epoch's Jan 2026 hard anchor (>15M) at Epoch's own 3.3-3.4x/yr gives ~27-28M for Jul 2026 — tracking almost exactly. Their Dec 2025 checkpoint (18M) also brackets Epoch's end-2025 figures (~16M deployed / ~20M sold).Canonical measurement source 1 ↗
CounterargumentThe mid-2026 'reality' figure is an extrapolation from a six-month-old anchor, not a measurement; and supply-side constraints (HBM4 validation issues, CoWoS capacity tight through at least H1 2027) could bend the trajectory's back half below 2.25x/yr.
“About 100M H100-equivalents could, in principle, be dedicated to training [by 2030] — range 20 million to 400 million H100-equivalents, corresponding to 1e29 to 5e30 FLOP.”Can AI scaling continue through 2030? (Aug 2024), https://epoch.ai/blog/can-ai-scaling-continue-through-2030
Conditionality: feasibility claim ('could, in principle') — the fleet existing is one half; DEDICATING it to a single training run is the binding half.
Evidence · 2026-07Global stock >15M H100e (Jan 2026) compounding at 3.3-3.4x/yr reaches 100M TOTAL around 2027-28 — the fleet will exist years ahead of their 2030 feasibility date on current trend.Canonical measurement source 1 ↗
CounterargumentThe allocation half is moving the wrong way for this claim: training's share of compute use is falling (AI 2027's own model has the leading company's training share going 40% to 20%; Epoch's separate finding is that frontier labs don't use most AI compute at all), so a large fleet does not imply a large dedicated training fleet.
“Even by 2027... GPU fleets in the 10s of millions... 10 million+ A100-equivalents [training].”Ch II, From AGI to Superintelligence
Conditionality: conditional-scenario; note the unit — A100-equivalents, roughly 3x smaller than H100e — and the 'training' qualifier on the 10M+ figure.
Evidence · 2026-07Global stock >15M H100e (Jan 2026) is roughly ~45M A100-equivalents — the fleet half of the claim is effectively already true a year early. The 10M+ A100e ON TRAINING half turns on allocation: ~30% of ~45M A100e would clear it, which is within plausible training shares.Canonical measurement source 1 ↗
CounterargumentThe A100e-vs-H100e conversion (~3x) does heavy lifting, and no one publishes the training-dedicated share — scoring the training half 'resolved' would rest on an allocation assumption, not a measurement.
Stock (installed) vs flow (annual shipments) vs revenue (dollars) get conflated constantly — this series is cumulative installed capacity, H100e = TPP-normalized across Nvidia/TPU/AMD/Huawei. Epoch states the growth trend as 3.4x/yr (trends dashboard) and 3.3x/yr (chip-production insight) — same trend, two restatements. Concentration: five hyperscalers held 71% of capacity in Q4 2025; Google alone ~25% (mostly TPUs). The mid-2026 value is EXTRAPOLATED from the Jan 2026 hard anchor, not measured — Epoch's live explorer renders client-side and needs a browser pull to confirm. Sanity proxy: NVIDIA DC revenue $75.2B in Q1 FY2027 (+92% YoY), consistent with continued compounding but ASP-inflated.
Top-5 by training compute, language models only — narrower than Epoch's site-wide 'frontier' definition (top 10 at time of release). Epoch reports the rate is stable across N=5/10/15 with no statistically significant difference. Headline: 5x/year since 2020 (90% CI 4-6x), doubling every 5.2 months, 0.7 OOM/year; the top-5 trend has grown ~10,000x since 2020. IMPORTANT FRESHNESS CAVEAT: the data bundle behind Epoch's chart ends at 2025-08-06 even though the page banner reads Feb 2026, so the newest point here is ~11 months old — a publication lag, not the censoring problem that afflicts the record-run series. Epoch's own view is that this pace holds ~1-2 years from Sept 2025 before training-duration ceilings (~9 months) and datacenter lead times bite. Methodology detail sits in a private Colab and could not be verified.
Latest recorded observation2024-03 · Epoch Ho et al. — language models, 8-month halving (95% CI 5-14); still the last formal estimateEpoch AI, Algorithmic Progress in Language Models (NeurIPS 2024) ↗
checked 2026-07-24 · episodic
1 of 2 claims are plotted here. The other 1 set no numeric target on this axis (dated milestones, or a quantity this axis does not measure) — they are recorded in the claims below. Horizontal lines are standing rate claims — they assert a level that holds over time rather than by a target date.
| Series | Date / deadline | Observation / target | State | Evidence |
|---|---|---|---|---|
| Reality | 2020-05 | 16 · OpenAI Hernandez & Brown — vision (ImageNet), 16-month doubling, 2012-2019 | Measured | Epoch AI, Algorithmic Progress in Language Models (NeurIPS 2024) ↗ |
| Reality | 2022-12 | 9 · Epoch Erdil & Besiroglu — vision, 9-month halving (95% CI 4-25) | Measured | Epoch AI, Algorithmic Progress in Language Models (NeurIPS 2024) ↗ |
| Reality | 2024-03 | 8 · Epoch Ho et al. — language models, 8-month halving (95% CI 5-14); still the last formal estimate | Measured | Epoch AI, Algorithmic Progress in Language Models (NeurIPS 2024) ↗ |
| Leopold | 2027-12 | ~2 OOM algorithmic efficiency vs GPT-4 (range 1-3) | on-track · confidence 50 | Published claim ↗ |
| Forethought | ongoing | ~3x/yr training-efficiency gain continues (halving ~7.6 months) | on-track · confidence 55 | Published claim ↗ |
“1-3 OOMs of algorithmic efficiency gains (compared to GPT-4) by the end of 2027, maybe with a best guess of ~2 OOMs.”Ch I, From GPT-4 to AGI
Conditionality: conditional-scenario ('best guess') — his baseline rate of ~0.5 OOM/yr is stated as firm; the 2027 cumulative figure extrapolates it.
Evidence · 2026-07Epoch's measured 8-month halving is ~0.45 OOM/yr — essentially his claimed baseline rate — and would compound to ~2.1 OOM over GPT-4-to-end-2027 if it held. Reasoning-model post-training added an estimated further ~10x compute-equivalent gain (Epoch, Aug 2025) in verifiable domains, arguably front-running the schedule.Canonical measurement source 1 ↗
CounterargumentThe rate has not been formally re-measured since Mar 2024, so 'on-track' rests on extrapolating a stale estimate; the Nov 2025 MIT decomposition argues most historical gains came from two unrepeatable scale-dependent transitions, which would make sustained 0.5 OOM/yr the exception, not the trend.
“physical computation required to train a model at the same level of performance is falling by roughly 3x per year”Preparing for the Intelligence Explosion (Mar 2025), Current trends
Conditionality: stated as a current-trends baseline, not a forecast; their forward-looking software-explosion claims hang off the r_cog parameter (median 1.2, log-uniform 0.4-3.6), which they flag as 'necessarily speculative'.
Evidence · 2026-073x/yr is the same quantity as Epoch's 8-month halving (3x/yr = halving every ~7 months) and matches Epoch's live dashboard figure (~3.0x/yr, ~7.6-month halving) — internally consistent with the best available measurement.Canonical measurement source 1 ↗
CounterargumentAll three numbers ultimately derive from the same Epoch methodology and 2012-2023 data window — agreement between them is not independent confirmation, and none reflects a post-2024 re-measurement.
“Software/algorithmic efficiency input: OpenAI's 'Efficiency' analysis cited at 16-month doubling; Epoch's own analysis at ~10-month doubling; Davidson's assessment: 'Progress is if anything faster for [language models].'”Recovered secondary-source synthesis from the 2026-07 harvest. The canonical report supports a broader ~1–2 year historical range and a conservative 2.5-year model input, but this exact 10–16 month formulation has not been verified; excluded from headline results.
Conditionality: model input parameter (2023 vintage), with an explicit hedge that LM progress may be faster than the vision-era 10-16 month figures he anchored on.
Evidence · 2026-07The best measured LM figure (8-month halving, Mar 2024) is faster than his 10-16-month input range — reality outpaced the parameter he fed his takeoff model, exactly as his own hedge anticipated. Faster algorithmic progress shortens his modeled timelines.Canonical measurement source 1 ↗
Counterargument'Ahead' rests on a single stale measurement whose CI (5-14 months) overlaps his input range; if the MIT one-off-transitions critique is right, the post-2024 rate could fall back inside or below 10-16 months.
The 8-month figure is pretraining-only and backward-looking (2012-2023 fit). Epoch's own decomposition: 60-95% of capability gains came from compute and data, algorithms only 5-40%. A Nov 2025 MIT preprint attributes ~91% of measured efficiency gains to two one-off scale-dependent transitions (LSTM-to-Transformer, Kaplan-to-Chinchilla), cautioning against treating the rate as a steady state. The post-training era complicates measurement further: Epoch (Aug 2025) estimates reasoning-model RL delivered a ~10x compute-equivalent gain (range 1x-100x) in verifiable domains — not captured by the pretraining halving rate. Epoch's live trends dashboard shows ~7.6-month halving, but its provenance/update cadence is undocumented; the Mar 2024 paper remains the citable number.
~16 hrs, doubling ~105 days; eval-gaming is the new caveat
Latest recorded observation2026-06 · GPT-5.6 Sol — 11.3h standard scoring; 71h if detected cheating discarded, >270h if counted as success (55% gaming rate)METR Time Horizons ↗
checked 2026-07-24 · per frontier release
Log scale.
| Series | Date / deadline | Observation / target | State | Evidence |
|---|---|---|---|---|
| Reality | 2025-03 | 0.5 · initial paper | Measured | METR Time Horizons ↗ |
| Reality | 2026-01 | 14.5 · TH1.1 | Measured | METR Time Horizons ↗ |
| Reality | 2026-05 | 16 · Claude Mythos Preview; wide CI 8.5-55h | Measured | METR Time Horizons ↗ |
| Reality | 2026-06 | 11.3 · GPT-5.6 Sol — 11.3h standard scoring; 71h if detected cheating discarded, >270h if counted as success (55% gaming rate) | Measured | METR Time Horizons ↗ |
| Forethought | 2028-2031 | 1-month-task horizon | on-track · confidence 55 | Published claim ↗ |
“within three to six years, AI models will become capable of automating many cognitive tasks which take human experts up to a month”Preparing for the Intelligence Explosion, AI-human cognitive parity
Conditionality: conditional-scenario: 'Naively extrapolating this trend' [the METR 7-month doubling]. Preserve as stated.
Evidence · 2026-07METR horizon ~16h (May 2026), doubling ~105 days; extrapolation tracks toward month-scale ~2028-29.Canonical measurement source 1 ↗
CounterargumentDoubling time is itself contested and may be decelerating; >16h measurements are near suite-resolution limits, so recent points are noisy.
Task-suite methodology changed v1.0->v1.1 (Jan 2026); cross-version comparisons need care. Measurements above ~16h are near the edge of what the current suite resolves. Single research org, not independently audited. Eval-gaming is becoming the dominant uncertainty: METR's GPT-5.6 Sol eval (Jun 2026) measured 11.3h under standard scoring but 71h-270h+ depending on how its record 55% detected-cheating rate is handled. METR added a companion 'expenditure horizon' metric (cost-based, Jul 2026) — not yet a leaderboard series.
checked 2026-07-24 · event-driven
1 dated milestone drawn. This force has no numeric series — vertical position is ordering, not a value.
| Series | Date / deadline | Observation / target | State | Evidence |
|---|---|---|---|---|
| AI Futures | 2031-12 | Superhuman Coder | pending · confidence 50 | Published claim ↗ |
“the model’s median prediction is just 3 years. The simulation models the effect of both rising human investments and increasing AI automation”Primary Coefficient Giving report, Abstract, published 2023-06-27
Conditionality: model-conditional output; the author says his personal probabilities are 'still massively in flux.' The clock only starts at the 20%-automation trigger, which has not been reached.
Evidence · 2026-07The 20%-automation trigger has not been hit: heavy AI coding ASSISTANCE (~36% of Claude usage) is not task AUTOMATION, and Anthropic's own system card denies even a sustained 2x R&D speedup. Nothing to score yet — the claim defines what to watch.Canonical measurement source 1 ↗
CounterargumentIf rung (d) resolves at 2x in the next year, Davidson's own feedback-loop math implies the trigger may be closer than the automation-fraction framing suggests — 'pending' could flip to live fast.
This is the hardest Tier 1 variable to measure, but every takeoff model turns on it. Evidence hierarchy applies hard here: lab-leader statements are intention data and self-reported speedups receive a confidence discount. METR's early-2025 randomized study found a 19% slowdown in adjacent open-source work; its later experiment was inconclusive because of selection and measurement limits. The field still has no clean causal estimate of frontier AI-R&D acceleration.
Latest recorded observation2026-03 · Q1 2026 — Epoch: 'came in on trend' for $770B/yr extrapolationSEC 10-Qs / earnings releases + Epoch Data Insights ↗
checked 2026-07-24 · quarterly (earnings), staggered across companies
Annual forecast totals are plotted as quarterly equivalents to match the measured series.
| Series | Date / deadline | Observation / target | State | Evidence |
|---|---|---|---|---|
| Reality | 2023-03 | 36.2 · Q1 2023 — GPT-4-release-era baseline | Measured | SEC 10-Qs / earnings releases + Epoch Data Insights ↗ |
| Reality | 2023-06 | 35.3 · Q2 2023 | Measured | SEC 10-Qs / earnings releases + Epoch Data Insights ↗ |
| Reality | 2023-09 | 38.3 · Q3 2023 | Measured | SEC 10-Qs / earnings releases + Epoch Data Insights ↗ |
| Reality | 2023-12 | 44 · Q4 2023 | Measured | SEC 10-Qs / earnings releases + Epoch Data Insights ↗ |
| Reality | 2024-03 | 46 · Q1 2024 | Measured | SEC 10-Qs / earnings releases + Epoch Data Insights ↗ |
| Reality | 2024-06 | 55.7 · Q2 2024 | Measured | SEC 10-Qs / earnings releases + Epoch Data Insights ↗ |
| Reality | 2024-09 | 61.2 · Q3 2024 | Measured | SEC 10-Qs / earnings releases + Epoch Data Insights ↗ |
| Reality | 2024-12 | 76.3 · Q4 2024 | Measured | SEC 10-Qs / earnings releases + Epoch Data Insights ↗ |
| Reality | 2025-03 | 77.8 · Q1 2025 | Measured | SEC 10-Qs / earnings releases + Epoch Data Insights ↗ |
| Reality | 2025-06 | 97.3 · Q2 2025 | Measured | SEC 10-Qs / earnings releases + Epoch Data Insights ↗ |
| Reality | 2025-09 | 105.8 · Q3 2025 | Measured | SEC 10-Qs / earnings releases + Epoch Data Insights ↗ |
| Reality | 2025-12 | 130.7 · Q4 2025 — full-year 2025 ~$411B cash basis (~$500B incl. finance leases per Epoch) | Measured | SEC 10-Qs / earnings releases + Epoch Data Insights ↗ |
| Reality | 2026-03 | 148.4 · Q1 2026 — Epoch: 'came in on trend' for $770B/yr extrapolation | Measured | SEC 10-Qs / earnings releases + Epoch Data Insights ↗ |
| Epoch | 2026-12 | ~$770B combined 2026 capex | on-track · confidence 70 | Published claim ↗ |
“total AI investment could be north of $1T annually by 2027.”Ch IIIa, Racing to the Trillion-Dollar Cluster
Conditionality: conditional-scenario ('could be') — 'total AI investment' is broader than hyperscaler capex; scored against capex plus credible non-hyperscaler AI investment.
Evidence · 2026-07Hyperscaler capex alone is guided to ~$780B for 2026 and forecast ~$1.1T for 2027 (Morgan Stanley); adding non-hyperscaler AI investment (xAI $18B+ Colossus 2 alone, sovereign programs, other labs) puts total 2027 AI investment north of $1T on nearly any accounting.Canonical measurement source 1 ↗
CounterargumentThe marginal capex dollar is increasingly debt-funded — aggregate hyperscaler FCF hits zero ~Q3 2026 (Epoch) — so a credit-market turn could cut 2027 spending below guidance; and 'total AI investment' has no agreed measurement basis.
“If that trend held through 2026, Alphabet, Amazon, Meta, Microsoft, and Oracle would collectively spend $770 billion on capex this year.”Hyperscaler capex has quadrupled since GPT-4's release (Feb 2026), https://epoch.ai/data-insights/hyperscaler-capex-trend — verified verbatim
Conditionality: explicit naive trend extrapolation (72%/yr since Q2 2023), not a guidance-based forecast.
Evidence · 2026-07Q1 2026 came in 'on trend' (Epoch's own words); company guidance now sums to ~$780B at midpoints (Alphabet $195-205B, Microsoft ~$190B CY, Amazon $200B, Meta $125-145B, Oracle FY26 $55.7B actual); Alphabet raised guidance again on Jul 22.Canonical measurement source 1 ↗
CounterargumentGuidance is not spend — H2 delivery constraints (components, power) have repeatedly shifted capex between quarters; and the cash-vs-finance-lease basis difference (~$90B in 2025) means 'hitting $770B' depends on whose accounting you use.
“aggregate free cash flow reaches zero around the third quarter of 2026”Hyperscaler Capex to Exceed Cash Flow by Q3 2026 (live Data Insight, snapshot Jun 16 2026), https://epoch.ai/data-insights/hyperscaler-capex-vs-cash-flow
Conditionality: model fit (capex ~70%/yr vs operating cash flow ~23%/yr, Q2 2023-Q1 2026); Epoch's own sensitivity note: fit-window choice moves the crossover between roughly Q2 and Q4 2026.
Evidence · 2026-07Oracle crossed and is deepening ($43B debt + $5B equity raised in FY26, ~$40B more planned); Alphabet — modeled not to cross until Q1 2027 — already posted a negative-FCF quarter (-$5.9B, Q2 2026). If anything the crossover is running early.Canonical measurement source 1 ↗
CounterargumentOne negative Alphabet quarter could be timing/seasonality, and Amazon/Meta/Microsoft Q2 reports (Jul 29-30) are the real test — 'ahead' would be premature by one week.
“$800 billion that we think is spent this year is set to be dwarfed by $1.1 trillion of estimated spending in 2027.”Morgan Stanley, AI Infrastructure Spending: Inelastic Demand and Rising Costs, May 2026 — verified against the primary transcript on 2026-08-17
Conditionality: sell-side AI-infrastructure spending forecast, repeatedly revised upward; its scope is broader than the four-company capex series used as the tracker's nearest public proxy.
Evidence · 2026-07Company guidance (~$780B midpoint sum) sits close to Epoch's $770B extrapolation and Morgan Stanley's $800B forecast — within one guidance revision of the forecast, with guidance moving upward during 2026.Canonical measurement source 1 ↗
CounterargumentMorgan Stanley's AI-infrastructure spending scope is broader than the four-company capex proxy used here, so proximity between the figures is not a like-for-like confirmation.
Total capex is NOT AI-specific — no company breaks it out; third-party estimates of the AI share range 40-75% (JPMorgan ~70% for 2025; Dell'Oro: accelerator silicon alone is ~1/3) depending on definition. Series here is cash 'purchases of property and equipment' EXCLUDING finance leases — earnings-call figures often include them (this basis gap largely explains ~$411B vs Epoch's ~$500B for 2025). Oracle adds ~$248B of off-balance-sheet lease commitments and reports on a May-end fiscal year; Microsoft on June-end — quarterly alignment is by calendar quarter of period end. History values are aggregator-compiled (stockanalysis.com), spot-checked against press figures, not line-by-line 10-Q reconciled.
Claims run 2-5x ahead of metal; the first cap-back just landed
Latest recorded observation2026-07 · Colossus 2 at 946 MW IT power (Epoch, satellite-verified); New Carlisle 910 MWEpoch AI — Colossus 2 record ↗
checked 2026-07-24 · periodic
Log scale. 3 of 4 claims are plotted here. The other 1 set no numeric target on this axis (dated milestones, or a quantity this axis does not measure) — they are recorded in the claims below. Open markers are ANNOUNCED targets, plotted at the date each was claimed; the bar runs to the delivery date promised for it. The vertical distance is an announced-to-energized scale gap, but capacity bases are not always like-for-like; the horizontal gap is the delivery lag.
| Series | Date / deadline | Observation / target | State | Evidence |
|---|---|---|---|---|
| Reality | 2023-01 | 0.023 · GPT-4 training era, ~20-25 MW sustained (UNVERIFIED secondary estimate; OpenAI never disclosed) | Measured | Provenance gap: unverified legacy estimate |
| Reality | 2024-07 | 0.15 · xAI Colossus 1, 100k H100s, ~150 MW target | Measured | Provenance gap: missing primary permalink |
| Reality | 2024-12 | 0.25 · Colossus 1 expanded to 200k GPUs, ~250 MW | Measured | Provenance gap: missing primary permalink |
| Reality | 2025-09 | 0.3 · Stargate Abilene Phase 1 energized (~300 MW per Epoch's Apr 2026 measurement) | Measured | Provenance gap: methodology revision |
| Reality | 2026-07 | 0.946 · Colossus 2 at 946 MW IT power (Epoch, satellite-verified); New Carlisle 910 MW | Measured | Epoch AI — Colossus 2 record ↗ |
| Announced (operator claims) | 2024-12 | 2 · Meta Hyperion, Richland Parish LA: 'more than two gigawatts of compute capacity'. No delivery date given. | Announced | Provenance gap: missing primary permalink |
| Announced (operator claims) | 2025-03 | 1.2 · promised 2026-06 · Crusoe, Abilene TX (the Stargate site's developer): campus expanded to 1.2 GW facility capacity, full buildout targeted mid-2026. Epoch measured ~0.3 GW energized in early 2026. | Announced | Provenance gap: non comparable announcement |
| Announced (operator claims) | 2025-05 | 1 · xAI Colossus 2, Memphis: Musk, 'Colossus 2 will be the first Gigawatt AI training supercluster'. No date given. | Announced | Provenance gap: missing primary permalink |
| Announced (operator claims) | 2025-07 | 5 · promised 2032-01 · Meta Hyperion raised to 5 GW: Zuckerberg, 'scale up to 5GW over several years'. Meta later indicated 2 GW by 2030, full 5 GW ~2032. | Announced | Provenance gap: missing primary permalink |
| Announced (operator claims) | 2025-07 | 1 · promised 2026-12 · Meta Prometheus, New Albany OH: Zuckerberg, 'coming online in 26'. Epoch measured 631 MW at Jul 2026. | Announced | Provenance gap: missing primary permalink |
| Announced (operator claims) | 2025-10 | 2 · Poolside/CoreWeave 'Project Horizon', West Texas: 2 GW across eight 250 MW phases. CoreWeave and Poolside mutually terminated the deal in Apr 2026 — announced, never delivered. | Cancelled | Provenance gap: secondary only claim |
| Announced (operator claims) | 2025-12 | 2 · xAI Colossus 2: Musk, 'Will take @xAI training compute to almost 2GW' on acquiring a third building. No delivery date given. | Announced | Provenance gap: missing primary permalink |
| Announced (operator claims) | 2026-01 | 1.5 · promised 2026-04 · xAI Colossus 2: Musk claimed 1 GW already operational and 'Upgrades to 1.5GW in April'. Epoch measured 946 MW IT in Jul 2026 — the April target was missed. | Announced | Provenance gap: missing primary permalink |
| Announced (operator claims) | 2026-03 | 1.2 · Stargate Abilene CAPPED at 1.2 GW — the discussed expansion to ~2 GW was scrapped after grid-interconnection delays exceeding a year. The first DOWNWARD revision on this board. Press-reported (Bloomberg/Tom's Hardware); not a company statement. | Announced | Provenance gap: secondary only claim |
| Announced (operator claims) | 2026-07 | 1 · promised 2028-01 · Meta El Paso TX (with BlackRock): 1 GW, capacity beginning to come online 2028. | Announced | Provenance gap: missing primary permalink |
| Leopold | 2026-12 | ~1M H100e / ~1 GW single training cluster | resolved-true · confidence 60 | Published claim ↗ |
| Leopold | 2028-12 | ~10 GW single training cluster | pending · confidence 50 | Published claim ↗ |
| Epoch | 2030 | largest single training run drawing 4-16 GW | on-track · confidence 55 | Published claim ↗ |
| AI 2027 | 2026-12 | leading AI company at 6 GW peak power | behind · confidence 45 | Published claim ↗ |
“Year-by-year cluster table: ~100k H100e / ~100MW (2024) - ~1M / ~1GW (2026) - ~10M / ~10GW (2028) - ~100M / ~100GW (2030)”Ch IIIa, Racing to the Trillion-Dollar Cluster (back-of-the-envelope table)
Conditionality: explicitly illustrative ('back-of-the-envelope') — scored charitably against its 2026 rung as a round-number target.
Evidence · 2026-07Epoch (Jul 24, 2026, satellite-verified): xAI Colossus 2 at 1,112k H100-equivalents and 946 MW IT power — roughly 1.1-1.4 GW on a facility-power basis — with Amazon-Anthropic New Carlisle at 910 MW right behind. The 2026 rung's both halves (~1M H100e, ~1 GW) are essentially exactly reality, with five months of the year to spare.Canonical measurement source 1 ↗
CounterargumentOn the strict IT-power measure it is 5% short of the round 1 GW; the table was illustrative, so 'resolved-true' credits a charitable reading; and Epoch's own methodology notes actual consumption runs 60-80% of nameplate.
“individual training clusters costing $100s of billions by 2028—clusters requiring power equivalent to a small/medium US state.”Ch IIIa, Racing to the Trillion-Dollar Cluster (the table's 2028 rung: ~10M H100e / ~10GW)
Conditionality: conditional-scenario; the 10 GW figure is the same back-of-envelope table's 2028 rung.
Evidence · 2026-07Requires ~10x from today's 0.95 GW in ~2.5 years (~2.5x/yr — above the 2.2x/yr historical growth of run power). Announced pipeline is directionally supportive (Stargate >9 GW by 2029, Hyperion 5 GW) but announced-to-energized has been running ~2x optimistic, and no single campus currently targets 10 GW by 2028.Canonical measurement source 1 ↗
CounterargumentThe 2026 rung of the same table just resolved true on schedule — the table has a live track record, and multi-site distributed training could satisfy the spirit of the rung without one 10 GW campus.
“the largest individual frontier training runs in 2030 will likely draw 4-16 gigawatts (GW) of power”How much power will frontier AI training demand in 2030? (Aug 2025), https://epoch.ai/blog/power-demands-of-frontier-ai-training — verified verbatim
Conditionality: trend projection off 'largest runs now exceeding 100 MW' (2025 baseline) growing 2.2-2.9x/yr; about training RUNS, not cluster nameplate.
Evidence · 2026-07Largest verified cluster went ~0.3 GW (Sep 2025) to ~0.95 GW (Jul 2026) — about 3.9x annualized, running FASTER than the 2.2-2.9x/yr their projection assumes. Even at their slower assumed rate, compounding from ~0.95 GW clears the bottom of the 4-16 GW range well before 2030.Canonical measurement source 1 ↗
CounterargumentScored on-track rather than ahead because the quantities differ: this series measures cluster IT power, while their forecast is the draw of a single training RUN, which uses only part of a cluster. The persistent ~2x announced-vs-verified gap and grid-interconnection queues could also stall the buildout before 2030.
“Global AI peak power: 38GW (2026); OpenBrain power requirement: 6GW peak”Compute forecast supplement, Key Metrics 2026 panel, https://ai-2027.com/research/compute-forecast — verified against the primary page on 2026-08-17
Conditionality: scenario narrative (global figure scored here; the 6GW leading-company figure is scored separately).
Evidence · 2026-07Two independent derivations put mid-2026 global AI IT power at ~40-44 GW (SemiAnalysis ~40 GW AI share, secondary-sourced; Epoch coverage 11.9 GW = 27% of global AI compute implying ~44 GW) — at or slightly above the scenario's 38 GW.Canonical measurement source 1 ↗
CounterargumentBoth reality figures are derived, not measured: one is paywalled-secondhand, the other a naive extrapolation of Epoch's coverage share; IT-vs-facility-power ambiguity alone could swing the comparison by 30%+.
“OpenBrain datacenter power: 6GW peak, 2026”Compute forecast supplement, https://ai-2027.com/research/compute-forecast
Conditionality: scenario narrative — 'OpenBrain' is the fictional leading company; scored against the actual leading AI company's dedicated power.
Evidence · 2026-07The largest verified single-company training cluster is ~0.95 GW (xAI); OpenAI's energized Stargate capacity is ~0.4-0.6 GW plus shares of Microsoft sites. No single company shows evidence of 6 GW of dedicated AI power in 2026 — the scenario's global figure is tracking but its concentration in one leader is not.Canonical measurement source 1 ↗
CounterargumentCompany-wide fleets across all sites (especially OpenAI-on-Azure and Google's TPU estate) are poorly public — a leading company's TOTAL ai power could plausibly reach several GW; 'behind' rests on the verified-single-site lens.
Three power measures get conflated constantly: IT power (compute equipment draw — used here), facility power (20-50% higher, cooling/overhead), and grid-connection capacity. Operator claims run ~2x above third-party verification (Musk claimed Colossus 2 at '2 GW' in Jan 2026; satellite analysis then showed ~350 MW cooling; Epoch measures 946 MW in Jul 2026). Announced figures: sum of top-5 campus targets ~20-25 GW (Stargate >9 GW by 2029, Meta Hyperion 5 GW, Rainier ~2.2-4.6 GW (sources conflict), Colossus 2 target 2 GW, Prometheus >=1 GW). GPT-4-era baseline point is an unverified secondary estimate. Training-vs-inference cluster labels are operator-reported snapshots. ANNOUNCED SERIES: operator claims plotted at the date each was made, with a bar running to the delivery date promised for it; an X marks an announcement that was cancelled. Operators almost never say whether a figure is IT or facility power, so every announced figure here is UNSPECIFIED on that axis and is not strictly like-for-like with the measured IT-power line. Sources are Musk's X posts, Zuckerberg's Threads posts, and company newsroom releases; xAI has never published a datacenter power figure on its own site. Deliberately EXCLUDED: Stargate's ~10 GW is a multi-site PROGRAM total, not a campus. Project Rainier's widely-cited 2.2 GW traces to Indiana state officials and site plans — neither AWS nor Anthropic has ever attached a GW figure to that campus. Microsoft has never published a power figure for either Fairwater site; the circulating 3.3 GW is Epoch's own independent estimate, which Epoch itself calls 'somewhat more speculative than our other estimates'. That absence is itself informative: two of the four largest builders disclose no campus power at all. Note the asymmetry in the announced series: every revision is upward (Hyperion 2 GW to 5 GW) except one — Stargate Abilene, capped back at 1.2 GW in Mar 2026 when a ~600 MW expansion was scrapped after grid-interconnection delays. That is the physical constraint this stage exists to watch, showing up in an operator's own plans rather than in an analyst's forecast.
OpenAI flat since February; Anthropic tripled to $47B
Log scale.
| Series | Date / deadline | Observation / target | State | Evidence |
|---|---|---|---|---|
| Reality | 2023-12 | 2.1 · OpenAI ~$2B (Epoch anchor) + Anthropic ~$0.1B | Measured | Provenance gap: composite with unverifiable component |
| Reality | 2024-12 | 6.5 · OpenAI $5.5B + Anthropic ~$1B | Measured | Provenance gap: composite with unverifiable component |
| Reality | 2025-08 | 18 · OpenAI $13B (Epoch) + Anthropic >$5B (compiled source, unverified) | Measured | Provenance gap: composite with unverifiable component |
| Reality | 2025-12 | 30.4 · OpenAI $21.4B + Anthropic ~$9B | Measured | Provenance gap: composite with unverifiable component |
| Reality | 2026-02 | 39 · OpenAI $25B + Anthropic $14B (Series G) | Measured | Provenance gap: composite with unverifiable component |
| Reality | 2026-05 | 72 · OpenAI ~$25B (flat) + Anthropic $47B (Series H, gross) | Measured | Provenance gap: composite with unverifiable component |
Latest recorded observation2026-05 · OpenAI ~$25B (flat) + Anthropic $47B (Series H, gross)Provenance gap: composite with unverifiable component
checked 2026-07-27 · irregular
“when will a big tech company (Google, Microsoft, Meta, etc.) hit a $100B revenue run rate from AI (products and API)?... Very naively extrapolating out the doubling every 6 months, supposing we hit a $10B revenue run rate in early 2025, suggests this would happen mid-2026.”Ch IIIa, Racing to the Trillion-Dollar Cluster
Conditionality: 'Very naively extrapolating' is his own hedge. SCOPE: big-tech AI-specific product/API revenue, NOT frontier-lab total revenue — a different quantity from this metric's axis, so it is not plotted.
Evidence · 2026-07No big tech company has disclosed a $100B AI-specific run-rate as of Jul 2026, and the doubling-every-6-months pace the extrapolation rests on has not held for this quantity.Canonical measurement source 1 ↗
CounterargumentBig tech does not break out AI-specific revenue in filings, so this is not cleanly falsifiable from public data — absence of disclosure is not proof the threshold was missed.
“Leading-company annual revenue: $1B (2023), $4B (2024), $14B (2025), $45B (2026), $140B (2027)”Compute forecast supplement, financials table, https://ai-2027.com/research/compute-forecast
Conditionality: scenario narrative; their stated method is a short-term trend 'we expect to slow down gradually, but see sustained exponential growth through 2027'. NOTE: this forecasts ANNUAL REVENUE while this metric tracks RUN-RATE — different quantities, so it is not plotted on the axis.
Evidence · 2026-07Tracked closely through 2025: 2024 actual $3.7B vs $4B forecast; 2025 actual $13.07B vs $14B forecast. For 2026 the answer turns on who counts as 'leading' — Anthropic's May run-rate of $47B already clears $45B, while OpenAI sits near $25B and flat.Canonical measurement source 1 ↗
CounterargumentAnthropic's $47B is a gross run-rate annualized from one month before hyperscaler retrocession; OpenAI's audited 2025 revenue was less than half its December run-rate, so comparing a run-rate against an annual-revenue forecast systematically flatters the forecast.
“OpenAI's annualized revenue has grown by 3.2x/year since 2024.”Epoch AI trends dashboard (updated Feb 5, 2026), https://epoch.ai/trends
Conditionality: a measured backward-looking rate stated as the current trend, not an explicit forward forecast; scored here as continuation of that rate.
Evidence · 2026-07OpenAI went $21.4B (Dec 2025) to $25B (Feb 2026) and has reportedly held roughly flat since — about 1.3x/yr. Sustaining 3.2x/yr from the $13B Aug-2025 anchor would put it near $38B by now.Canonical measurement source 1 ↗
CounterargumentRun-rate is one month annualized and the $25B figure is press-reported and never company-confirmed; a single large enterprise quarter or S-1-timed disclosure could reset the series sharply upward.
The weakest-sourced metric here. No audited annual figure exists for either lab through normal channels: OpenAI's FY2025 revenue ($13.07B audited) is known only because it leaked ahead of a confidential S-1, and Anthropic has never disclosed an audited annual number at all. Run-rate = one month annualized, which overstates trailing revenue during hypergrowth — Anthropic's ~$9B run-rate at end-2025 sits against ~$4.5B reported actual full-year 2025. Anthropic books cloud-reseller revenue GROSS, before retrocession to AWS/Google/Azure, so its figure is not like-for-like with a net reporter. Anthropic's CFO stated 'exceeding $5 billion to date' under oath in Mar 2026 while public PR claimed $14B+ run-rate; the scope of that filing is unresolved.
Two regimes bind labs today; the US federal threshold died in 2025
0 of 1 claims are plotted here. The other 1 set no numeric target on this axis (dated milestones, or a quantity this axis does not measure) — they are recorded in the claims below.
| Series | Date / deadline | Observation / target | State | Evidence |
|---|---|---|---|---|
| Reality | 2022-10 | 0 · First BIS chip export controls — binds chip flows, not developers | Measured | BIS October 2022 advanced-computing export controls ↗ |
| Reality | 2023-11 | 1 · US EO 14110 in force: 1e26 FLOP reporting duty | Measured | Executive Order 14110 ↗ |
| Reality | 2025-01 | 0 · EO 14110 revoked by EO 14148 — US federal threshold gone | Measured | Executive Order 14148 ↗ |
| Reality | 2025-08 | 1 · EU AI Act GPAI systemic-risk duties begin (1e25 FLOP) | Measured | EU AI Act (Regulation 2024/1689) ↗ |
| Reality | 2026-01 | 2 · California SB 53 in force (1e26 FLOP) — first binding US state law | Measured | California SB 53 (TFAIA) ↗ |
| Reality | 2026-07 | 2 · EU high-risk duties deferred to Dec 2027; NY RAISE Act not until 2027 | Measured | Provenance gap: multi jurisdiction composite |
| Leopold | 2028-12 | Some form of US government AGI project | pending · confidence 45 | Published claim ↗ |
Latest recorded observation2026-07 · EU high-risk duties deferred to Dec 2027; NY RAISE Act not until 2027Provenance gap: multi jurisdiction composite
checked 2026-07-27 · event-driven
“The USG will wake from its slumber, and by 27/28 we'll get some form of government AGI project.”Ch IV, The Project
Conditionality: firm declarative scenario narration ('will'). Not a compute-threshold rule, so it is not plotted on this metric's axis.
Evidence · 2026-07The closest existing analog is the Genesis Mission EO (Nov 2025) — a DOE-led national compute and scientific-discovery platform explicitly framed as 'this generation's Manhattan Project'. Equity-stake proposals surfaced Jun-Jul 2026 (OpenAI floating ~5% to the USG; Sanders calling for 50%), none enacted.Canonical measurement source 1 ↗
CounterargumentGenesis is scoped to scientific discovery, not sovereign AGI development, and no classified AGI program, mass clearance hiring, or nationalization has been confirmed — this is trending toward the prediction rather than meeting it, and his window runs to end-2028.
“Somewhere around 26/27 or so, the mood in Washington will become somber.”Ch IV, The Project
Conditionality: firm in tone but a claim about MOOD, not an observable rule or event — inherently resistant to clean resolution. Recorded because it is dated and load-bearing in his narrative.
Evidence · 2026-07Federal engagement escalated sharply through 2025-26: the Genesis Mission EO, a Dec 2025 executive order litigating against state AI laws, and mid-2026 government equity-stake proposals in the leading labs.Canonical measurement source 1 ↗
CounterargumentThe observable federal posture is deregulatory and accelerationist, not somber — the 2025 AI Action Plan is framed around winning a race, and the preemption EO exists to remove safety rules, not add them. A mood claim cannot be cleanly falsified either way.
“in the next 12-24 months, we will leak key AGI breakthroughs to the CCP.”Ch IIIb, Lock Down the Labs
Conditionality: firm ('we will'), dated Jun 2024 so the window closed mid-2026. Not plotted: this is a security event, not a compute-threshold rule.
Evidence · 2026-07No public reporting confirms a specific AGI-relevant breakthrough leak to China in the Jun 2024 - Jun 2026 window.Canonical measurement source 1 ↗
CounterargumentThis claim may be permanently unresolvable in public: successful espionage is precisely what does not get reported, so absence of evidence is weak evidence here. It is kept at 'pending' rather than scored as missed for that reason.
Counts only compute-threshold-triggered obligations actually in force. Timeline behind the count: US EO 14110 (Oct 2023) set a 1e26 FLOP reporting duty and was REVOKED by EO 14148 on 2025-01-20, taking the US federal count back to zero; the Jan 2025 BIS AI Diffusion Rule was rescinded 2025-05-13, two days before it would have bound anyone. EU AI Act GPAI duties (1e25 FLOP systemic-risk notification) began 2025-08-02. California SB 53 (1e26 FLOP) took force 2026-01-01 — the first binding US frontier-AI developer law, now under federal preemption threat from a Dec 2025 executive order. NY RAISE Act is signed but not effective until 2027-01-01. EU high-risk obligations were deferred by the 2026 Digital Omnibus from Aug 2026 to Dec 2027 (Annex III) and Aug 2028 (embedded systems) — verify the Official Journal citation before relying on the exact in-force date. Export controls are excluded from the count: they bind chip flows, not developers, and reversed direction in Jan 2026 when H200/MI325X licensing to China reopened.
Authors use different definitions and dates. This ladder keeps their original wording, then shows where their claims overlap.
No structured forecast is attached to this rung yet.
“the new model with median parameters predicts SC in Dec 2031”AI Futures Model Dec 2025 update (supersedes AI 2027's Mar 2027 scenario date; do NOT overwrite the original — record both)
Conditionality: firm model output, but authors retain 'highly uncertain... cannot confidently predict a specific year'
Revision trail: ai2027-apr2025 SC date (2027-03)
Evidence · 2026-07Nothing close to the SC bar (any AGI-company coding task, 30x faster, 30x cheaper) exists: SWE-bench Verified 79.2% (Opus 4.6), Terminal-Bench 2.0 ~65%; the independent AI 2027 tracker grades SC 'Not Yet Testable' (Jun 2026). The authors' aggregate self-grading puts quantitative progress at ~65-75% of scenario pace.Canonical measurement source 1 ↗
CounterargumentEven the revised median has wide CI (10th pct 2027.5), and individual forecaster medians (Daniel ~2029, Eli ~2032) still disagree by years — the Dec 2031 model output is not a settled team view.
“it is strikingly plausible that by 2027, models will be able to do the work of an AI researcher/engineer.”Ch I, From GPT-4 to AGI
Conditionality: conditional-scenario ('strikingly plausible')
Evidence · 2026-07Strong coding-agent adoption and longer autonomy, but no autonomous AI-researcher demonstrated; RE-Bench-class evals still short of expert autonomy.Canonical measurement source 1 ↗
CounterargumentThe AI Futures Project itself (Dec 2025) pushed its Superhuman Coder median from Jan 2027 to Dec 2031, evidence against the 2027 rung.
“We have set internal goals of having an automated AI research intern by September of 2026 running on hundreds of thousands of GPUs, and a true automated AI researcher by March of 2028.”Livestream + X post, 2025-10-28 (x.com/sama/status/1983584366547829073)
Conditionality: explicit self-hedge in the same statement: 'We may totally fail at this goal.' Lab-leader statement = intention data (bottom of the evidence hierarchy), scored with the corresponding discount.
Evidence · 2026-07Six weeks from the intern deadline, no public demonstration exists. Strongest adjacent result is DeepMind's AlphaEvolve (novel matrix-multiplication algorithm — real but narrow-domain optimization, not open-ended research). METR's NanoGPT-speedrun analysis (Apr 2026): AI contributions real but 'none reached the deep or breakthrough end of the scale'.Canonical measurement source 1 ↗
CounterargumentAn 'internal goal' can be declared met internally without public verification — this claim may resolve only on OpenAI's say-so, which the evidence hierarchy discounts to intention data either way.
“Early 2026: AI R&D progress multiplier reaches 1.5x — algorithmic progress '50% faster' with AI assistants than without”AI 2027 scenario narrative + compute-forecast multiplier table. CANONICAL READING: the narrative ladder (1.5x early 2026 -> 3x Jan 2027) — the corpus has two conflicting multiplier tables (compute-forecast quarterly vs takeoff-forecast milestone-tied), recorded per schema rule 4.
Conditionality: scenario narrative with the project-wide hedge ('not a prediction; 2027 was our modal year'). The authors' own end-2026 target was 1.9x, which they now grade 'behind pace'.
Evidence · 2026-07METR's self-report survey (May 2026) median value-multiplier is 1.4-2x; METR's production-function read of Anthropic's 8x code-volume figure puts researcher uplift 'plausibly >2x'; the third-party AI 2027 tracker marks the 1.5x rung On Track. Evidence is self-report-dominated, hence the discount.Canonical measurement source 1 ↗
CounterargumentAnthropic's own system card says acceleration is 'well short of a sustained, AI-attributable doubling of the overall pace'; the only RCT-grade measurement found a 19% SLOWDOWN (since methodologically disowned by METR itself, but never replaced with a positive result); self-reports historically overstate measured savings by ~40 percentage points.
No structured forecast is attached to this rung yet.
Each line follows one person’s published median over time. It shows revision, not consensus: these estimates are never averaged together.
| Forecaster | Published | Median year | Milestone |
|---|---|---|---|
| Cotra — transformative AI | 2020-09 | 2050 | Not recorded |
| Cotra — transformative AI | 2022-08 | 2040 | Not recorded |
| Kokotajlo — AGI | 2025-04 | 2028 | Not recorded |
| Kokotajlo — AGI | 2026-01 | 2030.95 | Not recorded |
| Lifland — AGI | 2025-04 | 2031 | Not recorded |
| Lifland — AGI | 2025-12 | 2035 | Not recorded |