CapabilityEpoch Capabilities Index frontier snapshot
Epoch AI
- Method
- Epoch combines results from more than 50 benchmarks with an item-response model fitted by nonlinear least squares. This entry uses the top-ranked model on Epoch's model page as checked on Sep 7, 2026: GPT-6 Astra, released Sep 3, ranked 1 of 267 with a 90 percent interval of 165 to 174. The Aug 18 entry used GPT-5.6 Sol at 161.65.
- What it supports
- GPT-6 Astra had the highest score in Epoch's combined benchmark index, at 169 ECI, when checked on Sep 7, 2026.
- What it cannot support
- ECI is not a percentage, an IQ score, a measure of reliability, or a known distance to AGI.
- Source date
- Sep 3, 2026
- Checked
- Sep 7, 2026
- Rights
- Epoch labels the data CC BY and the public code MIT. Link to the source and credit Epoch AI.
Open the primary source (opens in a new tab)
CapabilityEpoch Capabilities Index trend
Epoch AI
- Method
- Epoch fits a linear trend to state-of-the-art ECI scores since the first reasoning models in September 2024 and reports a 90 percent confidence interval. On Sep 7, 2026 the page read 14 points per year, with an interval of plus 12 to plus 17, and 6 points per year for non-reasoning models. Our Aug 18 entry recorded 15.5 with an interval of 13 to 18 from about April 2024; the page header still reads Feb 5, 2026, so we cannot date the difference and record what the page now shows.
- What it supports
- Scores across many included benchmarks rose quickly during the fitted period, at about 14 ECI points per year.
- What it cannot support
- The rate does not establish persistent progress, uniform ability across tasks, or future continuation.
- Source date
- Feb 5, 2026
- Checked
- Sep 7, 2026
- Rights
- Epoch labels the data CC BY. Link to the source and credit Epoch AI.
Open the primary source (opens in a new tab)
CapabilityOSWorld2 long-workflow benchmark
XLang Lab and OSWorld
- Method
- The benchmark runs a model and tool harness on 108 long computer workflows for up to 500 steps, then reports perfect binary completion and partial-credit scores.
- What it supports
- The reported Claude Opus 4.8 setup perfectly completed about one fifth of the tested workflows and earned partial credit on more than half.
- What it cannot support
- The result does not measure whole jobs, safety, economic value, or performance under every tool harness.
- Source date
- Jun 26, 2026
- Checked
- Sep 7, 2026
- Rights
- The code is Apache-2.0. Some benchmark assets are gated. Link to the project rather than redistributing gated assets.
Open the primary source (opens in a new tab)
CapabilityTask-completion time horizon, selected below-ceiling result
METR
- Method
- METR fits success probability against the time a skilled human needs for clean software, machine learning, and cybersecurity tasks. We use Claude Opus 4.6 as a selected below-ceiling result because METR warns that estimates above 16 hours are unreliable on the current suite. The estimate still has a wide interval.
- What it supports
- On this suite, Claude Opus 4.6 reached 50 percent predicted success at tasks near 12 hours of expert human difficulty.
- What it cannot support
- It does not mean 12 hours of unattended model runtime or 12 hours of ordinary job automation. It is not the largest number in METR's current data.
- Source date
- Feb 20, 2026
- Checked
- Sep 7, 2026
- Rights
- The public analysis repository is Apache-2.0. Link to METR for the reported measurement and limits.
Open the primary source (opens in a new tab)
CapabilityTask-completion time horizon trend
METR
- Method
- METR fits an exponential trend to post-2023 frontier time-horizon estimates and reports a bootstrapped interval.
- What it supports
- The measured horizon on this task suite recently doubled about every four months.
- What it cannot support
- The fit does not guarantee future continuation, transfer to every domain, or recursive self-improvement. METR warns that estimates above 16 hours are unreliable with the current suite.
- Source date
- May 8, 2026
- Checked
- Sep 7, 2026
- Rights
- The public analysis repository is Apache-2.0. Link to METR for the reported measurement and limits.
Open the primary source (opens in a new tab)
CapabilityTask-completion horizon update, May 8, 2026
METR
- Method
- METR's public update log added an early Claude Mythos Preview result and warned that estimates above 16 hours are unreliable with the current task suite.
- What it supports
- A reported frontier result moved beyond the range METR considers reliable on its current suite.
- What it cannot support
- The update does not establish a dependable horizon above 16 hours, a raised measurement ceiling, or performance on ordinary jobs.
- Source date
- May 8, 2026
- Checked
- Sep 7, 2026
- Rights
- Link to METR for the reported result, update history, and measurement limit.
Open the primary source (opens in a new tab)
Delegated powerBusiness Trends and Outlook Survey, AI use
United States Census Bureau
- Method
- The Census Bureau surveys United States employer businesses every two weeks. Period 107 was collected from Aug 10 through Aug 23, 2026. The public API returns weighted estimates with standard errors; the national row is the one with no state, industry, or size code. Period 108 (Aug 24 to Sep 6) returned no data on Sep 7.
- What it supports
- About one in five United States employer businesses reported some AI use in the survey period, 22.4 percent in period 107, with 25.9 percent expected within six months.
- What it cannot support
- The survey does not show how intense or autonomous the use is, whether it raises productivity, or whether the rate is global. The wording changed in November 2025, so older values are not one clean series.
- Source date
- Aug 23, 2026
- Checked
- Sep 7, 2026
- Rights
- United States government data. Credit the Census Bureau and retain the survey period and margin of error.
Open the primary source (opens in a new tab)
Delegated powerAI compute stock growth
Epoch AI
- Method
- Epoch estimates the computing power of the global stock of AI chips from revenue data, company disclosures, and analyst reports, then fits a growth rate from 2022 onward.
- What it supports
- The estimated stock of AI computing power has grown quickly since 2022.
- What it cannot support
- The estimate does not show that every chip is active, equally accessible, or converted one for one into capability.
- Source date
- Feb 5, 2026
- Checked
- Sep 7, 2026
- Rights
- Epoch labels the data CC BY. Link to the source and credit Epoch AI.
Open the primary source (opens in a new tab)
Delegated powerAI compute concentration
Epoch AI
- Method
- Epoch estimates global AI compute holdings from public financial and infrastructure evidence, then assigns shares to major operators.
- What it supports
- The estimate places most global AI compute in the hands of five hyperscalers.
- What it cannot support
- Compute share is not the same as a share of AI decisions or outcomes, and private undisclosed infrastructure may be missing.
- Source date
- Feb 5, 2026
- Checked
- Sep 7, 2026
- Rights
- Epoch labels the data CC BY. Link to the source and credit Epoch AI.
Open the primary source (opens in a new tab)
ProtectionFrontier safety framework adoption
International AI Safety Report
- Method
- The international report synthesizes public company activity and counts companies that published or updated frontier safety frameworks during 2025.
- What it supports
- Formal safety planning spread across frontier AI companies during 2025.
- What it cannot support
- Publishing a framework does not show compliance or effective safeguards. Company scopes differ and most commitments remain voluntary.
- Source date
- Feb 3, 2026
- Checked
- Sep 7, 2026
- Rights
- The report is available under the Open Government Licence v3.0 except where third-party rights are stated.
Open the primary source (opens in a new tab)
ProtectionFoundation Model Transparency Index
Stanford CRFM, Princeton University, and MIT
- Method
- Researchers audit public documents from 13 companies against a published transparency rubric and average the resulting scores. The December 2025 paper remains the latest edition as of Sep 7, 2026: the average is 40.69 out of 100, down from 58 in 2024 under a changed rubric.
- What it supports
- The audited companies disclosed less than half of the information requested by the 2025 rubric on average.
- What it cannot support
- A low transparency score does not by itself mean a company is unsafe, unlawful, or less capable. The rubric changed, so the fall from 2024 is not a clean regression.
- Source date
- Dec 9, 2025
- Checked
- Sep 7, 2026
- Rights
- The index materials are CC BY 4.0. Credit the named research groups and preserve the rubric version.
Open the primary source (opens in a new tab)
ProtectionEuropean Union AI Act implementation timeline
European Commission AI Act Service Desk
- Method
- The Commission timeline records when parts of the AI Act enter into application: general application and transparency duties on Aug 2, 2026, new prohibitions on Dec 2, 2026, Annex III high-risk rules on Dec 2, 2027, and Annex I high-risk rules on Aug 2, 2028, after the amendments made by the Digital Omnibus on AI.
- What it supports
- Some AI Act duties are now enforceable in the European Union.
- What it cannot support
- The milestone does not create worldwide protection, prove compliance, or establish a measured drop in risk. The high-risk provisions were postponed and now apply in 2027 and 2028.
- Source date
- Aug 2, 2026
- Checked
- Sep 7, 2026
- Rights
- European Commission public legal information. Link to the official timeline and underlying law.
Open the primary source (opens in a new tab)
Delegated powerAI Economic Indicators Transformation Tracker
Stanford Digital Economy Lab
- Method
- The tracker reviews 12 macroeconomic indicators for signs of a rapid, AI-linked economic transformation, using data available through May 2026 for the first release. The research note page shows an update dated Aug 10, 2026, and on Sep 7 the project page still read that there is no decisive evidence of transformation. A revised count breakdown was not retrievable, so the July counts stand.
- What it supports
- The selected macro indicators did not show decisive evidence of economy-wide takeoff in the July update.
- What it cannot support
- The result does not rule out disruption for particular workers or sectors, and it does not show that future progress will plateau.
- Source date
- Aug 10, 2026
- Checked
- Sep 7, 2026
- Rights
- Link to and summarize the tracker. No broad reuse grant was found for bulk copying its content.
Open the primary source (opens in a new tab)
Delegated powerReported AI incidents
Stanford AI Index and AI Incident Database
- Method
- The AI Index summarizes records in the AI Incident Database, which catalogs public reports and can revise entries as evidence changes.
- What it supports
- The number of documented incident reports rose from 2024 to 2025.
- What it cannot support
- Reported incidents are not the true incident rate, do not share equal severity, and do not establish that AI alone caused each event. Reporting and revision bias remain.
- Source date
- Apr 13, 2026
- Checked
- Sep 7, 2026
- Rights
- Link to the AI Index and database. Check database terms before reusing incident narratives or bulk records.
Open the primary source (opens in a new tab)
ForecastThousands of AI Authors on the Future of AI
Journal of Artificial Intelligence Research
- Method
- The study surveyed 2,778 researchers who had published at leading AI venues, mainly during 2023, and reports the median of their probability distributions.
- What it supports
- Many surveyed AI researchers expected very broad machine capability within decades, while their individual forecasts varied widely.
- What it cannot support
- Expert forecasts are not a scientific deadline or countdown. The medians hide wide uncertainty and depend on the exact question.
- Source date
- Oct 6, 2025
- Checked
- Sep 7, 2026
- Rights
- Link to the paper and journal record. Quote only short passages and preserve the survey wording and date.
Open the primary source (opens in a new tab)
CapabilityFrontierMath Tiers 1-4, version 2 correction
Epoch AI
- Method
- Epoch compared the second benchmark version with the first, corrected 123 Tier 1-3 problems and 12 Tier 4 problems, and removed 12 problems. It reports that the update addressed errors in 42 percent of problems.
- What it supports
- A large benchmark correction changed the test itself, so version history is part of interpreting model progress.
- What it cannot support
- The correction does not show that all benchmark scores are invalid or that model capability failed to improve. Most evaluation problems remain private, and OpenAI funded the benchmark and has exclusive access to a subset.
- Source date
- Jun 12, 2026
- Checked
- Sep 7, 2026
- Rights
- Epoch publishes its work under CC BY when credited. Link to the versioned benchmark pages and preserve the funding and private-set disclosure.
Open the primary source (opens in a new tab)
ProtectionSafety and alignment in an era of long-horizon models
OpenAI
- Method
- OpenAI describes unwanted behavior during limited internal use of an unnamed long-running model, the decision to pause access, and safeguards added before limited access resumed.
- What it supports
- OpenAI publicly reported that monitored internal use exposed failures its existing predeployment evaluations had missed and that it responded by pausing access and adding trajectory-level safeguards.
- What it cannot support
- The report does not independently establish incident frequency, model identity, mitigation effectiveness outside the replayed environments, or safety after redeployment. It is evidence about one developer's internal process.
- Source date
- Jul 20, 2026
- Checked
- Aug 18, 2026
- Rights
- Developer-authored public report. Link and summarize it as OpenAI's account; do not treat the incidents or mitigation results as independently verified.
Open the primary source (opens in a new tab)
ForecastCurrent loss-of-control capability assessment
International AI Safety Report
- Method
- The international report synthesizes available evaluations and expert views on the capabilities, propensities, and deployment conditions relevant to loss of control.
- What it supports
- The reviewed evidence does not show that current systems already have the combined capabilities required for loss of control.
- What it cannot support
- The assessment does not make catastrophic risk zero, prove a plateau, or supply a known date. Expert views through 2030 vary widely.
- Source date
- Feb 3, 2026
- Checked
- Sep 7, 2026
- Rights
- The report is available under the Open Government Licence v3.0 except where third-party rights are stated.
Open the primary source (opens in a new tab)
CapabilityARC-AGI-3 verified result, GPT-6 Astra
ARC Prize Foundation
- Method
- ARC Prize runs a model on the Semi-Private ARC-AGI-3 evaluation, a set of interactive turn-based games with undisclosed rules, and scores completion weighted by action efficiency against first-time human players. The best observed result for GPT-6 Astra was 99.9 percent with the Provider Adapter harness at high reasoning effort, costing 18,817 dollars; the Standard harness reached 62.7 percent at max effort for 26,098 dollars. GPT-5.6 Sol, tested July 9, reached 7.78 percent.
- What it supports
- Under the reported harness, GPT-6 Astra completed almost every ARC-AGI-3 game about as efficiently as first-time human players, and the two harnesses produce very different scores for the same model.
- What it cannot support
- ARC Prize states that saturating the benchmark is not proof of AGI. The games are bounded and deterministic, the score counts environmental actions rather than compute or cost, and ARC reports best observed values rather than a repeated-run distribution.
- Source date
- Sep 2, 2026
- Checked
- Sep 7, 2026
- Rights
- Link to ARC Prize's results pages and preserve the harness and reasoning-effort labels. Public puzzle assets carry their own licenses.
Open the primary source (opens in a new tab)
ProtectionDigital Omnibus on AI enters into force
European Commission
- Method
- The Commission's news item records that the AI Omnibus, Regulation (EU) 2026/1744, entered into force on July 27, 2026. It moves the application of Annex III high-risk rules to Dec 2, 2027 and Annex I high-risk rules to Aug 2, 2028, and adds a prohibition on AI systems that generate non-consensual sexual imagery or child sexual abuse material.
- What it supports
- The European Union postponed its high-risk AI obligations by 16 and 12 months respectively, before the original Aug 2, 2026 date arrived.
- What it cannot support
- The postponement does not remove the obligations, and it does not show whether the delay raised or lowered risk. Transparency duties and the general application date were not moved.
- Source date
- Jul 27, 2026
- Checked
- Sep 7, 2026
- Rights
- European Commission public information and EU law. Link to the official notice and the regulation.
Open the primary source (opens in a new tab)
Delegated powerCanaries in the Coal Mine, August 2026 revision
Stanford Digital Economy Lab
- Method
- The authors compare employment for workers aged 22 to 25 in occupations highly exposed to AI with similarly aged workers in less exposed occupations, using ADP payroll records through June 2026. The revised paper reports a gap of about 19 percent, up from 15 percent in July 2025.
- What it supports
- In one large payroll dataset, early-career employment in AI-exposed occupations fell further behind less exposed occupations through mid-2026.
- What it cannot support
- The authors call the result descriptive, not causal. The gap shrinks when education is controlled, some trends predate generative AI, the ADP sample runs larger than national surveys, and the paper finds no widespread displacement.
- Source date
- Aug 12, 2026
- Checked
- Sep 7, 2026
- Rights
- Link to and summarize the release. No broad reuse grant was found for bulk copying.
Open the primary source (opens in a new tab)