Superalignment

AI Progress Tracker / Method

How we know what changed

We check sources often. We change the public view only when the evidence earns a change.

Sources checked
Editorial view changed
Next review
Status
Public prototype

The contract

A daily watch with an editorial brake

  1. Watch.

    Check each named source for a new value, method, correction, or release.

  2. Compare.

    Keep the source date, our checked date, and the last real change separate.

  3. Classify.

    Mark the item as observed, modeled, forecast, or claim before writing the summary.

  4. Review.

    Verify the unit, scope, uncertainty, and plain-language limit before publishing.

  5. Preserve.

    Publish a new dated snapshot. Do not rewrite an old briefing in place.

Automation may tell us that a page changed. It may not decide what the change means.

Evidence labels

Four labels, used exactly

Observed

A source reports a measurement, event, count, or public record. The label does not make the record complete or causal.

Modeled

A source derives an estimate by fitting, combining, or extending observations. The method and uncertainty belong with the number.

Forecast

A source states an expectation about a future outcome. It is not a present fact.

Claim

An identified group makes an assessment that cannot be reduced to one direct measurement.

Why there is no score

Capability, power, and protection use different rulers

A benchmark score, a business adoption rate, and a legal duty do not share a unit. Adding them would create a precise-looking number with no stable meaning. The tracker keeps the three lanes side by side so disagreement remains visible.

A countdown has the same problem. Trend lines can test a stated measure. They cannot name the date of a singularity without assumptions that dominate the answer.

Source registry

21 monitored sources

Download JSON
Capability

Epoch Capabilities Index frontier snapshot

Epoch AI

Method
Epoch combines results from more than 50 benchmarks with an item-response model fitted by nonlinear least squares. This entry uses the top-ranked model on Epoch's model page as checked on Sep 7, 2026: GPT-6 Astra, released Sep 3, ranked 1 of 267 with a 90 percent interval of 165 to 174. The Aug 18 entry used GPT-5.6 Sol at 161.65.
What it supports
GPT-6 Astra had the highest score in Epoch's combined benchmark index, at 169 ECI, when checked on Sep 7, 2026.
What it cannot support
ECI is not a percentage, an IQ score, a measure of reliability, or a known distance to AGI.
Source date
Sep 3, 2026
Checked
Sep 7, 2026
Rights
Epoch labels the data CC BY and the public code MIT. Link to the source and credit Epoch AI.

Open the primary source (opens in a new tab)

Capability

Epoch Capabilities Index trend

Epoch AI

Method
Epoch fits a linear trend to state-of-the-art ECI scores since the first reasoning models in September 2024 and reports a 90 percent confidence interval. On Sep 7, 2026 the page read 14 points per year, with an interval of plus 12 to plus 17, and 6 points per year for non-reasoning models. Our Aug 18 entry recorded 15.5 with an interval of 13 to 18 from about April 2024; the page header still reads Feb 5, 2026, so we cannot date the difference and record what the page now shows.
What it supports
Scores across many included benchmarks rose quickly during the fitted period, at about 14 ECI points per year.
What it cannot support
The rate does not establish persistent progress, uniform ability across tasks, or future continuation.
Source date
Feb 5, 2026
Checked
Sep 7, 2026
Rights
Epoch labels the data CC BY. Link to the source and credit Epoch AI.

Open the primary source (opens in a new tab)

Capability

OSWorld2 long-workflow benchmark

XLang Lab and OSWorld

Method
The benchmark runs a model and tool harness on 108 long computer workflows for up to 500 steps, then reports perfect binary completion and partial-credit scores.
What it supports
The reported Claude Opus 4.8 setup perfectly completed about one fifth of the tested workflows and earned partial credit on more than half.
What it cannot support
The result does not measure whole jobs, safety, economic value, or performance under every tool harness.
Source date
Jun 26, 2026
Checked
Sep 7, 2026
Rights
The code is Apache-2.0. Some benchmark assets are gated. Link to the project rather than redistributing gated assets.

Open the primary source (opens in a new tab)

Capability

Task-completion time horizon, selected below-ceiling result

METR

Method
METR fits success probability against the time a skilled human needs for clean software, machine learning, and cybersecurity tasks. We use Claude Opus 4.6 as a selected below-ceiling result because METR warns that estimates above 16 hours are unreliable on the current suite. The estimate still has a wide interval.
What it supports
On this suite, Claude Opus 4.6 reached 50 percent predicted success at tasks near 12 hours of expert human difficulty.
What it cannot support
It does not mean 12 hours of unattended model runtime or 12 hours of ordinary job automation. It is not the largest number in METR's current data.
Source date
Feb 20, 2026
Checked
Sep 7, 2026
Rights
The public analysis repository is Apache-2.0. Link to METR for the reported measurement and limits.

Open the primary source (opens in a new tab)

Capability

Task-completion time horizon trend

METR

Method
METR fits an exponential trend to post-2023 frontier time-horizon estimates and reports a bootstrapped interval.
What it supports
The measured horizon on this task suite recently doubled about every four months.
What it cannot support
The fit does not guarantee future continuation, transfer to every domain, or recursive self-improvement. METR warns that estimates above 16 hours are unreliable with the current suite.
Source date
May 8, 2026
Checked
Sep 7, 2026
Rights
The public analysis repository is Apache-2.0. Link to METR for the reported measurement and limits.

Open the primary source (opens in a new tab)

Capability

Task-completion horizon update, May 8, 2026

METR

Method
METR's public update log added an early Claude Mythos Preview result and warned that estimates above 16 hours are unreliable with the current task suite.
What it supports
A reported frontier result moved beyond the range METR considers reliable on its current suite.
What it cannot support
The update does not establish a dependable horizon above 16 hours, a raised measurement ceiling, or performance on ordinary jobs.
Source date
May 8, 2026
Checked
Sep 7, 2026
Rights
Link to METR for the reported result, update history, and measurement limit.

Open the primary source (opens in a new tab)

Delegated power

Business Trends and Outlook Survey, AI use

United States Census Bureau

Method
The Census Bureau surveys United States employer businesses every two weeks. Period 107 was collected from Aug 10 through Aug 23, 2026. The public API returns weighted estimates with standard errors; the national row is the one with no state, industry, or size code. Period 108 (Aug 24 to Sep 6) returned no data on Sep 7.
What it supports
About one in five United States employer businesses reported some AI use in the survey period, 22.4 percent in period 107, with 25.9 percent expected within six months.
What it cannot support
The survey does not show how intense or autonomous the use is, whether it raises productivity, or whether the rate is global. The wording changed in November 2025, so older values are not one clean series.
Source date
Aug 23, 2026
Checked
Sep 7, 2026
Rights
United States government data. Credit the Census Bureau and retain the survey period and margin of error.

Open the primary source (opens in a new tab)

Delegated power

AI compute stock growth

Epoch AI

Method
Epoch estimates the computing power of the global stock of AI chips from revenue data, company disclosures, and analyst reports, then fits a growth rate from 2022 onward.
What it supports
The estimated stock of AI computing power has grown quickly since 2022.
What it cannot support
The estimate does not show that every chip is active, equally accessible, or converted one for one into capability.
Source date
Feb 5, 2026
Checked
Sep 7, 2026
Rights
Epoch labels the data CC BY. Link to the source and credit Epoch AI.

Open the primary source (opens in a new tab)

Delegated power

AI compute concentration

Epoch AI

Method
Epoch estimates global AI compute holdings from public financial and infrastructure evidence, then assigns shares to major operators.
What it supports
The estimate places most global AI compute in the hands of five hyperscalers.
What it cannot support
Compute share is not the same as a share of AI decisions or outcomes, and private undisclosed infrastructure may be missing.
Source date
Feb 5, 2026
Checked
Sep 7, 2026
Rights
Epoch labels the data CC BY. Link to the source and credit Epoch AI.

Open the primary source (opens in a new tab)

Protection

Frontier safety framework adoption

International AI Safety Report

Method
The international report synthesizes public company activity and counts companies that published or updated frontier safety frameworks during 2025.
What it supports
Formal safety planning spread across frontier AI companies during 2025.
What it cannot support
Publishing a framework does not show compliance or effective safeguards. Company scopes differ and most commitments remain voluntary.
Source date
Feb 3, 2026
Checked
Sep 7, 2026
Rights
The report is available under the Open Government Licence v3.0 except where third-party rights are stated.

Open the primary source (opens in a new tab)

Protection

Foundation Model Transparency Index

Stanford CRFM, Princeton University, and MIT

Method
Researchers audit public documents from 13 companies against a published transparency rubric and average the resulting scores. The December 2025 paper remains the latest edition as of Sep 7, 2026: the average is 40.69 out of 100, down from 58 in 2024 under a changed rubric.
What it supports
The audited companies disclosed less than half of the information requested by the 2025 rubric on average.
What it cannot support
A low transparency score does not by itself mean a company is unsafe, unlawful, or less capable. The rubric changed, so the fall from 2024 is not a clean regression.
Source date
Dec 9, 2025
Checked
Sep 7, 2026
Rights
The index materials are CC BY 4.0. Credit the named research groups and preserve the rubric version.

Open the primary source (opens in a new tab)

Protection

European Union AI Act implementation timeline

European Commission AI Act Service Desk

Method
The Commission timeline records when parts of the AI Act enter into application: general application and transparency duties on Aug 2, 2026, new prohibitions on Dec 2, 2026, Annex III high-risk rules on Dec 2, 2027, and Annex I high-risk rules on Aug 2, 2028, after the amendments made by the Digital Omnibus on AI.
What it supports
Some AI Act duties are now enforceable in the European Union.
What it cannot support
The milestone does not create worldwide protection, prove compliance, or establish a measured drop in risk. The high-risk provisions were postponed and now apply in 2027 and 2028.
Source date
Aug 2, 2026
Checked
Sep 7, 2026
Rights
European Commission public legal information. Link to the official timeline and underlying law.

Open the primary source (opens in a new tab)

Delegated power

AI Economic Indicators Transformation Tracker

Stanford Digital Economy Lab

Method
The tracker reviews 12 macroeconomic indicators for signs of a rapid, AI-linked economic transformation, using data available through May 2026 for the first release. The research note page shows an update dated Aug 10, 2026, and on Sep 7 the project page still read that there is no decisive evidence of transformation. A revised count breakdown was not retrievable, so the July counts stand.
What it supports
The selected macro indicators did not show decisive evidence of economy-wide takeoff in the July update.
What it cannot support
The result does not rule out disruption for particular workers or sectors, and it does not show that future progress will plateau.
Source date
Aug 10, 2026
Checked
Sep 7, 2026
Rights
Link to and summarize the tracker. No broad reuse grant was found for bulk copying its content.

Open the primary source (opens in a new tab)

Delegated power

Reported AI incidents

Stanford AI Index and AI Incident Database

Method
The AI Index summarizes records in the AI Incident Database, which catalogs public reports and can revise entries as evidence changes.
What it supports
The number of documented incident reports rose from 2024 to 2025.
What it cannot support
Reported incidents are not the true incident rate, do not share equal severity, and do not establish that AI alone caused each event. Reporting and revision bias remain.
Source date
Apr 13, 2026
Checked
Sep 7, 2026
Rights
Link to the AI Index and database. Check database terms before reusing incident narratives or bulk records.

Open the primary source (opens in a new tab)

Forecast

Thousands of AI Authors on the Future of AI

Journal of Artificial Intelligence Research

Method
The study surveyed 2,778 researchers who had published at leading AI venues, mainly during 2023, and reports the median of their probability distributions.
What it supports
Many surveyed AI researchers expected very broad machine capability within decades, while their individual forecasts varied widely.
What it cannot support
Expert forecasts are not a scientific deadline or countdown. The medians hide wide uncertainty and depend on the exact question.
Source date
Oct 6, 2025
Checked
Sep 7, 2026
Rights
Link to the paper and journal record. Quote only short passages and preserve the survey wording and date.

Open the primary source (opens in a new tab)

Capability

FrontierMath Tiers 1-4, version 2 correction

Epoch AI

Method
Epoch compared the second benchmark version with the first, corrected 123 Tier 1-3 problems and 12 Tier 4 problems, and removed 12 problems. It reports that the update addressed errors in 42 percent of problems.
What it supports
A large benchmark correction changed the test itself, so version history is part of interpreting model progress.
What it cannot support
The correction does not show that all benchmark scores are invalid or that model capability failed to improve. Most evaluation problems remain private, and OpenAI funded the benchmark and has exclusive access to a subset.
Source date
Jun 12, 2026
Checked
Sep 7, 2026
Rights
Epoch publishes its work under CC BY when credited. Link to the versioned benchmark pages and preserve the funding and private-set disclosure.

Open the primary source (opens in a new tab)

Protection

Safety and alignment in an era of long-horizon models

OpenAI

Method
OpenAI describes unwanted behavior during limited internal use of an unnamed long-running model, the decision to pause access, and safeguards added before limited access resumed.
What it supports
OpenAI publicly reported that monitored internal use exposed failures its existing predeployment evaluations had missed and that it responded by pausing access and adding trajectory-level safeguards.
What it cannot support
The report does not independently establish incident frequency, model identity, mitigation effectiveness outside the replayed environments, or safety after redeployment. It is evidence about one developer's internal process.
Source date
Jul 20, 2026
Checked
Aug 18, 2026
Rights
Developer-authored public report. Link and summarize it as OpenAI's account; do not treat the incidents or mitigation results as independently verified.

Open the primary source (opens in a new tab)

Forecast

Current loss-of-control capability assessment

International AI Safety Report

Method
The international report synthesizes available evaluations and expert views on the capabilities, propensities, and deployment conditions relevant to loss of control.
What it supports
The reviewed evidence does not show that current systems already have the combined capabilities required for loss of control.
What it cannot support
The assessment does not make catastrophic risk zero, prove a plateau, or supply a known date. Expert views through 2030 vary widely.
Source date
Feb 3, 2026
Checked
Sep 7, 2026
Rights
The report is available under the Open Government Licence v3.0 except where third-party rights are stated.

Open the primary source (opens in a new tab)

Capability

ARC-AGI-3 verified result, GPT-6 Astra

ARC Prize Foundation

Method
ARC Prize runs a model on the Semi-Private ARC-AGI-3 evaluation, a set of interactive turn-based games with undisclosed rules, and scores completion weighted by action efficiency against first-time human players. The best observed result for GPT-6 Astra was 99.9 percent with the Provider Adapter harness at high reasoning effort, costing 18,817 dollars; the Standard harness reached 62.7 percent at max effort for 26,098 dollars. GPT-5.6 Sol, tested July 9, reached 7.78 percent.
What it supports
Under the reported harness, GPT-6 Astra completed almost every ARC-AGI-3 game about as efficiently as first-time human players, and the two harnesses produce very different scores for the same model.
What it cannot support
ARC Prize states that saturating the benchmark is not proof of AGI. The games are bounded and deterministic, the score counts environmental actions rather than compute or cost, and ARC reports best observed values rather than a repeated-run distribution.
Source date
Sep 2, 2026
Checked
Sep 7, 2026
Rights
Link to ARC Prize's results pages and preserve the harness and reasoning-effort labels. Public puzzle assets carry their own licenses.

Open the primary source (opens in a new tab)

Protection

Digital Omnibus on AI enters into force

European Commission

Method
The Commission's news item records that the AI Omnibus, Regulation (EU) 2026/1744, entered into force on July 27, 2026. It moves the application of Annex III high-risk rules to Dec 2, 2027 and Annex I high-risk rules to Aug 2, 2028, and adds a prohibition on AI systems that generate non-consensual sexual imagery or child sexual abuse material.
What it supports
The European Union postponed its high-risk AI obligations by 16 and 12 months respectively, before the original Aug 2, 2026 date arrived.
What it cannot support
The postponement does not remove the obligations, and it does not show whether the delay raised or lowered risk. Transparency duties and the general application date were not moved.
Source date
Jul 27, 2026
Checked
Sep 7, 2026
Rights
European Commission public information and EU law. Link to the official notice and the regulation.

Open the primary source (opens in a new tab)

Delegated power

Canaries in the Coal Mine, August 2026 revision

Stanford Digital Economy Lab

Method
The authors compare employment for workers aged 22 to 25 in occupations highly exposed to AI with similarly aged workers in less exposed occupations, using ADP payroll records through June 2026. The revised paper reports a gap of about 19 percent, up from 15 percent in July 2025.
What it supports
In one large payroll dataset, early-career employment in AI-exposed occupations fell further behind less exposed occupations through mid-2026.
What it cannot support
The authors call the result descriptive, not causal. The gap shrinks when education is controlled, some trends predate generative AI, the ADP sample runs larger than national surveys, and the paper finds no widespread displacement.
Source date
Aug 12, 2026
Checked
Sep 7, 2026
Rights
Link to and summarize the release. No broad reuse grant was found for bulk copying.

Open the primary source (opens in a new tab)

Review and corrections

What is reviewed

This snapshot is a public prototype prepared with AI assistance for source retrieval, structure, first-pass prose, and implementation. Each listed value was checked against the linked source on the date shown. A named independent human review is not yet recorded.

Corrections should identify the item, old value, new value, reason, and date. Send corrections to [email protected].