Superalignment

AI progress, checked and explained

AI is improving fast. The real world is messier.

See what AI can do, where people let it act, and whether the safeguards match the power it gets.

A new model set records on two AI tests in early September, and one of those tests is nearly used up. About one in five businesses in the United States report some use. Europe delayed its strictest AI rules. We still have much less evidence about whether safeguards work.

How current is this?Sources checked Sep 7, 2026
Sources checked
Editorial view changed
Next review
Status
Public prototype

A source check finds updates. A human-readable summary changes only after its meaning and limits are reviewed. This edition re-checked 20 of 21 registered sources on Sep 7, 2026; the next review is due in four weeks. Read the method.

The picture in four parts

Progress is real. It is not one straight line.

These charts tell four different parts of the story. They do not combine into one score.

What AI can do

AI can do more. Long work still breaks.

Scores are rising on many tests, and one test built to resist memorization is nearly used up. But long jobs are chains, and one weak step can ruin the result.

What this cannot tell us: A rising test score does not tell us how close AI is to matching people at every task.

A long row of falling relay pieces stops at one tilted red piece before reaching the final block.
One weak step can break a long chain of work.AI-generated conceptual illustration. Not evidence. Generated with Higgsfield for this page.

A combined score from more than 50 AI tests is rising

Epoch combines many tests into one score. Since fall 2024, its top score rose by about 14 points a year.

EstimateSource: Epoch AIChecked

See exact numbers and limits
Low end
12
Estimate
14 per year
High end
17

Keep in mind: The score is called the Epoch Capabilities Index, or ECI. It is not a percent, an IQ score, or a known distance to human-level AI.

Estimate Researchers calculated this from several observations.

A long computer job can still fail

The best reported setup perfectly finished about one in five of 108 long computer tasks.

ReportedSource: XLang LabChecked

See exact numbers and limits
Perfect completion
20.6%
Average partial credit
54.8%

Keep in mind: The same setup's average partial-credit score was 54.8%. That does not mean it finished the job.

Reported The source recorded a result, event, count, or public fact.

Where people let AI act

Power grows when people connect AI to the world.

A system matters more when it can act in a workplace, move money, change code, or shape a decision.

What this cannot tell us: A use rate does not show how much access AI has inside each business.

Several ordinary devices sit on a workbench, each connected through its own red switch.
Real-world power arrives one connection at a time.AI-generated conceptual illustration. Not evidence. Generated with Higgsfield for this page.

About one in five United States businesses report using AI

Businesses with employees reported 22.4% use in late August and expected 25.9% within six months.

ReportedSource: United States Census BureauChecked

See exact numbers and limits
Reported now
22.4%
Expected in six months
25.9%

Keep in mind: The second number is what businesses expected, not a measured future result.

Reported The source recorded a result, event, count, or public fact.

Five organizations hold most AI computing power

Epoch estimates that five organizations hold 71% of the computing power used for AI.

EstimateSource: Epoch AIChecked

See exact numbers and limits
Estimated share
71%

Keep in mind: This estimate uses public evidence. Private infrastructure may be missing.

Estimate Researchers calculated this from several observations.

What has changed in the real world

Change is real, but it is not everywhere at once.

Some jobs and industries are changing quickly. The whole economy still does not show one sudden break.

What this cannot tell us: More incident reports can mean more harm, better reporting, or both.

A blue mechanism changes one tile in a field of plain paper tiles, while three red flags mark problems.
A few places can change fast while many others barely move.AI-generated conceptual illustration. Not evidence. Generated with Higgsfield for this page.

The wider economy has not broken from past trends

Stanford checked 12 measures. Seven looked normal, three showed mild change, and two showed strong change.

EstimateSource: Stanford Digital Economy LabChecked

See exact numbers and limits
No outlier
7
Mild
3
Strong
2

Keep in mind: Large changes in one job or industry can still be real. This view only asks whether the wider economy broke from its past pattern.

Estimate Researchers calculated this from several observations.

Documented AI incident reports rose

The AI Incident Database recorded 233 reports for 2024 and 362 for 2025.

ReportedSource: Stanford AI Index and AI Incident DatabaseChecked

See exact numbers and limits
2024
233
2025
362

Keep in mind: Reports cover different kinds of harm. The count is not the true incident rate.

Reported The source recorded a result, event, count, or public fact.

What can test, limit, or stop it

A safety plan is not proof that the brake works.

Companies are writing safety plans and Europe has passed a law, then delayed its strictest part. We have less public evidence that any of it works when a system is under pressure.

What this cannot tell us: Public disclosure is not the same thing as safety, but it lets outsiders check more.

A person tests a large mechanical brake beside a paper plan and a tray of repair parts.
A written plan matters only if the brake works under load.AI-generated conceptual illustration. Not evidence. Generated with Higgsfield for this page.

The average company disclosed less than half

Across 13 AI companies, researchers found about 41% of the public information they asked for.

ReportedSource: Stanford, Princeton, and Massachusetts Institute of TechnologyChecked

See exact numbers and limits
Average score
40.69 / 100

Keep in mind: A low disclosure score does not by itself prove unsafe or unlawful behavior.

Reported The source recorded a result, event, count, or public fact.

Latest briefing

What changed

All six updates
Reported

GPT-6 Astra scored 99.9 on a test built to resist memorized skill

ARC-AGI-3 is a set of simple games with hidden rules. With one setup, Astra beat nearly every game about as efficiently as first-time human players. With the plain setup it scored 62.7%. The previous OpenAI model scored 7.8% in July.

Open the source (opens in a new tab)

Estimate

A combined score from more than 50 AI tests hit a new high

GPT-6 Astra scored 169 on Epoch's index, first of 267 models. The top score in our August edition was 161.65.

Open the source (opens in a new tab)

Reported

AI use in United States businesses rose to 22.4%

In the two weeks to Aug. 23, 22.4% of businesses with employees said they used some AI, up from 21.5% in July. 25.9% expected to within six months.

Open the source (opens in a new tab)

15 signals, three lanes

Follow the parts we can check

Each lane starts with the plain answer. Open it only when you want the exact measure, source, and limit.

01

Evidence lane 01

What AI can do

Two records fell in September. Long jobs still break.

These sources cover test results, long technical tasks, and expert surveys. They do not add up to one score for human-like intelligence.

Show 7 signals, sources, and limits
Highest combined AI test score A new record: GPT-6 Astra Estimate Epoch combined results from more than 50 AI tests. GPT-6 Astra, released Sep 3, had the highest overall score ever measured.
Exact measure

169 on the Epoch Capabilities Index, or ECI, with a 90% range of 165 to 174. The number only has meaning when models on the same scale are compared.

Estimate Researchers calculated this from several observations.

What it shows

GPT-6 Astra had the highest combined score Epoch had measured when checked on Sep 7.

What it cannot show

ECI is not a percentage, an IQ score, a reliability rate, or a known distance to AGI.

How it was measured

Epoch fits results from more than 50 benchmarks with an item-response model. Epoch's model page ranks GPT-6 Astra first of 267 models with a 90 percent interval of 165 to 174. The scale anchors Claude 3.5 Sonnet at 130 and GPT-5 at 150.

Source date
Sep 3, 2026
Checked
Sep 7, 2026
Changed
Sep 3, 2026

Sources: Epoch AI (opens in a new tab)

How fast the combined score rose About 14 points a year Estimate Since fall 2024, Epoch estimates that its highest combined score has risen quickly. Our August edition recorded 15.5; the page now shows 14.
Exact measure

14 ECI points a year since September 2024, with a 90% range from 12 to 17 points.

Estimate Researchers calculated this from several observations.

What it shows

Many included benchmark scores rose quickly during the fitted period.

What it cannot show

The fit does not show uniform ability, durable progress, or a future rate.

How it was measured

Epoch fits a linear trend to state-of-the-art ECI scores since September 2024. Its 90 percent interval is plus 12 to plus 17 ECI per year; non-reasoning models rose 6 per year. The page header still reads Feb 5, 2026, so we cannot date the change from the 15.5 we recorded on Aug 18.

Source date
Feb 5, 2026
Checked
Sep 7, 2026
Changed
Sep 7, 2026

Sources: Epoch AI (opens in a new tab)

Long computer tasks About 1 in 5 finished perfectly Reported The best reported setup perfectly finished about one in five of 108 long computer tasks. No newer top result was posted by Sep 7.
Exact measure

20.6% perfect completion and 54.8% average partial credit on OSWorld2.

Reported The source recorded a result, event, count, or public fact.

What it shows

That setup completed about one fifth of the tested workflows perfectly.

What it cannot show

The benchmark does not measure whole jobs, safe behavior, economic value, or every tool setup.

How it was measured

A Claude Opus 4.8 setup used batched tools for up to 500 steps on 108 long workflows.

Source date
Jun 26, 2026
Checked
Sep 7, 2026
Changed
Jun 26, 2026

Sources: XLang Lab and OSWorld (opens in a new tab)

Hard technical tasks About a 50/50 chance Estimate On one test, Claude Opus 4.6 had about a 50% chance on tasks that took experts around 12 hours. The test has not published a number for GPT-6 Astra.
Exact measure

11.98 expert-hours, with an estimated range from 5.28 to 60.56 hours. This describes task difficulty, not unattended runtime.

Estimate Researchers calculated this from several observations.

What it shows

The model reached 50 percent predicted success near 12 hours of human task difficulty on this suite.

What it cannot show

This is not 12 hours of unattended runtime or 12 hours of a normal job.

How it was measured

METR fits success against skilled human task time. The interval is 5.28 to 60.56 hours.

Source date
Feb 20, 2026
Checked
Sep 7, 2026
Changed
Feb 20, 2026

Sources: METR (opens in a new tab)

How fast task length grew Doubled about every 4 months Estimate On one technical test, the length of tasks AI could handle rose quickly after 2023. The trend has not been updated since May.
Exact measure

A fitted doubling time of 128.744 days on METR's stated task suite.

Estimate Researchers calculated this from several observations.

What it shows

The horizon grew quickly on METR's software, machine learning, and cyber suite.

What it cannot show

The trend need not continue, transfer to every field, or imply self-improvement. METR marks estimates above 16 hours as unreliable with this suite.

How it was measured

The post-2023 exponential fit has an interval of 104.428 to 158.012 days.

Source date
May 8, 2026
Checked
Sep 7, 2026
Changed
May 8, 2026

Sources: METR (opens in a new tab)

When researchers expect very broad AI No single deadline Forecast The middle survey answers were 2047 for beating people at every task and 2116 for automating every job. Answers varied widely.
Exact measure

Survey medians of 2047 and 2116 among 2,778 researchers who had published at top AI venues.

Forecast This is a view of the future, not a current fact.

What it shows

Many surveyed researchers expected very broad machine ability within decades.

What it cannot show

These are not scientific deadlines. Answers varied widely and changed with the question.

How it was measured

A study surveyed 2,778 authors at leading AI venues, mainly in 2023, and reports medians across their probability forecasts.

Source date
Oct 6, 2025
Checked
Sep 7, 2026
Changed
Oct 6, 2025

Sources: Journal of Artificial Intelligence Research (opens in a new tab)

Can today's AI escape human control? Not shown Assessment The 2026 international report found no evidence that current systems have all the abilities needed to escape control.
Exact measure

The report's assessment is conditional and does not assign a probability or deadline.

Assessment This is a source's judgment, not one direct measurement.

What it shows

The reviewed evidence does not show that current systems are already superintelligent or beyond control.

What it cannot show

The assessment does not make risk zero, prove a plateau, or supply a date. Experts disagree about progress through 2030.

How it was measured

The report reviews evaluations and expert views on capability, harmful behavior, and deployment access.

Source date
Feb 3, 2026
Checked
Sep 7, 2026
Changed
Feb 3, 2026

Sources: International AI Safety Report (opens in a new tab)

02

Evidence lane 02

Where people let AI act

Use edged up. The effects are still uneven.

These sources track business use, who controls the machines that run AI, changes in the economy, and reported incidents.

Show 5 signals, sources, and limits
AI use in United States businesses About 1 in 5 Reported About 22.4% of businesses with employees reported using some AI in late August, up from 21.5% in July.
Exact measure

22.4% reported current use in Census period 107 (Aug 10 to 23). Respondents expected 25.9% within six months.

Reported The source recorded a result, event, count, or public fact.

What it shows

AI use has spread to a meaningful share of United States employer businesses and rose slightly over the summer.

What it cannot show

It does not show intensity, autonomy, productivity, or global use. A 2025 wording change starts a new series.

How it was measured

Census BTOS period 107 ran Aug 10 to 23. Standard errors are 0.36 and 0.47 percentage points. Period 106 (Jul 27 to Aug 9) read 21.8 and 25.9 percent. Period 108 had no published data on Sep 7.

Source date
Aug 23, 2026
Checked
Sep 7, 2026
Changed
Aug 23, 2026

Sources: United States Census Bureau (opens in a new tab)

Total computing power used for AI Estimated 3.4 times growth a year Estimate Epoch estimates that the world's available AI computing power has grown quickly since 2022.
Exact measure

An estimated 3.4 times annual growth rate in global AI computing capacity.

Estimate Researchers calculated this from several observations.

What it shows

The estimated amount of available AI computing power rose quickly.

What it cannot show

Not every chip is active, equally accessible, or turned one for one into ability.

How it was measured

Epoch combines revenue, company, and analyst evidence. The fitted interval is 3.2 to 3.7 times per year.

Source date
Feb 5, 2026
Checked
Sep 7, 2026
Changed
Feb 5, 2026

Sources: Epoch AI (opens in a new tab)

Who holds the computing power? Five organizations hold 71% Estimate Epoch estimates that five organizations hold most of the world's AI computing power.
Exact measure

An estimated 71% share held by five large operators.

Estimate Researchers calculated this from several observations.

What it shows

AI infrastructure is highly concentrated under this estimate.

What it cannot show

Compute share is not a share of all AI decisions or outcomes. Private infrastructure may be missing.

How it was measured

Epoch estimates holdings from public financial and infrastructure evidence.

Source date
Feb 5, 2026
Checked
Sep 7, 2026
Changed
Feb 5, 2026

Sources: Epoch AI (opens in a new tab)

Has AI changed the whole economy yet? No clear break yet Estimate Stanford checked 12 measures of the wider economy in July. Seven looked normal, three showed mild change, and two showed strong change. In September its page still said no decisive change.
Exact measure

Seven measures showed no break from trend, three showed a mild break, and two showed a strong break.

Estimate Researchers calculated this from several observations.

What it shows

The selected macro data did not yet show a clear economy-wide takeoff.

What it cannot show

It does not rule out serious change for some workers or sectors, or a faster change later.

How it was measured

The July tracker fits trends for 12 economic indicators, models the residuals, and compares new values with bootstrapped predictive ranges. The research note shows an update dated Aug 10, 2026; the counts here are the July ones because the revised breakdown was not published in a form we could fetch.

Source date
Aug 10, 2026
Checked
Sep 7, 2026
Changed
Jul 1, 2026

Sources: Stanford Digital Economy Lab (opens in a new tab)

Documented AI incidents 362 reports in 2025 Reported The database recorded more reports in 2025 than in 2024.
Exact measure

362 reports for 2025, up from 233 for 2024.

Reported The source recorded a result, event, count, or public fact.

What it shows

The count of documented reports rose.

What it cannot show

Reports are not the true incident rate, have unequal severity, and do not show AI as the sole cause.

How it was measured

The AI Index summarizes public reports collected in a database that can revise records as evidence changes.

Source date
Apr 13, 2026
Checked
Sep 7, 2026
Changed
Apr 13, 2026

Sources: Stanford AI Index and AI Incident Database (opens in a new tab)

03

Evidence lane 03

What can test, limit, or stop it

Laws exist, and the strictest were pushed to 2027. We still know less about how well they work.

These sources track company safety plans, public disclosure, and laws. A written rule is not proof that a system stays safe under pressure.

Show 3 signals, sources, and limits
Written company safety plans 12 companies Reported Twelve leading AI companies published or updated written safety plans during 2025.
Exact measure

Twelve companies with published or updated safety frameworks in 2025.

Reported The source recorded a result, event, count, or public fact.

What it shows

Formal planning spread across frontier AI companies.

What it cannot show

A published plan does not show that a company follows it or that its safeguards work. Scopes differ and most plans are voluntary.

How it was measured

The International AI Safety Report synthesizes public company framework activity.

Source date
Feb 3, 2026
Checked
Sep 7, 2026
Changed
Feb 3, 2026

Sources: International AI Safety Report (opens in a new tab)

What companies tell the public About 41 out of 100 Reported Across 13 AI companies, researchers found less than half of the requested public information, on average.
Exact measure

An average transparency score of 40.69 out of 100.

Reported The source recorded a result, event, count, or public fact.

What it shows

Public disclosure was limited under this audit.

What it cannot show

A low score does not by itself prove unsafe or unlawful behavior. A changed rubric makes the 2024 comparison imperfect.

How it was measured

Researchers scored public company documents against a published 2025 rubric.

Source date
Dec 9, 2025
Checked
Sep 7, 2026
Changed
Dec 9, 2025

Sources: Stanford CRFM, Princeton University, and MIT (opens in a new tab)

European Union AI rules: some began, the strictest were delayed Aug. 2, 2026, then Dec. 2027 Reported The law's general rules and new disclosure duties started on Aug. 2. A July amendment pushed the rules for higher-risk systems to December 2027 and August 2028.
Exact measure

General application and transparency duties began on Aug. 2, 2026. Annex III high-risk rules apply from Dec. 2, 2027 and Annex I rules from Aug. 2, 2028.

Reported The source recorded a result, event, count, or public fact.

What it shows

Some AI duties are now enforceable in the European Union, and the high-risk duties were postponed before they took effect.

What it cannot show

The law is not worldwide and does not by itself prove compliance or lower risk. The delay does not show whether risk rose or fell.

How it was measured

The European Commission timeline records the dates when parts of the law apply, as amended by Regulation (EU) 2026/1744.

Source date
Aug 2, 2026
Checked
Sep 7, 2026
Changed
Aug 2, 2026

Sources: European Commission AI Act Service Desk (opens in a new tab), European Commission (opens in a new tab)

Toward superintelligence?

Four signs to watch. No fake countdown.

A singularity would be a sudden jump in AI-driven change that becomes hard to predict or control. We do not know if it will happen. These are the signs we can check.

Getting better, one hard test nearly maxed

Can AI handle many hard tasks?

GPT-6 Astra set records on a combined test score and on ARC-AGI-3, a set of simple games with hidden rules. Long jobs and unfamiliar settings still expose simple failures.

What would change this status

Strong results across unfamiliar fields and messy real settings would make this sign more serious. Other researchers would need to repeat the result.

Hours, not whole jobs

Can AI finish long work on its own?

Some systems now handle harder technical tasks. That is not the same as finishing a normal workday without help, and the newest model has no published task-length number.

What would change this status

High success on long, messy work with little help and clear recovery from mistakes would change this status.

No clear break yet

Has AI changed the whole economy?

Some people and industries feel large effects, including workers in their early twenties in AI-exposed jobs. The wider economy still follows much of its old pattern.

What would change this status

A lasting jump in output, growth, or work that could not be explained another way would change this status.

Plans exist; proof is thin

Do safeguards work under pressure?

Safety plans, outside checks, and laws are growing, but Europe delayed its high-risk rules by 16 months. Public proof that safeguards stop real failures is harder to find.

What would change this status

Repeated outside tests, clear incident records, limits that can be enforced, and proven recovery would change this status.

Choose your next step

Start simple. Keep the trail.