A new AI model nearly maxed out one test. Europe delayed its rules.
Six changes since Aug. 18 show what AI can do, where people use it, and whether safeguards can keep up.
6 evidence changes
AI Progress Briefing
Each issue separates what changed, what it means, and what the evidence cannot show.
The wire
A combined score from more than 50 AI tests hit a new high (epoch.ai)
GPT-6 Astra scored 169 on Epoch's index, first of 267 models. The top score in our August edition was 161.65.
Frontier benchmark composite: 169 ECI (epoch.ai)
GPT-6 Astra, released Sep 3, holds the top score in Epoch's index, up from 161.65 for GPT-5.6 Sol in our Aug 18 snapshot.
GPT-6 Astra scored 99.9 on a test built to resist memorized skill (arcprize.org)
ARC-AGI-3 is a set of simple games with hidden rules. With one setup, Astra beat nearly every game about as efficiently as first-time human players. With the plain setup it scored 62.7%. The previous OpenAI model scored 7.8% in July.
AI use in United States businesses rose to 22.4% (census.gov)
In the two weeks to Aug. 23, 22.4% of businesses with employees said they used some AI, up from 21.5% in July. 25.9% expected to within six months.
United States business AI use: 22.4% now (census.gov)
About one in five employer businesses reported some AI use in the two weeks to Aug 23, up from 21.5 percent in July. The expected six-month rate was 25.9 percent.
Young workers in AI-exposed jobs fell 19% behind (digitaleconomy.stanford.edu)
Stanford researchers found that employment for 22 to 25 year olds in jobs most exposed to AI sat about 19% below similar workers in less exposed jobs, up from 15% a year earlier, in one payroll dataset through June.
Economy-wide transformation: 7 no, 3 mild, 2 strong (digitaleconomy.stanford.edu)
Stanford found no decisive sign of an economy-wide AI takeoff across 12 macro indicators. Its project page still said so on Sep 7, and no newer count breakdown was retrievable.
European Union legal milestone: Applied Aug 2; high-risk delayed to 2027 (ai-act-service-desk.ec.europa.eu)
General application and transparency duties began on Aug 2, 2026. The Digital Omnibus, in force since July 27, moved the Annex III high-risk rules to Dec 2, 2027 and Annex I rules to Aug 2, 2028.
Europe delayed its high-risk AI rules by 16 months (digital-strategy.ec.europa.eu)
An amendment in force since July 27 moved the rules for higher-risk AI to December 2027 and August 2028. General rules and disclosure duties still began on Aug. 2.
OpenAI paused a long-running model after monitored use exposed new failures (openai.com)
OpenAI says monitored internal use exposed problems that tests before release had missed. It added checks across the full chain of actions and later restored limited access. We found no later update.
Long computer workflows: 20.6% perfect (osworld-v2.xlang.ai)
The best reported setup fully completed about one in five OSWorld2 workflows and received 54.8 percent partial credit. The leaderboard showed no newer top result on Sep 7.
FrontierMath Tiers 1-4, version 2 correction (epoch.ai)
A large benchmark correction changed the test itself, so version history is part of interpreting model progress.
Task-horizon doubling trend: 128.744 days (metr.org)
METR's post-2023 fit says the measured horizon doubled about every four months. The fit has not been updated since May 8, 2026.
Task-completion horizon update, May 8, 2026 (metr.org)
A reported frontier result moved beyond the range METR considers reliable on its current suite.
Documented AI incidents: 362 in 2025 (hai.stanford.edu)
The AI Incident Database recorded 362 reports for 2025, up from 233 for 2024. The AI Index has published no newer annual count.
50 percent task horizon: 11.98 hours (metr.org)
On a clean technical task suite, Claude Opus 4.6 reached half success near tasks that took experts about 12 hours. METR's page last changed on May 8 and lists no result for GPT-6 Astra.
Frontier benchmark trend: +14 ECI per year (epoch.ai)
Epoch's fitted frontier score has risen about 14 points a year since the first reasoning models in September 2024. Our Aug 18 edition recorded 15.5; the page now shows 14.
AI compute stock growth: 3.4 times per year (epoch.ai)
Epoch estimates that the world's stock of AI computing power has grown fast since 2022. The trends page still carries its Feb 5 figure.
Compute held by five hyperscalers: 71% estimated (epoch.ai)
Epoch estimates that five large operators hold most global AI compute. The figure was unchanged on Sep 7.
Current loss-of-control capability: Not established (internationalaisafetyreport.org)
The 2026 international report says current systems do not yet have the combined abilities needed for loss of control. Its publications page listed nothing newer on Sep 7.
Company safety frameworks: 12 companies (internationalaisafetyreport.org)
Twelve companies published or updated frontier safety frameworks during 2025. The report has not been updated since February.
Company transparency: 40.69 out of 100 (crfm.stanford.edu)
Thirteen AI companies disclosed less than half of the information requested by the 2025 audit, on average. The December 2025 edition is still the latest.
Researcher forecast medians: 2047 and 2116 (arxiv.org)
The survey median was 2047 for machines beating humans at every task, and 2116 for all jobs becoming fully automatable. No newer edition of the survey has been published.
Archive
Six changes since Aug. 18 show what AI can do, where people use it, and whether safeguards can keep up.
6 evidence changesSix changes show what AI can do, where people use it, and whether safeguards can keep up.
6 evidence changes