Superalignment

AI Progress Briefing /

A new AI model nearly maxed out one test. Europe delayed its rules.

Six changes since Aug. 18 show what AI can do, where people use it, and whether safeguards can keep up.

This is not a countdown. GPT-6 Astra set a record on a combined score from more than 50 tests and nearly maxed out ARC-AGI-3, a test built to resist memorized skill. AI use in United States businesses rose by less than a point. Europe pushed its strictest AI rules to 2027. Long computer jobs still fail often.

Six changes

The evidence that moved

Each item says whether it was reported, estimated, forecast, or assessed. Open the source before carrying the number into another argument.

Reported

GPT-6 Astra scored 99.9 on a test built to resist memorized skill

ARC-AGI-3 is a set of simple games with hidden rules. With one setup, Astra beat nearly every game about as efficiently as first-time human players. With the plain setup it scored 62.7%. The previous OpenAI model scored 7.8% in July.

Why it matters: A test designed to separate practiced skill from new learning has almost no gap left. Its makers say that is not proof of human-level AI, and the games are small and fixed.

Open the source (opens in a new tab)

Estimate

A combined score from more than 50 AI tests hit a new high

GPT-6 Astra scored 169 on Epoch's index, first of 267 models. The top score in our August edition was 161.65.

Why it matters: The combined score keeps rising at about its recent pace. It is a scale for comparing models, not a percent or a distance to human-level AI.

Open the source (opens in a new tab)

Reported

AI use in United States businesses rose to 22.4%

In the two weeks to Aug. 23, 22.4% of businesses with employees said they used some AI, up from 21.5% in July. 25.9% expected to within six months.

Why it matters: Use is spreading, but slowly. The survey does not measure how much a business lets AI decide.

Open the source (opens in a new tab)

Reported

Young workers in AI-exposed jobs fell 19% behind

Stanford researchers found that employment for 22 to 25 year olds in jobs most exposed to AI sat about 19% below similar workers in less exposed jobs, up from 15% a year earlier, in one payroll dataset through June.

Why it matters: The whole economy shows no sudden break, but one group does. The authors say the pattern is descriptive, not proof of cause, and they find no wide job loss.

Open the source (opens in a new tab)

Reported

Europe delayed its high-risk AI rules by 16 months

An amendment in force since July 27 moved the rules for higher-risk AI to December 2027 and August 2028. General rules and disclosure duties still began on Aug. 2.

Why it matters: The strictest part of the largest AI law now arrives after the systems it was written for.

Open the source (opens in a new tab)

Assessment

OpenAI paused a long-running model after monitored use exposed new failures

OpenAI says monitored internal use exposed problems that tests before release had missed. It added checks across the full chain of actions and later restored limited access. We found no later update.

Why it matters: A system can look safe one step at a time and still fail across a long job. OpenAI reported this; no outside group verified the failure or the fix.

Open the source (opens in a new tab)

What this adds up to

Faster tests. Slow spread. Safeguards moved later.

Two simple stories both fail. A test built to catch memorization has almost nothing left to catch, yet one in five businesses use AI and the wider economy shows no break. Safeguards moved the other way: the strictest rules in the largest AI law now arrive after the systems they were written for. Ask a smaller question: when people give an AI system a new ability or a new kind of access, what check was built for that exact use, and has it worked under pressure?

Which safeguard was built for a system that learns new tasks about as efficiently as a person, and has anyone tested it?