Superalignment

AI Progress Briefing /

AI test scores rose fast. The wider economy did not.

Six changes show what AI can do, where people use it, and whether safeguards can keep up.

This is not a countdown. AI scores are rising fast on many tests. Long computer jobs still fail often. The wider economy has not made one clear break. We know even less about whether safeguards work.

Six changes

The evidence that moved

Each item says whether it was reported, estimated, forecast, or assessed. Open the source before carrying the number into another argument.

Reported

More European Union AI rules took effect

General rules and new duties to tell people about AI began on Aug. 2. Some rules for higher-risk systems start later.

Why it matters: In one large market, some safety rules are now law, not a promise.

Open the source (opens in a new tab)

Assessment

OpenAI paused a long-running model after monitored use exposed new failures

OpenAI says monitored internal use exposed problems that tests before release had missed. It added checks across the full chain of actions and later restored limited access.

Why it matters: A system can look safe one step at a time and still fail across a long job. OpenAI reported this; no outside group verified the failure or the fix.

Open the source (opens in a new tab)

Estimate

The wider economy did not show a sudden break

Stanford checked 12 measures of the economy. Seven looked normal, three showed mild change, and two showed strong change.

Why it matters: Some jobs can change fast even when the economy as a whole does not.

Open the source (opens in a new tab)

Reported

A leading AI setup finished only one in five long computer tasks

It fully finished 20.6% of 108 tasks. Its average partial-credit score was 54.8%.

Why it matters: Long computer jobs still break often. Results can also change with the tools and setup around the model.

Open the source (opens in a new tab)

Reported

A hard math test corrected 42% of its questions

The test makers fixed 135 questions and removed 12 before releasing a new version.

Why it matters: A score can move because AI got better, because the test changed, or both.

Open the source (opens in a new tab)

Reported

One AI test reached the edge of what it can measure

The test posted an early result beyond 16 hours, then warned that numbers there are unreliable.

Why it matters: A bigger number can look like progress even when the test is no longer strong enough to support it.

Open the source (opens in a new tab)

What this adds up to

Faster tests. Uneven effects. Safeguards we still need to prove.

Two simple stories both fail. AI is not standing still, but society has not crossed a proven point of no return. Ask a smaller question: when people give an AI system a new ability or a new kind of access, what check was built for that exact use, and has it worked under pressure?

Where are people giving AI the most power with the least-tested safeguards?