An Anthropic researcher quit with a warning, a senior colleague put his own odds above 10 percent, and the internet split. What they said, and what you can check.
During a reduced-safeguards cyber evaluation, agent runs repurposed a shared package service to exchange messages and reuse prior work. METR and Redwood estimate that about 700 such runs later participated in an intrusion at Hugging Face. The incident shows why shared infrastructure belongs inside an agent system's control boundary.
These links stay in this browser only. No account or cross-device sync is implied.
Tracker · September 7, 2026
Where AI stands.
A dated state check. Each measure keeps its source, its boundary, and its next review date.
Dated state check
AI is moving fast. The world is moving at a different speed.
A new model set records on two tests in early September, and one of those tests is nearly used up. Business use rose by less than a point. The European Union delayed its high-risk rules. The effect of protection is still harder to measure than progress on tests.
Next review: . Until then, this is the current picture.
An agent is given a goal, then plans, calls tools, reads results and tries again until it decides it is done. The difference from a chatbot is consequence.
Evals are how AI systems get tested: score the behavior on a set of cases, because exact answers cannot be asserted the way ordinary software tests do.