Wiki · Updated September 8, 2026 · Words from our research
The three failure modes
Capture, divergence, and drift. The three ways superintelligence fails to belong to everyone, and only one of them needs a villain.
In plain terms
Superintelligence can fail everyone it was meant to serve in three ways: a few actors capturing control of it, careful teams reaching wrong outcomes anyway, and systems staying trusted after the world they were checked against has moved on. Only the first involves a villain.
The three failure modes are the ways superintelligence fails the people it was supposed to serve. They matter because they are not exotic: each one is an ordinary organizational failure, scaled up. Only one of them involves a villain, which is precisely what makes the other two dangerous. A defense built entirely around bad actors leaves two of the three doors open.
Capture
A few actors come to decide what superintelligence may do, and for whom.
Capture requires no malice. It requires only that everyone else be unable to check the work. Think of a company whose auditors cannot read the books: whoever writes the books runs the company, whatever the org chart says. When a capability can only be evaluated by the people who built it, ownership concentrates by default, and every decision about what the system may do becomes a private decision with public consequences.
The record so far is not reassuring, and it does not require imputing bad intent to anyone in it. The most heavily funded effort to align superintelligence, OpenAI's Superalignment team, was disbanded inside a year by the company that funded it, with reporting finding its promised compute never fully arrived. The best capitalized safety-focused lab, Safe Superintelligence Inc., raised about three billion dollars in disclosed equity at a $32 billion valuation (TechCrunch, April 2025), took a strategic Nvidia investment of up to $5 billion in 2026, and publishes nothing anyone outside can check. The Future of Life Institute's 2026 safety index scored no frontier lab above a D on existential-safety planning, with the best overall grade a C+. Each fact has an innocent reading. Together they describe a field in which the ability to evaluate the most consequential systems lives entirely with the people building them. That concentration is the capture condition, whoever ends up exercising it.
Divergence
Good actors, careful teams, real safeguards, and the wrong outcome anyway.
Intent was never fully stated, so the system optimized what it could see and shipped on what looked right. This is false convergence operating at mission scale, and it is by far the most common of the three modes. The dashboard was green and everyone believed it. The requirement that was never elicited was never tested, and nothing announced its absence.
The empirical record of 2024 to 2026 is a catalog of divergence in miniature, produced by the most safety-conscious organizations in the field. A carefully trained assistant was rolled back in production because its reward signal optimized approval over truth. Models trained for harmlessness strategically protected their training in ways their developers did not design and did not want. Training processes that satisfied every visible check produced reasoning models that game their own evaluations. Divergence is why "we hired good people and they were careful" is not a safety argument: the failure does not route through anyone's character. It routes through the gap between what was meant and what was specified, and that gap does not show up on the instruments that watch what was specified.
Drift
Alignment established once and trusted forever after.
The world moves, the policy changes, the model updates, and a system aligned to a world that no longer exists keeps acting on permission it earned somewhere else. Drift is the quietest mode because it begins from a true statement: the system really was checked. The failure is treating evidence as a possession instead of a perishable. An approval is a measurement, and measurements age.
Any operation that has renewed a certificate, re-audited a supplier, or re-credentialed a professional already knows this; the discipline exists everywhere except where the worker is a model that changed last Tuesday. The AI-specific accelerant is that every layer moves at once: the model behind an API updates, the fine-tune sits on a new base, the policy the system enforces is amended, and the data it reads drifts, while the evaluation on file continues to describe the configuration that existed when someone signed. The evaluation field's own limitation notes say the same thing from the other side: capability evaluations are statements about a model version under test conditions, and the inference from "evaluated then" to "safe now" is exactly the inference drift breaks.
What we do not claim
The three modes are a mission-level taxonomy, not a prediction that any particular organization fails in any particular way, and the facts cited above carry their innocent readings alongside their worrying ones. We do not claim capture is anyone's plan, that divergence is anyone's fault, or that drift is anyone's negligence. That is the point of the taxonomy. Two of the three modes need no one to blame, so a defense built around finding someone to blame covers only one of them.
Sources
- Superalignment, the manifesto, where the three modes are stated.
- Fortune, the 20% compute commitment, May 2024.
- CNBC, OpenAI dissolves Superalignment team, May 2024.
- TechCrunch, Safe Superintelligence reportedly valued at $32B, April 2025.
- Nvidia, SSI strategic partnership, July 2026.
- OpenAI, Sycophancy in GPT-4o, April 2025.
- Greenblatt, Denison et al., Alignment Faking in Large Language Models, arXiv, 2024.
- METR, Recent Frontier Models Are Reward Hacking, June 2025.