Blog · Proposed taxonomy · August 6, 2026 · Failure modes · Updated August 15, 2026
Only one of the three failures needs a villain
A proposed taxonomy of how superintelligence fails the people it was meant to serve. Capture has a villain, divergence and drift do not, and defenses built around bad actors leave two doors open.
In brief
Superintelligence can fail everyone it was meant to serve in at least three ways. Capture is a few actors deciding what it may do. Divergence is good actors, careful teams, and the wrong outcome anyway. Drift is alignment established once and trusted forever while the world moves. Only capture involves a villain, and security thinking aimed only at villains leaves the other two doors open. This is a proposed grouping of outcomes, not a proven partition.
Scope: this piece is about outcomes at the frontier, where a system may be more capable than the people supervising it. The evidence available today is about ordinary deployments, so treat the argument as a structural one rather than as a finding, and treat the examples as illustrations of a mechanism rather than measurements of a rate.
Ask someone what could go wrong with superintelligence and the answer usually has a villain in it: a rogue lab, a hostile state, a bad actor with the weights. The villain story is not wrong. It is one part of the problem, and structurally it is the most comfortable part, because villains can be identified, sanctioned, and kept out. The other two have no one to arrest.
What kind of list this is
Three doors, proposed. This is a grouping of outcomes, of ways the technology ends up failing the people it was supposed to serve, and not a list of causal mechanisms or of governance failures. That distinction matters, because a single mechanism can open more than one door and a single door can be reached by several mechanisms.
We do not claim the set is exhaustive, and we have no exhaustiveness argument to offer. What we claim is narrower: these three are reached by different defenses, so a program that addresses one and calls it safety has a gap it can name.
The field already has vocabulary here, and most of it maps cleanly. Misuse and concentration of power are capture. Specification gaming, reward hacking, and the parts of loss of control that route through misspecified objectives are divergence. Distribution shift, model updating, and value drift are drift. Where our grouping differs from the standard lists is that it cuts by who can close the door rather than by what went wrong inside the system, which is why capture and divergence land in separate buckets even when the same deployment produces both. A reader who prefers the established categories loses nothing important by translating.
Capture: the door with a villain, sometimes
Capture is a few actors coming to decide what superintelligence may do, and for whom. The villain version is familiar. The version that should worry you more requires no malice at all: capture happens by default whenever a capability can only be checked by the people who built it. A company whose auditors cannot read the books is run by whoever writes the books, whatever the org chart says. Concentration follows from unverifiability, not from evil intent, which is why "we are the good guys" is not an answer to it.
Divergence: good actors, wrong outcome
Divergence is careful teams, real safeguards, and the wrong result anyway. Intent was never fully stated, so the system optimized what it could see and shipped on what looked right. This is false convergence operating at mission scale. Nobody lied. The dashboard was green. The requirement nobody elicited was never tested, and nothing announced its absence.
We used to write that divergence is by far the most common of the three. That claim cannot be supported: nobody has a sample of superintelligence outcomes, and the deployment failures we can observe today are evidence about ordinary systems rather than about the frontier. What we will say is that divergence is the door our own work keeps arriving at, and that it is the one screening for trustworthy people cannot close, because the failure does not route through anyone's character. It routes through the gap between what was meant and what was specified.
Drift: alignment with an expired date
Drift is alignment established once and trusted forever after. The world moves, the policy changes, the model updates, and a system aligned to a world that no longer exists keeps acting on permission it earned somewhere else. Drift begins from a true statement, that the system really was checked, and fails through the tense of the verb.
The full case is a separate essay: why every other certification regime treats evidence as perishable, what the three clocks are, and what re-verification costs. Note the vocabulary collision before you go: drift here is the governance outcome, not the machine-learning sense of data or concept drift, though the second is one way to reach the first.
One missing part behind three doors
Line the three up and they share a structure. Capture is uncheckable capability concentrating. Divergence is unchecked intent diverging. Drift is a check that was never repeated. Three doors, one missing part: nobody independent can check the work, keep checking it, and be believed.
This is why our position is that verification, not weight distribution, is what democratizes superintelligence. The steelman of the distribution view is real: open weights genuinely diffuse capability, enable outside scrutiny of the artifact, and prevent some forms of capture. Meta's Llama releases and the open-weight ecosystem around them are the strongest version of that case, and the scrutiny they enabled is a real public good. But handing out weights opens none of the three doors' locks by itself. It multiplies the number of deployments that can diverge and drift, and access to weights is not the same as the ability to check whether a particular deployment does what an organization requires. You do not democratize superintelligence by giving everyone an uncheckable thing. You democratize it by making readiness checkable by people who should not have to trust the builder.
The governance proposals we have read cluster into three families, and our reading, which we would like corrected, is that each mostly answers one door. Risk-tiered regulation of the kind the EU AI Act establishes, and the lab risk policies that set capability thresholds, are aimed most squarely at capture and at catastrophic misuse. Evaluation and red-teaming regimes are aimed at divergence, and inherit its hardest problem, which is that an evaluation cannot test a requirement nobody stated. Post-deployment monitoring commitments are the only family that takes drift seriously, and they are the least specified of the three. We know of no proposal that carries all three, and we would genuinely like to be shown one.
What would change our mind: evidence that broad weight access, by itself, measurably reduces divergence and drift failures rather than multiplying their surface; a demonstrated capture of a system whose verification was genuinely public; or a governance proposal that answers all three doors, which would make this taxonomy's central complaint obsolete. Any of them would force a rewrite, and we would write it.
The open question is the one none of the three doors answers alone: who checks, how often, and why should the rest of us believe them?
FAQ
What are the three failure modes of superintelligence?
Capture, divergence, and drift, as proposed here. Capture is a few actors deciding what superintelligence may do. Divergence is good actors with careful safeguards reaching the wrong outcome because stated objectives did not carry full intent. Drift is a system aligned once and trusted forever while the world it was aligned to changes. Only capture involves a bad actor, and the set is a proposed grouping rather than a proven partition.
How do these relate to misuse, specification gaming, and distribution shift?
They are the same territory cut differently. Misuse and concentration of power fall under capture, specification gaming and reward hacking under divergence, and distribution shift and value drift under drift. The grouping here cuts by which defense closes the door rather than by what went wrong inside the system.
Does open-sourcing AI models prevent these failures?
Open weights diffuse capability, enable outside scrutiny, and complicate some capture scenarios, but they do not by themselves make a given deployment checkable by the organization relying on it, and they multiply the deployments exposed to divergence and drift. The position argued here is that verification is the part that democratizes.
Corrections
August 15, 2026. The article claimed divergence is "by far the most common of the three doors." No sample of superintelligence outcomes exists, so the claim has been withdrawn and replaced with what we can say.
August 15, 2026. It also claimed that "every governance proposal we have read is an answer to at most one door" without naming a single proposal. That sentence now names the families it is talking about and is labeled as our reading, open to correction.
August 15, 2026. The piece now states that the three-door set is a proposed grouping of outcomes rather than an exhaustive partition, maps it to the field's established categories, and flags the collision between drift in this sense and drift in the machine-learning sense. The drift section is shortened, because Alignment is an orbit, not a proof is now the full treatment.