Superalignment
Watercolour of a small sailboat holding its heading in a winding estuary channel under a vast evening sky, its wake tracing a long continuous curve between sandbanks.

In brief

Elevators, aircraft, auditors, and surgeons operate under evidence that expires on a schedule, because the world their certification described keeps moving. AI deployments run the same risk faster, through model updates, policy changes and drifting data, and the approvals we have seen carry no expiry at all. Alignment is not a property you establish once; it is a relationship between intent, system, and world that must be re-verified as all three move, an orbit held by continuous correction rather than a theorem proved once.

Scope: this is the full treatment of drift, one of three ways a system can fail the people it was meant to serve. The argument is about deployments under ordinary oversight; where the reasoning reaches toward frontier systems, it is doing so by structural analogy and says so.

There is a placard in every elevator you have ever ridden, and it has a date on it. The machine did not stop being an elevator when the old date passed. The inspection expired, so the permission expired, and someone came and looked again. Nobody finds this strange. The same ritual runs through aircraft, boilers, auditors, surgeons, food lines and fire doors, each on its own clock: a pilot's flight review every 24 calendar months, an elevator's full-load safety test every five years on top of its annual one, an audit every year with the lead partner rotated off after five, a physician's certification renewed on a five or ten year cycle. Evidence ages, so certification is a subscription, never a purchase.

Now look at how organizations treat an AI approval. The system passed review in March. It is October. The model behind the API has been updated twice, the refund policy was amended in June, the customer mix shifted when the new market launched, and the approval from March is still the document everyone points to. The fastest-moving system in the building is the only one working on a permanent license.

Two honest qualifications, because the contrast is easy to overstate. Plenty of software assurance outside safety-critical industries never expires either, so AI approvals are not uniquely careless; they are joining a bad habit rather than inventing one. And we have not surveyed how organizations date AI approvals, so read the claim as what we have seen in the deployments we have worked on and in the review documents that get shown to us, which is a small and self-selected sample.

Three clocks, all running

An AI deployment sits under three clocks at once. The system changes: models update silently behind APIs, prompts get edited, tools swap versions. The rules change: policies are amended, regulations arrive, the org reshuffles who may approve what. And the world changes: the data drifts, the customers change, the adversaries adapt. Evidence gathered in March describes the March system under March rules in the March world. Every clock tick since is unverified inference.

None of this requires anything to go wrong. That is what makes drift the quietest failure mode: it begins from a true statement. The system really was checked. The failure is treating the check as a possession instead of a measurement, and measurements age.

The research field has been converging on the same conclusion from its own side. Evaluations are statements about a model version under test conditions, and the strongest evaluation organizations publish that limitation themselves. The UK AI Safety Institute writes that independent evaluations "can provide a snapshot of a given system's capabilities and vulnerabilities" and that the science is too nascent for them to act as a certification function. METR, reporting on a frontier model, wrote that its pre-deployment evaluations were not sufficient to rule out large-scale risks and likely underestimated the model's capabilities. What we add is the organizational half: the expiry is not a defect in the evidence, it is a property of all evidence, and an operation that does not date its approvals has decided not to know.

The orbit

Here is the picture we keep coming back to, and it is not a metaphor we chose for decoration. Three bodies in orbit: what you actually intend, what the system understood, and what the world requires. The three-body problem has no closed-form solution. You cannot solve it once, write down the answer, and walk away. You integrate it, step after step, correcting as the bodies move, or you lose it. Small perturbations compound. The elegant configurations are the fragile ones.

Alignment is that kind of object. Not because the mathematics is mystical, but because the thing being aligned is a relationship among three moving parts, and relationships do not stay proved. A system aligned to a world that no longer exists keeps acting on permission it earned somewhere else, and every quarter of silence widens the gap between the evidence on file and the system in production.

Not a proof to be found, but an orbit to be held.

What keeping it takes

The practices are unglamorous, which is a point in their favor, because every one of them already exists in some older discipline. Date every piece of evidence and declare its horizon when it is gathered, not when it is questioned. Treat an expired approval as no approval, the way the elevator inspector does. Re-verify on the clocks that actually tick, model updates, policy changes, distribution shifts, rather than on the calendar alone. And build the re-checking into the system's permission to act, so that autonomy is spent against current evidence and renewed by fresh observation, instead of drawn indefinitely against a measurement nobody remembers making.

None of this solves alignment, here or in the limit. That is the argument. The problems worth having are the ones you hold, with instruments, attention, and the humility to keep looking at a thing you already looked at. The harmony is not something you find. It is something you keep.

Prior art, and what would change our mind

Nothing in the practice section is ours. Recertification and surveillance audits are how professional and safety certification has worked for a century. Site reliability engineering already spends permission against measured evidence in the form of error budgets. Control theory has the receding horizon, where a plan is re-solved every step rather than executed to completion. Continuous monitoring in security assurance replaced the annual audit for the same reason we are arguing here. What we add is the mapping onto AI deployment, where the three clocks tick faster than the review cycle and the evidence tends to be dated implicitly or not at all.

What would change our mind: a demonstration that AI approvals hold up without re-verification, meaning deployments whose original review still predicts behavior after substantial model, policy and population change. A study that tracked approved deployments over a year and found that staleness did not predict incidents would take the argument apart, and we know of no one running it.

Sources

Corrections

August 15, 2026. The certification intervals in the opening are now specific and sourced, and the claim that evaluation organizations publish their own limitation now quotes one.

August 15, 2026. The article opened by claiming that every certification regime treats evidence as perishable except AI. That is too strong: much ordinary software assurance does not expire either. The claim is now scoped, and the observation about undated AI approvals is labeled as what we have seen rather than as a survey finding. This piece also now carries the drift argument in full, which Only one of the three failures needs a villain previously duplicated.