Superalignment

In plain terms

Four constraints on AI readiness work like walls, not budgets: no amount of model capability buys through them. Competent engineering means building with them, the way network engineers route around the speed of light instead of waiting for it to improve.

The four structural limits are engineering constraints on AI readiness, established in Convergence Programming. They are not curiosities about today's models. Read as constraints, they say what no amount of model capability will buy you; read constructively, they prescribe an architecture.

The plain version

Some limits in engineering are budgets: spend more, get more. Others are walls: no spend gets through them, and competent engineering means building with the wall rather than against it. You cannot out-spend the speed of light on a trading link; you route around it. These four are walls. "Wait for a better model" therefore buys nothing here: model capability has improved on a measured curve for years, and each of these limits is indifferent to it, because each is a fact about evidence rather than about models.

The four limits

  1. Specification debt. A prompt is an information bottleneck. Constraints neither encoded in it nor revealed by later interaction are simply absent, and no amount of model capability turns unobserved world structure into evidence. The hard requirements are the idiosyncratic ones: this organization, this policy, this situation.
  1. No free readiness. If two possible worlds produce the same observation history but demand different safety decisions, no rule that sees only that history can certify both. Safety stays reachable, and it is reached by making the distinguishing observation: elicit it, monitor it, or carry it openly as risk.
  1. Critical blockers do not average. One unresolved critical blocker changes the type of decision rather than lowering a score. A large enough pile of routine successes will hide a single critical failure under any average-gap threshold. Readiness is a checklist, and a checklist is read line by line.
  1. Autonomy spends evidence. Every accepted action draws down a finite budget of residual uncertainty. A system safe enough to draft one email may be unsafe to run a support desk for a month without reopening the loop. Long-lived agency is a sequence of evidence-bounded windows, not one standing approval. Permission recedes. It does not accumulate.

The limits in the field

Each limit is ours in name and formulation; each has recognizable cousins in the research literature, which is part of the argument that they are structural rather than parochial.

Specification debt is the working face of the alignment problem itself: intent does not survive being written down, and the remainder is silently absent rather than visibly missing. Machine learning has its own formal version in underspecification (D'Amour et al., 2020): pipelines whose training and test constraints leave the deployment-relevant behavior undetermined, so indistinguishable models diverge where nobody was looking.

No free readiness is an identifiability statement: observation histories that cannot distinguish two worlds cannot license different decisions about them. It is the organizational cousin of the eliciting latent knowledge worst case, where training signals that cannot distinguish a truthful reporter from a plausible one cannot select for truth, and its practical corollary is everywhere in the evaluation literature: evaluations demonstrate capability, never its absence, which is why safety-case inability arguments are the ones that expire.

Critical blockers do not average restates, for AI readiness, a rule every safety-critical discipline already enforces: worst-case properties are not purchasable with good averages. Aviation does not certify an aircraft whose failures are rare but catastrophic on the strength of its excellent mean performance, and a security review that finds one exploitable path does not net it against the paths that held. The limit matters for AI because benchmark culture reports means, and a deployment's real question is about the tail.

Autonomy spends evidence is the readiness version of the fact that evidence dates. Certificates expire, audits recur, and credentials are renewed on a schedule, because the world the evidence described moves; in AI deployments the same clock runs faster, through model updates, policy changes, and drift. The limit's specific content is that action itself consumes the budget: every accepted action is a draw against residual uncertainty, so even an unchanged world spends down an approval.

What follows from them

Each limit names the failure that appears when it is ignored. Ignore specification debt and you get the requirement that was never elicited, so never tested. Ignore no-free-readiness and you certify on a history that could not distinguish safe from unsafe. Ignore the third and a single critical failure hides under a good average. Ignore the fourth and last year's approval is still authorizing this year's actions. All four routes end in false convergence.

The constructive reading is an architecture, and the prescription is unusually literal: record assumptions as tentative anchors, actively seek observations that could falsify readiness, turn every discovered anchor into a standing regression obligation, and gate action by the scope of its evidence rather than by model confidence or a user's approval. Remove any one component and the corresponding failure mode returns. That architecture, made operational, is Verity.

On verifier independence

One corollary is worth stating on its own: verifier multiplicity is not verifier independence. Shared model families, shared prompts, shared training data and shared evaluative frames impose a floor. Past that floor, adding critics improves the evidence dashboard and changes nothing about the correct decision. Verifiers that share a blind spot vote as one.

The corollary has an experimental pedigree older than language models. The classic study of N-version programming (Knight and Leveson, 1986) tested the assumption that independently written programs fail independently and rejected it: separate teams, working apart, produced correlated faults, because they shared the hard parts of the problem and the human tendencies for getting them wrong. Model-based judging adds its own correlations on top, including evaluators measurably favoring outputs resembling their own (Panickssery et al., 2024). The implication is not that model verifiers are useless; it is that the verification budget must buy at least one check that does not share the system's view of the world, an execution, a ground-truth lookup, a human with different incentives, before multiplicity means anything.

Convergence infrastructure does not eliminate uncertainty. It prevents uncertainty from being mistaken for readiness.

What we do not claim

The formal statements and their assumptions live in the paper, which states its scope plainly: bounded, executable task worlds with a defined action surface. The limits are results about what observation histories can certify, not predictions about model behavior, and they do not say automation is unsafe. They say readiness claims are bounded by evidence, and they locate exactly where the evidence has to come from.

Sources