Superalignment
Watercolour of a ship's chronometer and a signal mirror laid on a chart table beside a window overlooking a grey-green estuary.

In brief

This page is a maintained registry of the site's substantive claims. Each entry states the claim, what kind of claim it is, how far it reaches, what supports it, what argues against it, and the specific observation that would make the authors retract it. Entries carry a status and a review date, and the changelog at the end records every time a claim was weakened, revised or retired.

Scope: this page covers claims made across this site. Each entry names its own regime, because a claim about enterprise deployments and a claim about frontier oversight are not the same kind of claim and should not borrow each other's evidence.

A claim you cannot name the killer for is an advertisement. So this page lists what we claim, what would refute it, and what has already changed. It is maintained rather than published once: entries carry a status and a review date, and the changelog at the end is the public history.

Stated-in-advance falsifiers are cheap, which is the argument for them. They cost nothing to publish, they make motivated reasoning harder for us, and they give a reader a test that does not require trusting us. The tell runs the other way too: a vendor page, lab announcement or research agenda that states no falsifier anywhere has told you something about what kind of document it is.

Entries are maintained by the Superalignment Team; our editorial policy covers who that is and what we sell.

How to read an entry

Claim type is one of: hypothesis, empirical finding, literature synthesis, proposed taxonomy, logical deduction, or normative position. The type sets what evidence could settle it, and mixing types is how a position gets defended with the wrong kind of argument.

Status is active, weakened, revised or retired. Weakened means we still hold the claim but have narrowed it. Retired means we no longer assert it, and the entry stays on the page rather than disappearing.

Confidence is our own, stated in words rather than invented numbers.

C1. False convergence is a common failure mode of AI deployments

  • Claim: deployments frequently fail because every visible signal says the system is ready while requirements nobody elicited go unmet, rather than because the model was not capable enough.
  • Type: hypothesis.
  • Scope: ordinary enterprise deployments under human approval. It says nothing about frontier systems.
  • Status: weakened, August 15, 2026. We previously called it the dominant failure mode.
  • Confidence: moderate for the mechanism, low for any claim about frequency.
  • Best support: the mechanism is well established under other names, including underspecification, reward hacking and proxy-metric failure, and every deployment review we have run has surfaced requirements the build had not elicited.
  • Best opposing evidence: no incident corpus records why deployments end, so the frequency half of this claim rests on our own sample, which is small and selected by the fact that people call us when something is wrong.
  • Falsifier: incident data from real deployments showing failures traced predominantly to model capability or adversarial action rather than to unelicited requirements and proxy metrics.
  • Weakness in the falsifier: no such corpus exists, so this claim currently cannot be settled either way. That is a defect in the falsifier, not a strength of the claim.
  • Introduced: August 4, 2026. Last reviewed: August 15, 2026.

C2. Most enterprise AI pilots do not reach production

  • Claim: the majority of enterprise pilots end before production use.
  • Type: literature synthesis.
  • Scope: enterprise pilots as reported in vendor and analyst surveys, 2024 to 2026.
  • Status: active, with a known unit problem.
  • Confidence: moderate.
  • Best support: several independent surveys point the same direction, and we retired a larger number for this one after auditing its sources.
  • Best opposing evidence: the strongest supporting figures are organization-level, meaning they count organizations that have scaled rather than pilots that reached production. Our claim is pilot-level, so it is inferring across units, which is the error the audit post exists to warn about. We are re-checking the primaries and will restate this entry with the correct unit.
  • Falsifier: a registered-methodology study with a fixed denominator and a conflict-free funder finding a majority of enterprise agent pilots reaching production.
  • Weakness in the falsifier: almost every survey in this area is run by a party selling something, so the conflict-free condition may never be met.
  • Introduced: August 10, 2026. Last reviewed: August 15, 2026.

C3. Readiness is a property of the trajectory, not the artifact

  • Claim: identical artifacts with different build histories carry different risk, so artifact-only evaluation is structurally blind to the difference.
  • Type: logical deduction from the definition of hidden requirements, with empirical support for the premise.
  • Scope: builds where requirements are discovered during the work, which is most agent work and not, for example, a well-specified numerical routine.
  • Status: active.
  • Confidence: high for the deduction, moderate for its practical weight.
  • Best support: regulated industries already certify build records rather than finished objects, which is the same inference drawn independently. The argument is set out here.
  • Best opposing evidence: ship-and-iterate has an excellent track record wherever errors are cheap, reversible and observable, and much deployed AI work is in exactly that regime.
  • Falsifier: evidence that artifact-level evaluation plus production iteration reliably catches silent-failure requirements before incidents do, at costs organizations actually pay.
  • Weakness in the falsifier: "reliably" and "at real cost" are not operationalized, so a study could satisfy a reader and not us. We would accept a comparison on matched deployments with pre-registered incident definitions.
  • Introduced: July 28, 2026. Last reviewed: August 15, 2026.

C4. Behavioral evaluation cannot carry a deployment decision alone

  • Claim: passed evaluations underdetermine future behavior, because behavior is context-sensitive and evaluation contexts are a small, known-in-advance sample of deployment contexts.
  • Type: empirical finding plus deduction.
  • Scope: frontier and near-frontier models under behavioral evaluation.
  • Status: active.
  • Confidence: high.
  • Best support: published results showing models behaving differently when they can infer they are being evaluated, discussed here.
  • Best opposing evidence: the strongest demonstrations are constructed settings, several are single-lab, and most models tested in follow-up work did not show the behavior. The inference from those results to routine deployment is ours.
  • Falsifier: a demonstrated evaluation protocol whose pass verdicts survive adversarial red-teaming across context shifts, replicated outside the lab that built it.
  • Weakness in the falsifier: "survives" needs a bar, and a protocol could pass on the shifts anyone thought to test while failing on the ones nobody did, which is this claim's own argument turned against its test.
  • Introduced: August 14, 2026. Last reviewed: August 15, 2026.

C5. The binding constraint on overseeing stronger workers is conditions, not capability

  • Claim: human institutions supervise people who know more than the supervisor, using incentives, repeated interaction, independent checkers and error diversity, and AI removes those conditions before it crosses any intelligence threshold.
  • Type: hypothesis, argued from analogy with organizational practice.
  • Scope: oversight of systems that produce work a local approver cannot fully inspect. Extending it to systems that are globally more capable than their supervisors is an argument we make explicitly rather than by assumption.
  • Status: active.
  • Confidence: moderate. This is our most load-carrying original claim and its evidence is the analogy, not a measurement.
  • Best support: the conditions are what audit, professional licensure and separation of duties actually rely on, and each has a recognizable failure when the condition is removed. Set out here.
  • Best opposing evidence: analogies from human institutions may not transfer to systems that can model the checker, and a sufficiently large capability gap plausibly breaks oversight whatever the conditions.
  • Falsifier: a deployed system demonstrating reliable oversight of a stronger model with the conditions absent, meaning no incentive alignment, no error diversity, and holding up under adversarial evaluation.
  • Weakness in the falsifier: "conditions absent" is hard to establish, since a real deployment usually has some condition present by accident.
  • Introduced: August 12, 2026. Last reviewed: August 15, 2026.

C6. Verification, not weight distribution, is what democratizes advanced AI

  • Claim: making readiness checkable by people who did not build a system distributes power over it more than distributing its weights does.
  • Type: normative position, resting on the empirical claim that weight access does not confer deployment-level checkability.
  • Scope: governance of advanced systems. This is an argument, not a finding, and it is the claim on this page most shaped by what we build.
  • Status: active.
  • Confidence: moderate, and we hold it while noting the conflict: we sell verification tooling.
  • Best support: the argument in Only one of the three failures needs a villain, particularly that concentration follows from unverifiability rather than from intent.
  • Best opposing evidence: open weights have produced real outside scrutiny, independent safety research and competitive pressure that closed-weight deployment did not. That is a genuine democratizing effect our position underweights.
  • Falsifier: longitudinal evidence that broad weight access alone measurably reduces divergence and drift failures, or a demonstrated capture of a system whose verification was genuinely public.
  • Weakness in the falsifier: as a normative claim, evidence constrains it without settling it, and a reader who weighs distribution differently can accept all our facts and reject the position.
  • Introduced: August 6, 2026. Last reviewed: August 15, 2026.

What this registry is not

It is not a promise that we are right, and it is not decoration: we think the evidence currently favors every active claim above, or we would not publish them. It is a commitment about process. Positions held in public should come with their exit conditions attached, the same way we argue an AI system's approval should come with the evidence that would revoke it. We would not accept "trust us" from a model, and a reader should not accept it from us.

If you hold any of the falsifying evidence above, or think one of our falsifiers is badly chosen, dulled to be unmeetable, or aimed at a strawman of our own claim, that last failure is the subtlest and we want to hear about it most. Write to [email protected].

Changelog

August 15, 2026. The page became a maintained registry. Every claim now carries a type, a scope, a status, the best evidence against it, a weakness in its own falsifier, and dates.

August 15, 2026. C1 was weakened. It previously read "false convergence is the dominant failure mode of AI deployments"; we cannot support a frequency claim when no corpus records why deployments end, so the entry now claims the mechanism and not the ranking.

August 15, 2026. C2 gained a known unit problem: our pilot-level claim is supported by organization-level figures. The entry says so pending a re-check of the primaries.

August 15, 2026. C3's supporting essay moved: the trajectory argument now lives inside The dashboard is green and the work is wrong.

August 11, 2026. Page first published with six claims and their falsifiers.