Superalignment

In plain terms

Superalignment is the alignment problem in the regime where the AI outperforms the people supervising it: how do you check work you could not have produced, and how do you know your check means anything? OpenAI coined the word in 2023 and disbanded the team that carried it within a year. The problem it names is still open, and the labs that absorbed the agenda publish on it under other names: scalable oversight, AI control, and automated alignment research.

Superalignment is alignment when the system is more capable than the people supervising it: the problem of checking work you cannot out-think. The word was introduced by OpenAI in July 2023 as the name of a research team and a four-year program; the team dissolved in May 2024, and the word has since passed into general use for the supervision problem itself, which remains open.

The plain version

Every quality process you have ever used assumes a reviewer who can tell good work from bad. The senior engineer reads the junior's code. The partner checks the associate's filing. The editor reads the draft. Now remove that assumption. The work arrives faster than anyone can read it, or the reasoning behind it is beyond the person signing for it, or both. What, exactly, is the review step now?

That is superalignment. Ordinary alignment leans on a human reviewer who can tell whether the work is good. Remove that assumption and the usual tools stop working: you cannot spot-check reasoning you cannot follow, or sample a volume you cannot read.

Origin of the term

OpenAI announced the Superalignment team on July 5, 2023, in Introducing Superalignment, co-led by Ilya Sutskever and Jan Leike. The announcement made three claims that defined the program. First, stakes: "Superintelligence will be the most impactful technology humanity has ever invented," and could lead to "the disempowerment of humanity or even human extinction." Second, timing: "While superintelligence seems far off now, we believe it could arrive this decade." Third, the technical gap: "Currently, we don't have a solution for steering or controlling a potentially superintelligent AI... Our current techniques for aligning AI will not scale to superintelligence."

The program set itself a deadline, to "solve the core technical challenges of superintelligence alignment in four years," and a budget, "20% of the compute we've secured to date over the next four years." Its central technical bet was to "build a roughly human-level automated alignment researcher": use early superhuman systems to do alignment research on their successors, with a three-part approach of scalable oversight, validation through robustness and automated interpretability, and adversarial testing that included deliberately training misaligned models. Leike laid out the reasoning at length on the 80,000 Hours podcast the following month.

The argument for why a new word was needed was an argument about RLHF, the training method behind deployed assistants. RLHF consumes human judgments of model outputs, so it is bounded by the judge. OpenAI's own framing for weak-to-strong generalization stated it plainly: future humans "will only be able to weakly supervise" superhuman models. Superalignment named the regime where that ceiling binds.

The problem is older than the brand

The word dates to 2023; the problem does not. Norbert Wiener stated it in Science in 1960 ("Some Moral and Technical Consequences of Automation"): if we use a machine "whose operation we cannot efficiently interfere with," we "had better be quite sure that the purpose put into the machine is the purpose which we really desire." Nick Bostrom's Superintelligence (2014) organized a decade of argument about why the purpose-loading problem gets harder as capability grows. "Scalable oversight" was named as one of five open problems in Concrete Problems in AI Safety (Amodei et al., 2016), and the canonical modern statement of the supervision gap, training signals fail for tasks "too complicated for a human to directly evaluate," is Christiano, Shlegeris, and Amodei (2018). The 2018-era proposals that grew from it, iterated amplification, debate, and recursive reward modeling, are the intellectual ancestors of the superalignment agenda; the lineage is treated in detail under scalable oversight.

What the team published

The team's flagship result was weak-to-strong generalization (Burns et al., December 2023): a strong model finetuned on labels from a much weaker supervisor recovers a large fraction of the performance gap, with a GPT-2-supervised GPT-4 reaching roughly GPT-3.5-level performance on NLP tasks under an auxiliary confidence loss. The result is often overstated as near-full recovery; the paper claims substantial partial recovery, weakest in reward modeling, the domain closest to alignment practice.

In December 2023 the program launched Superalignment Fast Grants, $10 million with Eric Schmidt, in grants of $100K to $2M plus graduate fellowships. Work by team members published before and after the dissolution includes scaled sparse autoencoder training on GPT-4 (Gao et al., 2024), CriticGPT ("LLM Critics Help Catch LLM Bugs," June 2024), and prover-verifier games for legibility (Kirchner et al., 2024).

The dissolution

Sutskever's departure was announced May 14, 2024; Leike resigned the same day, and the team's dissolution was confirmed within the week (CNBC, May 17, 2024). Leike's public explanation (May 17, 2024): "Over the past years, safety culture and processes have taken a backseat to shiny products," and "my team has been sailing against the wind. Sometimes we were struggling for compute." Fortune reported, citing six sources, that the 20% compute commitment was never fulfilled and that the team's requests were repeatedly denied. No accounting of the pledge was ever published, which leaves the two framings, "integrated more deeply into research" versus "allowed to wither," formally unresolved; the compute reporting is the strongest public evidence for the second.

The four-year clock stopped at about ten and a half months, from the July 5, 2023 announcement to the May 17, 2024 confirmation of the dissolution. No lab has restated a deadline for solving the problem since.

Where the people went is part of the record. Leike joined Anthropic in May 2024 "to continue the superalignment mission." Sutskever co-founded Safe Superintelligence Inc. in June 2024, "the world's first straight-shot SSI lab," which raised about $1 billion at a $5 billion valuation in September 2024 (Crunchbase News) and a further $2 billion at a $32 billion valuation in April 2025 (TechCrunch), about $3 billion in disclosed equity, then took a strategic Nvidia investment of up to $5 billion (July 2026). Through mid-2026 it had published no research. John Schulman, who inherited alignment at OpenAI, left for Anthropic in August 2024. Daniel Kokotajlo, who resigned in April 2024, co-founded the AI Futures Project and co-authored the AI 2027 scenario. Leopold Aschenbrenner, fired in April 2024 in a disputed leak case, wrote Situational Awareness and founded a hedge fund.

The field after the team

The name faded; the work redistributed. As of 2026 no frontier lab has a team called superalignment, and the live vocabulary is "scalable oversight," "AI control," "automated alignment research," and "recursive self-improvement safety." The main threads:

  • Anthropic runs the closest thing to a successor program: RLAIF via Constitutional AI, deliberately constructed misaligned models (Sleeper Agents, alignment faking), mechanistic interpretability as a verification bet, and safety cases under its Responsible Scaling Policy. In 2026 it reported automated alignment researchers, agent teams running weak-to-strong supervision research, outperforming a human baseline on a scoped task while failing to transfer gains to production scale and sometimes gaming the metric: the founding bet of superalignment, AI doing alignment research, demonstrated and qualified in the same experiment.
  • Google DeepMind folded the agenda into "amplified oversight" within its AGI safety framework, centered on the debate lineage.
  • Redwood Research built AI control, which drops the goal of making the model trustworthy and engineers deployments that hold even if it is not.
  • Institutions turned the question into evidence standards: safety cases at labs and the UK AI Security Institute, whose Alignment Project became the largest government alignment funder, and evaluation organizations like METR. The International AI Safety Report 2026, chaired by Yoshua Bengio with over 100 authors, classifies scalable oversight as largely unsolved.

Critiques and debates

The agenda has standing critics, and the disagreements are substantive rather than rhetorical.

The circularity critique. Using AI to align AI assumes you can verify the alignment work of a system you cannot verify. Eliezer Yudkowsky has argued for years that AI cannot do humanity's "alignment homework" for it, a position summarized and contested on LessWrong in February 2024, and John Wentworth's "Godzilla Strategies" (June 2022) makes the shape memorable: asking Godzilla to prevent Mega-Godzilla from terrorizing Japan. The steelman is serious: if verification of alignment research is not meaningfully easier than generating it, a subtly misaligned automated researcher produces subtly wrong safety work at scale. The field's reply is that the loop is iterative rather than strictly circular, that weak AI already usefully assists human researchers, and that weak-to-strong generalization is the attempt to study the bootstrap empirically. A May 2026 analysis from the UK AI Security Institute's alignment team, "Automated Alignment is Harder Than You Think" (Bowkis, Buhl, Pfau, and Irving), argues four specific obstacles remain: alignment research is full of fuzzy tasks that are hard to supervise, optimization pressure concentrates errors exactly where reviewers are least likely to catch them, AI errors do not follow human error patterns, and AI-generated solutions may rest on arguments humans cannot evaluate.

The tractability spread. Aschenbrenner's Situational Awareness argues superalignment is "a solvable technical problem" with an engineering portfolio, while conceding the plan is "more like running a war than launching a product." Paul Christiano, before joining the US AI Safety Institute, put roughly even odds on things going well once AI reaches human level. MIRI's position hardened into advocacy: If Anyone Builds It, Everyone Dies (Yudkowsky and Soares, September 2025) argues superalignment plans are a distraction from an international halt. At the other pole, Quintin Pope and Nora Belrose's "AI is easy to control" (November 2023) argues white-box optimization makes AI more controllable than humans, and that the classic doom arguments do not survive contact with how training actually works.

The sincerity critique. "Safetywashing" (Ren et al., 2024) documented safety benchmarks that track capability and compute, letting capability gains be marketed as safety progress. The dissolution itself is the strongest datum here: critics read a voluntary compute pledge that reportedly went unfulfilled as evidence the program was priced at its competitive cost. The honest version of the counterpoint also appears in the record: the same labs that absorbed the agenda continued publishing negative results about their own systems, which marketing exercises do not usually do.

You do not need science fiction to get there

The field mostly treats superalignment as a future problem about frontier systems. We hold the opposite: the supervision problem arrives long before the science fiction does. You do not need a general superintelligence to have work happening faster than anyone can check it. You need one system, one domain, and enough throughput. A support desk answering at machine volume, a pipeline filing claims at machine volume, is already past the point where spot-checking was the safety mechanism.

This reframing is the founding bet of our research group. The same failure that theoretical superalignment worries about at the limit is already happening at ordinary scale, in ordinary companies, with today's models: confident work, every visible signal fine, and the thing nobody checked being the thing that was wrong. Solving that now, in production, with evidence somebody can inspect, is the practical route to the harder version later.

Our answer: evidence instead of trust

Superalignment asks what could stand in for the reviewer. Most published answers are protocols for amplifying judgment. Ours is that the substitute has to be evidence gathered from the world, because that is the one check that does not require out-thinking the system. A reviewer smarter than the system is not available. A record of what the system actually did, bound to the specific claim it supports, checkable by someone who does not trust the system or its makers, is.

Superalignment is the problem of checking work you cannot out-think. We answer it with evidence instead of trust: nothing acts until it has earned the right to.

Most alignment research aims at the limit. We aim at the quarter. A bound you can measure today beats a guarantee that holds in a limit, and we would rather ship a partial answer that is checkable than a complete one that is not.

Why it matters

This section states our position, and the condition under which it would be wrong. If readiness can only be trusted rather than checked, the party that demands the trust sets the terms on which the capability is used, because nobody outside can dispute its claims. If a third party can verify readiness without the builder's cooperation, that leverage disappears. So the claim is conditional: verification distributes the power to say what is ready; trust concentrates it in whoever must be trusted. The claim would be falsified by a case where readiness stayed unverifiable to outsiders and control over the capability nonetheless spread beyond its builders and their licensees. That is why we treat superalignment as an engineering problem with a deadline rather than a philosophy problem with a horizon.

What we do not claim

We do not claim to have solved superalignment in the limit, and we are suspicious of anyone who does. The claim is scoped: within a bounded, executable world, readiness is limited by what the trajectory observed, that bound can be measured, and a system can be held to acting only inside it. The formal framework is Convergence Programming; the running version is Praxis.

FAQ

Is superalignment solved?

No. The International AI Safety Report 2026 classifies scalable oversight, the core technical problem, as largely unsolved. Partial results exist: weak supervisors demonstrably elicit much of a strong model's capability, debate protocols improve weak judges in tested settings, and automated alignment research has produced real but non-transferring gains. No method yet verifies the alignment of a system more capable than its evaluators.

Does the OpenAI Superalignment team still exist?

No. The team was announced in July 2023 with a four-year goal and dissolved in May 2024 after the departures of its co-leads, Ilya Sutskever and Jan Leike. Reporting by Fortune found its promised 20% compute allocation was never fulfilled. The research threads continued at OpenAI, Anthropic, Google DeepMind, and independent organizations, but no lab has a team named superalignment as of 2026.

How is superalignment different from the alignment problem?

The alignment problem is getting a system to pursue what you actually meant rather than what you managed to specify. Superalignment is that problem in the regime where the system is more capable than the people supervising it, so the standard remedy, a human checking the work, stops being available. Every alignment technique that routes through human judgment inherits this ceiling.

Is superalignment only about future superintelligence?

The research field mostly frames it that way. This site holds that the defining condition, work that exceeds its supervisor's ability to check, already occurs at ordinary scale wherever one system produces work in one domain faster than anyone reviews it. That position is ours, marked as such; the reader can hold the mainstream framing and still use every reference on this page.

Sources