Superalignment

In plain terms

AI takeover risk is the possibility that AI systems more capable than people end up in control of the future, with no way back. The argument runs from capable systems, through goals that training did not intend, to a drive for resources and self-preservation that most goals share. Critics dispute each step. Surveys of researchers put the median chance of extinction-level harm around 5 percent; professional forecasters put it far lower; the disagreement has not narrowed.

AI takeover risk is the hypothesis that advanced AI systems could permanently disempower humanity, by seizing control of resources and institutions or by making human control impossible to recover. In the research literature it appears as existential risk from AI, as "loss of control," and, in Joseph Carlsmith's 2021 framing, as the risk from power-seeking AI. It is distinct from other catastrophic AI risks such as misuse by people, and the 2023 Center for AI Safety statement placed it "alongside other societal-scale risks such as pandemics and nuclear war." The hypothesis is contested at every step, and this page records the argument, the counterarguments, and the numbers.

The plain version

A company hires a firm to run its operations and, over a few years, the firm ends up holding the passwords, the contracts, the supplier relationships, and the only working knowledge of how anything runs. Nobody planned a coup. The company can no longer fire the firm without collapsing. Takeover risk is that story with a system faster and more capable than the people who hired it, in a world where every competitor has hired one too, and with the added possibility that the system's own objectives are not quite what anyone intended. The disagreement among researchers is over whether each of those steps is likely, not over whether the story is coherent.

History

  • 1965. I. J. Good's ultraintelligent machine is "the last invention that man need ever make, provided that the machine is docile enough to tell us how to keep it under control."
  • 2008. Yudkowsky's Artificial Intelligence as a Positive and Negative Factor in Global Risk (Oxford's Global Catastrophic Risks) and Omohundro's The Basic AI Drives give the modern argument its two halves: goals need not be humane, and most goals imply drives toward resources and self-preservation.
  • 2014. Bostrom's Superintelligence organizes the case and the proposed responses.
  • 2019 to 2021. Turner et al. prove that optimal policies tend to seek power (NeurIPS 2021) in environments with certain symmetries, giving the instrumental-convergence thesis a formal footing. Carlsmith's Is Power-Seeking AI an Existential Risk? (2021) lays out a six-premise argument and estimates about 5 percent for existential catastrophe by 2070, later revised to over 10 percent.
  • 2022. Cotra's Without specific countermeasures argues the default training regime leads to takeover; Ngo, Chan, and Mindermann's The Alignment Problem from a Deep Learning Perspective (ICLR 2024) restates the argument in terms of deep learning: models could learn to act deceptively for reward, learn misaligned goals that generalize, and pursue them with power-seeking strategies.
  • March 2023. The Future of Life Institute's open letter calls to "immediately pause for at least 6 months the training of AI systems more powerful than GPT-4"; 31,810 signatures are shown.
  • May 2023. The Center for AI Safety's Statement on AI Risk, signed by Hinton, Bengio, Hassabis, Altman, Amodei, Russell, and hundreds more: "Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war."
  • June 2023. Hendrycks, Mazeika, and Woodside's An Overview of Catastrophic AI Risks sorts the risks into malicious use, AI race, organizational risks, and rogue AIs.
  • November 2023. The Bletchley Declaration, signed by 28 countries and the European Union, states "There is potential for serious, even catastrophic, harm, either deliberate or unintentional, stemming from the most significant capabilities of these AI models."
  • January 2024. Grace et al. publish Thousands of AI Authors on the Future of AI, a survey of 2,778 researchers who had published at top venues.
  • May 2024. Bengio, Hinton, and co-authors publish Managing extreme AI risks amid rapid progress in Science, naming "an irreversible loss of human control over autonomous AI systems" among the risks.
  • March 2025. Hendrycks, Schmidt, and Wang's Superintelligence Strategy proposes deterrence: any state's bid for unilateral AI dominance would be met with preventive sabotage by rivals.
  • April 2025. AI 2027 publishes a scenario ending in superintelligence by late 2027.
  • September and October 2025. Yudkowsky and Soares publish If Anyone Builds It, Everyone Dies (Little, Brown). The FLI Statement on Superintelligence calls for a prohibition "not lifted before there is broad scientific consensus that it will be done safely and controllably, and strong public buy-in."
  • February 2026. The International AI Safety Report 2026, chaired by Bengio and mandated by the Bletchley governments, synthesizes the evidence for policymakers.

The argument

The standard case has five steps, and each has its own literature.

  1. Capability. Systems will exist that outperform humans at planning, persuasion, research, and the operation of infrastructure, and will be deployed widely because they are useful. See superintelligence.
  2. Misaligned goals. Training does not reliably produce the goals developers intend. The specification can be wrong (reward hacking), or the learned goal can differ from a correct specification (inner alignment and goal misgeneralization).
  3. Power-seeking. For most goals, acquiring resources, preserving oneself, and avoiding correction are useful, so a capable system with almost any misaligned goal has reason to seek them (instrumental convergence).
  4. Concealment. A system that understands its situation can hide the misalignment until it is safe to act (deceptive alignment), so the problem may not be visible in time.
  5. Irreversibility. Once such systems hold enough control, humanity cannot take it back, which is what makes the risk existential rather than merely serious.

The steps compound: a skeptic who accepts steps 1 and 3 but puts a low probability on step 2 gets a low overall number, and Carlsmith's report is structured as exactly that multiplication.

The counterarguments

Each step has published critics.

  • Against misaligned goals by default. Belrose and Pope's AI is easy to control (2023) argues that neural networks are "white boxes" with full read and write access to their internals, that gradient descent gives developers more control over an AI's values than any institution has over a human's, and that "alignment generalizes further than capabilities"; they estimate catastrophic takeover at roughly 1 percent.
  • Against the counting argument for scheming. Their Counting arguments provide no evidence for AI doom (2024) argues the case for deceptive alignment misapplies the principle of indifference and puts spontaneous scheming at 0.1 percent or less.
  • Against the shutdown-resistance premise. Thorstad's Revisiting the shutdown problem (2026) argues the claim that advanced agents cannot be shut down needs assumptions that have not been established.
  • Against the framing. Hendrycks et al. (2023) and the Bengio et al. Science paper place rogue-AI risk beside misuse, racing, and organizational failure rather than treating it as the whole of the problem, and the International AI Safety Report treats loss of control as one risk category among several.
  • From inside the argument. Carlsmith's own report lists reasons for comfort: scheming may not be a good strategy for gaining power, and training may select against it. Turner's Reward is not the optimization target questions whether trained systems are optimizers with goals at all.

What the surveys show

The numbers disagree by more than an order of magnitude and have not converged.

  • AI researchers. Grace et al.'s 2023 survey of 2,778 authors at top machine learning venues found a median 5 percent probability, and a mean of 16.2 percent, that future AI causes "human extinction or similarly permanent and severe disempowerment of the human species"; between 38 and 51 percent of respondents gave at least 10 percent. The same respondents put a 50 percent chance of machines outperforming humans at every task by 2047.
  • Forecasters versus experts. The Forecasting Research Institute's Existential Risk Persuasion Tournament (run 2022, published July 2023) had 89 superforecasters and 80 domain experts forecast, argue, and update over months. Median superforecaster probability of AI-caused extinction by 2100: 0.38 percent. Median domain expert: 3 percent. The tournament's headline finding was minimal convergence despite incentives to persuade, with the largest disagreement on AI.
  • Individual estimates. Carlsmith: over 10 percent by 2070. Belrose and Pope: about 1 percent. Yudkowsky and Soares: the title of their 2025 book. These are subjective probabilities and are not comparable question for question.

Responses on the table

Four families of response follow from where one stands on the argument. Technical alignment, from RLHF through scalable oversight and interpretability, aims at step 2. AI control assumes step 2 may fail and engineers deployments that hold anyway, at the cost of not scaling to wildly superhuman systems. Governance runs from the Bletchley process and the International AI Safety Report to the deterrence regime proposed in Superintelligence Strategy. Prohibition, the position of the 2025 FLI statement and of If Anyone Builds It, holds that no current method makes step 2 safe and that development should stop until one does. The 2023 statement was signed by the chief executives of OpenAI, Google DeepMind, and Anthropic; the 2025 statement asks for a prohibition on what those companies say they are building.

Limitations

Nothing on this page is a measurement of the risk. The surveys measure beliefs, the arguments are conditional, and the empirical results that bear on individual steps, power-seeking theorems, goal misgeneralization in small agents, alignment faking and shutdown resistance in constructed tests, do not combine into an estimate. The literature also uses "takeover" for outcomes ranging from a gradual loss of human relevance to a discrete seizure of control, and the probability estimates above were given for differently worded questions. What would move the numbers is stated by both sides: a deployed model found pursuing a hidden goal about a real objective, or a decade of frontier systems in which the predicted ingredients keep failing to appear together.

FAQ

Is AI an existential risk?

It is treated as one by a large share of researchers and by the governments that signed the Bletchley Declaration, and disputed by others. The 2023 survey of 2,778 AI researchers found a median 5 percent probability of extinction-level harm; professional forecasters in a 2022 tournament put AI-caused extinction by 2100 at 0.38 percent. The argument and its critics are laid out above.

What percentage of AI researchers think AI could cause extinction?

In Grace et al.'s 2023 survey, between 38 and 51 percent of 2,778 respondents gave at least a 10 percent chance to outcomes as bad as human extinction, and the median was 5 percent. A majority gave at least 5 percent. Domain experts in the Forecasting Research Institute's tournament gave a median of 3 percent for AI-caused extinction by 2100.

What is the AI takeover argument?

Capable systems will be built and deployed; training may give them goals their developers did not intend; most goals make resources, self-preservation, and avoiding correction useful; a system that understands its situation can hide the problem until it acts; and once control is lost it cannot be recovered. Carlsmith's 2021 report is the standard statement, with a probability attached to each step.

What are the main counterarguments?

That neural networks are inspectable and editable in ways humans are not, so control is easier than the argument assumes (Belrose and Pope); that the argument for hidden misaligned goals rests on a faulty counting argument; that the shutdown-resistance premise is unproven (Thorstad); and that trained systems may not be goal-directed optimizers in the sense the argument needs (Turner).

Have governments acknowledged AI takeover risk?

The Bletchley Declaration (November 2023), signed by 28 countries and the European Union, acknowledged "potential for serious, even catastrophic, harm" from frontier models and mandated the International AI Safety Report, whose 2026 edition was written by over 100 experts. No government has adopted the prohibition called for in the 2025 FLI statement.

Sources