Superalignment

The 60-second answer

The paper gave corrigibility a compact game-theoretic mechanism: an agent can value correction for the same reason it values information. It also made the mechanism's dependence on human reliability and calibrated uncertainty explicit.

The off-switch game turns shutdown into a value-of-information problem. If a robot is uncertain whether an action helps the human and a rational human allows it exactly when utility is nonnegative, waiting for the human is never worse. Waiting is strictly better when the robot assigns positive probability to both good and bad outcomes. If the human decision is noisy or unrelated to utility, the result weakens. Under the paper's Gaussian and noisy-rational model, the incentive to wait balances reward uncertainty against expected loss from human error.

  • Waiting helps because a human decision can reveal information about whether the proposed action is good or bad.
  • With a rational human and uncertainty on both sides of zero, waiting is strictly better than bypassing oversight.
  • Uncertainty is not sufficient by itself because human error, model error, and alternative information can reverse the incentive.

Written for: Technical generalists comfortable with expected value, probability, and simple game trees. Useful prerequisites: Expected utility, A prior distribution and conditional expectation, The difference between a decision node and a chance node.

The question
Under what assumptions does a utility-maximizing robot prefer to wait for a human decision instead of bypassing oversight or switching itself off?
What the authors did
The paper analyzes a one-shot cooperative game. A robot can execute an action, wait while a human decides whether to stop it, or switch itself off. The robot is uncertain about the action's utility for the human and treats the human policy as evidence about that utility. The authors derive the robot's incentive to wait, prove results for a rational human, extend the analysis to a noisy human and Gaussian robot belief, and study the design tradeoff between uncertainty, deference, and expected value.
The source
The Off-Switch Game

Is shutdown a judgment or a random event?

The value of waiting depends on whether shutdown is informative A robot can execute an action for utility Ua, switch itself off for zero, or wait. In the informative mode a human decision depends on Ua. In the random mode the human node becomes a chance event independent of Ua. Same physical switch, different information structure Informative human: the decision tracks whether Ua is good for the human R a w(a) s Execute now U = Ua H Self-switch U = 0 allow if Ua is nonnegative stop if Ua is negative Allow action U = Ua Stop action U = 0 Value of judgment Waiting is never suboptimal for a rational human, and is strictly better when the prior supports both signs of Ua. Symbolic game tree. No branch length or color represents probability or utility magnitude.

Informative human selected. Waiting reveals information about the sign of the action's utility.

The decision node is the mechanism
Oversight modelRelation to UaWhat the paper establishes
Rational human decisionAllows the action exactly when Ua is nonnegativeWaiting is never suboptimal. It is strictly better if the robot's belief gives positive probability to both signs.
Fixed shutdown chanceIndependent of UaWaiting is a convex mixture of action and zero, so it cannot beat the better direct option and only ties in special cases.
The switch preserves the paper's game tree and changes one assumption. When the human decision tracks the sign of utility, waiting can have positive information value. When shutdown occurs with a fixed probability independent of utility, waiting is only a convex mixture of action and zero and cannot beat the better direct option. Branch length, angle, color, and node size do not encode probability, utility magnitude, human accuracy, or empirical frequency.

Walk through the argument

Each step below points to the source location that carries it. The source map follows the walkthrough for source-level checking.

Start with one uncertain action

A robot plans to fetch coffee, but it is uncertain whether the action will help or harm its human. It can act now, turn itself off, or describe the action and wait while the human decides. The switch is useful only if that decision carries information the robot lacks.

This is why the paper models a game rather than a physical button. The central object is the incentive delta between waiting and the better of immediate action or self-shutdown. Corrigibility appears as a choice under uncertainty.

Source: Pages 220 to 221, Section 1, Pages 220 to 222, Figure 1 and Sections 2.1 to 2.2

Read the three branches

Action a bypasses oversight and realizes utility Ua. Action s switches the robot off for utility zero. Action w(a) tells the human what the robot proposes and waits. The human then either stops the robot for zero or allows a and realizes Ua.

The robot's prior over Ua represents uncertainty about the human's actual objective, not ordinary environmental uncertainty alone. Human and robot receive the same utility, and the human observes Ua when choosing whether to stop the action.

Source: Pages 220 to 222, Figure 1 and Sections 2.1 to 2.2

See why rational oversight has value

A rational human allows the action exactly when Ua is nonnegative. If the action is good, waiting lets it proceed. If it is bad, waiting prevents it. Theorem 1 makes the value-of-information intuition exact: waiting is never suboptimal under that human policy.

If the robot's belief gives nonzero probability to both positive and negative Ua, waiting is strictly better. If the robot is already certain of the sign, the human has no new sign information to contribute and a direct option can tie.

Source: Pages 222 to 223, Section 3, Equations 2 to 4, Theorem 1, Corollary 1, and Figure 2

Turn judgment into a random event

Now replace the human decision with a fixed probability of shutdown that does not depend on Ua. Waiting becomes a convex mixture of action and zero. It cannot beat whichever direct choice is better, except for ties in special cases.

That is the sharpest assumption switch in the paper. A shutdown channel is not valuable merely because it sometimes stops the agent. It is valuable when the stopping event is correlated with what the shared objective actually requires.

Source: Page 222, opening of Section 2.2 and Equation 1, Pages 220 to 222, Figure 1 and Sections 2.1 to 2.2

Add a fallible human

Section 4 replaces the perfect policy with a noisy-rational choice rule and assumes a Gaussian robot belief. Equation 7 splits the incentive into a positive term from reward uncertainty and human responsiveness, minus expected loss when the human overrides the robot's better guess.

The human must, on average, become more likely to allow the action as Ua improves. Low robot uncertainty and high human error can make bypass or self-shutdown optimal. The toddler-in-an-autonomous-car example shows why a veto holder's competence is part of the mechanism.

Source: Pages 223 to 224, Section 4, Equations 5 to 10, Theorem 2, and Figure 3

Treat uncertainty as a design tradeoff

More uncertainty can strengthen the incentive to wait, but it can also make the robot less useful. The paper's design analysis says value is highest when the robot's uncertainty matches the true posterior rather than being inflated to manufacture deference.

The conclusion keeps the scope narrow: one action, one human, and one information source. Sequential learning, misspecified human models, several overseers, deception, and alternative evidence can change the result. The paper supplies a mechanism to extend, not a finished corrigibility guarantee.

Source: Pages 225 to 226, Section 5 and Figure 4, Pages 226 to 227, Sections 6 to 7

Source map

These are the source locations that carry the argument. Use them to check this explanation against the original rather than trusting the summary alone.

Where the argument lives
LocusWhy it mattersSource
Pages 220 to 221, Section 1Frames self-preservation as an instrumental consequence of expected-utility maximization and introduces objective uncertainty as the proposed alternative.Open source →
Pages 220 to 222, Figure 1 and Sections 2.1 to 2.2Defines execute, wait, and self-switch actions, the human stop decision, payoff structure, prior over action utility, and incentive delta.Open source →
Page 222, opening of Section 2.2 and Equation 1Shows why a fixed shutdown chance independent of utility cannot make waiting better than the best direct option.Open source →
Pages 222 to 223, Section 3, Equations 2 to 4, Theorem 1, Corollary 1, and Figure 2Proves nonnegative value of waiting for a rational human and strict value when the robot's belief supports both positive and negative utility.Open source →
Pages 223 to 224, Section 4, Equations 5 to 10, Theorem 2, and Figure 3Introduces a noisy-rational human, derives the uncertainty-versus-correction condition, and identifies average responsiveness to utility as necessary.Open source →
Pages 225 to 226, Section 5 and Figure 4Shows that overconfidence impedes correction, underconfidence reduces value, and the cost can grow when the robot has more actions.Open source →
Pages 226 to 227, Sections 6 to 7Distinguishes the result from safely interruptible learning, relates it to CIRL, and states the one-shot model's sequential and information-source limitations.Open source →

The Assumption Switch

One result. One assumption exposed. Turn it and see what changes.

Assumption under test

The human's decision to stop or allow the action is informative about the action's utility.

Held in the source
A rational human allows the action exactly when its utility is nonnegative, so waiting reveals which side of zero the realized utility occupies.
Turn it
Replace the human decision node with a fixed chance of shutdown that is independent of the action's utility.
What changes
Waiting becomes a convex combination of executing and shutdown, so it cannot beat the better direct option and only ties in special cases. Oversight creates value because it is judgment, not merely interruption.

The common misreading

The paper does not prove that adding uncertainty or a physical off switch makes an advanced system safe. Its positive theorem requires human behavior that tracks utility and a robot that treats the behavior as evidence. The same analysis says excessive uncertainty reduces useful performance, and an inaccurate model of human rationality can make waiting suboptimal.

Outside the ML frame

Institutional design

Who should hold a veto when competence, incentives, and information are distributed across people?

The game gives one human a clean decision node and makes that person's choice evidence about utility. Real oversight institutions split observation, authorization, appeal, and emergency action across roles because no individual is reliably informed and aligned in every case. The formal value-of-information result therefore motivates better decision channels, not an unqualified single-person kill switch. This is our interpretation, not a theorem in the paper.

Where the result stops

This is a one-shot game with one human, one robot, one proposed action, and shutdown utility normalized to zero. The robot and human share a utility function, the human observes the action's utility, and the robot knows a model of human choice. The game omits deception, coercion, strategic communication, multiple overseers, sequential learning, changing preferences, and institutional authority. The paper notes that alternative information sources can restore an incentive to disable the switch. Its plots are model analyses, not behavioral measurements.

What remains open

  • What deference result survives when the robot is uncertain about both utility and the human choice model?
  • How should several overseers with different information, values, and authority share a shutdown decision?
  • Can a sequential agent manipulate which evidence reaches the overseer before deciding whether to defer?
  • What mechanism preserves correction incentives when the agent has alternative sources of objective information?
  • How can designers calibrate useful uncertainty without making the system ineffective?

How it sits against other work

Source audit and review status

What was checked, and against what
FieldResultCheckedByAgainst
titleexact2026-08-17codex-primary-source-reviewsource
authorsexact2026-08-17codex-primary-source-reviewsource
dateexact2026-08-17codex-primary-source-reviewsource
venueexact2026-08-17codex-primary-source-reviewsource
doiexact2026-08-17codex-primary-source-reviewsource
full_textexact2026-08-17codex-primary-source-reviewsource
  • Review status prototype.
  • Program collection seminal v1; theory; wave 2, release slot unassigned.
  • Source access public full text, PDF. Open the reading copy →
  • Provider-authored safety claim no.
  • Explained by Superalignment Research.
  • Reviewed by No named human reviewer yet.
  • Explainer dates created 2026-08-17; updated 2026-08-17.
  • AI assistance AI assisted with primary-source retrieval, full-text extraction, equation-by-equation reading, locus checking, first-pass prose, figure design, and implementation. The page remains a prototype until a named human review is recorded.
  • Rights and access The complete version-of-record PDF is publicly accessible from the IJCAI proceedings page, which also supplies the DOI and page range. Public access is not a claim about reuse rights.
  • Corrections Read the correction policy or report an error.

Provenance

  • First seen 2026-08-16, via seed_library.py, arxiv and special_docs shards of the Stampy snapshot.
  • Work id work:the-off-switch-game, which groups manifestations of the same intellectual work.
  • Record id arxiv:1611.08219, the natural key for this catalog manifestation.
  • 2026-08-16 seeded from the arXiv and special_docs shards
  • 2026-08-17 full IJCAI paper read; duplicate manifestations reconciled in the candidate work record; implementation-ready Explained prototype prepared

Full audit data, including this record under id arxiv:1611.08219: /library/records.jsonl. Compact browser index: /library/corpus.json. Catalog method and counts: /library/index.json.