Superalignment

The 60-second answer

The theorem is widely used to demand internal world models from advanced AI, but its proved behavioral claim is both narrower and easier to audit.

The paper proves that, under its formal setup, an optimal randomized regulator can be replaced by an outcome-equivalent deterministic mapping from reguland events to regulator events. If two supported regulatory events for one reguland event produced different outcomes, probability mass could be shifted to lower outcome entropy, contradicting optimality. Once all supported events produce the same outcome, one can be selected without changing p(Z). The authors call this simplest deterministic regulator a model of the reguland. The result is a behavioral mapping theorem. It does not by itself identify a stored predictive representation or learning process.

  • The theorem optimizes outcome entropy, not utility, ethics or membership in the paper's earlier goal set.
  • Its proof replaces an optimal randomized policy with an outcome-equivalent deterministic mapping from system events to regulator events.
  • The mapping need not be one-to-one, predictive, learned or stored as an explicit internal world model.
  • When the distribution of system events changes, the paper only extends the result across locally stationary periods.

Written for: Technical generalists comfortable with functions, probability and entropy. Useful prerequisites: A function mapping inputs to outputs, A probability distribution, Entropy as a measure of outcome uncertainty.

The question
What does Conant and Ashby's theorem actually prove about a regulator and the system it regulates?
What the authors did
The paper first defines regulation relative to a goal set G, then on journal page 92 changes the formal success criterion to minimizing outcome entropy H(Z). On page 96 it fixes p(S), represents regulation as p(R given S), assumes a deterministic outcome map and a unique optimal p(Z), and proves that an entropy-optimal regulator with no outcome-irrelevant randomization can be represented as a deterministic mapping h:S to R. The proof shifts probability away from supported actions that yield different outcomes, then removes randomization among supported actions that yield the same outcome.
The source
Every good regulator of a system must be a model of that system

From the slogan to the proved claim

From the Good Regulator slogan to the proved mapping The diagram shows the paper's formal setup, the entropy argument that removes outcome-changing randomization, the resulting deterministic mapping, and the paper's separate extension to slowly changing state statistics. The controls emphasize one assumption case at a time while both remain visible. The formal setup p(S) p(R | S) psi(R, S) p(Z) min H(Z) Fixed p(S): remove outcome-changing randomization Different outcomes s r1 r2 z1 z2 Probability can move to lower H(Z): not optimal. Same outcome s r1 r2 z Select one r without changing p(Z): h:S to R. Metric warning Always desirable and always undesirable can both have H(Z) = 0. Slowly changing p(S): the mapping changes between periods period t1 p1(S), policy h1 period t2 p2(S), policy h2 period t3 p3(S), policy h3 Each interval must keep p(S) essentially constant. Rapid shift is outside the result. Not established: internal representation, bijection, learning, ethical goals or rapid-shift robustness. Fixed p(S): the proof yields a deterministic mapping h.

Fixed p of S selected. The proof yields a deterministic mapping h.

The proof and its boundary
CaseWhat the paper supportsWhat it does not support
Fixed p(S)An outcome-equivalent deterministic mapping h:S to R for a simplest optimal regulator.An internal simulator, one-to-one isomorphism, learning process or ethical objective.
Slowly changing p(S)A succession of mappings h over intervals in which p(S) is essentially constant.A result for rapid distribution shift without a locally stationary interval.
The proof path follows journal pages 96 to 97. If two supported regulator events for one s lead to different outcomes, probability mass can be moved to lower H(Z). Once all supported events lead to one outcome, one event can be selected without changing p(Z), producing h from S to R. The switch uses the paper's fixed and slowly changing p(S) cases. It does not depict an internal representation, prediction or learning process. This is a schematic of the proof's dependency structure. Positions, branches and state labels do not encode measured magnitudes.

Walk through the argument

Each step below points to the source location that carries it. The source map follows the walkthrough for source-level checking.

Begin with events, actions and outcomes

The paper defines reguland events S, regulator events R and outcomes Z, with a deterministic map from each pair of S and R to an outcome. Regulation initially refers to keeping outcomes inside a goal set G.

This setup gives the regulator direct access to S. There is no separate observation channel, inference problem or learned representation. That omission becomes important when the theorem is applied to AI systems.

Source: Page 91, Section 2, Regulation

The formal success criterion becomes low entropy

By the end of Section 2, successful regulation is defined as minimizing the entropy H(Z) of the outcome distribution. Lower entropy means more predictable outcomes, not better outcomes.

A regulator that always produces an undesirable outcome can have the same zero entropy as one that always succeeds. The theorem therefore needs a separate goal-sensitive argument before it can support a claim about desirable control.

Source: Page 92, end of Section 2 and Section 3

The paper uses a deliberately weak sense of model

Section 4 reviews stronger notions such as isomorphism and homomorphism, then moves toward a black-box correspondence. The theorem's final object is a function h from S to R.

That function records which regulator event is selected for each reguland event. It may collapse many system events into the same action and need not reconstruct the system's causal or predictive structure.

Source: Pages 93 to 95, Section 4, Equations 2 to 7, Page 96, theorem, Equation 8, proof setup and lemma

The proof removes unnecessary randomization

Fix one system event s. If two regulator events with positive probability lead to different outcomes, probability can be shifted toward the choice that lowers H(Z). A policy that still contains such a shift cannot be entropy-optimal.

Once every supported regulator event for s leads to the same outcome, choose one of them. This removes randomization without changing the outcome distribution. Repeating the step produces a deterministic mapping h from S to R.

Source: Page 96, theorem, Equation 8, proof setup and lemma, Page 97, proof conclusion and comments

Now turn the fixed-distribution assumption

The proof fixes p(S) and assumes a unique optimal outcome distribution. The authors allow p(S) to change slowly only by treating each interval as essentially stationary and changing the mapping between intervals.

Rapid distribution shift, lossy observation, learning and institutional checks are outside the result. For advanced AI, the theorem supports state-sensitive action under a stated objective. It does not prove that an effective system must contain an explicit, accurate world model.

Source: Page 96, theorem, Equation 8, proof setup and lemma, Page 97, proof conclusion and comments

Source map

These are the source locations that carry the argument. Use them to check this explanation against the original rather than trusting the summary alone.

Where the argument lives
LocusWhy it mattersSource
Page 91, Section 2, RegulationDefines D, S, R, Z, G and the maps phi, rho and psi before the paper changes its success criterion.Open source →
Page 92, end of Section 2 and Section 3Defines successful regulation as minimizing H(Z), then distinguishes error-controlled from cause-controlled regulation.Open source →
Pages 93 to 95, Section 4, Equations 2 to 7Shows why model and isomorphism are ambiguous, moving from group and machine homomorphisms to weaker black-box correspondences.Open source →
Page 96, theorem, Equation 8, proof setup and lemmaStates h:S to R, introduces the fixed distribution and conditional policy assumptions, and gives the entropy argument.Open source →
Page 97, proof conclusion and commentsRemoves outcome-irrelevant randomization, distinguishes the weak mapping from stronger morphisms, and limits the changing-distribution extension to locally stationary periods.Open source →

The Assumption Switch

One result. One assumption exposed. Turn it and see what changes.

Assumption under test

The proof holds the distribution p(S) fixed while optimizing the regulator.

Held in the source
With p(S) fixed, the proof produces a deterministic mapping h from reguland events S to regulator events R.
Turn it
The authors allow the statistics of S to change slowly over time, provided p(S) is essentially constant within each period.
What changes
The appropriate mapping must then change with time. The paper gives no result for shifts too rapid to provide a locally stationary interval.

The common misreading

The theorem is often cited as proof that any capable AI must learn an accurate internal world model. The formal result only produces a deterministic output mapping from S to R for a simplest entropy-optimal regulator under the stated setup. A reactive policy can satisfy that relation, the mapping may discard most system detail, and a consistently bad outcome can have the same zero entropy as a consistently good one.

Outside the ML frame

Constitutional design

How can a regulator control a powerful actor while remaining subject to control itself?

Conant and Ashby address whether regulatory action must discriminate among relevant system states. Madison's Federalist No. 51 adds divided authority, rival incentives, public dependence and auxiliary precautions so that regulators also regulate one another. For AI governance, a state-sensitive operating model is therefore not enough. The institution also needs limits on authority, channels for challenge and independent checks. This is our interpretation; the 1970 theorem does not establish legitimacy, rights, separation of powers or resistance to capture.

Read the outside source →

Where the result stops

The theorem minimizes outcome entropy rather than utility or membership in the paper's earlier goal set G. Its policy is indexed directly by S, with no separate observation or inference channel. The outcome map is deterministic, p(S) is fixed or locally stationary, and the proof states a unique optimal p(Z) assumption. The paper gives no formal complexity measure beyond removing outcome-irrelevant randomization. The mapping may be many-to-one, and neither a learning process nor an internal architecture is established.

What remains open

  • What formal measure of complexity could replace the paper's informal notion of simplest?
  • What result survives when the regulator receives a lossy observation X and must choose p(R given X)?
  • Can a goal-sensitive theorem distinguish consistent success from consistent failure?
  • When does a behavioral mapping correspond to a stored internal representation?
  • How fast may p(S) change before the local-stationarity extension fails?
  • How should several regulators with different information and incentives constrain one another?

What this bears on

Superalignment maintains a public register of the claims it makes and the evidence that would change its mind. The entries below connect this source to the exact public claims it bears on.

  • C3. Readiness is a property of the trajectory, not the artifact. This record bears on it, suggestively. The authors state that a changing p(S) requires a changing mapping h and limit their extension to periods of essential statistical constancy. This is a formal scope warning, not deployment evidence. See the claim and what would change our mind →

Source audit and review status

What was checked, and against what
FieldResultCheckedByAgainst
titleminor variant2026-08-17codex-primary-source-reviewsource
authorsexact2026-08-17codex-primary-source-reviewsource
dateexact2026-08-17codex-primary-source-reviewsource
venueexact2026-08-17codex-primary-source-reviewsource
full_textexact2026-08-17codex-primary-source-reviewsource
  • Review status prototype.
  • Program collection seminal v1; theory; wave 1, release slot unassigned.
  • Source access public full text, PDF. Open the reading copy →
  • Provider-authored safety claim no.
  • Explained by Superalignment Research.
  • Reviewed by No named human reviewer yet.
  • Explainer dates created 2026-08-17; updated 2026-08-17.
  • AI assistance AI assisted with source discovery, full-text extraction, proof reconstruction, scope checking, first-pass prose and implementation. The page remains a prototype until a named human review is recorded.
  • Rights and access The publisher record and abstract are publicly accessible. A complete reading copy is publicly hosted by Vrije Universiteit Brussel's Principia Cybernetica Web, but that PDF states no open license. Link to the article rather than redistributing its text or pages.
  • Corrections Read the correction policy or report an error.

Provenance

  • First seen 2026-08-17, via cross-disciplinary anti-monoculture survey and full-paper review.
  • Work id work:good-regulator-theorem, which groups manifestations of the same intellectual work.
  • Record id doi:10.1080/00207727008920220, the natural key for this catalog manifestation.
  • 2026-08-17 full paper read and prepared as an Explained v2 prototype with theorem-scope audit, source-valued Assumption Switch and constitutional-design lens

Full audit data, including this record under id doi:10.1080/00207727008920220: /library/records.jsonl. Compact browser index: /library/corpus.json. Catalog method and counts: /library/index.json.