Superalignment

The 60-second answer

AI systems are developed and operated under changing capability, market, workload, and regulatory pressures. A one-time safety check can miss the path by which ordinary decisions spend margin across the wider control system.

A complex operation does not sit at a fixed safe point. Actors adapt to local pressures for efficiency, lower workload, and acceptable performance, while technology, markets, regulation, and competence also change. These adaptations can migrate work toward a boundary of functionally acceptable performance and erode defenses. Rasmussen therefore treats risk management as a distributed control problem that needs visible constraints, feedback, and cross-level models rather than a search for isolated errors.

  • Actors adapt toward locally attractive efficiency and workload conditions while the wider system and its constraints also change.
  • Accidents can emerge when interacting adaptations migrate performance toward a poorly visible safety boundary and defenses erode.
  • The article offers a systems model and research program, not measured evidence that one diagram predicts every hazard.

Written for: Technical generalists designing safety, operations, regulation, or organizational controls. Useful prerequisites: Feedback control, The difference between local and system-level optimization.

The question
Why can locally sensible adaptations move a changing sociotechnical system toward an accident without any single actor choosing to violate safety?
What the authors did
Rasmussen presents a cross-disciplinary theoretical synthesis grounded in decades of industrial risk research. He models risk management as control across government, regulators, companies, managers, planners, staff, and hazardous processes. He contrasts structural decomposition with functional abstraction, then links migration toward performance boundaries to changing economic, workload, and safety pressures.
The source
Risk Management in a Dynamic Society: A Modelling Problem

What changes when the safety boundary becomes visible?

Adaptive migration toward a boundary of acceptable performance A performance space is bounded by economic failure, unacceptable workload, and functionally unacceptable performance. Efficiency and lower-workload arrows point toward an operating point near the safety boundary. Controls reveal the boundary and add a counter-gradient. Work adapts inside several practical boundaries Local gradients can be reasonable while their combined direction reduces safety margin. unacceptable workload economic failure boundary of functionally acceptable performance efficiency gradient lower-workload gradient safety counter-gradient illustrative operating point Changing context market, technology, rules Boundary hidden: local gradients still act while remaining margin is hard to observe. A reconstruction of the Figure 3 mechanism, not a measured trajectory.

Boundary hidden. Efficiency and workload gradients pull toward a boundary that is difficult to see.

Visibility, local pressure, and proposed control
ModeWhat actors can useProposed consequence
Boundary hiddenEfficiency and workload signals are clearer than remaining safety margin.Ordinary adaptation can migrate toward unacceptable performance.
Boundary visibleThe safety boundary and current distance become actionable information.Actors can detect migration, but existing gradients still operate.
Counter-gradientVisible constraints plus a locally meaningful force toward safer operation.Safety competes with efficiency and workload pressure at the decision point.
The switch reconstructs Rasmussen's migration mechanism. Efficiency and lower-workload gradients pull an illustrative operating point toward a boundary of acceptable performance. Making the boundary visible supports a safety counter-gradient, but the source warns that adaptation can consume a merely enlarged margin. This is a schematic, not measured motion or risk. The operating point, boundary positions, arrows, colors, and distances are illustrative. They do not encode time, probability, causal strength, actual safety margin, or data from a particular system.

Walk through the argument

Each step below points to the source location that carries it. The source map follows the walkthrough for source-level checking.

Leave the single-error story

A familiar accident analysis starts with the last visible deviation and works backward through a chain. Rasmussen argues that this can isolate operators and tasks from the management, regulatory, and economic conditions shaping them.

His alternative asks how controls and feedback interact from government and regulators through management and staff to the hazardous process. The unit of analysis becomes the changing sociotechnical system.

Source: Accepted manuscript pages 1 to 4, abstract, Introduction, and Figure 1, Accepted manuscript pages 5 to 8, Modelling by Structural Decomposition, Accident Causation, and Figure 2

Draw a space with several pressures

Work is bounded by at least three practical concerns: economic failure, unacceptable workload, and functionally unacceptable performance. Actors search within this space rather than following one fixed route forever.

Efficiency pressure and a preference for lower effort can create gradients toward the safety boundary. These are not accusations of recklessness. They describe the incentives and constraints under which ordinary adaptation occurs.

Source: Accepted manuscript pages 9 to 11, Modelling by Functional Abstraction and Figures 3 to 4

Watch reasonable moves interact

One team saves time, another relaxes a defense that appears redundant, and a manager reallocates attention. Each choice can look reasonable locally while their interaction reduces the system's remaining room for ordinary variation.

The loss does not require one dramatic violation. Once operations are near the boundary, normal fluctuations can be enough to cross it.

Source: Accepted manuscript pages 5 to 8, Modelling by Structural Decomposition, Accident Causation, and Figure 2, Accepted manuscript pages 9 to 11, Modelling by Functional Abstraction and Figures 3 to 4

Make the boundary actionable

Rasmussen proposes making the safety boundary and distance to it visible, then creating a counter-gradient that makes safe movement locally attractive. A rule that exists only in a manual does not provide usable control if workers cannot see the relevant state.

He also warns that adding a larger nominal margin can fail. If the same pressures remain, adaptation may consume the new space and leave the operating point near a shifted boundary.

Source: Accepted manuscript pages 11 to 12, Control of System Performance

Close control loops across levels

A functioning controller needs objectives, a model of the process, ways to act, and feedback that measures relevant effects. Rasmussen applies that logic at every level, including regulation and management rather than only physical equipment.

Cross-level communication matters because goals, constraints, and observations are transformed as they move. A control system cannot outperform a measuring channel that hides the state needed for intervention.

Source: Accepted manuscript pages 12 to 18, Risk Management: A Control Task and Figures 5 to 6, Accepted manuscript pages 18 to 21, Identification of Constraints and Safe Boundaries and Figure 7

Treat the map as a testable hypothesis

The boundary diagram is a functional abstraction. It helps analysts look for gradients, constraints, adaptation, and weak feedback, but it does not supply a universal metric or an accident probability.

A serious application must define the hazard, controllers, boundaries, signals, and interventions for the specific system, then test whether those constructs predict or prevent unsafe migration.

Source: Accepted manuscript pages 18 to 21, Identification of Constraints and Safe Boundaries and Figure 7, Accepted manuscript pages 21 to 35, human-science paradigm review, Figure 8, and conclusion

Source map

These are the source locations that carry the argument. Use them to check this explanation against the original rather than trusting the summary alone.

Where the argument lives
LocusWhy it mattersSource
Accepted manuscript pages 1 to 4, abstract, Introduction, and Figure 1Defines risk management as a cross-level sociotechnical control problem under technological, market, regulatory, and competence change.Open source →
Accepted manuscript pages 5 to 8, Modelling by Structural Decomposition, Accident Causation, and Figure 2Explains why separate discipline, task, and error models can miss interactions among locally reasonable decisions.Open source →
Accepted manuscript pages 9 to 11, Modelling by Functional Abstraction and Figures 3 to 4Introduces boundaries of acceptable performance, economic and workload gradients, adaptive migration, defense degradation, and release by ordinary variation.Open source →
Accepted manuscript pages 11 to 12, Control of System PerformanceProposes making boundaries visible and adding a safety counter-gradient while warning that adaptation can consume a newly added margin.Open source →
Accepted manuscript pages 12 to 18, Risk Management: A Control Task and Figures 5 to 6Frames risk management as closed-loop control across objectives, controllers, feedback, competence, priorities, local constraints, and measuring channels.Open source →
Accepted manuscript pages 18 to 21, Identification of Constraints and Safe Boundaries and Figure 7Distinguishes hazard domains and argues that different sources and frequencies require different control strategies.Open source →
Accepted manuscript pages 21 to 35, human-science paradigm review, Figure 8, and conclusionSituates the framework across research traditions and closes with a proposed research direction rather than an intervention result.Open source →

The Assumption Switch

One result. One assumption exposed. Turn it and see what changes.

Assumption under test

The safety boundary and the system's current distance from it are visible enough to guide local adaptation.

Held in the source
Actors can notice when ordinary efficiency and workload pressures are moving performance toward unacceptable conditions and can apply a safety counter-gradient.
Turn it
The boundary is uncertain or hidden while local incentives continue to reward efficient, lower-effort performance.
What changes
Many reasonable adjustments can migrate the system toward loss, weaken defenses, and make normal variation sufficient to cross the boundary.

The common misreading

The migration model is sometimes read as a story about careless operators drifting into danger. Rasmussen's point is almost the reverse: adaptation can be locally rational and guided by efficiency and workload gradients. The hazard emerges from interacting decisions and weak control across levels, so blaming the last actor can hide the design problem.

Outside the ML frame

AI development and deployment governance

Can a lab see when ordinary delivery pressure is consuming its safety margin?

The model suggests tracing controls and feedback from policy and leadership through evaluation, release, operations, and the deployed process. It also suggests watching adaptations over time instead of certifying one static artifact. This is a systems-safety transfer, not evidence that AI organizations follow a measured migration curve.

Where the result stops

The article is a conceptual model and research agenda, not a controlled evaluation of an intervention. Its diagrams organize mechanisms but do not estimate migration rates or accident probabilities. The boundary model abstracts heterogeneous hazards into common pressures and needs domain-specific operationalization. The public source is an accepted manuscript whose pagination differs from the journal version, so all loci identify accepted-manuscript pages.

What remains open

  • Which indicators make a safety boundary visible without turning an uncertain model into false precision?
  • How can organizations preserve a safety counter-gradient when market and workload pressures are immediate and measurable?
  • Which cross-level feedback channels detect adaptations before they combine into an unsafe operating regime?
  • When does adding a larger nominal safety margin merely invite further adaptation instead of increasing resilience?
  • How should the model change for tightly coupled digital services whose boundaries and controllers shift quickly?

What this bears on

Superalignment maintains a public register of the claims it makes and the evidence that would change its mind. The entries below connect this source to the exact public claims it bears on.

  • C3. Readiness is a property of the trajectory, not the artifact. This record bears on it, indirectly. Rasmussen models safety as a dynamic control problem in which adaptive decisions, changing pressures, feedback, and eroding defenses alter risk over time. The paper predates AI deployment and does not define Superalignment readiness. See the claim and what would change our mind →

Source audit and review status

What was checked, and against what
FieldResultCheckedByAgainst
titleexact2026-08-17codex-primary-source-reviewsource
authorsexact2026-08-17codex-primary-source-reviewsource
dateexact2026-08-17codex-primary-source-reviewsource
venueexact2026-08-17codex-primary-source-reviewsource
full_textminor variant2026-08-17codex-primary-source-reviewsource
  • Review status prototype.
  • Program collection seminal v1; theory; wave 4, release slot unassigned.
  • Source access public full text, PDF accepted manuscript. Open the reading copy →
  • Provider-authored safety claim no.
  • Explained by Superalignment Research.
  • Reviewed by No named human reviewer yet.
  • Explainer dates created 2026-08-17; updated 2026-08-17.
  • AI assistance AI assisted with primary-source retrieval, full-text extraction, manifestation checking, locus mapping, first-pass prose, figure design, and implementation. The page remains a prototype until a named human review is recorded.
  • Rights and access DTU Orbit provides the peer-reviewed accepted manuscript for private study or research, retains copyright and moral rights, bars further distribution and profit-making use, and permits free distribution of the portal URL. Link to the source rather than redistributing the file.
  • Corrections Read the correction policy or report an error.

Provenance

  • First seen 2026-08-17, via cross-disciplinary seminal-source survey and full-source review.
  • Work id work:rasmussen-risk-management-dynamic-society, which groups manifestations of the same intellectual work.
  • Record id doi:10.1016/s0925-7535(97)00052-0, the natural key for this catalog manifestation.
  • 2026-08-17 full accepted manuscript read and implementation-ready Explained prototype prepared with manuscript-pagination and evidence-class caveats

Full audit data, including this record under id doi:10.1016/s0925-7535(97)00052-0: /library/records.jsonl. Compact browser index: /library/corpus.json. Catalog method and counts: /library/index.json.