Superalignment

Five lenses, one supervision problem

Superalignment is the problem of checking work beyond your ability to evaluate directly. The technical model matters, but so do the evaluator's incentives, the organization's reporting lines, the institution that grants authority, the measurements that become legible and the people whose values enter the test. Each lens below has to change a prediction, a design or an experiment. Otherwise it is decoration.

Models and control
Training, scalable oversight, interpretability, permissions and runtime controls when alignment is imperfect.
Measures and evidence
Construct validity, stochastic evals, safety cases, evaluator independence and the envelope in which a result remains valid.
Incentives and games
How models, reviewers and institutions adapt to rewards, audits, observability and one another.
Organizations and institutions
Distributed knowledge, psychological safety, escalation, high reliability, legibility and legitimate authority.
Cultures and plural values
Who evaluates, how disagreement is aggregated and what disappears when one population stands in for everyone.

One evidence graph, six kinds of reading

SurfaceIts one job
CultureStory-led gateways that turn familiar films and live cultural pressures into testable alignment mechanisms.
ResearchOriginal claims, methods, experiments and active questions.
ExplainedVisual close readings that expose one assumption and show what changes.
Wiki and glossaryCanonical concepts, definitions, relationships and revision dates.
LibraryWorks, identifiers, provenance, source checks and reusable data.
BlogArguments, evidence briefs, field notes and changes of mind.

Superalignment supplies the frame. Praxis is where people and agents receive authority, do work, surface uncertainty and escalate. Verity binds claims to scenarios, observations and evidence from the running system. The public research program should make those connections testable rather than turn product architecture into proof.

Works

What the program has produced so far, and what each piece is for. A program is larger than any one of its works, the way scaling laws or learning from human feedback were each one work inside a larger vision. The first entry below is ours.

Paper · being prepared for release

Convergence Programming

Iterative, World-Grounded Alignment of Lossy Human Intent with Executable Behavior

Software can now be written from a sentence, by someone who cannot read the result, and put to work before anyone has established what it was for. The paper names the failure that follows (false convergence), moves the unit of programming from the artifact to the trajectory, proves four limits that no model capability removes, and reads them constructively as an architecture: a harness that earns permission to act.

Evidence instrument

The tracker

Dated, sourced indicators of what frontier AI can do, where it is used, and what the public can inspect. Measures and evidence, applied to the present.

Close readings

Explained

One result. One assumption exposed. Turn it. Visual readings of consequential outside work, each with a section where another discipline changes the interpretation.

Testbed

Praxis

The harness as a workspace: people, agents and business apps, with authority, uncertainty and escalation made visible. Where the program's claims get run against real work.

Open questions

The questions the program is organized around. Each is one we expect to change a design or an experiment, not a topic.

How long can an agent run before it must reopen the loop?
Autonomy spends evidence. What does a receding-horizon approval look like in practice, and what renews it?
When is verifier diversity real?
Critics that share a blind spot vote as one. What makes evidence independent in a world where most evaluators descend from the same few models?
Who gets to grant permission to act?
Authority, escalation and audit are institutional questions before they are technical ones. Which organizational structures make calibrated permission legible?
What do measurements hide?
Construct validity, stochastic evaluations and the envelope in which a result remains true. How do we keep a green dashboard from becoming the goal?
Whose values enter the test?
Evaluator composition, aggregation of disagreement, and what disappears when one population stands in for everyone.

What we are building toward

The target is not maximal uninterrupted autonomy. It is calibrated permission to act: discover the hidden constraints, repair and preserve them, bind every readiness claim to independent evidence, act inside the envelope that evidence covers, and reopen grounding when the budget runs out.

Reliable AI will require better generation. It will require something deeper as well: better convergence.

The Convergence Programming paper contributes a conceptual framework and an evaluation design, supported by a preliminary pilot. Measured benchmark and leaderboard results remain future work.