Superalignment

Five ways into the same problem

Alignment failures do not live only inside a model. They also emerge from measurements, incentives, organizations, governing institutions and disagreements about whose values count. These lenses organize the research program without pretending they are separate worlds.

Search runs in your browser: the catalog loads once on your first keystroke, then filters instantly.

Explained: where the reading happens

These are close readings, not summaries. Each page walks through the method, uses an original figure, turns one assumption, states the common misreading, imports one outside lens and publishes its review status. There are 38 so far because the standard is intentionally expensive.

Explained

Can a stronger model learn past a weak supervisor's mistakes?

seminal v1

This paper establishes an experimental apparatus, not a solution. Positive performance gap recovered is common in its model-to-model proxy, but the amount recovered changes with the task, supervisor size, training objective and the learnability of the supervisor's errors. The last variable is the most consequential: when the weak answer becomes trivial to copy, average recovery collapses.

Read the record, including what it cannot support →

Explained

What does the Good Regulator theorem actually prove?

seminal v1

The proof establishes an outcome-equivalent deterministic state-action map under a fixed entropy objective. It does not establish that every effective regulator contains an explicit world model, that the mapping is one-to-one, or that the stabilized outcome is desirable.

Read the record, including what it cannot support →

Explained

How can a reward model learn a goal from human comparisons?

seminal v1

The paper made learned reward models practical enough for contemporary deep reinforcement learning and established the core feedback loop later associated with RLHF. It did not show that pairwise preferences recover human values. It showed that a small, task-specific comparison channel could sometimes replace a much denser programmatic reward in simulated control and games.

Read the record, including what it cannot support →

Explained

When does a rational robot choose to keep its off-switch?

seminal v1

The result is not that uncertainty automatically makes an agent corrigible. Deference has value when the robot is uncertain in the right way and the human decision is informative about the objective. Replace that decision with a random interruption, or make the human sufficiently unreliable relative to the robot's confidence, and bypassing oversight can become optimal.

Read the record, including what it cannot support →

See the complete Explainer program, coverage boundary and release states →

One evidence graph, six reader jobs

The site uses one source graph for different kinds of reading. A fact should not acquire five slightly different versions because it appears in six sections.

38

Explained

The source was read in full. Claims, method, limits and open questions are traced in a visual explanation with an Assumption Switch, one outside discipline, and visible review status.

26

Verified

The title, authors and date were checked against the canonical source and recorded field by field. This verifies the citation, not the result.

10,554

Listed

Imported from a public archive and searchable here. The source metadata has not yet been checked by this publication.

Checked at the source

26 records where an agent opened the canonical page and confirmed the title, the authors and the date, then recorded each check. This supports the bibliography only. It is not an endorsement of the argument, method or result. Rows link to the original source; full check provenance is in the audit export.

How to trust a row

Every row wears one of three labels, and each label is a specific promise about how far we went. Read the label and you know exactly how much of our word you are taking.

  • Listed. Imported from a public archive and searchable here. The source metadata has not yet been checked by this publication. A listed row points at somebody else's page and links you straight there.
  • Verified. The title, authors and date were checked against the canonical source and recorded field by field. This verifies the citation, not the result. The record keeps the individual checks rather than a verdict, so you can see which field was checked, against which URL, on what date, by which agent. Verified rows do not get editorial pages. Their complete checks are published in the full audit export.
  • Explained. The source was read in full. Claims, method, limits and open questions are traced in a visual explanation with an Assumption Switch, one outside discipline, and visible review status. These are the only records on this site that carry our interpretation, and the only ones that get a page here. The validator requires the question, method, common misreading, Assumption Switch, outside lens, open questions, original figure and editorial provenance before this label can appear.

Labels move up when the work is done and down when a re-check fails. Every move remains in the record changelog. A prototype is labeled as a prototype until a named human review is recorded.

Where the 10,618 rows came from

The bulk is the StampyAI/alignment-research-dataset, reported Nov 2023 snapshot. General forums are keyword gated rather than imported whole, so the catalog stays about the research instead of filling with community meta posts. The gate is versioned in data/library/seed_filters.json and you can read exactly what it admits.

Records by source
SourceRecords
AI Alignment Forum2,839
LessWrong1,718
EA Forum1,712
arXiv preprint1,581
intelligence.org503
aiimpacts.org253
carado.moe246
cold-takes.com102
drive.google.com70
scholar.google.com51
Distill50
vkrakovna.wordpress.com47
Where the coverage thins out. That snapshot reports through November 2023, so the catalog is strongest on the historical canon and gets thinner after that date. Supervised frontier refreshes fill the recent end while the reproducible daily runner is repaired. We also left out video transcripts, one wiki mid migration, and a community FAQ, each with a reason recorded in the filter file. If you expected to find something and it is not here, tell us: that is a coverage hole, and it is useful.

Recently catalogued

The newest of the 10,554 listed records, straight from the archives, most recent first. Every row links to its source.

Browse all 10,618 records → 43 pages

Machine readable, and meant to be: /library/index.json carries the counts, the method and the tier definitions; /library/corpus.json is the compact browser index; /library/records.jsonl carries every complete record and check; and /library/keys.txt is the deployed natural-key set for rejecting exact identifier repeats. It does not merge alternate identifiers or manifestations. Superalignment-authored compilation and editorial material are CC BY 4.0. Titles, abstracts and other third-party source material in the full export retain their original rights.