Library · AI alignment, seen through more than AI
The evidence base for superalignment, mapped with every check exposed.
A maintained catalog of work on how stronger systems are supervised, governed, measured and trusted. It covers machine learning and the fields alignment research often leaves outside the room: game theory, organizational science, safety engineering, public administration and culture. Every label says exactly what we checked.
10,618 records · 38 Explained · 37-work seminal spine: 36 prototypes, 1 source gate · 26 Verified · updated September 8, 2026
One result. One assumption exposed. Turn it. Every Explained record traces the evidence, names the boundary and uses one outside discipline only where it changes what the result means.
Five ways into the same problem
Alignment failures do not live only inside a model. They also emerge from measurements, incentives, organizations, governing institutions and disagreements about whose values count. These lenses organize the research program without pretending they are separate worlds.
Search runs in your browser: the catalog loads once on your first keystroke, then filters instantly.
Explained: where the reading happens
These are close readings, not summaries. Each page walks through the method, uses an original figure, turns one assumption, states the common misreading, imports one outside lens and publishes its review status. There are 38 so far because the standard is intentionally expensive.
Can a stronger model learn past a weak supervisor's mistakes?
seminal v1
This paper establishes an experimental apparatus, not a solution. Positive performance gap recovered is common in its model-to-model proxy, but the amount recovered changes with the task, supervisor size, training objective and the learnability of the supervisor's errors. The last variable is the most consequential: when the weak answer becomes trivial to copy, average recovery collapses.
What does the Good Regulator theorem actually prove?
seminal v1
The proof establishes an outcome-equivalent deterministic state-action map under a fixed entropy objective. It does not establish that every effective regulator contains an explicit world model, that the mapping is one-to-one, or that the stabilized outcome is desirable.
How can a reward model learn a goal from human comparisons?
seminal v1
The paper made learned reward models practical enough for contemporary deep reinforcement learning and established the core feedback loop later associated with RLHF. It did not show that pairwise preferences recover human values. It showed that a small, task-specific comparison channel could sometimes replace a much denser programmatic reward in simulated control and games.
When does a rational robot choose to keep its off-switch?
seminal v1
The result is not that uncertainty automatically makes an agent corrigible. Deference has value when the robot is uncertain in the right way and the human decision is informative about the objective. Replace that decision with a random interruption, or make the human sufficiently unreliable relative to the robot's confidence, and bypassing oversight can become optimal.
See the complete Explainer program, coverage boundary and release states →
One evidence graph, six reader jobs
The site uses one source graph for different kinds of reading. A fact should not acquire five slightly different versions because it appears in six sections.
38
Explained
The source was read in full. Claims, method, limits and open questions are traced in a visual explanation with an Assumption Switch, one outside discipline, and visible review status.
26
Verified
The title, authors and date were checked against the canonical source and recorded field by field. This verifies the citation, not the result.
10,554
Listed
Imported from a public archive and searchable here. The source metadata has not yet been checked by this publication.
Checked at the source
26 records where an agent opened the canonical page and confirmed the title, the authors and the date, then recorded each check. This supports the bibliography only. It is not an endorsement of the argument, method or result. Rows link to the original source; full check provenance is in the audit export.
-
Does DiffusionGemma do latent reasoning?
Verified -
Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds
Verified -
Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
Verified -
Rules or Character? Scaling Laws for AI Safety Design
Verified -
AI swarms are starting to pose indirect takeover risk
Verified -
Introducing the Conceptual Reasoning Index
Verified -
GPT-Red: Automated Red Teaming via Self-Play at Scale
Verified -
How independent researchers could investigate AI propensities after misalignment incidents
Verified -
Agentic Misalignment in Summer 2026
Verified -
Modular Pretraining Enables Access Control
Verified -
Separating signal from noise in coding evaluations
Verified -
Verbalizable Representations Form a Global Workspace in Language Models
Verified -
Summary of METR's predeployment evaluation of GPT-5.6 Sol
Verified -
"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms
Verified -
Diffuse AI Control on Fuzzy Tasks
Verified -
Quantifying the Salience of Geo-Cultural Values for Pluralistic Safety Alignment
Verified -
Gram: Assessing sabotage propensities via automated alignment auditing
Verified -
Realistic honeypot evaluations for scheming propensity
Verified -
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
Verified -
Automated alignment is harder than you think
Verified -
Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
Verified -
Model Spec Midtraining: Improving How Alignment Training Generalizes
Verified - Verified
-
BashArena: A Control Setting for Highly Privileged AI Agents
Verified -
Imitation Learning is Probably Existentially Safe
Verified -
Ctrl-Z: Controlling AI Agents via Resampling
Verified
How to trust a row
Every row wears one of three labels, and each label is a specific promise about how far we went. Read the label and you know exactly how much of our word you are taking.
- Listed. Imported from a public archive and searchable here. The source metadata has not yet been checked by this publication. A listed row points at somebody else's page and links you straight there.
- Verified. The title, authors and date were checked against the canonical source and recorded field by field. This verifies the citation, not the result. The record keeps the individual checks rather than a verdict, so you can see which field was checked, against which URL, on what date, by which agent. Verified rows do not get editorial pages. Their complete checks are published in the full audit export.
- Explained. The source was read in full. Claims, method, limits and open questions are traced in a visual explanation with an Assumption Switch, one outside discipline, and visible review status. These are the only records on this site that carry our interpretation, and the only ones that get a page here. The validator requires the question, method, common misreading, Assumption Switch, outside lens, open questions, original figure and editorial provenance before this label can appear.
Labels move up when the work is done and down when a re-check fails. Every move remains in the record changelog. A prototype is labeled as a prototype until a named human review is recorded.
Where the 10,618 rows came from
The bulk is the StampyAI/alignment-research-dataset, reported Nov 2023 snapshot. General forums are keyword
gated rather than imported whole, so the catalog stays about the
research instead of filling with community meta posts. The gate is
versioned in data/library/seed_filters.json and you can
read exactly what it admits.
| Source | Records |
|---|---|
| AI Alignment Forum | 2,839 |
| LessWrong | 1,718 |
| EA Forum | 1,712 |
| arXiv preprint | 1,581 |
| intelligence.org | 503 |
| aiimpacts.org | 253 |
| carado.moe | 246 |
| cold-takes.com | 102 |
| drive.google.com | 70 |
| scholar.google.com | 51 |
| Distill | 50 |
| vkrakovna.wordpress.com | 47 |
Recently catalogued
The newest of the 10,554 listed records, straight from the archives, most recent first. Every row links to its source.
- Listed
- Listed
- Listed
-
Algorithmic learning in a random world
Listed -
An Introduction to Game Theory, Chapters 1-7,14,15
Listed - Listed
-
Artificial Intelligence Safety and Security
Listed - Listed
- Listed
-
Causality: models, reasoning, and inference
Listed - Listed
-
CEA's Existential Risk and the Far Future playlist
Listed - Listed
- Listed
- Listed
- Listed
-
Computability and Logic, Chapters 1-4, 8-20, 23, 25, and 27
Listed - Listed
- Listed
- Listed
- Listed
- Listed
- Listed
- Listed
-
Defining Human Values for Value Learners
Listed -
Distributed Representations: Composition & Superposition
Listed -
Do Androids Dream of Electric Sheep?
Listed -
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Listed - Listed
- Listed
-
Empirical examples of power-law CDFs
Listed -
Ethics Background (Introduction through “Absolute Rights or Prima Facie Duties”)
Listed -
Everything else Chris Olah has ever written
Listed -
Flatland: A Romance of Many Dimensions
Listed -
FLI's AI Safety Research Landscape
Listed -
Formalizing Convergent Instrumental Goals
Listed - Listed
- Listed
-
Game Theory: Analysis of Conflict
Listed - Listed
- Listed
-
Handbook of Model Checking (to appear soon)
Listed -
Harry Potter and the Methods of Rationality (#1 of 6)
Listed - Listed
- Listed
-
Information Theory, Inference, and Learning Algorithms Parts I-III
Listed - Listed
- Listed
-
Introduction to Artificial Intelligence
Listed -
Introduction to Automata Theory, Languages, and Computation, Chapters 1-10
Listed - Listed
-
Learning the Preferences of Bounded Agents
Listed -
Log-normal distributions (with comparisons to power laws)
Listed - Listed
- Listed
-
Machine Learning (lecture notes)
Listed -
Machine Learning (online course)
Listed -
Mathematical Logic : A course with exercises — Part I
Listed -
Mathematical Logic : A course with exercises — Part II, Chapters 5 and 6
Listed -
Maximizing a quantity while ignoring effect through some channel
Listed
Browse all 10,618 records → 43 pages
Machine readable, and meant to be:
/library/index.json
carries the counts, the method and the tier definitions;
/library/corpus.json
is the compact browser index;
/library/records.jsonl
carries every complete record and check; and
/library/keys.txt
is the deployed natural-key set for rejecting exact identifier repeats.
It does not merge alternate identifiers or manifestations.
Superalignment-authored compilation
and editorial material are CC BY 4.0. Titles, abstracts and
other third-party source material in the full export retain their
original rights.