Wiki
A maintained vocabulary for the alignment problem.
Canonical terms for models, evidence, incentives, organizations, institutions and values. Use this page as the glossary, then follow a term for the maintained entry and its sources.
Every term on the site, defined. Entries in bold carry a full treatment.
No terms match.
The big picture
Six words that get used as if everyone agrees what they mean.
When software starts acting
The vocabulary of the shift from advice to action.
Knowing whether it works
The measurement vocabulary, and why the measurements mislead.
The words we use for proof
The vocabulary of knowing rather than hoping. Several of these are ours, so they need defining rather than assuming.
Words from our research
Terms from the Convergence Paper and the frameworks around it. They are ours, so they need defining rather than assuming.
Words your engineers will use
Enough to follow the conversation without nodding along.
Words from the risk debate
The vocabulary of the argument over how dangerous advanced AI could be, and of the commitments labs make in response. Defining a word is not taking a side.
Alignment research
The field's named problems, methods and results, each treated at full length.
- AI control
- AI safety via debate
- Alignment faking
- Automated alignment research
- Constitutional AI
- Corrigibility
- Deceptive alignment
- Eliciting latent knowledge
- Inner alignment
- Instrumental convergence
- Mechanistic interpretability
- Safety case
- Sandbagging
- Scalable oversight
- Sycophancy
- Weak-to-strong generalization