Superalignment

Building codes ask more of a tower than of a garden shed: more inspections, stronger materials, more ways out in a fire. A responsible scaling policy applies that idea to AI. The lab writes down in advance which dangerous capabilities it will test for, which safeguards each level of capability requires, and what it will do if those safeguards are not ready in time. Anthropic published the first policy under this name in 2023. Other labs publish similar frameworks under other names.

Why it matters

It turns a general promise to be careful into specific commitments outsiders can check. Unlike a building code, the lab writes it, largely checks its own compliance, and can revise it.