Superalignment

Banks pay people to try to break into their own vaults, and armies run exercises where one side plays the enemy. Red-teaming applies the same idea to AI. Testers try to make a system do what it should not: give dangerous instructions, leak private data, ignore its own rules. They report what worked so it can be fixed. It finds failures ordinary tests miss, because the testers are looking for them on purpose. It cannot prove a system is safe, only that these testers did not find a way through.

Why it matters

A clean result describes what the testers tried, not what the next attacker will try.