The catalog, page 1
Records 1 to 250 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Alignment Faking in Large Language Models
Explained -
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Explained -
AI Control: Improving Safety Despite Intentional Subversion
Explained -
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Explained -
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
Explained -
Scaling Laws for Reward Model Overoptimization
Explained -
Constitutional AI: Harmlessness from AI Feedback
Explained - Explained
-
Underspecification Presents Challenges for Credibility in Modern Machine Learning
Explained -
Goal Misgeneralization in Deep Reinforcement Learning
Explained -
The Impact of Network Connectivity on Collective Learning
Explained -
Training language models to follow instructions with human feedback
Explained -
Optimal Policies Tend To Seek Power
Explained -
Eliciting latent knowledge: How to tell if your eyes deceive you
Explained -
Algorithmic Monoculture and Social Welfare
Explained - Explained
-
Risks from Learned Optimization in Advanced Machine Learning Systems
Explained -
Supervising strong learners by amplifying weak experts
Explained - Explained
-
Deep Reinforcement Learning from Human Preferences
Explained - Explained
-
Concrete Problems in AI Safety
Explained -
Cooperative Inverse Reinforcement Learning
Explained - Explained
-
Engineering a Safer World: Systems Thinking Applied to Safety
Explained -
The Weirdest People in the World?
Explained -
Beyond Markets and States: Polycentric Governance of Complex Economic Systems
Explained -
The Communication Structure of Epistemic Communities
Explained -
Psychological Safety and Learning Behavior in Work Teams
Explained -
Risk Management in a Dynamic Society: A Modelling Problem
Explained -
Multitask Principal-Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design
Explained -
The Viable System Model: Its Provenance, Development, Methodology and Pathology
Explained -
Strategic Information Transmission
Explained -
Assessing the Impact of Planned Social Change
Explained -
On the Folly of Rewarding A, While Hoping for B
Explained -
Every good regulator of a system must be a model of that system
Explained -
Requisite Variety and Its Implications for the Control of Complex Systems
Explained - Explained
-
Does DiffusionGemma do latent reasoning?
Verified -
Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds
Verified -
Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
Verified -
Rules or Character? Scaling Laws for AI Safety Design
Verified -
AI swarms are starting to pose indirect takeover risk
Verified -
Introducing the Conceptual Reasoning Index
Verified -
GPT-Red: Automated Red Teaming via Self-Play at Scale
Verified -
How independent researchers could investigate AI propensities after misalignment incidents
Verified -
Agentic Misalignment in Summer 2026
Verified -
Modular Pretraining Enables Access Control
Verified -
Separating signal from noise in coding evaluations
Verified -
Verbalizable Representations Form a Global Workspace in Language Models
Verified -
Summary of METR's predeployment evaluation of GPT-5.6 Sol
Verified -
"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms
Verified -
Diffuse AI Control on Fuzzy Tasks
Verified -
Quantifying the Salience of Geo-Cultural Values for Pluralistic Safety Alignment
Verified -
Gram: Assessing sabotage propensities via automated alignment auditing
Verified -
Realistic honeypot evaluations for scheming propensity
Verified -
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
Verified -
Automated alignment is harder than you think
Verified -
Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
Verified -
Model Spec Midtraining: Improving How Alignment Training Generalizes
Verified - Verified
-
BashArena: A Control Setting for Highly Privileged AI Agents
Verified -
Imitation Learning is Probably Existentially Safe
Verified -
Ctrl-Z: Controlling AI Agents via Resampling
Verified - Listed
- Listed
- Listed
-
Algorithmic learning in a random world
Listed -
An Introduction to Game Theory, Chapters 1-7,14,15
Listed - Listed
-
Artificial Intelligence Safety and Security
Listed - Listed
- Listed
-
Causality: models, reasoning, and inference
Listed - Listed
-
CEA's Existential Risk and the Far Future playlist
Listed - Listed
- Listed
- Listed
- Listed
-
Computability and Logic, Chapters 1-4, 8-20, 23, 25, and 27
Listed - Listed
- Listed
- Listed
- Listed
- Listed
- Listed
- Listed
-
Defining Human Values for Value Learners
Listed -
Distributed Representations: Composition & Superposition
Listed -
Do Androids Dream of Electric Sheep?
Listed -
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Listed - Listed
- Listed
-
Empirical examples of power-law CDFs
Listed -
Ethics Background (Introduction through “Absolute Rights or Prima Facie Duties”)
Listed -
Everything else Chris Olah has ever written
Listed -
Flatland: A Romance of Many Dimensions
Listed -
FLI's AI Safety Research Landscape
Listed -
Formalizing Convergent Instrumental Goals
Listed - Listed
- Listed
-
Game Theory: Analysis of Conflict
Listed - Listed
- Listed
-
Handbook of Model Checking (to appear soon)
Listed -
Harry Potter and the Methods of Rationality (#1 of 6)
Listed - Listed
- Listed
-
Information Theory, Inference, and Learning Algorithms Parts I-III
Listed - Listed
- Listed
-
Introduction to Artificial Intelligence
Listed -
Introduction to Automata Theory, Languages, and Computation, Chapters 1-10
Listed - Listed
-
Learning the Preferences of Bounded Agents
Listed -
Log-normal distributions (with comparisons to power laws)
Listed - Listed
- Listed
-
Machine Learning (lecture notes)
Listed -
Machine Learning (online course)
Listed -
Mathematical Logic : A course with exercises — Part I
Listed -
Mathematical Logic : A course with exercises — Part II, Chapters 5 and 6
Listed -
Maximizing a quantity while ignoring effect through some channel
Listed -
Mechanistic Interpretability, Variables, and the Importance of Interpretable Bases
Listed - Listed
-
Moral Machines: Teaching robots right from wrong
Listed -
Motivated Value Selection for Artificial Agents
Listed -
Multiplicative processes produce log normals
Listed -
Multiplicative processes produce power laws
Listed - Listed
- Listed
- Listed
-
Numerous power laws for cities
Listed - Listed
-
On Explainability in Machine Learning
Listed - Listed
-
Pattern Recognition and Machine Learning
Listed -
Probability: Theory and Examples, Chapters 1-6
Listed -
Quantilizers Limited Optimization
Listed - Listed
-
Review of properties of power laws
Listed -
Robert Miles discuss AI on Computerphile
Listed -
Robert Miles's own YouTube channel
Listed - Listed
- Listed
-
Superforecasting – Philip Tetlock
Listed - Listed
- Listed
- Listed
- Listed
- Listed
- Listed
-
The Dark Forest (#2 of Three Body Problem)
Listed - Listed
- Listed
-
The Righteous Mind: Why Good People Are Divided by Politics and Religion
Listed - Listed
-
The Singularity: A Philosophical Analysis
Listed -
The Structure of Normative Ethics
Listed -
The Unilateralist’s Curse: The Case for a Principle of Conformity
Listed -
Theory and applications of Robust Optimization
Listed - Listed
- Listed
-
Universal Artificial Intelligence, Chapters 2-5
Listed - Listed
-
What happens when our computers get smarter than we are?
Listed - Listed
-
Workshop On Safety And Control For Artificial Intelligence
Listed -
Experiences and learnings from both sides of the AI safety job market
Listed - Listed
- Listed
-
New report: "Scheming AIs: Will AIs fake alignment during training in order to get power?"
Listed -
Betting on what is un-falsifiable and un-verifiable
Listed -
Is Interpretability All We Need?
Listed -
Is there Work on Embedded Agency in Cellular Automata Toy Models?
Listed -
Would this be Progress in Solving Embedded Agency?
Listed -
AISC Project: Benchmarks for Stable Reflectivity
Listed -
AISC Project: Modelling Trajectories of Language Models
Listed -
Optionality approach to ethics
Listed - Listed
-
The Science Algorithm AISC Project
Listed -
Theories of Change for AI Auditing
Listed -
Why small phenomenons are relevant to morality
Listed -
AISC project: SatisfIA – AI that satisfies without overdoing it
Listed -
Control Symmetry: why we might want to start investigating asymmetric alignment interventions
Listed -
Game Theory without Argmax [Part 1]
Listed -
Game Theory without Argmax [Part 2]
Listed -
Open Phil releases RFPs on LLM Benchmarks and Forecasting
Listed -
The Existential Risk of Speciesist Bias in AI
Listed -
The Top AI Safety Bets for 2023: GiveWiki’s Latest Recommendations
Listed - Listed
-
Artefacts generated by mode collapse in GPT-4 Turbo serve as adversarial attacks.
Listed -
EA Poland is facing an existential risk
Listed -
GPT-2030 and Catastrophic Drives: Four Vignettes
Listed -
Munk Debate on AI: a few observations and opinions
Listed -
Update on the UK AI Summit and the UK's Plans
Listed -
We have promising alignment plans with low taxes
Listed -
ACI#6: A Non-Dualistic ACI Model
Listed - Listed
-
Learning-theoretic agenda reading list
Listed -
Polysemantic Attention Head in a 4-Layer Transformer
Listed -
What we're missing: the case for structural risks from AI
Listed -
Open-ended/Phenomenal Ethics (TLTR)
Listed -
Alignment Frame/Exercise: Building The Puzzle of Alignment
Listed -
Growth and Form in a Toy Model of Superposition
Listed - Listed
-
Open-ended ethics of phenomena (a desiderata with universal morality)
Listed -
Tall Tales at Different Scales: Evaluating Scaling Trends For Deception In Language Models
Listed -
What’s going on? LLMs and IS-A sentences
Listed -
AI Alignment Research Engineer Accelerator (ARENA): call for applicants
Listed -
Announcing Athena - Women in AI Alignment Research
Listed - Listed
- Listed
- Listed
- Listed
-
Please, someone make a dataset of supposed cases of "tech panic"
Listed -
Scalable And Transferable Black-Box Jailbreaks For Language Models Via Persona Modulation
Listed -
Scalable And Transferable Black-Box Jailbreaks For Language Models Via Persona Modulation
Listed -
Scalable And Transferable Black-Box Jailbreaks For Language Models Via Persona Modulation
Listed - Listed
-
20+ tips, tricks, lessons and thoughts on hosting hackathons
Listed -
AI Fables Writing Contest Winners!
Listed -
An illustrative model of backfire risks from pausing AI research
Listed - Listed
-
Governance of AI, Breakfast Cereal, Car Factories, Etc.
Listed -
On running a city-wide university group
Listed -
Tips, tricks, lessons and thoughts on hosting hackathons
Listed -
Why building ventures in AI Safety is particularly challenging
Listed -
Why building ventures in AI Safety is particularly challenging
Listed - Listed
-
Disentangling four motivations for acting in accordance with UDT
Listed -
Eric Schmidt on recursive self-improvement
Listed -
xAI announces Grok, beats GPT-3.5
Listed -
[Linkpost] Concept Alignment as a Prerequisite for Value Alignment
Listed -
Despair about AI progressing too slowly
Listed -
Genetic fitness is a measure of selection strength, not the selection target
Listed -
The 6D effect: When companies take risks, one email can be very powerful.
Listed -
Untrusted smart models and trusted dumb models
Listed -
We are already in a persuasion-transformed world and must take precautions
Listed -
If AGI is imminent, why can’t I hail a robotaxi?
Listed -
Paul Christiano on Dwarkesh Podcast
Listed -
Sam Altman: "safety and capabilities are not these two separate things"
Listed - Listed
- Listed
-
Why is learning economics, psychology, sociology important for preventing AI risks?
Listed -
A Critique of The Evidentialist's Wager
Listed -
Mech Interp Challenge: November - Deciphering the Cumulative Sum Model
Listed -
Still no strong evidence that LLMs increase bioterrorism risk
Listed -
[Congressional Hearing] Oversight of A.I.: Legislating on Artificial Intelligence
Listed