The catalog, page 12
Records 2,751 to 3,000 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Two contrasting models of “intelligence” and future growth
Listed - Listed
-
Against a General Factor of Doom
Listed -
Announcing AI safety Mentors and Mentees
Listed -
Conjecture Second Hiring Round
Listed -
Conjecture: a retrospective after 8 months of work
Listed -
Human-level Diplomacy was my fire alarm
Listed -
Injecting some numbers into the AGI debate - by Boaz Barak
Listed -
Notes on an Experiment with Markets
Listed -
Simulators, constraints, and goal agnosticism: porbynotes vol. 1
Listed -
What is the best source to explain short AI timelines to a skeptical person?
Listed -
A Walkthrough of In-Context Learning and Induction Heads (w/ Charles Frye) Part 1 of 2
Listed -
AI will change the world, but won’t take it over by playing “3-dimensional chess”.
Listed -
Announcing AI Alignment Awards: $100k research contests about goal misgeneralization & corrigibility
Listed -
Benchmarking the next generation of never-ending learners
Listed -
Brute-forcing the universe: a non-standard shot at diamond alignment
Listed -
Epoch is hiring a Research Data Analyst
Listed -
Human-level Full-Press Diplomacy (some bare facts).
Listed -
imitation: Clean Imitation Learning Implementations
Listed - Listed
-
Meta AI announces Cicero: Human-Level Diplomacy play (with dialogue)
Listed -
Toby Ord's new report on lessons from the development of the atomic bomb
Listed -
What is the best article to introduce someone to AI safety for the first time?
Listed -
[Hebbian Natural Abstractions] Introduction
Listed -
Benefits/Risks of Scott Aaronson's Orthodox/Reform Framing for AI Alignment
Listed -
Beyond Simple Existential Risk: Survival in a Complex Interconnected World
Listed -
Pre-Announcing the 2023 Open Philanthropy AI Worldviews Contest
Listed -
Review: What We Owe The Future
Listed -
ARC paper: Formalizing the presumption of independence
Listed - Listed
-
Decision Theory but also Ghosts
Listed -
let's stick with the term "moral patient"
Listed -
A Short Dialogue on the Meaning of Reward Functions
Listed -
By Default, GPTs Think In Plain Sight
Listed - Listed
-
Update to Mysteries of mode collapse: text-davinci-002 not RLHF
Listed -
wonky but good enough alignment schemes
Listed -
"humans aren't aligned" and "human values are incoherent"
Listed -
Artificial Intelligence and Nuclear Command, Control, & Communications: The Risks of Integration
Listed -
Cognitive science and failed AI forecasts
Listed -
Distillation of "How Likely Is Deceptive Alignment?"
Listed -
Don't design agents which exploit adversarial inputs
Listed -
Engineering Monosemanticity in Toy Models
Listed - Listed
-
The Disastrously Confident And Inaccurate AI
Listed - Listed
- Listed
-
LLMs may capture key components of human agency
Listed -
Massive Scaling Should be Frowned Upon
Listed -
Results from the interpretability hackathon
Listed -
The Ground Truth Problem (Or, Why Evaluating Interpretability Methods Is Hard)
Listed -
Current themes in mechanistic interpretability research
Listed -
Disagreement with bio anchors that lead to shorter timelines
Listed -
Questions about Value Lock-in, Paternalism, and Empowerment
Listed -
Unpacking "Shard Theory" as Hunch, Question, Theory, and Insight
Listed -
Graph of % of tasks AI is superhuman at?
Listed -
If FTX is liquidated, who ends up controlling Anthropic?
Listed - Listed
-
The economy as an analogy for advanced AI systems
Listed -
The limited upside of interpretability
Listed -
Training for Good - Update & Plans for 2023
Listed -
Value Formation: An Overarching Model
Listed -
Winners of the AI Safety Nudge Competition
Listed - Listed
- Listed
- Listed
-
Will we run out of ML data? Evidence from projecting dataset size trends
Listed -
a safer experiment than quantum suicide
Listed -
A short critique of Vanessa Kosoy's PreDCA
Listed -
Decision making under model ambiguity, moral uncertainty, and other agents with free will?
Listed -
The Alignment Community Is Culturally Broken
Listed -
Will AI Worldview Prize Funding Be Replaced?
Listed -
fully aligned singleton as a solution to everything
Listed - Listed
-
Vanessa Kosoy's PreDCA, distilled
Listed - Listed
-
Apply now for the EU Tech Policy Fellowship 2023
Listed -
Instrumental convergence is what makes general intelligence possible
Listed -
What are some low-cost outside-the-box ways to do/fund alignment research?
Listed -
Why I'm Working On Model Agnostic Interpretability
Listed -
Adversarial Priors: Not Paying People to Lie to You
Listed -
I there a demo of "You can't fetch the coffee if you're dead"?
Listed -
Is full self-driving an AGI-complete problem?
Listed -
[ASoT] Instrumental convergence is useful
Listed -
AI Safety groups should imitate career development clubs
Listed -
Restricting brain organoid research to slow down AGI
Listed -
Trying to Make a Treacherous Mesa-Optimizer
Listed -
A first success story for Outer Alignment: InstructGPT
Listed -
Applying superintelligence without collusion
Listed -
Applying superintelligence without collusion
Listed -
Inverse scaling can become U-shaped
Listed - Listed
-
People care about each other even though they have imperfect motivational pointers?
Listed -
Some advice on independent research
Listed -
4 Key Assumptions in AI Safety
Listed -
A philosopher's critique of RLHF
Listed - Listed
-
AI Safety Unconference NeurIPS 2022
Listed - Listed
-
How could we know that an AGI system will have good consequences?
Listed -
How does one find out their AGI timelines?
Listed - Listed
-
Examining the Differential Risk from High-level Artificial Intelligence and the Question of Control
Listed -
Has anyone increased their AGI timelines?
Listed -
Longevity research as AI X-risk intervention
Listed -
You won’t solve alignment without agent foundations
Listed -
"AGI timelines: ignore the social factor at their peril" (Future Fund AI Worldview Prize submission)
Listed -
"AI predictions" (Future Fund AI Worldview Prize submission)
Listed - Listed
-
Instead of technical research, more people should focus on buying time
Listed -
Is AI forecasting a waste of effort on the margin?
Listed -
My summary of “Pragmatic AI Safety”
Listed -
My summary of “Pragmatic AI Safety”
Listed -
Recommend HAIST resources for assessing the value of RLHF-related alignment research
Listed -
Takeaways from a survey on AI alignment resources
Listed -
The Slippery Slope from DALLE-2 to Deepfake Anarchy
Listed -
The Slippery Slope from DALLE-2 to Deepfake Anarchy
Listed -
A new place to discuss cognitive science, ethics and human alignment
Listed -
A newcomer’s guide to the technical AI safety field
Listed -
Applications are now open for Intro to ML Safety Spring 2023
Listed -
Are alignment researchers devoting enough time to improving their research capacity?
Listed -
Don't you think RLHF solves outer alignment?
Listed -
For ELK truth is mostly a distraction
Listed -
How to store human values on a computer
Listed -
Measuring Progress on Scalable Oversight for Large Language Models
Listed - Listed
-
A Mystery About High Dimensional Concept Encoding
Listed -
A Theologian's Response to Anthropogenic Existential Risk
Listed -
Further considerations on the Evidentialist's Wager
Listed -
Liability regimes in the age of AI: a use-case driven analysis of the burden of proof
Listed -
Mechanistic Interpretability as Reverse Engineering (follow-up to "cars and elephants")
Listed -
Why do we post our AI safety plans on the Internet?
Listed -
AI Safety Needs Great Product Builders
Listed -
AI X-risk >35% mostly based on a recent peer-reviewed argument
Listed -
Announcing: What Future World? - Growing the AI Governance Community
Listed -
Humans do acausal coordination all the time
Listed -
WFW?: Opportunity and Theory of Impact
Listed -
a casual intro to AI doom and alignment
Listed -
a casual intro to AI doom and alignment
Listed -
Adversarial Policies Beat Professional-Level Go AIs
Listed -
AI X-Risk: Integrating on the Shoulders of Giants
Listed -
All AGI Safety questions welcome (especially basic ones) [~monthly thread]
Listed -
Auditing games for high-level interpretability
Listed -
Caution when interpreting Deepmind's In-context RL paper
Listed - Listed
-
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Listed -
ML Safety Scholars Summer 2022 Retrospective
Listed - Listed
-
Real-Time Research Recording: Can a Transformer Re-Derive Positional Info?
Listed -
Should AI focus on problem-solving or strategic planning? Why not both?
Listed -
Threat Model Literature Review
Listed -
"Cars and Elephants": a handwavy argument/analogy against mechanistic interpretability
Listed -
[Book] Interpretable Machine Learning: A Guide for Making Black Box Models Explainable
Listed -
Announcing The Most Important Century Writing Prize
Listed - Listed
-
Embedding safety in ML development
Listed -
My (naive) take on Risks from Learned Optimization
Listed -
publishing alignment research and exfohazards
Listed -
Superintelligent AI is necessary for an amazing future, but far from sufficient
Listed -
Teacher-student curriculum learning for reinforcement learning
Listed -
What sorts of systems can be deceptive?
Listed -
Instrumental ignoring AI, Dumb but not useless.
Listed -
Me (Steve Byrnes) on the “Brain Inspired” podcast
Listed -
«Boundaries», Part 3a: Defining boundaries as directed Markov blankets
Listed - Listed
-
Is there a news-tracker about GPT-4? Why has everything become so silent about it?
Listed - Listed
-
aisafety.community - A living document of AI safety communities
Listed -
Join the interpretability research hackathon
Listed -
Prizes for ML Safety Benchmark Ideas
Listed -
Prizes for ML Safety Benchmark Ideas
Listed - Listed
-
Resources that (I think) new alignment researchers should know about
Listed -
Some Lessons Learned from Studying Indirect Object Identification in GPT-2 small
Listed - Listed
- Listed
- Listed
-
counterfactual computations in world models
Listed - Listed
-
Intent alignment should not be the goal for AGI x-risk reduction
Listed - Listed
-
Paper: In-context Reinforcement Learning with Algorithm Distillation [Deepmind]
Listed -
Summary of "Technology Favours Tyranny" by Yuval Noah Harari
Listed -
Why some people believe in AGI, but I don't.
Listed -
A Brief Summary Of The Most Important Century
Listed -
A Walkthrough of A Mathematical Framework for Transformer Circuits
Listed - Listed
-
Maps and Blueprint; the Two Sides of the Alignment Equation
Listed -
Mechanism Design for AI Safety - Reading Group Curriculum
Listed -
What does it take to defend the world against out-of-control AGIs?
Listed -
A Barebones Guide to Mechanistic Interpretability Prerequisites
Listed -
Call to action: Read + Share AI Safety / Reinforcement Learning Featured in Conversation
Listed -
Emergent world representations: Exploring a sequence model trained on a synthetic task
Listed - Listed
-
POWERplay: An open-source toolchain to study AI power-seeking
Listed -
The optimal timing of spending on AGI safety work; why we should probably be spending more now
Listed -
Empowerment is (almost) All We Need
Listed -
QACI: question-answer counterfactual intervals
Listed -
Newsletter for Alignment Research: The ML Safety Updates
Listed -
Simple question about corrigibility and values in AI.
Listed -
Intelligent behaviour across systems, scales and substrates
Listed - Listed
-
Learning societal values from law as part of an AGI alignment strategy
Listed -
Notes on "Can you control the past"
Listed -
Scaling Laws for Reward Model Overoptimization
Listed -
Task Phasing: Automated Curriculum Learning from Demonstrations
Listed -
The heritability of human values: A behavior genetic critique of Shard Theory
Listed -
The heritability of human values: A behavior genetic critique of Shard Theory
Listed - Listed
-
What Does AI Alignment Success Look Like?
Listed -
Governments pose larger risks than corporations: a brief response to Grace
Listed -
Response to Katja Grace's AI x-risk counterarguments
Listed -
Scaling laws for reward model overoptimization
Listed -
Should we push for requiring AI training data to be licensed?
Listed -
[Link post] AI could fuel factory farming—or end it
Listed -
A conversation about Katja's counterarguments to AI risk
Listed -
An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers
Listed -
Decision theory does not imply that we get to have nice things
Listed -
Distilled Representations Research Agenda
Listed -
Infinite Possibility Space and the Shutdown Problem
Listed -
Metaculus is building a team dedicated to AI forecasting
Listed -
Science of Deep Learning - a technical agenda
Listed -
‘Dissolving’ AI Risk – Parameter Uncertainty in AI Future Forecasting
Listed - Listed
-
A.I. Robustness: a Human-Centered Perspective on Technological Challenges and Opportunities
Listed -
AI Safety Ideas: A collaborative AI safety research platform
Listed - Listed
-
Is interest in alignment worth mentioning for grad school applications?
Listed -
Maximal lotteries for value learning
Listed -
Why not to solve alignment by making superintelligent humans?
Listed -
Best resource to go from "typical smart tech-savvy person" to "person who gets AGI risk urgency"?
Listed -
Toward Next-Generation Artificial Intelligence: Catalyzing the NeuroAI Revolution
Listed -
[Job]: AI Standards Development Research Assistant
Listed -
Another problem with AI confinement: ordinary CPUs can work as radio transmitters
Listed -
Counterarguments to the basic AI risk case
Listed -
Counterarguments to the basic AI x-risk case
Listed -
Counterarguments to the basic AI x-risk case
Listed -
Instrumental convergence: scale and physical interactions
Listed -
The US expands restrictions on AI exports to China. What are the x-risk effects?
Listed -
The Vitalik Buterin Fellowship in AI Existential Safety is open for applications!
Listed -
Cataloguing Priors in Theory and Practice
Listed -
CNAS report: 'Artificial Intelligence and Arms Control'
Listed -
Contra shard theory, in the context of the diamond maximizer problem
Listed -
Greed Is the Root of This Evil
Listed -
Misalignment-by-default in multi-agent systems
Listed - Listed
- Listed
-
Sixty years after the Cuban Missile Crisis, a new era of global catastrophic risks
Listed -
You are better at math (and alignment) than you think
Listed -
[MLSN #6]: Transparency survey, provable robustness, ML models that predict the future
Listed