The catalog, page 5
Records 1,001 to 1,250 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Towards Measuring the Representation of Subjective Global Opinions in Language Models
Listed -
When do "brains beat brawn" in Chess? An experiment
Listed -
AISC team report: Soft-optimization, Bayes and Goodhart
Listed - Listed
-
An overview of the points system
Listed -
Catastrophic Risks from AI #5: Rogue AIs
Listed -
Catastrophic Risks from AI #6: Discussion and FAQ
Listed -
ML4G Germany - AI Alignment Camp
Listed -
"Safety Culture for AI" is important, but isn't going to be easy
Listed -
AI Safety Field Building vs. EA CB
Listed -
Catastrophic Risks from AI #4: Organizational Risks
Listed -
Deceptive AI vs. shifting instrumental incentives
Listed - Listed
-
Let’s set new AI safety actors up for success
Listed -
Looking for Canadian summer co-op position in AI Governance
Listed -
The EU AI Act: A Simple Explanation - A Stanford Study Reveals the gaps of ChatGPT and 9 more
Listed -
The fraught voyage of aligned novelty
Listed -
Where on the continuum of pure EA to pure AIS should you be? (Uni Group Organizers Focus)
Listed -
Did Bengio and Tegmark lose a debate about AI x-risk against LeCun and Mitchell?
Listed -
Map of maps of interesting fields
Listed -
Would a super-intelligent AI necessarily support its own existence?
Listed -
Democratic AI Constitution: Round-Robin Debate and Synthesis
Listed -
DSLT 4. Phase Transitions in Neural Networks
Listed -
Announcing the AIPolicyIdeas.com Database
Listed -
Catastrophic Risks from AI #3: AI Race
Listed -
On the compute governance era and what has to come after (Lennart Heim on The 80,000 Hours Podcast)
Listed -
OpenAI's grant program for democratic process for deciding what rules AI systems should follow
Listed -
Slaying the Hydra: toward a new game board for AI
Listed -
Thoughts about AI safety field-building in LMIC
Listed -
What should I ask Ezra Klein about AI policy proposals?
Listed -
Announcing the EA Project Ideas Database
Listed -
Catastrophic Risks from AI #1: Summary
Listed -
Catastrophic Risks from AI #2: Malicious Use
Listed -
RP’s AI Governance & Strategy team - June 2023 interim overview
Listed -
The Hubinger lectures on AGI safety: an introductory lecture series
Listed -
US public perception of CAIS statement and the risk of extinction
Listed - Listed
-
Yip Fai Tse on animal welfare & AI safety and long termism
Listed -
20 concrete projects for reducing existential risk
Listed - Listed
-
An Overview of Catastrophic AI Risks
Listed -
EU AI Act passed Plenary vote, and X-risk was a main topic
Listed -
Join the Virtual AI Safety Unconference (VAISU)!
Listed -
Upcoming speaker series on emerging tech, national security & US policy careers
Listed -
Using Claude to convert dialog transcripts into great posts?
Listed -
A Friendly Face (Another Failure Story)
Listed -
Ban development of unpredictable powerful models?
Listed -
Causality: A Brief Introduction
Listed -
Corrigibility test #1: Shutdown activations in a Virus Research Lab
Listed -
DSLT 3. Neural Networks are Singular
Listed -
Lightning Post: Things people in AI Safety should stop talking about
Listed -
LPP Summer Research Fellowship in Law & AI 2023: Applications Open
Listed -
Simulating Shutdown Code Activations in an AI Virus Lab
Listed -
Summary of the AI Bill of Rights and Policy Implications
Listed -
A Multidisciplinary Approach to Alignment (MATA) and Archetypal Transfer Learning (ATL)
Listed -
Experiments in Evaluating Steering Vectors
Listed -
Mode collapse in RL may be fueled by the update equation
Listed -
New reference standard on LLM Application security started by OWASP
Listed -
Principles for AI Welfare Research
Listed - Listed
-
The Multidisciplinary Approach to Alignment (MATA) and Archetypal Transfer Learning (ATL)
Listed -
DSLT 2. Why Neural Networks obey Occam's Razor
Listed -
My lab's small AI safety agenda
Listed -
UK Foundation Model Task Force - Expression of Interest
Listed -
A summary of current work in AI governance
Listed -
Partial Simulation Extrapolation: A Proposal for Building Safer Simulators
Listed -
The AI governance gaps in developing countries
Listed -
[Replication] Conjecture's Sparse Coding in Small Transformers
Listed -
Conjecture: A standing offer for public debates on AI
Listed -
Critiques of non-existent AI safety labs: Yours
Listed - Listed
-
DSLT 0. Distilling Singular Learning Theory
Listed -
DSLT 1. The RLCT Measures the Effective Dimension of Neural Networks
Listed -
LLMs Sometimes Generate Purely Negatively-Reinforced Text
Listed -
Safety evaluations and standards for AI | Beth Barnes | EAG Bay Area 23
Listed -
Scaffolded LLMs: Less Obvious Concerns
Listed -
Updates from Campaign for AI Safety
Listed - Listed
-
What would it look like for AIS to no longer be neglected?
Listed -
Aligned Objectives Prize Competition
Listed -
AXRP Episode 22 - Shard Theory with Quintin Pope
Listed -
Brief thoughts on Data, Reporting, and Response for AI Risk Mitigation
Listed -
EU AI Act passed vote, and x-risk was a main topic
Listed -
human intelligence may be alignment-limited
Listed -
Inverse Scaling: When Bigger Isn’t Better
Listed -
PhD student and postdoc positions philosophy of AI in Erlangen (Germany)
Listed -
Report: Artificial Intelligence Risk Management in Spain
Listed -
UN Secretary-General recognises existential threat from AI
Listed -
Why "AI alignment" would better be renamed into "Artificial Intention research"
Listed -
a short chat about realityfluid
Listed -
AI Safety Strategy - A new organization for better timelines
Listed -
Anthropic | Charting a Path to AI Accountability
Listed - Listed
-
Instrumental Convergence? [Draft]
Listed -
Linkpost: Dwarkesh Patel interviewing Carl Shulman
Listed -
<$750k grants for General Purpose AI Assurance/Safety Research
Listed -
Aptitudes for AI governance work
Listed -
Epoch and FRI Mentorship Program Summer 2023
Listed - Listed
-
MetaAI: less is less for alignment.
Listed -
Raising the voices that actually count
Listed -
Some talent needs in AI governance
Listed -
TASRA: A Taxonomy and Analysis of Societal-Scale Risks from AI
Listed -
There is only one goal or drive - only self-perpetuation counts
Listed -
Tony Blair Institute AI Safety Work
Listed -
What's the exact way you predict probability of AI extinction?
Listed -
A Manifold Market "Leaked" the AI Extinction Statement and CAIS Wanted it Deleted
Listed -
ARC is hiring theoretical researchers
Listed -
ARC is hiring theoretical researchers
Listed -
Contingency: A Conceptual Tool from Evolutionary Biology for Alignment
Listed -
Critiques of prominent AI safety labs: Conjecture
Listed -
Critiques of prominent AI safety labs: Conjecture
Listed - Listed
-
If you are too stressed, walk away from the front lines
Listed -
Import AI 332: Mini-AI; safety through evals; Facebook releases a RLHF dataset
Listed -
Introduction to Towards Causal Foundations of Safe AGI
Listed - Listed
-
TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI
Listed -
What can superintelligent ANI tell us about superintelligent AGI?
Listed -
Higher Dimension Cartesian Objects and Aligning ‘Tiling Simulators’
Listed -
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
Listed - Listed
-
an Evangelion dialogue explaining the QACI alignment plan
Listed -
Are we confident that superintelligent artificial intelligence disempowering humans would be bad?
Listed -
formalizing the QACI alignment formal-goal
Listed -
Goal-misgeneralization is ELK-hard
Listed -
Using Consensus Mechanisms as an approach to Alignment
Listed - Listed
-
A plea for solutionism on AI safety
Listed - Listed
-
an Evangelion dialogue explaining the QACI alignment plan
Listed -
Announcement: You can now listen to the “AI Safety Fundamentals” courses
Listed -
formalizing the QACI alignment formal-goal
Listed -
How biosafety could inform AI standards
Listed -
How does AI progress affect other EA cause areas?
Listed -
Improvement on MIRI's Corrigibility
Listed -
[Linkpost] Scaling laws for language encoding models in fMRI
Listed -
A comparison of causal scrubbing, causal abstractions, and related methods
Listed -
A potentially high impact differential technological development area
Listed -
A survey of concrete risks derived from Artificial Intelligence
Listed -
Beware popular discussions of AI "sentience"
Listed -
Takeaways from the Mechanistic Interpretability Challenges
Listed -
Transformative AI is a process
Listed -
UK government to host first global summit on AI Safety
Listed -
Wild Animal Welfare Scenarios for AI Doom
Listed -
A note of caution about recent AI risk coverage
Listed -
An Exercise to Build Intuitions on AGI Risk
Listed -
Article Summary: Current and Near-Term AI as a Potential Existential Risk Factor
Listed -
Could AI accelerate economic growth?
Listed -
Proposal: Tune LLMs to Use Calibrated Language
Listed -
Rethink Priorities is hiring a Compute Governance Researcher or Research Assistant
Listed -
The current alignment plan, and how we might improve it | EAG Bay Area 23
Listed -
Understanding how hard alignment is may be the most important research direction right now
Listed - Listed
-
[Linkpost] Given Extinction Worries, Why Don’t AI Researchers Quit? Well, Several Reasons
Listed -
A Playbook for AI Risk Reduction (focused on misaligned AI)
Listed -
Agentic Mess (A Failure Story)
Listed -
AISN #9: Statement on Extinction Risks, Competitive Pressures, and When Will AI Reach Human-Level?
Listed -
Algorithmic Improvement Is Probably Faster Than Scaling Now
Listed - Listed
-
Stampy's AI Safety Info - New Distillations #3 [May 2023]
Listed -
The Sharp Right Turn: sudden deceptive alignment as a convergent goal
Listed -
Tim Cook was asked about extinction risks from AI
Listed -
Transformative AGI by 2043 is <1% likely
Listed - Listed
-
AISafety.info "How can I help?" FAQ
Listed -
Moral Spillover in Human-AI Interaction
Listed - Listed
-
AI Safety Fundamentals: An Informal Cohort Starting Soon!
Listed -
AI Safety Fundamentals: An Informal Cohort Starting Soon! (cross-posted to lesswrong.com)
Listed -
Decomposing alignment to take advantage of paradigms
Listed - Listed
-
How to Think About Activation Patching
Listed -
One implementation of regulatory GPU restrictions
Listed -
Details on how an IAEA-style AI regulator would function?
Listed - Listed
-
Terry Tao is hosting an "AI to Assist Mathematical Reasoning" workshop
Listed -
The AGI Race Between the US and China Doesn’t Exist.
Listed -
Unfaithful Explanations in Chain-of-Thought Prompting
Listed -
Upcoming AI regulations are likely to make for an unsafer world
Listed -
[Replication] Conjecture's Sparse Coding in Toy Models
Listed -
Advice for Entering AI Safety Research
Listed -
Applications open for AI Safety Fundamentals: Governance Course
Listed -
Catastrophic Risks from Unsafe AI: Navigating a Tightrope Scenario (Ben Garfinkel, EAG London 2023)
Listed -
Co-found an incubator for independent AI Safety researchers (rolling applications)
Listed - Listed
-
Proposal: labs should precommit to pausing if an AI argues for itself to be improved
Listed -
Some thoughts on "AI could defeat all of us combined"
Listed -
The Control Problem: Unsolved or Unsolvable?
Listed -
Think carefully before calling RL policies "agents"
Listed -
AI Manufactured Crisis (don't trust AI to protect us from AI)
Listed -
An explanation of decision theories
Listed -
Four levels of understanding decision theory
Listed - Listed
-
Open Source LLMs Can Now Actively Lie
Listed -
Outreach success: Intro to AI risk that has been successful
Listed - Listed
- Listed
- Listed
-
Uncertainty about the future does not imply that AGI will go well
Listed -
Update from Campaign for AI Safety
Listed - Listed
-
A compute-based framework for thinking about the future of AI
Listed -
A moral backlash against AI will probably slow down AGI development
Listed -
A push towards interactive transformer decoding
Listed -
Considerations on transformative AI and explosive growth from a semiconductor-industry perspective
Listed -
Contrast Pairs Drive the Empirical Performance of Contrast Consistent Search (CCS)
Listed -
Cosmopolitan values don't come free
Listed -
Exponential AI takeoff is a myth
Listed -
Improving mathematical reasoning with process supervision
Listed - Listed
-
Limiting factors to predict AI take-off speed
Listed - Listed
-
Neuroevolution, Social Intelligence, and Logic
Listed - Listed
- Listed
-
Unpredictability and the Increasing Difficulty of AI Alignment for Increasingly Intelligent AI
Listed -
Advice for new alignment people: Info Max
Listed -
AI Doom and David Hume: A Defence of Empiricism in AI Safety
Listed - Listed
- Listed
-
Boomerang - protocol to dissolve some commitment races
Listed -
Implications of AGI on Subjective Human Experience
Listed -
LIMA: Less Is More for Alignment
Listed -
PaLM-2 & GPT-4 in "Extrapolating GPT-N performance"
Listed -
Statement on AI Extinction - Signed by AGI Labs, Top Academics, and Many Other Notable Figures
Listed -
Statement on AI Extinction - Signed by AGI Labs, Top Academics, and Many Other Notable Figures
Listed - Listed
-
The bullseye framework: My case against AI doom
Listed - Listed
-
The case for removing alignment and ML research from the training dataset
Listed - Listed
-
Aligning an H-JEPA agent via training on the outputs of an LLM-based "exemplary actor"
Listed -
An LLM-based “exemplary actor”
Listed - Listed
-
Language Agents Reduce the Risk of Existential Catastrophe
Listed -
List of Masters Programs in Tech Policy, Public Policy and Security (Europe)
Listed - Listed
-
What are some of the best introductions/breakdowns of AI existential risk for those unfamiliar?
Listed -
Wikipedia as an introduction to the alignment problem
Listed -
Without a trajectory change, the development of AGI is likely to go badly
Listed -
Language Agents Reduce the Risk of Existential Catastrophe
Listed -
My AI Alignment Research Agenda and Threat Model, right now (May 2023)
Listed - Listed
-
TinyStories: Small Language Models That Still Speak Coherent English
Listed -
Why and When Interpretability Work is Dangerous
Listed -
[Linkpost] Longtermists Are Pushing a New Cold War With China
Listed - Listed
-
Diminishing Returns in Machine Learning Part 1: Hardware Development and the Physical Frontier
Listed -
Hands-On Experience Is Not Magic
Listed