The catalog, page 4
Records 751 to 1,000 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
EU’s AI ambitions at risk as US pushes to water down international treaty (linkpost)
Listed -
How to find AI alignment researchers to collaborate with?
Listed -
If AIs had subcortical brain simulation, would that solve the alignment problem?
Listed - Listed
-
Is there any existing term summarizing non-scalable oversight methods in outer alignment?
Listed -
Open Problems and Fundamental Limitations of RLHF
Listed -
The “no sandbagging on checkable tasks” hypothesis
Listed -
The “no sandbagging on checkable tasks” hypothesis
Listed -
Thoughts on sharing information about language model capabilities
Listed -
Trading off compute in training and inference (Overview)
Listed -
Watermarking considered overrated?
Listed - Listed
-
Shutting down AI Safety Support
Listed - Listed
-
Announcing the ITAM AI Futures Fellowship
Listed -
Introductory Textbook to Vision Models Interpretability
Listed -
Mech Interp Puzzle 2: Word2Vec Style Embeddings
Listed -
Reducing sycophancy and improving honesty via activation steering
Listed -
Reducing sycophancy and improving honesty via activation steering
Listed -
US Congress introduces CREATE AI Act for establishing National AI Research Resource
Listed -
Visible loss landscape basins don't correspond to distinct algorithms
Listed -
Visit Mexico City in January & February to interact with the AI Futures Fellowship
Listed -
When can we trust model evaluations?
Listed -
Animal Advocacy in the Age of AI
Listed -
AXRP Episode 23 - Mechanistic Anomaly Detection with Mark Xu
Listed -
AXRP Episode 24 - Superalignment with Jan Leike
Listed -
Discussing AI-Human Collaboration Through Fiction: The Story of Laika and GPT-∞
Listed -
Partial Transcript of Recent Senate Hearing Discussing AI X-Risk
Listed -
Preference Aggregation as Bayesian Inference
Listed -
AGI Takeoff dynamics - Intelligence vs Quantity explosion
Listed -
Apply to CEEALAR to do AGI moratorium work
Listed -
EleutherAI's Thoughts on the EU AI Act
Listed - Listed
-
Existential risk from AI and what DC could do about it (Ezra Klein on the 80,000 Hours Podcast)
Listed - Listed
- Listed
-
"The Universe of Minds" - call for reviewers (Seeds of Science)
Listed -
[Linkpost] My attempt at trying to summarize 'Intro to ML Safety'
Listed -
AI Safety Hub Serbia Soft Launch
Listed -
AI Safety Hub Serbia Soft Launch
Listed - Listed
- Listed
-
How LLMs are and are not myopic
Listed -
Should you work at a leading AI lab? (including in non-safety roles)
Listed -
Summary of posts on XPT forecasts on AI risk and timelines
Listed -
Task decomposition for scalable oversight (AGISF Distillation)
Listed -
Towards evidence gap-maps for AI safety
Listed -
[Crosspost] An AI Pause Is Humanity's Best Bet For Preventing Extinction (TIME)
Listed -
[Crosspost] An AI Pause Is Humanity's Best Bet For Preventing Extinction (TIME)
Listed -
[link post] AI Should Be Terrified of Humans
Listed -
Asterisk Magazine Issue 03: AI
Listed -
Open problems in activation engineering
Listed -
Slowing down AI progress is an underexplored alignment strategy
Listed -
XPT forecasts on (some) biological anchors inputs
Listed -
My favorite AI governance research this year so far
Listed -
QAPR 5: grokking is maybe not *that* big a deal?
Listed -
Supplementary Alignment Insights Through a Highly Controlled Shutdown Incentive
Listed -
AI-Relevant Regulation: Insurance in Safety-Critical Industries
Listed -
Compute Thresholds: proposed rules to mitigate risk of a “lab leak” accident during AI training runs
Listed -
Could someone help me understand why it's so difficult to solve the alignment problem?
Listed -
Examples of Prompts that Make GPT-4 Output Falsehoods
Listed -
Australians call for AI safety to be taken seriously
Listed -
BCIs and the ecosystem of modular minds
Listed - Listed
-
GPT-2's positional embedding matrix is a helix
Listed -
Linkpost: 7 A.I. Companies Agree to Safeguards After Pressure From the White House
Listed - Listed
-
Priorities for the UK Foundation Models Taskforce
Listed -
Reward Hacking from a Causal Perspective
Listed - Listed
-
What do XPT forecasts tell us about AI timelines?
Listed -
All AGI Safety questions welcome (especially basic ones) [July 2023]
Listed - Listed
-
Epoch is hiring an ML Hardware Researcher
Listed -
Even Superhuman Go AIs Have Surprising Failure Modes
Listed -
Should we nationalize AI development?
Listed - Listed
-
The Dilemma of Ultimate Technology
Listed - Listed
-
Alignment Grantmaking is Funding-Limited Right Now
Listed -
All AGI Safety questions welcome (especially basic ones) [July 2023]
Listed -
An Introduction to Critiques of prominent AI safety organizations
Listed - Listed
- Listed
-
Incident reporting for AI safety
Listed -
Thoughts on yesterday’s UN Security Council meeting on AI
Listed -
Updates from Campaign for AI Safety
Listed -
Using predictors in corrigible systems
Listed -
What do XPT forecasts tell us about AI risk?
Listed -
AI Impacts Quarterly Newsletter, Apr-Jun 2023
Listed -
Five Years of Rethink Priorities: Impact, Future Plans, Funding Needs (July 2023)
Listed -
I'm interviewing Jan Leike, co-lead of OpenAI's new Superalignment project. What should I ask him?
Listed -
Measuring and Improving the Faithfulness of Model-Generated Reasoning
Listed -
Meta announces Llama 2; "open sources" it for commercial use
Listed -
Simple alignment plan that maybe works
Listed -
Still no Lie Detector for LLMs
Listed -
Tiny Mech Interp Projects: Emergent Positional Embeddings of Words
Listed -
Train for incorrigibility, then reverse it (Shutdown Problem Contest Submission)
Listed -
A fictional AI law laced w/ alignment theory
Listed -
A fictional AI law laced w/ alignment theory
Listed -
AutoInterpretation Finds Sparse Coding Beats Alternatives
Listed -
Developing reliable AI tools for healthcare
Listed -
Eliciting responses to Marc Andreessen's "Why AI Will Save the World"
Listed -
New career review: AI safety technical research
Listed -
The shape of AGI: Cartoons and back of envelope
Listed -
Thoughts on “Process-Based Supervision”
Listed -
What we can learn from stress testing for AI regulation
Listed -
A simple way of exploiting AI's coming economic impact may be highly-impactful
Listed -
Activation adding experiments with llama-7b
Listed -
An upcoming US Supreme Court case may impede AI governance efforts
Listed -
Embracing the automated future
Listed -
Even briefer summary of ai-plans.com
Listed -
Less activations can result in high corrigibility?
Listed -
Mech Interp Puzzle 1: Suspiciously Similar Embeddings in GPT-Neo
Listed -
Runaway Optimizers in Mind Space
Listed -
Scaling and Sustaining Standards: A Case Study on the Basel Accords
Listed - Listed
- Listed
-
Cambridge AI Safety Hub is looking for full- or part-time organisers
Listed -
Introducción al Riesgo Existencial de Inteligencia Artificial
Listed -
Only a hack can solve the shutdown problem
Listed -
Robustness of Model-Graded Evaluations and Automated Interpretability
Listed -
Simplified bio-anchors for upper bounds on AI timelines
Listed -
Why was the AI Alignment community so unprepared for this moment?
Listed -
AI Risk and Survivorship Bias - How Andreessen and LeCun got it wrong
Listed -
Gearing Up for Long Timelines in a Hard World
Listed -
New DeepMind report on institutions for global AI governance
Listed - Listed
-
Instrumental Convergence to Complexity Preservation
Listed - Listed
-
What criterion would you use to select companies likely to cause AI doom?
Listed -
What new psychology research could best promote AI safety & alignment research?
Listed -
Winners of AI Alignment Awards Research Contest
Listed -
[Linkpost] NY Times Feature on Anthropic
Listed -
A transcript of the TED talk by Eliezer Yudkowsky
Listed -
AISN#14: OpenAI’s ‘Superalignment’ team, Musk’s xAI launches, and developments in military AI use
Listed -
Alignment Megaprojects: You're Not Even Trying to Have Ideas
Listed -
An Overview of the AI Safety Funding Situation
Listed -
Announcing the AI Fables Writing Contest!
Listed - Listed
-
Could unions be an underrated driver for AI safety policy?
Listed -
Eric Michaud on the Quantization Model of Neural Scaling, Interpretability and Grokking
Listed -
Goal-Direction for Simulated Agents
Listed -
How I Learned To Stop Worrying And Love The Shoggoth
Listed -
Report on modeling evidential cooperation in large worlds
Listed - Listed
-
Towards Developmental Interpretability
Listed -
What does the launch of x.ai mean for AI Safety?
Listed -
(How) Is technical AI Safety research being evaluated?
Listed - Listed
-
Disincentivizing deception in mesa optimizers with Model Tampering
Listed -
How to regulate cutting-edge AI models (Markus Anderljung on The 80,000 Hours Podcast)
Listed -
OpenAI Launches Superalignment Taskforce
Listed -
What is the most convincing article, video, etc. making the case that AI is an X-Risk
Listed -
Arguments against existential risk from AI, part 2
Listed -
Consciousness as a conflationary alliance term
Listed -
Consider Joining the UK Foundation Model Taskforce
Listed -
Cost-effectiveness of professional field-building programs for AI safety research
Listed -
Cost-effectiveness of student programs for AI safety research
Listed -
Do you think the probability of future AI sentience(suffering) is >0.1%? Why?
Listed -
GPT-7: The Tale of the Big Computer (An Experimental Story)
Listed -
Import AI 334: Better distillation; the UK's AI taskforce; money and AI
Listed -
Incentives from a causal perspective
Listed -
Infographics report risk management of Artificial Intelligence in Spain
Listed -
Is the Endowment Effect Due to Incomparability?
Listed -
Modeling the impact of AI safety field-building programs
Listed - Listed
- Listed
-
“Reframing Superintelligence” + LLMs + 4 years
Listed - Listed
-
Some basics of the hypercompetence theory of government
Listed -
"Concepts of Agency in Biology" (Okasha, 2023) - Brief Paper Summary
Listed -
Announcing AI Alignment workshop at the ALIFE 2023 conference
Listed -
Announcing AI Alignment workshop at the ALIFE 2023 conference
Listed -
Constructive Discussion and Thinking Methodology for Severe Situations including Existential Risks
Listed -
Continuous Adversarial Quality Assurance: Extending RLHF and Constitutional AI
Listed -
Minetester: A fully open RL environment built on Minetest
Listed -
Really Strong Features Found in Residual Stream
Listed -
Seven Strategies for Tackling the Hard Part of the Alignment Problem
Listed -
Views on when AGI comes and on strategy to reduce existential risk
Listed -
Views on when AGI comes and on strategy to reduce existential risk
Listed -
What Does LessWrong/EA Think of Human Intelligence Augmentation as of mid-2023?
Listed -
What is everyone doing in AI governance
Listed -
Will the vast majority of technological progress happen in the longterm future?
Listed -
Announcing the Existential InfoSec Forum
Listed - Listed
-
Internal independent review for language model agent alignment
Listed - Listed
-
A Defense of Work on Mathematical AI Safety
Listed -
Concrete open problems in mechanistic interpretability: a technical overview
Listed -
Empirical Evidence Against "The Longest Training Run"
Listed -
Frontier AI regulation: Managing emerging risks to public safety
Listed -
Jesse Hoogland on Developmental Interpretability and Singular Learning Theory
Listed -
Localizing goal misgeneralization in a maze-solving policy network
Listed - Listed
-
€200k in European AI & Society Fund grants
Listed -
(tentatively) Found 600+ Monosemantic Features in a Small LM Using Sparse Autoencoders
Listed -
[Linkpost] Introducing Superalignment
Listed - Listed
- Listed
-
Exploring Functional Decision Theory (FDT) and a modified version (ModFDT)
Listed -
Know a grad student studying AI's economic impacts?
Listed -
OpenAI is starting a new "Superintelligence alignment" team and they're hiring
Listed -
Optimized for Something other than Winning or: How Cricket Resists Moloch and Goodhart's Law
Listed -
Washington Post article about EA university groups
Listed -
What did AI Safety’s specific funding of AGI R&D labs lead to?
Listed -
[linkpost] Ten Levels of AI Alignment Difficulty
Listed -
AI labs' statements on governance
Listed -
Animal Weapons: Lessons for Humans in the Age of X-Risk
Listed -
The Neural Net Tank Urban Legend
Listed -
Ways I Expect AI Regulation To Increase Extinction Risk
Listed -
[Job] Managing Director at the Cooperative AI Foundation ($5000 Referral Bonus)
Listed -
Douglas Hoftstadter concerned about AI xrisk
Listed - Listed
-
Ten Levels of AI Alignment Difficulty
Listed -
(Intro/1) - My Understandings of Mechanistic Interpretability Notebook
Listed -
Apply to fall policy internships (we can help)
Listed - Listed
-
Quantitative cruxes in Alignment
Listed -
Sources of evidence in Alignment
Listed -
Using (Uninterpretable) LLMs to Generate Interpretable AI Code
Listed - Listed
-
Elements of Computational Philosophy, Vol. I: Truth
Listed -
Agency from a causal perspective
Listed - Listed
-
George Hotz on AI safety: ~"centralized power is bad"
Listed -
Inherently Interpretable Architectures
Listed -
Introducing EffiSciences’ AI Safety Unit
Listed -
Introducing EffiSciences’ AI Safety Unit
Listed - Listed
-
Little attention seems to be on discouraging hardware progress
Listed - Listed
- Listed
-
Three camps in AI x-risk discussions: My personal very oversimplified overview
Listed -
AI Safety without Alignment: How humans can WIN against AI
Listed -
Anthropically Blind: the anthropic shadow is reflectively inconsistent
Listed -
Biosafety Regulations (BMBL) and their relevance for AI
Listed -
Biosafety Regulations (BMBL) and their relevance for AI
Listed -
Challenge proposal: smallest possible self-hardening backdoor for RLHF
Listed - Listed
-
Updates from Campaign for AI Safety
Listed - Listed
-
A "weak" AGI may attempt an unlikely-to-succeed takeover
Listed -
AGI x Animal Welfare: A High-EV Outreach Opportunity?
Listed -
AI & Drug Discovery - Security and Risks
Listed - Listed
- Listed
- Listed
-
Carl Shulman on The Lunar Society (7 hour, two-part podcast)
Listed -
My research agenda in agent foundations
Listed