The catalog, page 6
Records 1,251 to 1,500 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Is Deontological AI Safe? [Feedback Draft]
Listed -
Project Idea: Challenge Groups for Alignment Researchers
Listed -
[Job Ad] SERI MATS is hiring for our summer program
Listed - Listed
-
Bandgaps, Brains, and Bioweapons: The limitations of computational science and what it means for AGI
Listed -
Before smart AI, there will be many mediocre or specialized AIs
Listed - Listed
-
Conditional Prediction with Zero-Sum Training Solves Self-Fulfilling Prophecies
Listed - Listed
-
Some thoughts on automating alignment research
Listed - Listed
-
[Linkpost] OpenAI leaders call for regulation of "superintelligence" to reduce existential risk.
Listed -
An early warning system for novel AI risks
Listed -
DeepMind: Model evaluation for extreme risks
Listed -
Exploiting Newcomb's Game Show
Listed -
Is behavioral safety "solved" in non-adversarial conditions?
Listed -
Requirements for a STEM-capable AGI Value Learner (my Case for Less Doom)
Listed -
Solving the Mechanistic Interpretability challenges: EIS VII Challenge 2
Listed -
The Genie in the Bottle: An Introduction to AI Alignment and Risk
Listed -
Two ideas for alignment, perpetual mutual distrust and induction
Listed -
Will AI end everything? A guide to guessing | EAG Bay Area 23
Listed -
[Linkpost] Interpretability Dreams
Listed -
AGI Catastrophe and Takeover: Some Reference Class-Based Priors
Listed -
Aligned AI via monitoring objectives in AutoGPT-like systems
Listed -
Diagram with Commentary for AGI as an X-Risk
Listed -
New s-risks audiobook available now
Listed -
October 2022 AI Risk Community Survey Results
Listed -
Rishi Sunak mentions "existential threats" in talk with OpenAI, DeepMind, Anthropic CEOs
Listed -
What projects and efforts are there to promote AI safety research?
Listed -
'Fundamental' vs 'applied' mechanistic interpretability research
Listed -
[Linkpost] The AGI Show podcast
Listed -
A Different Approach to Community Building: The Spiral Path to Impact
Listed - Listed
- Listed
-
AI self-improvement is possible
Listed -
Data and "tokens" a 30 year old human "trains" on
Listed -
How I learned to stop worrying and love skill trees
Listed -
How I learned to stop worrying and love skill trees
Listed -
Is "brittle alignment" good enough?
Listed - Listed
- Listed
-
Will Artificial Superintelligence Kill Us?
Listed -
đ¶Safetensors audited as really safe and becoming the default
Listed -
[Linkpost] "Governance of superintelligence" by OpenAI
Listed -
Activation additions in a small residual network
Listed - Listed
- Listed
-
Conjecture internal survey: AGI timelines and probability of human extinction from advanced AI
Listed -
Distillation of Neurotech and Alignment Workshop January 2023
Listed - Listed
- Listed
-
Import AI 330: Palantir's AI-War future; BLOOMChat; and more money for distributed AI training
Listed -
When will digital compute match the human brain?
Listed -
Former Israeli Prime Minister Speaks About AI X-Risk
Listed -
Effective Altruism Florida's AI Expert Panel - Recording and Slides Available
Listed -
G7 SummitâCooperation on AI Policy
Listed -
Mr. Meeseeks as an AI capability tripwire
Listed - Listed
-
âThe Race to the End of Humanityâ â Structural Uncertainty Analysis in AI Risk Models
Listed -
A recent write-up of the case for AI (existential) risk
Listed -
AI #12:The Quest for Sane Regulations
Listed -
Asking for online resources why AI now is near AGI
Listed - Listed
-
Some background for reasoning about dual-use alignment research
Listed - Listed
-
Thread: Reflections on the AGI Safety Fundamentals course?
Listed -
We Shouldn't Expect AI to Ever be Fully Rational
Listed -
AI Alignment in The New Yorker
Listed -
Creating a self-referential system prompt for GPT-4
Listed -
Eisenhower's Atoms for Peace Speech
Listed -
GPT-4 implicitly values identity preservation: a study of LMCA identity management
Listed -
Lessons on project management from âHow Big Things Get Doneâ
Listed -
Letâs use AI to harden human defenses against AI manipulation
Listed -
Play Regrantor: Move up to $250,000 to Your Top High-Impact Projects!
Listed -
Some quotes from Tuesday's Senate hearing on AI
Listed -
Why AGI systems will not be fanatical maximisers (unless trained by fanatical humans)
Listed -
$500 Bounty/Prize Problem: Channel Capacity Using "Insensitive" Functions
Listed -
A Mechanistic Interpretability Analysis of a GridWorld Agent-Simulator (Part 1 of N)
Listed -
AI Risk & Policy Forecasts from Metaculus & FLI's AI Pathways Workshop
Listed -
AI Risk & Policy Forecasts from Metaculus & FLI's AI Pathways Workshop
Listed - Listed
- Listed
-
AI Will Not Want to Self-Improve
Listed -
Decision Theory with the Magic Parts Highlighted
Listed -
Evaluating Language Model Behaviours for Shutdown Avoidance in Textual Scenarios
Listed -
My current workflow to study the internal mechanisms of LLM
Listed -
OpenAI CEO Sam Altman testifies during Senate hearing on AI oversight â 05/16/23
Listed -
Oversight of A.I.: Rules for Artificial Intelligence
Listed -
Proposal: we should start referring to the risk from unaligned AI as a type of *accident risk*
Listed - Listed
-
Accidentally teaching AI models to deceive us (Ajeya Cotra on The 80,000 Hours Podcast)
Listed -
AI policy & governance in Australia: notes from an initial discussion
Listed -
Can we learn much by studying the behaviour of RL policies?
Listed -
Catastrophic Regressional Goodhart: Appendix
Listed -
EA and AI Safety Schism: AGI, the last tech humans will (soon*) build
Listed -
GovAI: Towards best practices in AGI safety and governance: A survey of expert opinion
Listed -
Import AI 329: Compute IS data; don't build AI agents; AI needs a precautionary principle
Listed -
Reward is the optimization target (of capabilities researchers)
Listed -
Simple experiments with deceptive alignment
Listed -
Some Summaries of Agent Foundations Work
Listed -
The Lightcone Theorem: A Better Foundation For Natural Abstraction?
Listed -
Un-unpluggability - can't we just unplug it?
Listed -
Why don't quantilizers also cut off the upper end of the distribution?
Listed -
A strong mind continues its trajectory of creativity
Listed -
Asking for online calls on AI s-risks discussions
Listed -
CEA Should Invest in Helping Altruists Navigate Advanced AI
Listed -
Difficulties in making powerful aligned AI
Listed -
How much do markets value Open AI?
Listed -
Simpler explanations of AGI risk
Listed - Listed
- Listed
-
Are there enough opportunities for AI safety specialists?
Listed - Listed
-
PCAST Working Group on Generative AI Invites Public Input
Listed -
Steering GPT-2-XL by adding an activation vector
Listed -
Aggregating Utilities for Corrigible AI [Feedback Draft]
Listed -
Aggregating Utilities for Corrigible AI [Feedback Draft]
Listed -
Infinite-width MLPs as an "ensemble prior"
Listed -
Input Swap Graphs: Discovering the role of neural network components at scale
Listed -
Towards Measures of Optimisation
Listed -
Turning off lights with model editing
Listed -
US public opinion of AI policy and risk
Listed - Listed
-
A more grounded idea of AI risk
Listed -
A request to keep pessimistic AI posts actionable.
Listed -
Alignment, Goals, & The Gut-Head Gap: A Review of Ngo. et al
Listed -
Is Infra-Bayesianism Applicable to Value Learning?
Listed -
Notes on the importance and implementation of safety-first cognitive architectures for AI
Listed -
A Corrigibility Metaphore - Big Gambles
Listed -
AGI-Automated Interpretability is Suicide
Listed -
AI interpretability could be harmful?
Listed -
Continuous doesnât mean slow
Listed -
Crises Reveal Centralisation (Stefan Schubert)
Listed -
How much of a concern are open-source LLMs in the short, medium and long terms?
Listed -
New OpenAI Paper - Language models can explain neurons in language models
Listed -
Roadmap for a collaborative prototype of an Open Agency Architecture
Listed -
You don't need to be a genius to be in AI safety research
Listed -
A note of caution on believing things on a gut level
Listed -
A Search for More ChatGPT / GPT-3.5 / GPT-4 "Unspeakable" Glitch Tokens
Listed - Listed
- Listed
-
Announcing âKey Phenomena in AI Riskâ (facilitated reading group)
Listed -
Announcing âKey Phenomena in AI Riskâ (facilitated reading group)
Listed -
Chilean AIS Hackathon Retrospective
Listed -
Language models can explain neurons in language models
Listed -
Result Of The Bounty/Contest To Explain Infra-Bayes In The Language Of Game Theory
Listed -
Solving the Mechanistic Interpretability challenges: EIS VII Challenge 1
Listed -
Stampy's AI Safety Info - New Distillations #2 [April 2023]
Listed -
Stopping dangerous AI: Ideal lab behavior
Listed -
Stopping dangerous AI: Ideal US behavior
Listed -
When is Goodhart catastrophic?
Listed -
Why "just make an agent which cares only about binary rewards" doesn't work.
Listed -
A technical note on bilinear layers for interpretability
Listed -
Acausal trade naturally results in the Nash bargaining solution
Listed -
All AGI Safety questions welcome (especially basic ones) [May 2023]
Listed -
Annotated reply to Bengio's "AI Scientists: Safe and Useful AI?"
Listed -
H-JEPA might be technically alignable in a modified form
Listed -
How "AGI" could end up being many different specialized AI's stitched together
Listed -
How quickly AI could transform the world (Tom Davidson on The 80,000 Hours Podcast)
Listed -
How The EthiSizer Almost Broke `Story'
Listed -
Import AI 328: Cheaper StableDiffusion; sim2soccer; AI refinement
Listed -
Inference Speed is Not Unbounded
Listed -
Is EDT correct? Does "EDT" == "logical EDT" == "logical CDT"?
Listed - Listed
-
Predictable updating about AI risk
Listed -
Reminder: AI Worldviews Contest Closes May 31
Listed - Listed
-
What does it take to ban a thing?
Listed -
Against sacrificing AI transparency for generality gains
Listed - Listed
-
An artificially structured argument for expecting AGI ruin
Listed -
Corrigibility, Much more detail than anyone wants to Read
Listed -
Graphical Representations of Paul Christiano's Doom Model
Listed - Listed
-
On the Loebner Silver Prize (a Turing test)
Listed -
Residual stream norms grow exponentially over the forward pass
Listed - Listed
-
How much do you believe your results?
Listed -
Is "red" for GPT-4 the same as "red" for you?
Listed -
My preferred framings for reward misspecification and goal misgeneralisation
Listed -
Rank best universities for AI Saftey
Listed -
An Update On The Campaign For AI Safety Dot Org
Listed -
Intro to ML Safety virtual program: 12 June - 14 August
Listed -
Introducing the AI Objectives Institute's Research: Differential Paths toward Safe and Beneficial AI
Listed - Listed
-
Orthogonal's Formal-Goal Alignment theory of change
Listed -
Regulate or Compete? The China Factor in U.S. AI Policy (NAIR #2)
Listed -
Transcript of a presentation on catastrophic risks from AI
Listed -
[Link Post: New York Times] White House Unveils Initiatives to Reduce Risks of A.I.
Listed -
AI risk/reward: A simple model
Listed - Listed
- Listed
-
Most Leading AI Experts Believe That Advanced AI Could Be Extremely Dangerous to Humanity
Listed -
Trying to measure AI deception capabilities using temporary simulation fine-tuning
Listed -
White House Announces "New Actions to Promote Responsible AI Innovation"
Listed -
Alignment Research @ EleutherAI
Listed -
Finding Neurons in a Haystack: Case Studies with Sparse Probing
Listed -
How CISA can Support the Security of Large AI Models Against Theft [Grad School Assignment]
Listed -
How much do personal biases in risk assessment affect assessment of AI risks?
Listed -
My choice of AI misalignment introduction for a general audience
Listed -
Prizes for matrix completion problems
Listed -
«Boundaries/Membranes» and AI safety compilation
Listed -
A Case for the Least Forgiving Take On Alignment
Listed - Listed
- Listed
- Listed
- Listed
-
An Impossibility Proof Relevant to the Shutdown Problem and Corrigibility
Listed -
Avoiding xrisk from AI doesn't mean focusing on AI xrisk
Listed -
AXRP Episode 21 - Interpretability for Engineers with Stephen Casper
Listed -
Finding Neurons in a Haystack: Case Studies with Sparse Probing
Listed -
Owain Evans on LLMs, Truthful AI, AI Composition, and More
Listed -
P(doom|AGI) is high: why the default outcome of AGI is doom
Listed -
Simulating a possible alignment solution in GPT2-medium using Archetypal Transfer Learning
Listed -
Systems that cannot be unsafe cannot be safe
Listed -
[Linkpost] âThe Godfather of A.I.â Leaves Google and Warns of Danger Ahead
Listed -
Call for Pythia-style foundation model suite for alignment research
Listed - Listed
-
Exploring Metaculusâs AI Track Record
Listed - Listed
-
Import AI 327: Stable Diffusion on phones; GPT-Hacker; UK launches a ÂŁ100m AI taskforce
Listed -
List of AI safety newsletters and other resources
Listed -
My current take on existential AI risk [FB post]
Listed -
Retrospective on recent activity of Riesgos CatastrĂłficos Globales
Listed -
Safety standards: a framework for AI regulation
Listed -
Shah (DeepMind) and Leahy (Conjecture) Discuss Alignment Cruxes
Listed - Listed
- Listed
-
A small update to the Sparse Coding interim research report
Listed -
Call for submissions: Choice of Futures survey questions
Listed -
Career uncertainty: Medicine vs. AI
Listed -
Connectomics seems great from an AI x-risk perspective
Listed -
Discussion about AI Safety funding (FB transcript)
Listed - Listed
- Listed
-
[SEE EDIT] No, *You* Need to Write Clearer
Listed -
A Guide to Forecasting AI Science Capabilities
Listed -
A Guide to Forecasting AI Science Capabilities
Listed -
Research agenda: Supervising AIs improving AIs
Listed -
AI safety logo design contest, due end of May (extended)
Listed -
New open letter on AI â "Include Consciousness Research"
Listed - Listed
-
Towards Automated Circuit Discovery for Mechanistic Interpretability
Listed -
AI doom from an LLM-plateau-ist perspective
Listed - Listed
-
Infrafunctions and Robust Optimization
Listed -
Proposals for the AI Regulatory Sandbox in Spain
Listed -
The AI guide I'm sending my grandparents
Listed -
What are the limits of superintelligence?
Listed -
A simple presentation of AI risk arguments
Listed