The catalog, page 17
Records 4,001 to 4,250 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Causal confusion as an argument against the scaling hypothesis
Listed -
Digital Sentience Requires Solving the Boundary Problem
Listed -
How to become more agentic, by GPT-EA-Forum-v1
Listed -
Key Papers in Language Model Safety
Listed -
On corrigibility and its basin
Listed - Listed
-
[linkpost] Christiano on agreement/disagreement with Yudkowsky's "List of Lethalities"
Listed -
Let's See You Write That Corrigibility Tag
Listed -
Modeling Transformative AI Risks (MTAIR) Project -- Summary Report
Listed -
On Deference and Yudkowsky's AI Risk Estimates
Listed -
Where I agree and disagree with Eliezer
Listed - Listed
-
Do yourself a FAVAR: security mindset
Listed -
Scott Aaronson is joining OpenAI to work on AI safety
Listed -
‘Force multipliers’ for EA research
Listed - Listed
-
anthropic reasoning coordination
Listed -
Apply to the Machine Learning For Good bootcamp in France
Listed -
Evolution through large models
Listed -
generalized computation interpretability
Listed -
Pivotal outcomes and pivotal processes
Listed -
Pivotal outcomes and pivotal processes
Listed -
Quantifying General Intelligence
Listed -
solonomonoff induction, time penalty, the universal program, and deism
Listed -
The Unified Theory of Normative Ethics
Listed -
Unlocking High-Accuracy Differentially Private Image Classification through Scale
Listed -
Value extrapolation vs Wireheading
Listed - Listed
-
A transparency and interpretability tech tree
Listed -
Breaking Down Goal-Directed Behaviour
Listed -
Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models
Listed -
Humans are very reliable agents
Listed -
Interaction-Grounded Learning with Action-inclusive Feedback
Listed -
Refer the Cooperative AI Foundation’s New COO, Receive $5000
Listed -
Ten experiments in modularity, which we'd like you to run!
Listed -
Towards Gears-Level Understanding of Agency
Listed -
A central AI alignment problem: capabilities generalization, and the sharp left turn
Listed -
Alignment Risk Doesn't Require Superintelligence
Listed - Listed
-
What are all the AI Alignment and AI Safety Communication Hubs?
Listed -
Blake Richards on Why he is Skeptical of Existential Risk from AI
Listed -
Catholic theologians and priests on artificial intelligence
Listed -
Expected impact of a career in AI safety under different opinions
Listed -
Investigating causal understanding in LLMs
Listed -
Resources I send to AI researchers about AI safety
Listed -
Slow motion videos as AI risk intuition pumps
Listed -
Steering AI to care for animals, and soon
Listed -
Vael Gates: Risks from Advanced AI (June 2022)
Listed -
AI-written critiques help humans notice flaws
Listed - Listed
-
Contra EY: Can AGI destroy us without trial & error?
Listed -
Towards Autonomous Grading In The Real World
Listed - Listed
- Listed
-
EA AI/Emerging Tech Orgs Should Be Involved with Patent Office Partnership
Listed -
Grokking “Semi-informative priors over AI timelines”
Listed -
Grokking “Semi-informative priors over AI timelines”
Listed -
Why all the fuss about recursive self-improvement?
Listed -
AGI Ruin: A List of Lethalities
Listed -
AGI Safety Communications Initiative
Listed -
ELK Proposal - Make the Reporter care about the Predictor’s beliefs
Listed - Listed
-
Steganography and the CycleGAN - alignment failure case study
Listed -
[linkpost] The final AI benchmark: BIG-bench
Listed -
AI Could Defeat All Of Us Combined
Listed -
AI Twitter accounts to follow?
Listed -
Could Patent-Trolling delay AI timelines?
Listed -
Digital people could make AI safer
Listed -
Open Problems in AI X-Risk [PAIS #5]
Listed -
Open Problems in AI X-Risk [PAIS #5]
Listed -
Putting GPT-3's Creativity to the (Alternative Uses) Test
Listed -
why assume AGIs will optimize for fixed goals?
Listed - Listed
-
AI Could Defeat All Of Us Combined
Listed - Listed
-
How Do Selection Theorems Relate To Interpretability?
Listed -
If no near-term alignment strategy, research should aim for the long-term
Listed -
If there was a millennium equivalent prize for AI alignment, what would the problems be?
Listed -
outer alignment: politics & philosophy
Listed -
Towards Safe Reinforcement Learning via Constraining Conditional Value-at-Risk
Listed -
where are your alignment bits?
Listed -
You Only Get One Shot: an Intuition Pump for Embedded Agency
Listed -
Eliciting Latent Knowledge (ELK) - Distillation/Summary
Listed -
Is the time crunch for AI Safety Movement Building now?
Listed -
Research Questions from Stained Glass Windows
Listed -
Six Dimensions of Operational Adequacy in AGI Projects
Listed -
AGI Safety FAQ / all-dumb-questions-allowed thread
Listed -
CAISAR: A platform for Characterizing Artificial Intelligence Safety and Robustness
Listed -
Imitating Past Successes can be Very Suboptimal
Listed - Listed
-
Who models the models that model models? An exploration of GPT-3's in-context model fitting ability
Listed -
[Link] GCRI's Seth Baum reviews The Precipice
Listed -
A descriptive, not prescriptive, overview of current AI Alignment Research
Listed -
AGI Ruin: A List of Lethalities
Listed -
Epistemological Vigilance for Alignment
Listed -
Grokking “Forecasting TAI with biological anchors”
Listed -
Grokking “Forecasting TAI with biological anchors”
Listed -
Here are the finalists from FLI’s $100K Worldbuilding Contest
Listed -
Improving Model Understanding and Trust with Counterfactual Explanations of Model Confidence
Listed -
Reading the ethicists 2: Hunting for AI alignment papers
Listed -
Some ideas for follow-up projects to Redwood Research’s recent paper
Listed - Listed
-
AGI Ruin: A List of Lethalities
Listed -
New cooperation mechanism - quadratic funding without a matching pool
Listed -
Announcing the Alignment of Complex Systems Research Group
Listed -
Deep Learning Systems Are Not Less Interpretable Than Logic/Probability/Etc
Listed -
How to pursue a career in technical AI alignment
Listed -
How to pursue a career in technical AI alignment
Listed -
Towards a Formalisation of Returns on Cognitive Reinvestment (Part 1)
Listed -
Training a GPT model on EA texts: what data?
Listed - Listed
-
Data collection for AI alignment - Career review
Listed - Listed
-
I'm trying out "asteroid mindset"
Listed -
Intergenerational trauma impeding cooperative existential safety efforts
Listed - Listed
-
Adversarial training, importance sampling, and anti-adversarial training for AI whistleblowing
Listed -
Confused why a "capabilities research is good for alignment progress" position isn't discussed more
Listed -
Paradigms of AI alignment: components and enablers
Listed -
Paradigms of AI alignment: components and enablers
Listed -
Responsible/fair AI vs. beneficial/safe AI?
Listed - Listed
-
The prototypical catastrophic AI action is getting root access to its datacenter
Listed -
Appendix to Bridging Demonstration
Listed -
Contest: 250€ for translation of "longtermism" to German
Listed -
Elucidating the Design Space of Diffusion-Based Generative Models
Listed -
HYCEDIS: HYbrid Confidence Engine for Deep Document Intelligence System
Listed -
IDANI: Inference-time Domain Adaptation via Neuron-level Interventions
Listed -
Machines vs Memes Part 3: Imitation and Memes
Listed -
Advice on Pursuing Technical AI Safety Research
Listed -
Machines vs Memes Part 1: AI Alignment and Memetics
Listed -
Machines vs. Memes 2: Memetically-Motivated Model Extensions
Listed -
Paper: Teaching GPT3 to express uncertainty in words
Listed -
The Hard Intelligence Hypothesis and Its Bearing on Succession Induced Foom
Listed - Listed
-
Which possible AI impacts should receive the most additional attention?
Listed -
Multi-Game Decision Transformers
Listed -
Perform Tractable Research While Avoiding Capabilities Externalities [Pragmatic AI Safety #4]
Listed -
Perform Tractable Research While Avoiding Capabilities Externalities [Pragmatic AI Safety #4]
Listed - Listed
-
Six Dimensions of Operational Adequacy in AGI Projects
Listed -
Distilled - AGI Safety from First Principles
Listed - Listed
-
Multiple AIs in boxes, evaluating each other's alignment
Listed - Listed
-
The Problem With The Current State of AGI Definitions
Listed -
We should expect to worry more about speculative risks
Listed -
concentric rings of illiberalism
Listed -
Teaching models to express their uncertainty in words
Listed -
Understanding Selection Theorems
Listed -
Croesus, Cerberus, and the magpies: a gentle introduction to Eliciting Latent Knowledge
Listed -
Evaluating Multimodal Interactive Agents
Listed -
GALOIS: Boosting Deep Reinforcement Learning via Generalizable Logic Synthesis
Listed -
Infernal Corrigibility, Fiendishly Difficult
Listed - Listed
-
Personalized Algorithmic Recourse with Preference Elicitation
Listed - Listed
-
say "AI risk mitigation" not "alignment"
Listed -
Where Utopias Go Wrong, or: The Four Little Planets
Listed -
A Story of AI Risk: InstructGPT-N
Listed -
Dynamic language understanding: adaptation to new knowledge in parametric and semi-parametric models
Listed -
EA, Psychology & AI Safety Research
Listed -
How Could AI Governance Go Wrong?
Listed -
Infra-Bayesianism Distillation: Realizability and Decision Theory
Listed -
The pointers problem, distilled
Listed -
A Human-Centric Assessment Framework for AI
Listed -
autonomy: the missing AGI ingredient?
Listed -
RL with KL penalties is better seen as Bayesian inference
Listed -
The "Measuring Stick of Utility" Problem
Listed -
Complex Systems for AI Safety [Pragmatic AI Safety #3]
Listed -
Complex Systems for AI Safety [Pragmatic AI Safety #3]
Listed -
Explaining inner alignment to myself
Listed -
The No Free Lunch theorems and their Razor
Listed -
2022 Uehiro Lectures: Ethics and Artificial Intelligence
Listed -
AXRP Episode 15 - Natural Abstractions with John Wentworth
Listed -
Bits of Optimization Can Only Be Lost Over A Distance
Listed - Listed
-
RL with KL penalties is better viewed as Bayesian inference
Listed -
The Windfall Clause has a remedies problem
Listed - Listed
-
X-Risk Motivations for Safety Research Directions
Listed -
Adversarial attacks and optimal control
Listed -
implementing the platonic realm
Listed -
Responsible Artificial Intelligence -- from Principles to Practice
Listed -
SERI ML application deadline is extended until May 22.
Listed -
[Short version] Information Loss --> Basin flatness
Listed - Listed
-
Clarifying what ELK is trying to achieve
Listed -
Information Loss --> Basin flatness
Listed -
Scaling Laws and Interpretability of Learning from Repeated Data
Listed - Listed
-
How RL Agents Behave When Their Actions Are Modified? [Distillation post]
Listed -
Are you really in a race? The Cautionary Tales of Szilárd and Ellsberg
Listed -
[Fiction] Improved Governance on the Critical Path to AI Alignment by 2045.
Listed - Listed
-
A bridge to Dath Ilan? Improved governance on the critical path to AI alignment.
Listed -
Gato's Generalisation: Predictions and Experiments I'd Like to See
Listed -
generalized adding reality layers
Listed -
How to get into AI safety research
Listed -
Maxent and Abstractions: Current Best Arguments
Listed -
Mimicking Behaviors in Separated Domains
Listed -
predictablizing ethic deduplication
Listed -
We have achieved Noob Gains in AI
Listed -
[Intro to brain-like-AGI safety] 15. Conclusion: Open problems, how to help, AMA
Listed -
Actionable-guidance and roadmap recommendations for the NIST AI Risk Management Framework
Listed -
Actionable-guidance and roadmap recommendations for the NIST AI Risk Management Framework
Listed -
LW4EA: Some cruxes on impactful alternatives to AI policy work
Listed -
We Ran an AI Timelines Retreat
Listed -
“Intro to brain-like-AGI safety” series—just finished!
Listed -
AGI Risk: How to internationally regulate industries in non-democracies
Listed -
DeepMind’s generalist AI, Gato: A non-technical explainer
Listed - Listed
-
Emergent Bartering Behaviour in Multi-Agent Reinforcement Learning
Listed -
How Different Groups Prioritize Ethical Values for Responsible AI
Listed - Listed
-
Proxy misspecification and the capabilities vs. value learning race
Listed -
To what extent is your AGI timeline bimodal or otherwise "bumpy"?
Listed - Listed
- Listed
- Listed
-
What does the Project Management role look like in AI safety?
Listed -
"Tech company singularities", and steering them to reduce x-risk
Listed -
"Tech company singularities", and steering them to reduce x-risk
Listed - Listed
-
Agency As a Natural Abstraction
Listed - Listed
-
An observation about Hubinger et al.'s framework for learned optimization
Listed -
Clarifying the confusion around inner alignment
Listed -
cognitive biases regarding the evaluation of AI risk when doing AI capabilities work
Listed -
DeepMind is hiring for the Scalable Alignment and Alignment Teams
Listed -
Fermi estimation of the impact you might have working on AI safety
Listed -
Frame for Take-Off Speeds to inform compute governance & scaling alignment
Listed -
I'm interviewing Max Tegmark about AI safety and more. What shouId I ask him?
Listed - Listed
- Listed
-
A tentative dialogue with a Friendly-boxed-super-AGI on brain uploads
Listed -
Deepmind's Gato: Generalist Agent
Listed -
Interpretability’s Alignment-Solving Potential: Analysis of 7 Scenarios
Listed -
Introduction to the sequence: Interpretability Research for the Most Important Century
Listed - Listed
-
New series of posts answering one of Holden's "Important, actionable research questions"
Listed -
[Intro to brain-like-AGI safety] 14. Controlled AGI
Listed -
[Intro to brain-like-AGI safety] 14. Controlled AGI
Listed - Listed
- Listed
-
AI safety should be made more accessible using non text-based media
Listed - Listed
-
Rabbits, robots and resurrection
Listed -
The limits of AI safety via debate
Listed -
A Bird's Eye View of the ML Field [Pragmatic AI Safety #2]
Listed