The catalog, page 13
Records 3,001 to 3,250 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
A strange twist on the road to AGI
Listed - Listed
-
Article Review: Google's AlphaTensor
Listed -
Building a transformer from scratch - AI safety up-skilling challenge
Listed -
Instrumental convergence in single-agent systems
Listed - Listed
-
[Sketch] Validity Criterion for Logical Counterfactuals
Listed -
BenevolentAI - an effectively impactful company?
Listed -
Human-AI Coordination via Human-Regularized Search and Learning
Listed -
Power-Seeking AI and Existential Risk
Listed - Listed
- Listed
-
“Technological unemployment” AI vs. “most important century” AI: how far apart?
Listed -
Disentangling inner alignment failures
Listed -
Generating Executable Action Plans with Environmentally-Aware Language Models
Listed -
Lessons learned from talking to >100 academics about AI safety
Listed -
outer alignment: two failure modes and past-user satisfaction
Listed - Listed
-
When reporting AI timelines, be clear who you're deferring to
Listed - Listed
-
Good ontologies induce commutative diagrams
Listed -
Let’s talk about uncontrollable AI
Listed -
Uncontrollable AI as an Existential Risk
Listed -
Don't leave your fingerprints on the future
Listed -
Georgetown EA Fall 2022"Intro to AI" Reading Group
Listed -
Mutual Assured Destruction used against AGI
Listed -
SERI MATS Program - Winter 2022 Cohort
Listed -
Analysis: US restricts GPU sales to China
Listed - Listed
-
Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals
Listed -
How undesired goals can arise with correct rewards
Listed -
Knowledge-Grounded Reinforcement Learning
Listed -
More examples of goal misgeneralization
Listed -
Polysemanticity and Capacity in Neural Networks
Listed -
Public Explainer on AI as an Existential Risk
Listed -
What does it mean for an AGI to be 'safe'?
Listed -
A shot at the diamond-alignment problem
Listed -
AI Timelines via Cumulative Optimization Power: Less Long, More Short
Listed -
Analysing a 2036 Takeover Scenario
Listed -
confusion about alignment requirements
Listed -
More Recent Progress in the Theory of Neural Networks
Listed -
NAS-Bench-Suite-Zero: Accelerating Research on Zero Cost Proxies
Listed -
The probability that Artificial General Intelligence will be developed by 2043 is extremely low.
Listed -
Warning Shots Probably Wouldn't Change The Picture Much
Listed - Listed
-
confusion about alignment requirements
Listed -
Paper: Discovering novel algorithms with AlphaTensor [Deepmind]
Listed -
Reflection Mechanisms as an Alignment target: A follow-up survey
Listed -
Tracking Compute Stocks and Flows: Case Studies?
Listed -
What are the risks of an oracle AI?
Listed -
CHAI, Assistance Games, And Fully-Updated Deference [Scott Alexander]
Listed -
How are you dealing with ontology identification?
Listed -
Humans aren't fitness maximizers
Listed -
Paper+Summary: OMNIGROK: GROKKING BEYOND ALGORITHMIC DATA
Listed -
Polysemanticity and Capacity in Neural Networks
Listed - Listed
-
A review of the Bio-Anchors report
Listed -
Data for IRL: What is needed to learn human values?
Listed - Listed
-
my current outlook on AI risk mitigation
Listed -
Recall and Regurgitation in GPT2
Listed -
Tony Blair Institute - Compute for AI Index ( Seeking a Supplier)
Listed -
Against the weirdness heuristic
Listed -
Any further work on AI Safety Success Stories?
Listed -
Establishing Meta-Decision-Making for AI: An Ontology of Relevance, Representation and Reasoning
Listed - Listed
-
my current outlook on AI risk mitigation
Listed -
Paper: Large Language Models Can Self-improve [Linkpost]
Listed -
Questions on databases of AI Risk estimates
Listed -
Why does AGI occur almost nowhere, not even just as a remark for economic/political models?
Listed -
Announcing the AI Safety Nudge Competition to Help Beat Procrastination
Listed - Listed
-
Do anthropic considerations undercut the evolution anchor from the Bio Anchors report?
Listed -
Google could build a conscious AI in three months
Listed -
(Structural) Stability of Coupled Optimizers
Listed -
Carnegie Council MisUnderstands Longtermism
Listed -
EAG DC: Meta-Bottlenecks in Preventing AI Doom
Listed -
Eli's review of "Is power-seeking AI an existential risk?"
Listed -
Rethinking and Recomputing the Value of ML Models
Listed -
We all teach: here's how to do it better
Listed -
Builder/Breaker for Deconfusion
Listed -
Clarifying the Agent-Like Structure Problem
Listed -
Distribution Shifts and The Importance of AI Safety
Listed -
FDT is not directly comparable to CDT and EDT
Listed - Listed
-
It matters when the first sharp left turn happens
Listed - Listed
-
Repairing Bugs in Python Assignments Using Large Language Models
Listed -
Where I currently disagree with Ryan Greenblatt’s version of the ELK approach
Listed -
A Library and Tutorial for Factored Cognition with Language Models
Listed - Listed
-
Estimating the Current and Future Number of AI Safety Researchers
Listed -
How Open Source Machine Learning Software Shapes AI
Listed -
InFi: End-to-End Learning to Filter Input for Resource-Efficiency in Mobile-Centric Inference
Listed -
LOVE in a simbox is all you need
Listed -
Optimism, AI risk, and EA blind spots
Listed -
QAPR 3: interpretability-guided training of neural nets
Listed -
Strange Loops - Self-Reference from Number Theory to AI
Listed - Listed
-
Threat-Resistant Bargaining Megapost: Introducing the ROSE Value
Listed -
Why I think strong general AI is coming soon
Listed -
7 traps that (we think) new alignment researchers often fall into
Listed -
Collaborative Decision Making Using Action Suggestions
Listed -
Failure modes in a shard theory alignment plan
Listed -
Learning When to Advise Human Decision Makers
Listed -
Likelihood of an anti-AI backlash: Results from a preliminary Twitter poll
Listed -
My Thoughts on the ML Safety Course
Listed -
Why we're not founding a human-data-for-alignment org
Listed - Listed
- Listed
-
existential self-determination
Listed -
Inverse Scaling Prize: Round 1 Winners
Listed -
Lessons from Three Mile Island for AI Warning Shots
Listed - Listed
- Listed
-
Project Idea: The cost of Coccidiosis on Chicken farming and if AI can help
Listed -
Stress Externalities More in AI Safety Pitches
Listed -
surprise! you want what you want
Listed -
Understanding Hindsight Goal Relabeling from a Divergence Minimization Perspective
Listed -
You are Underestimating The Likelihood That Convergent Instrumental Subgoals Lead to Aligned AGI
Listed -
An Unexpected GPT-3 Decision in a Simple Gamble
Listed -
AI Risk Intro 2: Solving The Problem
Listed -
Brain-over-body biases, and the embodied value problem in AI alignment
Listed -
Papers to start getting into NLP-focused alignment research
Listed -
Two reasons we might be closer to solving alignment than it seems
Listed -
Two reasons we might be closer to solving alignment than it seems
Listed -
7 Learnings and a Detailed Description of an AI Safety Reading Group
Listed -
Announcing the Future Fund's AI Worldview Prize
Listed -
Interlude: But Who Optimizes The Optimizer?
Listed -
Interpreting Neural Networks through the Polytope Lens
Listed -
Shahar Avin On How To Regulate Advanced AI Systems
Listed -
Shahar Avin on How to Strategically Regulate Advanced AI Systems
Listed -
The Rival AI Deployment Problem: a Pre-deployment Agreement as the least-bad response
Listed -
Under what circumstances have governments cancelled AI-type systems?
Listed -
What are people's thoughts on working for DeepMind as a general software engineer?
Listed -
(My suggestions) On Beginner Steps in AI Alignment
Listed -
[Cause Exploration Prizes] Expanding communication about AGI risks
Listed -
AGI Battle Royale: Why “slow takeover” scenarios devolve into a chaotic multi-AGI fight to the death
Listed -
AI Risk Intro 2: Solving The Problem
Listed -
Crypto 'oracle protocols' for AI alignment with real-world data?
Listed -
Dath Ilan's Views on Stopgap Corrigibility
Listed -
Initial Thoughts on Dissolving "Couldness"
Listed -
Mathematical Circuits in Neural Networks
Listed -
Methodological Therapy: An Agenda For Tackling Research Bottlenecks
Listed -
Understanding Infra-Bayesianism: A Beginner-Friendly Video Series
Listed -
An issue with MacAskill's Evidentialist's Wager
Listed -
Announcing AISIC 2022 - the AI Safety Israel Conference, October 19-20
Listed -
EA’s brain-over-body bias, and the embodied value problem in AI alignment
Listed -
Establishing Oxford’s AI Safety Student Group: Lessons Learnt and Our Model
Listed - Listed
-
LCRL: Certified Policy Synthesis via Logically-Constrained Reinforcement Learning
Listed -
Nearcast-based "deployment problem" analysis
Listed -
Towards deconfusing wireheading and reward maximization
Listed - Listed
- Listed
- Listed
-
Doing oversight from the very start of training seems hard
Listed -
I'm Interviewing Kat Woods, EA Powerhouse. What Should I Ask?
Listed -
What Do AI Safety Pitches Not Get About Your Field?
Listed -
Why AGIs utility can't outweigh humans' utility?
Listed -
PIBBSS (AI alignment) is hiring for a Project Manager
Listed -
Quintin's alignment papers roundup - week 2
Listed -
Safety timelines: How long will it take to solve alignment?
Listed -
Safety timelines: How long will it take to solve alignment?
Listed -
Summaries: Alignment Fundamentals Curriculum
Listed - Listed
-
Updates on FLI'S Value Alignment Map?
Listed -
Aligning AI with Humans by Leveraging Legal Informatics
Listed -
Inner alignment: what are we pointing at?
Listed -
Leveraging Legal Informatics to Align AI
Listed -
Prize and fast track to alignment research at ALTER
Listed -
Summaries: Alignment Fundamentals Curriculum
Listed -
The Inter-Agent Facet of AI Alignment
Listed -
A Bite Sized Introduction to ELK
Listed -
Apply for mentorship in AI Safety field-building
Listed -
Prize and fast track to alignment research at ALTER
Listed -
Refine's Third Blog Post Day/Week
Listed -
Sparse trinary weighted RNNs as a path to better language model interpretability
Listed -
Takeaways from our robust injury classifier project [Redwood Research]
Listed -
[linkpost] When does technical work to reduce AGI conflict make a difference?: Introduction
Listed -
Katja Grace on Slowing Down AI, AI Expert Surveys And Estimating AI Risk
Listed - Listed
-
ordering capability thresholds
Listed -
Refine Blogpost Day #3: The shortforms I did write
Listed -
Representational Tethers: Tying AI Latents To Human Ones
Listed -
The heterogeneity of human value types: Implications for AI alignment
Listed -
The Pugwash Conferences and the Anti-Ballistic Missile Treaty as a case study of Track II diplomacy
Listed -
The religion problem in AI alignment
Listed -
'Artificial Intelligence Governance under Change' (PhD dissertation)
Listed - Listed
-
Black Box Investigations Research Hackathon
Listed -
Capability and Agency as Cornerstones of AI risk — My current model
Listed -
FDT defects in a realistic Twin Prisoners' Dilemma
Listed -
General advice for transitioning into Theoretical AI Safety
Listed -
How should DeepMind's Chinchilla revise our AI forecasts?
Listed -
ordering capability thresholds
Listed -
Why deceptive alignment matters for AGI safety
Listed -
Are Speed Superintelligences Feasible for Modern ML Techniques?
Listed - Listed
-
Coordinate-Free Interpretability Theory
Listed -
Emily Brontë on: Psychology Required for Serious™ AGI Safety Research
Listed -
Forecasting thread: How does AI risk level vary based on timelines?
Listed -
Future Matters #5: supervolcanoes, AI takeover, and What We Owe the Future
Listed -
Roodman's Thoughts on Biological Anchors
Listed -
Some ideas for epistles to the AI ethicists
Listed -
The Defender’s Advantage of Interpretability
Listed -
When does technical work to reduce AGI conflict make a difference?: Introduction
Listed -
When is intent alignment sufficient or necessary to reduce AGI conflict?
Listed -
When would AGIs engage in conflict?
Listed -
Why Do People Think Humans Are Stupid?
Listed -
Would a Misaligned SSI Really Kill Us All?
Listed -
Announcing an Empirical AI Safety Program
Listed -
Improving Language Model Prompting in Support of Semi-autonomous Task Learning
Listed -
New tool for exploring EA Forum, LessWrong and Alignment Forum - Tree of Tags
Listed -
Trying to find the underlying structure of computational systems
Listed -
[Linkpost] A survey on over 300 works about interpretability in deep networks
Listed -
Alignment via prosocial brain algorithms
Listed -
An experiment eliciting relative estimates for Open Philanthropy’s 2018 AI safety grants
Listed -
Differential technology development: preprint on the concept
Listed -
EA & LW Forums Weekly Summary (5 - 11 Sep 22’)
Listed -
Ideological Inference Engines: Making Deontology Differentiable*
Listed -
Resource Allocation to Agents with Restrictions: Maximizing Likelihood with Minimum Compromise
Listed -
What could an AI-caused existential catastrophe actually look like?
Listed -
Why do People Think Intelligence Will be "Easy"?
Listed -
AI Risk Intro 1: Advanced AI Might Be Very Bad
Listed -
AI Risk Intro 1: Advanced AI Might Be Very Bad
Listed -
AI Safety field-building projects I'd like to see
Listed -
Briefly thinking through some analogs of debate
Listed -
Join ASAP (AI Safety Accountability Programme) 🚀
Listed -
Path dependence in ML inductive biases
Listed -
Quintin's alignment papers roundup - week 1
Listed -
Unbounded utility functions and precommitment
Listed -
A California Effect for Artificial Intelligence
Listed -
Evaluations project @ ARC is hiring a researcher and a webdev/engineer
Listed -
Gatekeeper Victory: AI Box Reflection
Listed -
Markus Anderljung On The AI Policy Landscape
Listed -
Most People Start With The Same Few Bad Ideas
Listed -
Ought will host a factored cognition “Lab Meeting”
Listed -
Oversight Leagues: The Training Game as a Feature
Listed -
Samotsvety's AI risk forecasts
Listed - Listed
-
Understanding and avoiding value drift
Listed - Listed
- Listed
-
AI alignment with humans... but with which humans?
Listed -
All AGI safety questions welcome (especially basic ones) [Sept 2022]
Listed -
ethics and anthropics of homomorphically encrypted computations
Listed -
Linkpost: Github Copilot productivity experiment
Listed -
Monitoring for deceptive alignment
Listed -
Searching for Modularity in Large Language Models
Listed