The catalog, page 24
Records 5,751 to 6,000 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Humanities Research Ideas for Longtermists
Listed - Listed
-
AXRP Episode 8 - Assistance Games with Dylan Hadfield-Menell
Listed -
Big picture of phasic dopamine
Listed -
Conservative Agency with Multiple Stakeholders
Listed -
Curriculum Design for Teaching via Demonstrations: Theory and Applications
Listed -
Dangerous optimisation includes variance minimisation
Listed -
Definitions of intent suitable for algorithms
Listed -
Engines of Power: Electricity, AI, and General-Purpose Military Transformations
Listed -
Evan Hubinger on Homogeneity in Takeoff Speeds, Learned Optimization and Interpretability
Listed -
Game-theoretic Alignment in terms of Attainable Utility
Listed -
Provably Robust Detection of Out-of-distribution Data (almost) for free
Listed -
Supplement to "Big picture of phasic dopamine"
Listed -
Survey on AI existential risk scenarios
Listed -
Survey on AI existential risk scenarios
Listed - Listed
- Listed
-
There Is No Turning Back: A Self-Supervised Approach for Reversibility-Aware Reinforcement Learning
Listed -
Improving Social Welfare While Preserving Autonomy via a Pareto Mediator
Listed - Listed
-
Some AI Governance Research Ideas
Listed -
Speculations against GPT-n writing alignment papers
Listed -
Review of "Learning Normativity: A Research Agenda"
Listed -
The dumbest kid in the world (joke)
Listed - Listed
-
High Impact Careers in Formal Verification: Artificial Intelligence
Listed -
Search-in-Territory vs Search-in-Map
Listed - Listed
-
Finite Factored Sets: Introduction and Factorizations
Listed -
I'm creating a world simulation video game
Listed - Listed
-
SIA is basically just Bayesian updating on existence
Listed -
The underlying model of a morphism
Listed -
An Intuitive Guide to Garrabrant Induction
Listed -
Offline Reinforcement Learning as One Big Sequence Modeling Problem
Listed - Listed
-
Rogue AGI Embodies Valuable Intellectual Property
Listed -
Some AI Governance Research Ideas
Listed -
Towards a Mathematical Theory of Abstraction
Listed -
Thoughts on the Alignment Implications of Scaling Language Models
Listed -
Why Release a Large Language Model?
Listed -
"Existential risk from AI" survey results
Listed -
"Existential risk from AI" survey results
Listed - Listed
-
Final Report of the National Security Commission on Artificial Intelligence (NSCAI, 2021)
Listed -
Machine Learning and Cybersecurity
Listed - Listed
-
What Matters for Adversarial Imitation Learning?
Listed -
How much will pre-transformative AI speed up R&D?
Listed -
Institutionalising Ethics in AI through Broader Impact Requirements
Listed -
[Event] Weekly Alignment Research Coffee Time
Listed -
AI Safety Career Bottlenecks Survey Responses Responses
Listed -
AXRP Episode 7.5 - Forecasting Transformative AI from Biological Anchors with Ajeya Cotra
Listed -
Predict responses to the "existential risk from AI" survey
Listed -
Teaching ML to answer questions honestly instead of predicting human answers
Listed -
The blue-minimising robot and model splintering
Listed -
An Offline Risk-aware Policy Selection Method for Bayesian Markov Decision Processes
Listed -
Interactive Explanations: Diagnosis and Repair of Reinforcement Learning Based Agent Behaviors
Listed -
Long-Term Future Fund: May 2021 grant recommendations
Listed -
List of good AI safety project ideas?
Listed -
MDP models are determined by the agent architecture and the environmental dynamics
Listed - Listed
-
Decoupling deliberation from competition
Listed -
Knowledge is not just map/territory resemblance
Listed - Listed
-
Controlling Intelligent Agents The Only Way We Know How: Ideal Bureaucratic Structure (IBS)
Listed -
Evaluating Different Fewshot Description Prompts on GPT-3
Listed -
Finetuning Models on Downstream Tasks
Listed -
On the Sizes of OpenAI API Models
Listed -
Problems facing a correspondence theory of knowledge
Listed -
True Few-Shot Learning with Language Models
Listed - Listed
- Listed
-
[Event] Weekly Alignment Research Coffee Time (05/24)
Listed -
AI Safety Research Project Ideas
Listed -
Navigation Turing Test (NTT): Learning to Evaluate Human-Like Navigation
Listed -
Response to "What does the universal prior actually look like?"
Listed -
[AN #151]: How sparsity in the final layer makes a neural net debuggable
Listed -
A Tour of Emerging Cryptographic Technologies | GovAI
Listed - Listed
- Listed
- Listed
- Listed
- Listed
- Listed
-
Unifying Principles and Metrics for Safe and Assistive AI
Listed -
Knowledge Neurons in Pretrained Transformers
Listed -
Why should we *not* put effort into AI safety research?
Listed -
[Event] Weekly Alignment Research Coffee Time (05/17)
Listed - Listed
-
Saving The Client-Side Web: just WASM and the DOM
Listed - Listed
- Listed
-
AXRP Episode 7 - Side Effects with Victoria Krakovna
Listed -
Our all-time largest donation, and major crypto support from Vitalik Buterin
Listed -
Understanding the Lottery Ticket Hypothesis
Listed -
Agency in Conway’s Game of Life
Listed -
[AN #150]: The subtypes of Cooperative AI research
Listed -
Formal Inner Alignment, Prospectus
Listed -
Challenge: know everything that the best go bot knows about go
Listed - Listed
-
Leveraging Sparse Linear Layers for Debuggable Deep Networks
Listed -
Yampolskiy on AI Risk Skepticism
Listed -
Human priors, features and models, languages, and Solmonoff induction
Listed -
[Event] Weekly Alignment Research Coffee Time (05/10)
Listed -
Pre-Training + Fine-Tuning Favors Deception
Listed -
Finding the unicorn: Predicting early stage startup success through a hybrid intelligence method
Listed -
Life and expanding steerable consequences
Listed -
Using reinforcement learning to design an AI assistantfor a satisfying co-op experience
Listed -
Adversarial Reprogramming of Neural Cellular Automata
Listed -
Anthropics: different probabilities, different questions
Listed - Listed
-
Parsing Chris Mingard on Neural Networks
Listed -
[AN #149]: The newsletter's editorial policy
Listed - Listed
-
Hard Choices and Hard Limits for Artificial Intelligence
Listed -
Mundane solutions to exotic problems
Listed -
Mundane solutions to exotic problems
Listed -
Parsing Abram on Gradations of Inner Alignment Obstacles
Listed -
Consistencies as (meta-)preferences
Listed - Listed
-
RL-IoT: Reinforcement Learning to Interact with IoT Devices
Listed -
[Weekly Event] Alignment Researcher Coffee Time (in Walled Garden)
Listed - Listed
- Listed
-
Planning for Proactive Assistance in Environments with Partial Observability
Listed -
pyBKT: An Accessible Python Library of Bayesian Knowledge Tracing Models
Listed -
The Unsatisfactorily Far Reach Of Property
Listed - Listed
-
Contending Frames: Evaluating Rhetorical Dynamics in AI
Listed -
Cooperative AI: machines must learn to find common ground
Listed -
Machine Intelligence for Scientific Discovery and Engineering Invention
Listed -
Symmetry, Equilibria, and Robustness in Common-Payoff Games
Listed -
The Societal Implications of Deep Reinforcement Learning
Listed - Listed
- Listed
- Listed
-
25 Min Talk on MetaEthical.AI with Questions from Stuart Armstrong
Listed -
[AN #148]: Analyzing generalization across more axes than just accuracy or loss
Listed -
AMA: Paul Christiano, alignment researcher
Listed -
Draft report on existential risk from power-seeking AI
Listed -
Draft report on existential risk from power-seeking AI
Listed -
Why AI is Harder Than We Think - Melanie Mitchell
Listed -
Agents Over Cartesian World Models
Listed - Listed
- Listed
-
[Linkpost] Treacherous turns in the wild
Listed -
A new proposal for regulating AI in the EU
Listed -
Announcing the Alignment Research Center
Listed -
Axes for Sociotechnical Inquiry in AI Research
Listed -
FAQ: Advice for AI Alignment Researchers
Listed -
Why AI is Harder Than We Think
Listed -
Beware over-use of the agent model
Listed -
Causal Learning for Socially Responsible AI
Listed - Listed
-
Let's not generalize over people
Listed -
Probability theory and logical induction as lenses
Listed -
Is there anything that can stop AGI development in the near term?
Listed -
NTK/GP Models of Neural Nets Can't Learn Features
Listed -
[AN #147]: An overview of the interpretability landscape
Listed -
Where are intentions to be found?
Listed -
Gradations of Inner Alignment Obstacles
Listed -
Rotary Embeddings: A Relative Revolution
Listed -
International cooperation as a tool to reduce two existential risks.
Listed -
The Power of Scale for Parameter-Efficient Prompt Tuning
Listed -
Updating the Lottery Ticket Hypothesis
Listed -
Action Advising with Advice Imitation in Deep Reinforcement Learning
Listed -
Learning on a Budget via Teacher Imitation
Listed -
An EPIC way to evaluate reward functions
Listed -
Superrational Agents Kelly Bet Influence!
Listed -
Computing Natural Abstractions: Linear Approximation
Listed -
Gradient-based Adversarial Attacks against Text Transformers
Listed -
[AN #146]: Plausible stories of how we might fail to avert an existential catastrophe
Listed -
An Interpretability Illusion for BERT
Listed - Listed
- Listed
- Listed
-
Fiction relevant to AI futurism
Listed -
Is there evidence that recommender systems are changing users' preferences?
Listed - Listed
-
Working in Congress (Part #1): Background and some EA cause area analysis
Listed -
Adapting Language Models for Zero-shot Learning by Meta-tuning on Dataset and Prompt Collections
Listed -
[AN #145]: Our three year anniversary!
Listed -
A Framework for Ethical AI at the United Nations
Listed -
Identifiability Problem for Superrational Decision Theories
Listed -
My Current Take on Counterfactuals
Listed -
Opinions on Interpretable Machine Learning and 70 Summaries of Recent Papers
Listed -
Why unriggable *almost* implies uninfluenceable
Listed -
A possible preference algorithm
Listed -
AXRP Episode 6 - Debate and Imitative Generalization with Beth Barnes
Listed -
Could Advanced AI Drive Explosive Economic Growth?
Listed -
If you don't design for extrapolation, you'll extrapolate poorly - possibly fatally
Listed -
Learning What To Do by Simulating the Past
Listed -
Solving the whole AGI control problem, version 0.0001
Listed -
Voluntary safety commitments provide an escape from over-regulation in AI development
Listed - Listed
-
What do coherence arguments imply about the behavior of advanced AI?
Listed -
Alignment Newsletter Three Year Retrospective
Listed -
Another (outer) alignment failure story
Listed -
Scaling Scaling Laws with Board Games
Listed -
Which counterfactuals should an AI follow?
Listed -
Case studies of self-governance to reduce technology risk
Listed -
Coherence arguments imply a force for goal-directed behavior
Listed - Listed
-
Testing The Natural Abstraction Hypothesis: Project Intro
Listed -
The Many Faces of Infra-Beliefs
Listed - Listed
-
Risk Budgets vs. Basic Decision Theory
Listed -
How do scaling laws work for fine-tuning?
Listed -
"AI and Compute" trend isn't predictive of what is happening
Listed -
[AN #144]: How language models can also be finetuned for non-language tasks
Listed -
Artificial intelligence, human rights, democracy, and the rule of law: a primer
Listed -
GPT-3 on Coherent Extrapolated Volition
Listed - Listed
-
My take on Michael Littman on "The HCI of HAI"
Listed - Listed
- Listed
-
Ethics and Artificial Intelligence
Listed -
Formal Methods for the Informal Engineer: Workshop Recommendations
Listed - Listed
-
DEALIO: Data-Efficient Adversarial Learning for Imitation from Observation
Listed - Listed
-
What Multipolar Failure Looks Like, and Robust Agent-Agnostic Processes (RAAPs)
Listed - Listed
-
How do we prepare for final crunch time?
Listed -
Optimization, speculations on the X and only X problem.
Listed -
"Weak AI" is Likely to Never Become "Strong AI", So What is its Greatest Value for us?
Listed - Listed
-
A Bayesian Approach to Identifying Representational Errors
Listed - Listed
- Listed
- Listed
-
Inframeasures and Domain Theory
Listed -
Review of "Fun with +12 OOMs of Compute"
Listed - Listed
- Listed
-
Coherence arguments imply a force for goal-directed behavior
Listed -
Coherence arguments imply a force for goal-directed behavior
Listed -
Report on Semi-informative Priors for AI timelines (Open Philanthropy)
Listed -
Characterizing and Detecting Mismatch in Machine-Learning-Enabled Systems
Listed -
My AGI Threat Model: Misaligned Model-Based RL Agent
Listed -
Why 1-boxing doesn't imply backwards causation
Listed -
[AN #143]: How to make embedded agents that reason probabilistically about their environments
Listed -
Counterfactual Explanation with Multi-Agent Reinforcement Learning for Drug Target Prediction
Listed -
Toward Building Science Discovery Machines
Listed -
Toy model of preference, bias, and extra information
Listed - Listed
-
Against evolution as an analogy for how humans will create AGI
Listed -
AGI risk: analogies & arguments
Listed -
Assured Learning-enabled Autonomy: A Metacognitive Reinforcement Learning Framework
Listed