The catalog, page 31
Records 7,501 to 7,750 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
- Listed
-
Scaling Out-of-Distribution Detection for Real-World Settings
Listed -
Thoughts on Robin Hanson's AI Impacts interview
Listed -
Analysing: Dangerous messages from future UFAI via Oracles
Listed -
Ultra-simplified research agenda
Listed -
A Brief Intro to Domain Theory
Listed - Listed
-
ReMixMatch: Semi-Supervised Learning with Distribution Alignment and Augmentation Anchoring
Listed -
[AN #74]: Separating beneficial AI into competence, alignment, and coping with impacts
Listed -
AI safety scholarships look worth-funding (if other funding is sane)
Listed -
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Listed -
Planning with Goal-Conditioned Policies
Listed -
Impossible moral problems and moral authority
Listed -
Self-Fulfilling Prophecies Aren't Always About Self-Awareness
Listed - Listed
-
The new dot com bubble is here: it’s called online advertising
Listed - Listed
-
How common is it for one entity to have a 3+ year technological lead on its nearest competitor?
Listed -
I'm Buck Shlegeris, I do research and outreach at MIRI, AMA
Listed - Listed
-
[AN #73]: Detecting catastrophic failures by learning how agents tend to break
Listed -
Conversation with Robin Hanson
Listed -
Momentum Contrast for Unsupervised Visual Representation Learning
Listed - Listed
-
Robin Hanson on the futurist focus on AI
Listed -
A conversation with Rohin Shah
Listed - Listed
-
(When) Is Truth-telling Favored in AI Debate?
Listed - Listed
-
Operationalizing Newcomb's Problem
Listed -
Self-training with Noisy Student improves ImageNet classification
Listed - Listed
- Listed
- Listed
- Listed
-
A mechanistic model of meditation
Listed -
AI Alignment Research Overview (by Jacob Steinhardt)
Listed - Listed
-
Nonverbal Robot Feedback for Human Teachers
Listed -
On the Measure of Intelligence
Listed -
Computing Receptive Fields of Convolutional Neural Networks
Listed -
More variations on pseudo-alignment
Listed -
Will transparency help catch deception? Perhaps not
Listed -
But exactly how complex and fragile?
Listed -
“embedded self-justification,” or something like that
Listed -
AlphaStar: Impressive for RL progress, not for AGI progress
Listed -
Assessing the state of AI R&D in the US, China, and Europe – Part 1: Output indicators
Listed -
Chris Olah’s views on AGI safety
Listed -
Generating Justifications for Norm-Related Agent Decisions
Listed -
Positive-Unlabeled Reward Learning
Listed -
A Narration-based Reward Shaping Approach using Grounded Natural Language Commands
Listed -
Adversarial NLI: A New Benchmark for Natural Language Understanding
Listed - Listed
- Listed
-
Rohin Shah on reasons for AI optimism
Listed -
Rohin Shah on reasons for AI optimism
Listed -
[AN #71]: Avoiding reward tampering through current-RF optimization
Listed -
Network Classifiers With Output Smoothing
Listed - Listed
-
Doing Global Priorities or AI Policy research from remote location?
Listed - Listed
-
Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
Listed -
[AN #70]: Agents that help humans who are still learning about their own preferences
Listed -
Deliberation as a method to find the "actual preferences" of humans
Listed -
How can AI Automate End-to-End Data Science?
Listed - Listed
- Listed
-
An Alternative Surrogate Loss for PGD-based Adversarial Testing
Listed -
Collaborating with Humans Requires Understanding Them
Listed -
Multi-agent Hierarchical Reinforcement Learning with Dynamic Termination
Listed -
The Psychology of Existential Risk: Moral Judgments about Human Extinction
Listed -
The problem/solution matrix: Calculating the probability of AI safety "on the back of an envelope"
Listed -
[AN #69] Stuart Russell's new book on why we need to replace the standard model of AI
Listed - Listed
-
Summary of Stuart Russell's new book, "Human Compatible"
Listed -
Jeffrey Ding: Re-deciphering China’s AI dream
Listed -
Jesse Clifton: Open-source learning — a bargaining approach
Listed -
Ross Gruetzemacher: Defining and unpacking transformative AI
Listed -
Technical AGI safety research outside AI
Listed -
Technical AGI safety research outside AI
Listed -
Random Thoughts on Predict-O-Matic
Listed -
The Dualist Predict-O-Matic ($100 prize)
Listed -
Full toy model for preference learning
Listed - Listed
-
Impact measurement and value-neutrality verification
Listed -
Restoring ancient text using deep learning: a case study on Greek epigraphy
Listed -
Solving Logic Grid Puzzles with an Algorithm that Imitates Human Behavior
Listed -
The Parable of Predict-O-Matic
Listed -
[AN #68]: The attainable utility theory of impact
Listed - Listed
-
Using AI/ML to gain situational understanding from passive network observations
Listed - Listed
-
On the Utility of Learning about Humans for Human-AI Coordination
Listed -
Stabilizing Transformers for Reinforcement Learning
Listed -
Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
Listed -
Imitation Learning from Observations by Minimizing Inverse Dynamics Disagreement
Listed -
The Quest for Interpretable and Responsible Artificial Intelligence
Listed -
Thoughts on "Human-Compatible"
Listed -
Improving Generalization in Meta Reinforcement Learning using Learned Objectives
Listed - Listed
-
Minimization of prediction error as a foundation for human values in AI alignment
Listed -
Can We Distinguish Machine Learning from Human Learning?
Listed -
Characterizing Real-World Agents as a Research Meta-Strategy
Listed -
Detecting AI Trojans Using Meta Neural Analysis
Listed -
Human Compatible: Artificial Intelligence and the Problem of Control
Listed -
Misconceptions about continuous takeoff
Listed -
What's the dream for giving natural language commands to AI?
Listed -
[AN #67]: Creating environments in which to study inner alignment failures
Listed -
AI Alignment Writing Day Roundup #2
Listed - Listed
- Listed
-
Towards Deployment of Robust AI Agents for Human-Machine Partnerships
Listed -
AI Alignment Open Thread October 2019
Listed -
Debate on Instrumental Convergence between LeCun, Russell, Bengio, Zador, and More
Listed - Listed
-
Universality and model-based RL
Listed -
Can we make peace with moral indeterminacy?
Listed -
Formal Language Constraints for Markov Decision Processes
Listed -
Human instincts, symbol grounding, and the blank-slate neocortex
Listed -
Improving Sample Efficiency in Model-Free Reinforcement Learning from Images
Listed -
AI Alignment Research Overview
Listed -
Doing more with less: meta-reasoning and meta-learning in humans and machines
Listed - Listed
- Listed
-
World State is the Wrong Abstraction for Impact
Listed -
[AN #66]: Decomposing robustness into capability robustness and alignment robustness
Listed -
List of resolved confusions about IDA
Listed -
The Paths Perspective on Value Learning
Listed -
Christiano decision theory excerpt
Listed -
Gradient Descent: The Ultimate Optimizer
Listed -
Learning from Observations Using a Single Video Demonstration and Human Feedback
Listed -
UK policy and politics careers
Listed -
[Talk] Paul Christiano on his alignment taxonomy
Listed -
A Constructive Prediction of the Generalization Error Across Scales
Listed -
Attainable Utility Theory: Why Things Matter
Listed -
Automated curricula through setter-solver interactions
Listed - Listed
-
A simple environment for showing mesa misalignment
Listed -
Scaling data-driven robotics with reward sketching and batch reinforcement learning
Listed -
Toward Evaluating Robustness of Deep Reinforcement Learning with Continuous Control
Listed - Listed
-
[AN #65]: Learning useful skills by watching humans “play”
Listed -
Towards an empirical investigation of inner alignment
Listed - Listed
-
Scaled Autonomy: Enabling Human Operators to Control Robot Fleets
Listed -
Leveraging Human Guidance for Deep Reinforcement Learning Tasks
Listed -
What are the differences between all the iterative/recursive approaches to AI alignment?
Listed -
Meta-Inverse Reinforcement Learning with Probabilistic Context Variables
Listed - Listed
- Listed
-
How does the offense-defense balance scale?
Listed -
Fine-Tuning Language Models from Human Preferences
Listed -
The unexpected difficulty of comparing AlphaStar to humans
Listed -
The unexpected difficulty of comparing AlphaStar to humans
Listed -
Emergent Tool Use From Multi-Agent Autocurricula
Listed -
[AN #64]: Using Deep RL and Reward Uncertainty to Incentivize Preference Learning
Listed -
Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural Networks
Listed - Listed
-
The strategy-stealing assumption
Listed -
The strategy-stealing assumption
Listed -
VILD: Variational Imitation Learning with Diverse-quality Demonstrations
Listed -
A Critique of Functional Decision Theory
Listed -
Do Sufficiently Advanced Agents Use Logic?
Listed -
What You See Isn't Always What You Want
Listed -
Better AI through Logical Scaffolding
Listed -
Finding Generalizable Evidence by Learning to Convince Q&A Models
Listed -
Conversation with Paul Christiano
Listed -
Conversation with Paul Christiano
Listed -
Paul Christiano on the safety of future AI systems
Listed -
Soft takeoff can still lead to decisive strategic advantage
Listed - Listed
-
Counterfactual Oracles = online supervised learning with random selection of training episodes
Listed -
Is my result wrong? Maths vs intuition vs evolution in learning human preferences
Listed -
Meta-Learning with Implicit Gradients
Listed -
Relaxed adversarial training for inner alignment
Listed - Listed
-
Are minimal circuits deceptive?
Listed -
Concrete experiments in inner alignment
Listed -
How much EA analysis of AI safety as a cause area exists?
Listed - Listed
-
Implications of Quantum Computing for Artificial Intelligence alignment research (ABRIDGED)
Listed -
Logical Counterfactuals and Proposition graphs, Part 3
Listed -
Making Efficient Use of Demonstrations to Solve Hard Exploration Problems
Listed - Listed
- Listed
-
Achieving Verified Robustness to Symbol Substitutions via Interval Bound Propagation
Listed -
AI Forecasting Question Database (Forecasting infrastructure, part 3)
Listed -
Authoritarian Audiences, Rhetoric, and Propaganda in International Crises: Evidence from China
Listed -
Counterfactuals are an Answer, Not a Question
Listed -
LCA: Loss Change Allocation for Neural Network Training
Listed -
Making Efficient Use of Demonstrations to Solve Hard Exploration Problems
Listed - Listed
-
Shortcomings of the Bow Tie and Other Safety Tools Based on Linear Causality
Listed -
Logical Counterfactuals and Proposition graphs, Part 2
Listed - Listed
- Listed
-
AI Alignment Writing Day Roundup #1
Listed -
From the Internet of Information to the Internet of Intelligence
Listed - Listed
-
AI Forecasting Resolution Council (Forecasting infrastructure, part 2)
Listed -
[Link] Book Review: Reframing Superintelligence (SSC)
Listed - Listed
- Listed
- Listed
- Listed
- Listed
- Listed
-
End Times: A Brief Guide to the End of the World
Listed - Listed
-
Building The Castle vs Finding The Monolith • carado.moe
Listed -
Embedded Agency via Abstraction
Listed - Listed
-
Reversible changes: consider a bucket of water
Listed -
Gratification: a useful concept, maybe new
Listed -
Under a week left to win $1,000! By questioning Oracle AIs.
Listed -
Ernie Davis on the landscape of AI risks
Listed -
Release Strategies and the Social Impacts of Language Models
Listed - Listed
- Listed
-
Creating Environments to Design and Test Embedded Agents
Listed -
Does Agent-like Behavior Imply Agent-like Architecture?
Listed -
Existential risks: a philosophical analysis
Listed -
Formalising decision theory is hard
Listed - Listed
- Listed
-
Soft takeoff can still lead to decisive strategic advantage
Listed -
Tabooing 'Agent' for Prosaic Alignment
Listed - Listed
- Listed
-
Towards an Intentional Research Agenda
Listed - Listed
- Listed
-
Vaniver's View on Factored Cognition
Listed -
When do utility functions constrain?
Listed -
[AN #62] Are adversarial examples caused by real but imperceptible features?
Listed -
Announcement: Writing Day Today (Thursday)
Listed -
Computational Model: Causal Diagrams with Symmetry
Listed -
Implications of Quantum Computing for Artificial Intelligence Alignment Research
Listed -
Logical Counterfactuals and Proposition graphs, Part 1
Listed -
Markets are Universal for Logical Induction
Listed -
Towards a mechanistic understanding of corrigibility
Listed -
Call for contributors to the Alignment Newsletter
Listed -
Testing Robustness Against Unforeseen Adversaries
Listed - Listed
- Listed
-
Classifying specification problems as variants of Goodhart's Law
Listed -
Classifying specification problems as variants of Goodhart’s Law
Listed -
Goodhart's Curse and Limitations on AI Alignment
Listed -
Implications of Quantum Computing for Artificial Intelligence alignment research
Listed -
Problems in AI Alignment that philosophers could potentially contribute to
Listed