The catalog, page 36
Records 8,751 to 9,000 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Worrying about the Vase: Whitelisting
Listed -
Evolving simple programs for playing Atari games
Listed -
Scrutinizing and De-Biasing Intuitive Physics with Neural Stethoscopes
Listed - Listed
-
Weak arguments against the universal prior being malign
Listed -
Counterfactual Mugging Poker Game
Listed -
The IQ of Artificial Intelligence
Listed -
Understanding the Meaning of Understanding
Listed -
Resource-Efficient Neural Architect
Listed -
A general model of safety-oriented AI development
Listed -
Adaptive Mechanism Design: Learning to Promote Cooperation
Listed -
An Efficient, Generalized Bellman Update For Cooperative Inverse Reinforcement Learning
Listed -
Announcing the second AI Safety Camp
Listed - Listed
-
The Alignment Newsletter #10: 06/11/18
Listed -
Jade Leung and Seth Baum: The role of existing institutions in AI strategy
Listed - Listed
-
RFC: Meta-ethical uncertainty in AGI alignment
Listed - Listed
-
Simplifying Reward Design through Divide-and-Conquer
Listed -
The first AI Safety Camp & onwards
Listed - Listed
-
Resource-Limited Reflective Oracles
Listed -
The first AI Safety Camp & onwards
Listed - Listed
-
Learning to Understand Goal Specifications by Modelling Reward
Listed -
Measuring and avoiding side effects using relative reachability
Listed -
Prisoners' Dilemma with Costs to Modeling
Listed -
Relational Deep Reinforcement Learning
Listed -
Relational recurrent neural networks
Listed -
Penalizing side effects using stepwise relative reachability
Listed -
Relational inductive bias for physical construction in humans and machines
Listed -
Relational inductive biases, deep learning, and graph networks
Listed - Listed
-
The Alignment Newsletter #9: 06/04/18
Listed -
Between Progress and Potential Impact of AI: the Neglected Dimensions
Listed - Listed
- Listed
-
Special issue on learning for human–robot collaboration
Listed -
Agents and Devices: A Relative Definition of Agency
Listed -
Explaining Explanations: An Overview of Interpretability of Machine Learning
Listed -
Probabilistically Safe Robot Planning with Confidence-Based Human Predictions
Listed -
Managing Loss of Control as Many Militaries Pursue Technological Superiority
Listed -
Robustness May Be at Odds with Accuracy
Listed -
To Trust Or Not To Trust A Classifier
Listed - Listed
-
Playing hard exploration games by watching YouTube
Listed -
Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition
Listed -
Virtuously Safe Reinforcement Learning
Listed -
The Alignment Newsletter #8: 05/28/18
Listed -
The simple picture on AI safety
Listed -
Training verified learners with learned verifiers
Listed -
When is unaligned AI morally valuable?
Listed -
Decision theory and zero-sum game theory, NP and PSPACE
Listed -
A Psychopathological Approach to Safety Engineering in AI and AGI
Listed -
Do Better ImageNet Models Transfer Better?
Listed -
Towards the first adversarially robust neural network model on MNIST
Listed -
How To Solve Moral Conundrums with Computability Theory
Listed -
Maximum Causal Tsallis Entropy Imitation Learning
Listed -
Meta-Learning with Hessian-Free Approach in Deep Neural Nets Training
Listed -
Verifiable Reinforcement Learning via Policy Extraction
Listed -
A Framework and Method for Online Inverse Reinforcement Learning
Listed -
Constructing Unrestricted Adversarial Examples with Generative Models
Listed -
Hierarchical Reinforcement Learning with Hindsight
Listed -
Imitating Latent Policies from Observation
Listed -
Learning Safe Policies with Expert Guidance
Listed -
Learning What Information to Give in Partially Observed Domains
Listed -
The Alignment Newsletter #7: 05/21/18
Listed -
Constrained Policy Improvement for Safe and Efficient Reinforcement Learning
Listed -
Task-Agnostic Meta-Learning for Few-shot Learning
Listed -
Challenges to Christiano’s capability amplification proposal
Listed -
Challenges to Christiano’s capability amplification proposal
Listed -
Solving the Rubik's Cube Without Human Knowledge
Listed -
Unsupervised Learning of Neural Networks to Explain Neural Networks
Listed - Listed
-
The Blessings of Multiple Causes
Listed -
Trend in compute used in training for headline AI results
Listed -
Feedback-Based Tree Search for Reinforcement Learning
Listed -
RFC: Philosophical Conservatism in AI Alignment Research
Listed -
The Alignment Newsletter #6: 05/14/18
Listed -
Directions and desiderata for AI alignment
Listed -
Thoughts on "AI safety via debate"
Listed -
Automated Mechanism Design via Neural Networks
Listed -
Thoughts on AI Safety via Debate
Listed -
Problems integrating decision theory and inverse reinforcement learning
Listed -
The Alignment Newsletter #5: 05/07/18
Listed - Listed
-
Open question: are minimal circuits daemon-free?
Listed -
AGI Safety Literature Review (Everitt, Lea & Hutter 2018)
Listed -
Dynamic Control Flow in Large-Scale Machine Learning
Listed -
Everything I ever needed to know, I learned from World of Warcraft: Goodhart’s law
Listed -
Rigging is a form of wireheading
Listed -
Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Listed -
Soon: a weekly AI Safety prerequisites module on LessWrong
Listed -
The Alignment Newsletter #4: 04/30/18
Listed -
Issues with Iterated Distillation and Amplification
Listed - Listed
-
A Logic of Agent Organizations
Listed -
Double Cruxing the AI Foom debate
Listed - Listed
-
Reward Learning from Narrated Demonstrations
Listed -
Goertzel’s GOLEM implements evidential decision theory applied to policy choice
Listed -
The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation
Listed -
No Metrics Are Perfect: Adversarial Reward Learning for Visual Storytelling
Listed -
Realistic Evaluation of Deep Semi-Supervised Learning Algorithms
Listed -
The Alignment Newsletter #3: 04/23/18
Listed - Listed
-
Preventing Side-effects in Gridworlds
Listed -
Deep Probabilistic Programming Languages: A Qualitative Study
Listed -
Understanding Iterated Distillation and Amplification: Claims and Oversight
Listed -
Heuristic Approaches for Goal Recognition in Incomplete Domain Models
Listed -
On Gradient-Based Learning in Continuous Games
Listed -
The Alignment Newsletter #2: 04/16/18
Listed -
Adversarial Attacks Against Medical Deep Learning Systems
Listed - Listed
- Listed
-
Quantilal control for finite MDPs
Listed -
Capsules for Object Segmentation
Listed -
Emergent Communication through Negotiation
Listed - Listed
- Listed
-
First Experiments with a Flexible Infrastructure for Normative Reasoning
Listed -
Large scale distributed neural network training through online distillation
Listed -
The Alignment Newsletter #1: 04/09/18
Listed - Listed
- Listed
-
Programmatically Interpretable Reinforcement Learning
Listed - Listed
-
The tyranny of the god scenario
Listed -
Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks
Listed -
Specification gaming examples in AI
Listed - Listed
-
2018 research plans and predictions
Listed -
Artificial Intelligence and its Role in Near Future
Listed -
Can corrigibility be learned safely?
Listed -
Corrigible but misaligned: a superintelligent messiah
Listed -
My take on agent foundations: formalizing metaphilosophical competence
Listed -
Specification gaming examples in AI
Listed -
Adversarial Attacks and Defences Competition
Listed -
Iterative Learning with Open-set Noisy Labels
Listed -
Meta-Learning Update Rules for Unsupervised Representation Learning
Listed -
Opportunities for individual donors in AI safety
Listed -
Three wagers for multiverse-wide superrationality
Listed -
Brain wiring: The long and short of it
Listed -
Resolving human values, completely and adequately
Listed -
Reward hacking and Goodhart’s law by evolutionary algorithms
Listed -
Transmitting fibers in the brain: Total length and distribution of lengths
Listed -
Autonomous Intelligent Cyber-defense Agent (AICA) Reference Architecture. Release 2.0
Listed -
New paper: “Categorizing variants of Goodhart’s Law”
Listed -
Evaluating Existing Approaches to AGI Alignment
Listed - Listed
-
Non-Adversarial Goodhart and AI Risks
Listed - Listed
- Listed
-
Computational Power and the Social Impact of Artificial Intelligence
Listed -
Idea: Open Access AI Safety Journal
Listed -
Learning-based Model Predictive Control for Safe Exploration
Listed -
An Untrollable Mathematician Illustrated
Listed -
Generating Multi-Agent Trajectories using Programmatic Weak Supervision
Listed -
AI Alignment Prize: Super-Boxing
Listed - Listed
- Listed
-
Is the Star Trek Federation really incapable of building AI?
Listed -
A Dual Approach to Scalable Verification of Deep Networks
Listed - Listed
-
Jan Leike on how to become a machine learning alignment researcher
Listed - Listed
-
Using lying to detect human values
Listed -
Active Reinforcement Learning with Monte-Carlo Tree Search
Listed -
Categorizing Variants of Goodhart's Law
Listed -
Deep k-Nearest Neighbors: Towards Confident, Interpretable and Robust Deep Learning
Listed -
Fractal AI: A fragile theory of intelligence
Listed -
AI Alignment Prize: Round 2 due March 31, 2018
Listed -
Opportunities for individual donors in AI safety
Listed -
Brains and backprop: a key timeline crux
Listed - Listed
-
The Challenge of Crafting Intelligible Intelligence
Listed - Listed
- Listed
-
SentRNA: Improving computational RNA design by incorporating a prior of human design strategies
Listed -
A Brandom-ian view of Reinforcement Learning towards strong-AI
Listed -
Value Alignment, Fair Play, and the Rights of Service Robots
Listed -
The Building Blocks of Interpretability
Listed -
Iterated Distillation and Amplification
Listed -
Takeoff Speed: Simple Asymptotics in a Toy Model.
Listed -
AI impacts and Paul Christiano on takeoff speeds
Listed - Listed
-
Human-aligned artificial intelligence is a multiobjective problem
Listed -
Quick Nate/Eliezer comments on discontinuity
Listed -
Sam Harris and Eliezer Yudkowsky on “AI: Racing Toward the Brink”
Listed -
Beyond algorithmic equivalence: algorithmic noise
Listed -
Beyond algorithmic equivalence: self-modelling
Listed - Listed
-
Antifragility for Intelligent Autonomous Systems
Listed - Listed
-
Learning from Physical Human Corrections, One Feature at a Time
Listed -
More on the Linear Utility Hypothesis and the Leverage Prior
Listed -
Walkthrough of 'Formalizing Convergent Instrumental Goals'
Listed - Listed
- Listed
-
Self-regulation of safety in AI research
Listed -
The abruptness of nuclear weapons
Listed - Listed
- Listed
-
June 2012: 0/33 Turing Award winners predict computers beating humans at go within next 10 years.
Listed -
Likelihood of discontinuous progress around the development of AGI
Listed -
Don't Condition on no Catastrophes
Listed - Listed
-
Manipulating and Measuring Model Interpretability
Listed - Listed
-
The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation
Listed -
Using surrogate goals to deflect threats
Listed - Listed
-
Electrical efficiency of computing
Listed -
Learning Data-Driven Objectives to Optimize Interactive Systems
Listed -
Nordhaus hardware price performance dataset
Listed -
Toward a New Technical Explanation of Technical Explanation
Listed -
Adversarial Risk and the Dangers of Evaluating Against Weak Attacks
Listed -
The law of effect, randomization and Newcomb’s problem
Listed - Listed
-
2018 price of performance by Tensor Processing Units
Listed -
Examples of AI systems producing unconventional solutions
Listed -
Some conceptual highlights from “Disjunctive Scenarios of Catastrophic AI Risk”
Listed - Listed
-
More Robust Doubly Robust Off-policy Evaluation
Listed - Listed
-
Stable Pointers to Value II: Environmental Goals
Listed -
Goal Inference Improves Objective and Perceived Performance in Human-Robot Collaboration
Listed -
Shared Autonomy via Deep Reinforcement Learning
Listed - Listed
-
First-order Adversarial Vulnerability of Neural Networks and Input Dimension
Listed -
Learning from Richer Human Guidance: Augmenting Comparison-Based Learning with Feature Queries
Listed -
Factorio, Accelerando, Empathizing with Empires and Moderate Takeoffs
Listed -
Logical counterfactuals and differential privacy
Listed -
AI Safety Research Camp - Project Proposal
Listed -
Techniques for optimizing worst-case performance
Listed -
The Utility of Human Atoms for the Paperclip Maximizer
Listed -
Bias in AI: How we Build Fair AI Systems and Less-Biased Humans
Listed -
Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
Listed - Listed
-
Epiphenomenal Oracles Ignore Holes in the Box
Listed -
Sources of intuitions and data on AGI
Listed - Listed
-
Against Instrumental Convergence
Listed -
Is there a tradeoff between immediate and longer-term AI safety efforts?
Listed -
Strategy Nonconvexity Induced by a Choice of Potential Oracles
Listed -
Safe Exploration in Continuous Action Spaces
Listed -
Space races: Settling the universe Fast
Listed - Listed
-
AI alignment prize winners and next round [link]
Listed