The catalog, page 32
Records 7,751 to 8,000 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Clarifying some key hypotheses in AI alignment
Listed - Listed
-
Legible Normativity for AI Alignment: The Value of Silly Rules.
Listed -
Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning.
Listed -
On the Utility of Model Learning in HRI.
Listed -
Reward-rational (implicit) choice: A unifying formalism for reward learning.
Listed -
The Assistive Multi-Armed Bandit.
Listed - Listed
-
A Risk-Sensitive Finite-Time Reachability Approach for Safety of Stochastic Dynamic Systems.
Listed -
A unified framework for planning in adversarial and cooperative environments.
Listed - Listed
-
Adversarial Training with Voronoi Constraints.
Listed -
An Agent-Based Model of Financial Benchmark Manipulation.
Listed -
Approximate Causal Abstraction.
Listed -
Bayesian Robustness: A Nonasymptotic Viewpoint.
Listed -
Benchmarking Neural Network Robustness to Common Corruptions and Perturbations.
Listed -
Blameworthiness in Multi-Agent Settings.
Listed -
Bridging Hamilton-Jacobi Safety Analysis and Reinforcement Learning.
Listed - Listed
-
Causal Discovery in the Presence of Missing Data.
Listed -
Deep Anomaly Detection with Outlier Exposure.
Listed -
Distributed Protocols for Leader Election: A Game-Theoretic Perspective.
Listed -
Epistemic Therapy for Bias in Automated Decision-Making.
Listed -
Graphical Models for Processing Missing Data.
Listed -
Hierarchically Decoupled Imitation for Morphological Transfer.
Listed -
How You Act Tells a Lot: Privacy-Leaking Attack on Deep Reinforcement Learning.
Listed -
Human Compatible: Artificial Intelligence and The Problem of Control.
Listed -
Implementing Mediators with Asynchronous Cheap Talk.
Listed -
Incentivizing Collaboration in a Competition.
Listed -
Learning a Prior over Intent via Meta-Inverse Reinforcement Learning.
Listed -
Learning-Based Trading Strategies in the Face of Market Manipulation.
Listed - Listed
-
On the Existence of Nash Equilibrium in Games with Resource-Bounded Players.
Listed - Listed
- Listed
- Listed
-
Security in Asynchronous Interactive Systems.
Listed -
Sequential equilibrium in computational games.
Listed -
Strategic Classification is Causal Modeling in Disguise.
Listed -
The Computational Structure of Unintentional Meaning.
Listed - Listed
-
Using Machine Learning to Guide Cognitive Modeling: A Case Study in Moral Reasoning.
Listed -
Using Self-Supervised Learning Can Improve Model Robustness and Uncertainty.
Listed - Listed
-
Evidence against current methods leading to human level artificial intelligence
Listed - Listed
-
Attention in value-based choice as optimal sequential sampling.
Listed -
Bayesian Relational Memory for Semantic Visual Navigation.
Listed -
Combining reward information from multiple sources.
Listed -
Deception in finitely repeated security games.
Listed -
Demonstrating the Impact of Prior Knowledge in Risky Choice.
Listed -
Doing more with less: meta-reasoning and meta-learning in humans and machines.
Listed -
Human-robot interaction for truck platooning using hierarchical dynamic games.
Listed - Listed
-
Learning Causal Trees with Latent Variables via Controlled Experimentation.
Listed -
Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient.
Listed -
Scaling Out-of-Distribution Detection for Real-World Settings.
Listed -
The truth behind the myth of the folk theorem.
Listed - Listed
-
Using Pre-Training Can Improve Model Robustness and Uncertainty.
Listed -
Why Can’t You Do That, HAL? Explaining Unsolvability of Planning Tasks.
Listed -
Behaviour Suite for Reinforcement Learning
Listed -
AI Forecasting Dictionary (Forecasting infrastructure, part 1)
Listed -
Four Ways An Impact Measure Could Help Alignment
Listed - Listed
- Listed
-
Which of these five AI alignment research projects ideas are no good?
Listed - Listed
-
Self-Supervised Learning and AGI Safety
Listed -
Understanding Recent Impact Measures
Listed -
A Discussion of 'Adversarial Examples Are Not Bugs, They Are Features'
Listed - Listed
- Listed
- Listed
- Listed
- Listed
-
A Discussion of 'Adversarial Examples Are Not Bugs, They Are Features': Robust Feature Leakage
Listed - Listed
-
A Survey of Early Impact Measures
Listed -
An overview of arguments for concern about automation
Listed -
New paper: Corrigibility with Utility Preservation
Listed -
Project Proposal: Considerations for trading off capabilities and safety impacts of AI research
Listed -
[AN #61] AI policy and governance, from two people in the field
Listed -
AI Alignment Open Thread August 2019
Listed -
Improving Deep Reinforcement Learning in Minecraft with Action Advice
Listed -
Practical consequences of impossibility of value learning
Listed - Listed
- Listed
-
Contest: $1,000 for good questions to ask to an Oracle AI
Listed -
Towards a Theory of Intentions for Human-Robot Collaboration
Listed -
An upper bound for the background rate of human extinction
Listed - Listed
-
What does Optimization Mean, Again? (Optimizing and Goodhart Effects - Clarifying Thoughts, Part 2)
Listed - Listed
-
Some post- words for the future
Listed -
The Artificial Intentional Stance
Listed -
A Unified Bellman Optimality Principle Combining Reward Maximization and Empowerment
Listed -
Ought: why it matters and ways to help
Listed -
On the purposes of decision theory research
Listed -
Ought: why it matters and ways to help
Listed - Listed
-
IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL
Listed -
AI Safety Debate and Its Applications
Listed -
[AN #60] A new AI challenge: Minecraft agents that assist human players in creative mode
Listed - Listed
- Listed
-
A system of different layers of abstraction for artificial intelligence
Listed -
Why Build an Assistant in Minecraft?
Listed - Listed
-
Delegative Reinforcement Learning: learning to avoid traps with a little help
Listed -
Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery
Listed - Listed
-
Historic trends in land speed records
Listed -
Robust Multi-Agent Reinforcement Learning via Minimax Deep Deterministic Policy Gradient
Listed -
An Inductive Synthesis Framework for Verifiable Reinforcement Learning
Listed -
Jeff Hawkins on neuromorphic AGI within 20 years
Listed -
How Europe might matter for AI governance
Listed -
Grounding Value Alignment with Ethical Principles
Listed - Listed
- Listed
-
An Optimistic Perspective on Offline Reinforcement Learning
Listed -
The Role of Cooperation in Responsible AI Development
Listed -
Better-than-Demonstrator Imitation Learning via Automatically-Ranked Demonstrations
Listed -
[AN #59] How arguments for AI risk have changed over time
Listed - Listed
-
Some Comments on Stuart Armstrong's "Research Agenda v0.9"
Listed -
Musings on Cumulative Cultural Evolution and AI
Listed -
Learning biases and rewards simultaneously
Listed -
Quantifying the pathways to life using assembly spaces
Listed -
Learning a Behavioral Repertoire from Demonstrations
Listed -
On Inductive Biases in Deep Reinforcement Learning
Listed -
Adversarial Robustness through Local Linearization
Listed -
Large Scale Adversarial Representation Learning
Listed - Listed
-
Dynamics-Aware Unsupervised Discovery of Skills
Listed -
Generalizing from a few environments in safety-critical reinforcement learning
Listed - Listed
-
An Increasingly Manipulative Newsfeed
Listed - Listed
-
Detecting Spiky Corruption in Markov Decision Processes
Listed -
Aligning a toy model of optimization
Listed -
Artificial Intelligence Governance and Ethics: Global Perspectives
Listed -
Conceptual Problems with UDT and Policy Selection
Listed -
Self-confirming prophecies, and simplified Oracle designs
Listed -
Embedded Agency: Not Just an AI Problem
Listed -
Norms for Beneficial A.I.: A Computational Analysis of the Societal Value Alignment Problem
Listed -
Towards Empathic Deep Q-Learning
Listed -
Universal Litmus Patterns: Revealing Backdoor Attacks in CNNs
Listed -
Reinforcement Learning with Competitive Ensembles of Information-Constrained Primitives
Listed -
Research Agenda in reverse: what *would* a solution look like?
Listed -
[AN #58] Mesa optimization: what it is, and why we should care
Listed -
Confidence-aware motion prediction for real-time collision avoidance <sup>1</sup>
Listed -
Learning to Interactively Learn and Assist
Listed -
Machine Learning Projects on IDA
Listed -
An AGI with Time-Inconsistent Preferences
Listed -
On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inference
Listed -
"The Bitter Lesson", an article about compute vs human knowledge in AI
Listed -
Categorizing Wireheading in Partially Embedded Agents
Listed -
Modeling AGI Safety Frameworks with Causal Influence Diagrams
Listed -
Information security careers for GCR reduction
Listed -
Modeling AGI Safety Frameworks with Causal Influence Diagrams
Listed -
Unsupervised State Representation Learning in Atari
Listed -
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Listed - Listed
- Listed
- Listed
-
Research Agenda v0.9: Synthesising a human's preferences into a utility function
Listed -
Goal-conditioned Imitation Learning
Listed -
Stand-Alone Self-Attention in Vision Models
Listed -
The Al Does Not Hate You: Superintelligence, Rationality and the Race to Save the World
Listed -
Let's talk about "Convergent Rationality"
Listed -
Weight Agnostic Neural Networks
Listed -
Weight Agnostic Neural Networks
Listed -
A Survey of Reinforcement Learning Informed by Natural Language
Listed -
E-LPIPS: Robust Perceptual Image Similarity via Random Transformation Ensembles
Listed -
Self-Supervised Exploration via Disagreement
Listed -
Tackling Climate Change with Machine Learning
Listed - Listed
-
AGI will drastically increase economies of scale
Listed -
For the past, in some ways only, we are moral degenerates
Listed -
Likelihood Ratios for Out-of-Distribution Detection
Listed -
New paper: “Risks from learned optimization”
Listed -
Planning With Uncertain Specifications (PUnS)
Listed -
Risks from Learned Optimization: Conclusion and Related Work
Listed -
Risks from Learned Optimization: Conclusion and Related Work
Listed -
An Extensible Interactive Interface for Agent Design
Listed -
Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift
Listed -
Image Synthesis with a Single (Robust) Classifier
Listed -
Visualizing and Measuring the Geometry of BERT
Listed -
[AN #57] Why we should focus on robustness in AI safety, and the analogous problems in programming
Listed - Listed
- Listed
-
Methodology for discontinuous progress investigation
Listed -
Teaching AI to Explain its Decisions Using Embeddings and Multi-Task Learning
Listed - Listed
- Listed
-
Adversarial Robustness as a Prior for Learned Representations
Listed - Listed
- Listed
-
Learner-aware Teaching: Inverse Reinforcement Learning with Preferences and Constraints
Listed - Listed
- Listed
-
The Principle of Unchanged Optimality in Reinforcement Learning Generalization
Listed - Listed
-
AI Governance and the Policymaking Process: Key Considerations for Reducing AI Risk
Listed -
Conditions for Mesa-Optimization
Listed -
Conditions for Mesa-Optimization
Listed -
Multiparty Dynamics and Failure Modes for Machine Learning and Artificial Intelligence
Listed -
Risks from Learned Optimization: Introduction
Listed -
Risks from Learned Optimization: Introduction
Listed -
Better Future through AI: Avoiding Pitfalls and Guiding AI Towards its Full Potential
Listed -
Imitation Learning as $f$-Divergence Minimization
Listed -
Asymptotically Unambitious Artificial General Intelligence
Listed -
Defending Against Neural Fake News
Listed -
Learning Representations by Humans, for Humans
Listed -
SATNet: Bridging deep learning and logical reasoning using a differentiable satisfiability solver
Listed -
A shift in arguments for AI risk
Listed -
Causal Confusion in Imitation Learning
Listed - Listed
-
Cold Case: The Lost MNIST Digits
Listed -
On modelling the emergence of logical thinking
Listed -
AI-CARGO: A Data-Driven Air-Cargo Revenue Management System
Listed -
And the AI would have got away with it too, if...
Listed -
Cognitive Model Priors for Predicting Human Decisions
Listed -
Imitation Learning from Video by Leveraging Proprioception
Listed -
Where are people thinking and talking about global coordination for AI safety?
Listed -
[AN #56] Should ML researchers stop running experiments before making hypotheses?
Listed -
By default, avoid ambiguous distant situations
Listed -
Perceptual Values from Observation
Listed -
On Variational Bounds of Mutual Information
Listed - Listed
-
Jade Leung: Why companies should be leading on AI governance
Listed - Listed
-
Lie on the Fly: Strategic Voting in an Iterative Preference Elicitation Process
Listed - Listed
-
Coherent decisions imply consistent utilities
Listed -
Complex Behavior from Simple (Sub)Agents
Listed -
Integrating Artificial Intelligence into Weapon Systems
Listed - Listed
-
Training human models is an unsolved problem
Listed -
Aligning Recommender Systems as Cause Area
Listed -
Meta-learning of Sequential Strategies
Listed -
Toybox: A Suite of Environments for Experimental Evaluation of Deep Reinforcement Learning
Listed -
Adversarial Examples Are Not Bugs, They Are Features
Listed -
Deconstructing Lottery Tickets: Zeros, Signs, and the Supermask
Listed -
Value learning for moral essentialists
Listed -
[AN #55] Regulatory markets and international standards as a means of ensuring beneficial AI
Listed -
Deconstructing Lottery Tickets: Zeros, Signs, and the Supermask
Listed -
Meta-learners' learning dynamics are unlike learners'
Listed -
Oracles, sequence predictors, and self-confirming predictions
Listed