The catalog, page 33
Records 8,001 to 8,250 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
PRECOG: PREdiction Conditioned On Goals in Visual Multi-Agent Settings
Listed -
Self-confirming predictions can be arbitrarily bad
Listed -
Transfer of Adversarial Robustness Between Perturbation Types
Listed -
Bridging Hamilton-Jacobi Safety Analysis and Reinforcement Learning
Listed -
Nash equilibriums can be arbitrarily bad
Listed -
The relationship between Biological and Artificial Intelligence
Listed -
Challenges of Real-World Reinforcement Learning
Listed -
[AN #54] Boxing a finite-horizon AI system to keep it unambitious
Listed -
What are some good examples of incorrigibility?
Listed -
Regulating AI: do we need new tools?
Listed -
Knowing When to Stop: Evaluation and Verification of Conformity to Output-size Specifications
Listed -
Using Sub-Optimal Plan Detection to Identify Commitment Abandonment in Discrete Environments
Listed -
Ray Interference: a Source of Plateaus in Deep Reinforcement Learning
Listed -
Strategic implications of AIs' ability to coordinate at low cost, for example by merging
Listed -
New paper: “Delegative reinforcement learning”
Listed -
Long-Term Future Fund: April 2019 grant recommendations
Listed -
Risk Structures: Towards Engineering Risk-aware Autonomous Systems
Listed -
AI Alignment Problem: “Human Values” don’t Actually Exist
Listed - Listed
-
Optimization and Abstraction: A Synergistic Approach for Analyzing Neural Network Robustness
Listed -
The MineRL 2019 Competition on Sample Efficient Reinforcement Learning using Human Priors
Listed -
Any rebuttals of Christiano and AI Impacts on takeoff speeds?
Listed -
Generative Exploration and Exploitation
Listed -
Helen Toner on China, CSET, and AI
Listed - Listed
-
When is a Prediction Knowledge?
Listed -
Analysing Neural Network Topologies: a Game Theoretic Approach
Listed - Listed
-
Counterfactual Visual Explanations
Listed -
End-to-End Robotic Reinforcement Learning without Reward Engineering
Listed -
HARK Side of Deep Learning -- From Grad Student Descent to Automated Machine Learning
Listed -
Predicting human decisions with behavioral theories and machine learning
Listed -
Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations
Listed -
Corrigibility as Constrained Optimisation
Listed -
Alignment Newsletter One Year Retrospective
Listed -
Alignment Newsletter One Year Retrospective
Listed -
Best reasons for pessimism about impact of impact measures?
Listed -
Semantics: Primes and Universals, a book review • carado.moe
Listed -
Open Questions about Generative Adversarial Networks
Listed -
Extending planning knowledge using ontologies for goal opportunities
Listed -
Reinforcement learning with imperceptible rewards
Listed - Listed
- Listed
-
Defeating Goodhart and the "closest unblocked strategy" problem
Listed - Listed
-
A Visual Exploration of Gaussian Processes
Listed -
Are Query-Based Ontology Debuggers Really Helping Knowledge Engineers?
Listed -
Finding and Visualizing Weaknesses of Deep Reinforcement Learning Agents
Listed - Listed
-
Multitask Soft Option Learning
Listed -
New grants from the Open Philanthropy Project and BERI
Listed -
Informed Machine Learning -- A Taxonomy and Survey of Integrating Knowledge into Learning Systems
Listed - Listed
- Listed
- Listed
-
Historic trends in particle accelerator performance
Listed -
A Concrete Proposal for Adversarial IDA
Listed - Listed
- Listed
-
The LogBarrier adversarial attack: making effective use of decision boundary information
Listed -
Visualizing memorization in RNNs
Listed - Listed
-
Improving Safety in Reinforcement Learning Using Model-Based Architectures and Human Intervention
Listed - Listed
- Listed
-
Towards Characterizing Divergence in Deep Q-Learning
Listed - Listed
-
What's wrong with these analogies for understanding Informed Oversight and IDA?
Listed -
Semantic Image Synthesis with Spatially-Adaptive Normalization
Listed - Listed
-
Boeing 737 MAX MCAS as an agent corrigibility failure
Listed -
Comparison of decision theories (with a focus on logical-counterfactual decision theories)
Listed -
Algorithms for Verifying Deep Neural Networks
Listed -
Eric Drexler: Paretotopian goal alignment
Listed -
Humans aren't agents - what then for value learning?
Listed - Listed
-
Deep Reinforcement Learning with Feedback-based Exploration
Listed - Listed
-
Question: MIRI Corrigbility Agenda
Listed - Listed
- Listed
-
Applications are open for the MIRI Summer Fellows Program!
Listed -
Designing agent incentives to avoid side effects
Listed -
Example population ethics: ordered discounted utility
Listed - Listed
-
Alignment Research Field Guide
Listed -
CSER Advice to EU High-Level Expert Group on AI
Listed -
CSER and FHI advice to UN High-level Panel on Digital Cooperation
Listed -
Smoothmin and personal identity
Listed - Listed
-
Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples
Listed - Listed
- Listed
- Listed
-
Historical economic growth trends
Listed -
Learning Exploration Policies for Navigation
Listed -
Learning Latent Plans from Play
Listed -
Simplified preferences needed; simplified preferences sufficient
Listed -
Stabilizing the Lottery Ticket Hypothesis
Listed -
Three ways that "Sufficiently optimized agents appear coherent" can be false
Listed -
Using Natural Language for Reward Shaping in Reinforcement Learning
Listed -
A Strongly Asymptotically Optimal Agent in General Environments
Listed - Listed
-
Amanda Askell: AI safety needs social scientists
Listed - Listed
-
IRL 1/8: Inverse Reinforcement Learning and the problem of degeneracy
Listed -
Model Primitive Hierarchical Lifelong Reinforcement Learning
Listed -
Using Causal Analysis to Learn Specifications from Task Demonstrations
Listed -
Hacking Google reCAPTCHA v3 using Reinforcement Learning
Listed - Listed
-
Learning Robust Representations by Projecting Superficial Statistics Out
Listed -
Primates vs birds: Is one brain architecture better than the other?
Listed - Listed
-
The Ethics of AI Ethics -- An Evaluation of Guidelines
Listed - Listed
-
Diagnosing Bottlenecks in Deep Q-learning Algorithms
Listed -
How to get value learning and reference wrong
Listed - Listed
-
Understanding Agent Incentives using Causal Influence Diagrams. Part I: Single Action Settings
Listed -
Challenges for an Ontology of Artificial Intelligence
Listed - Listed
- Listed
-
Improving Robustness of Machine Translation with Synthetic Noise
Listed -
Verification of Non-Linear Specifications for Neural Networks
Listed -
Can HCH epistemically dominate Ramanujan?
Listed - Listed
- Listed
-
What type of Master's is best for AI policy work?
Listed -
Confused about AI research as a means of addressing AI risk
Listed -
FHI Report: Stable Agreements in Turbulent Times
Listed -
Quantifying Perceptual Distortion of Adversarial Examples
Listed - Listed
-
From Language to Goals: Inverse Reinforcement Learning for Vision-Based Instruction Following
Listed -
Meta-Weight-Net: Learning an Explicit Mapping For Sample Weighting
Listed - Listed
- Listed
-
AI Safety Needs Social Scientists
Listed -
Parenting: Safe Reinforcement Learning from Human Input
Listed -
Regularizing Black-box Models for Improved Interpretability
Listed -
STRIP: A Defence Against Trojan Attacks on Deep Neural Networks
Listed -
Was ist eine Professur fuer Kuenstliche Intelligenz?
Listed -
Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
Listed -
How does OpenAI's language model affect our AI timeline estimates?
Listed -
How the MtG Color Wheel Explains AI Safety
Listed - Listed
-
Fixed-point solutions to the regress problem in normative uncertainty
Listed -
Unsupervised Visuomotor Control through Distributional Planning Networks
Listed -
Three Biases That Made Me Believe in AI Risk
Listed -
Deep Reinforcement Learning from Policy-Dependent Human Feedback
Listed -
Learning preferences by looking at the world
Listed -
Nuances with ascription universality
Listed -
Preferences Implicit in the State of the World
Listed -
Coherent behaviour in the real world is an incoherent concept
Listed - Listed
-
Learning Preferences by Looking at the World
Listed - Listed
-
Would I think for ten thousand years?
Listed -
Some Thoughts on Metaphilosophy
Listed -
The Argument from Philosophical Difficulty
Listed -
Ben Garfinkel: How sure are we about this AI stuff?
Listed -
HCH is not just Mechanical Turk
Listed -
Reinforcement Learning in the Iterated Amplification Framework
Listed -
Ask Not What AI Can Do, But What AI Should Do: Towards a Framework of Task Delegability
Listed -
Certified Adversarial Robustness via Randomized Smoothing
Listed -
Evidence on good forecasting practices from the Good Judgment Project
Listed -
Evidence on good forecasting practices from the Good Judgment Project: an accompanying blog post
Listed -
Hybrid Models with Deep and Invertible Features
Listed - Listed
- Listed
-
Test Cases for Impact Regularisation Methods
Listed - Listed
-
PUTWorkbench: Analysing Privacy in AI-intensive Systems
Listed - Listed
-
(notes on) Policy Desiderata for Superintelligent AI: A Vector Field Approach
Listed -
Conclusion to the sequence on value learning
Listed - Listed
- Listed
-
How does Gradient Descent Interact with Goodhart?
Listed - Listed
- Listed
-
The Hanabi Challenge: A New Frontier for AI Research
Listed -
Human-Centered Artificial Intelligence and Machine Learning
Listed - Listed
-
The case for building expertise to work on US AI policy, and how to do it
Listed -
A Comparative Analysis of Expected and Distributional Reinforcement Learning
Listed -
Deconfusing Logical Counterfactuals
Listed -
Wireheading is in the eye of the beholder
Listed - Listed
-
Can there be an indescribable hellworld?
Listed -
How much can value learning be disentangled?
Listed -
Which textbook would you recommend to learn decision theory?
Listed -
Lyapunov-based Safe Policy Optimization for Continuous Control
Listed -
Techniques for optimizing worst-case performance
Listed -
Using Pre-Training Can Improve Model Robustness and Uncertainty
Listed -
Epistemic Therapy for Bias in Automated Decision-Making
Listed -
Specifying AI Objectives As a Human-AI Collaboration Problem
Listed -
Future directions for narrow value learning
Listed -
Forecasting Transformative AI: An Expert Survey
Listed - Listed
-
Is Agent Simulates Predictor a "fair" problem?
Listed -
Theoretically Principled Trade-off between Robustness and Accuracy
Listed -
Thoughts on reward engineering
Listed -
Thoughts on reward engineering
Listed -
Allowing a formal proof system to self improve while avoiding Lobian obstacles.
Listed -
Disentangling arguments for the importance of AI safety
Listed - Listed
-
S-Curves for Trend Forecasting
Listed - Listed
-
Disentangling arguments for the importance of AI safety
Listed -
Announcement: AI alignment prize round 4 winners
Listed - Listed
- Listed
- Listed
- Listed
-
Theory of Minds: Understanding Behavior in Groups Through Inverse Planning
Listed - Listed
-
Amplifying the Imitation Effect for Reinforcement Learning of UCAV's Mission Execution
Listed - Listed
- Listed
-
Debate AI and the Decision to Release an AI
Listed -
The reward engineering problem
Listed -
The reward engineering problem
Listed -
What AI Safety Researchers Have Written About the Nature of Human Values
Listed -
Artificial Intelligence and Robotization
Listed - Listed
-
Identifying and Correcting Label Bias in Machine Learning
Listed - Listed
-
Directions and desiderata for AI alignment
Listed - Listed
-
Towards formalizing universality
Listed - Listed
-
Ambitious vs. narrow value learning
Listed - Listed
-
Non-Consequentialist Cooperation?
Listed -
Towards formalizing universality
Listed -
A New Tensioning Method using Deep Reinforcement Learning for Surgical Pattern Cutting
Listed -
Universality and consequentialism within HCH
Listed -
What is narrow value learning?
Listed -
AlphaGo Zero and capability amplification
Listed - Listed
-
No surjection onto function space for manifold X
Listed - Listed
-
Reframing Superintelligence: Comprehensive AI Services as General Intelligence
Listed -
Risk-Aware Active Inverse Reinforcement Learning
Listed - Listed
-
AI safety without goal-directed behavior
Listed - Listed
-
Failures of UDT-AIXI, Part 1: Improper Randomizing
Listed -
Supervising strong learners by amplifying weak experts
Listed -
Hierarchical Reinforcement Learning via Advantage-Weighted Information Maximization
Listed