The catalog, page 34
Records 8,251 to 8,500 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Will humans build goal-directed agents?
Listed -
Two More Decision Theory Problems for Humans
Listed -
A Comprehensive Survey on Graph Neural Networks
Listed -
Logical inductors in multistable situations.
Listed -
Universality and security amplification
Listed - Listed
- Listed
- Listed
- Listed
-
Approval-directed agency and the decision theory of Newcomb-like problems
Listed -
Artificial Intelligence: American Attitudes and Trends
Listed -
Bridging near- and long-term concerns about AI
Listed -
Digital Authoritarianism: Evolving Chinese And Russian Models
Listed -
Failure Modes in Machine Learning - Security documentation
Listed -
Feasibility of Training an AGI using Deep RL: A Very Rough Estimate
Listed -
How Useful Is Quantilization For Mitigating Specification-Gaming?
Listed -
Learning Reward Machines for Partially Observable Reinforcement Learning
Listed -
Lessons for Artificial Intelligence from Other Global Risks
Listed -
Long-term trajectories of human civilization
Listed -
Machine Learning Projects for Iterated Distillation and Amplification
Listed -
ObjectNet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
Listed -
ObjectNet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
Listed -
Optimization Regularization through Time Penalty
Listed - Listed
-
Reframing Superintelligence: Comprehensive AI Services as General Intelligence
Listed - Listed
- Listed
-
Stable Agreements in Turbulent Times: A Legal Toolkit for Constrained Temporal Decision Transmission
Listed - Listed
-
Surveying Safety-relevant AI Characteristics
Listed - Listed
- Listed
-
The Role and Limits of Principles in AI Ethics: Towards a Focus on Tensions
Listed -
There is plenty of time at the bottom: the economics, risk and ethics of time compression
Listed - Listed
-
Why I expect successful (narrow) alignment
Listed -
Penalizing Impact via Attainable Utility Preservation
Listed -
Reconciling modern machine learning practice and the bias-variance trade-off
Listed -
An AI Race for Strategic Advantage: Rhetoric and Risks
Listed -
Learning Not to Learn: Training Deep Neural Networks with Biased Data
Listed - Listed
- Listed
-
Reinterpreting "AI and Compute"
Listed -
[Link] Center for the Governance of AI (GovAI) Annual Report 2018
Listed -
Anthropic probabilities and cost functions
Listed -
Human-AI Learning Performance in Multi-Armed Bandits
Listed -
The case for taking AI seriously as a threat to humanity
Listed -
Anthropic paradoxes transposed into Anthropic Decision Theory
Listed -
Reasons compute may not drive AI capabilities growth
Listed -
2018 AI Alignment Literature Review and Charity Comparison
Listed -
Reinterpreting “AI and Compute”
Listed - Listed
-
Announcing a new edition of “Rationality: From AI to Zombies”
Listed - Listed
-
The E-Coli Test for AI Alignment
Listed -
The limit of artificial intelligence: Can machines be rational?
Listed -
Two Neglected Problems in Human-AI Safety
Listed -
Scaling shared model governance via model splitting
Listed -
Critique of Superintelligence Part 1
Listed -
Critique of Superintelligence Part 2
Listed -
Critique of Superintelligence Part 3
Listed -
Critique of Superintelligence Part 4
Listed -
Critique of Superintelligence Part 5
Listed -
IRLAS: Inverse Reinforcement Learning for Architecture Search
Listed - Listed
- Listed
-
Linking Artificial Intelligence Principles
Listed -
Multi-agent predictive minds and AI alignment
Listed -
Assuming we've solved X, could we do Y...
Listed -
Quantum immortality: Is decline of measure compensated by merging timelines?
Listed - Listed
-
Feature Denoising for Improving Adversarial Robustness
Listed -
Photos from the first AI Safety Camp
Listed -
Photos from the second AI Safety Camp
Listed - Listed
-
Building Ethics into Artificial Intelligence
Listed -
Off-Policy Deep Reinforcement Learning without Exploration
Listed -
Toward the Engineering of Virtuous Machines
Listed -
Verification of deep probabilistic models
Listed - Listed
-
Truly Autonomous Machines Are Ethical
Listed -
Why we need a *theory* of human values
Listed - Listed
-
Learning from Extrapolated Corrections
Listed -
Rigorous Agent Evaluation: An Adversarial Approach to Uncover Catastrophic Failures
Listed -
Coherence arguments do not entail goal-directed behavior
Listed - Listed
-
Deep Learning Application in Security and Privacy -- Theory and Practice: A Position Paper
Listed -
Intuitions about goal-directed behavior
Listed -
Iterated Distillation and Amplification
Listed - Listed
-
Formal Open Problem in Decision Theory
Listed - Listed
-
Reflective oracles as a solution to the converse Lawvere problem
Listed -
The Ubiquitous Converse Lawvere Problem
Listed - Listed
-
MIRI’s newest recruit: Edward Kmett!
Listed - Listed
-
Exploring Restart Distributions
Listed - Listed
- Listed
- Listed
-
GAN Dissection: Visualizing and Understanding Generative Adversarial Networks
Listed -
Please Stop Explaining Black Box Models for High-Stakes Decisions
Listed -
Approval-directed bootstrapping
Listed - Listed
- Listed
- Listed
-
Hierarchical visuomotor control of humanoids
Listed -
Representer Point Selection for Explaining Deep Neural Networks
Listed -
Robustness via curvature regularization, and vice versa
Listed -
2018 Update: Our New Research Directions
Listed - Listed
-
Iteration Fixed Point Exercises
Listed -
Oversight of Unsafe Systems via Dynamic Safety Envelopes
Listed -
Some cruxes on impactful alternatives to AI policy work
Listed -
Time for AI to cross the human performance range in diabetic retinopathy
Listed -
New safety research agenda: scalable agent alignment via reward modeling
Listed - Listed
-
"Taking AI Risk Seriously" – Thoughts by Andrew Critch
Listed - Listed
-
Deeper Interpretability of Deep Networks
Listed -
Guiding Policies with Language via Meta-Learning
Listed -
Reinforcement Learning and Inverse Reinforcement Learning with System 1 and System 2
Listed -
Safely Probabilistically Complete Real-Time Planning and Exploration in Unknown Environments
Listed -
Scalable agent alignment via reward modeling: a research direction
Listed -
Diagonalization Fixed Point Exercises
Listed - Listed
- Listed
-
Topological Fixed Point Exercises
Listed -
Is Clickbait Destroying Our General Intelligence?
Listed - Listed
-
Economics of Human-AI Ecosystem: Value Bias and Lost Utility in Multi-Dimensional Gaps
Listed -
Embedded Agency (full-text version)
Listed -
Guiding the One-to-one Mapping in CycleGAN via Optimal Transport
Listed -
Reward learning from human preferences and demonstrations in Atari
Listed -
Switching hosting providers today, there probably will be some hiccups
Listed -
Woulda, Coulda, Shoulda: Counterfactually-Guided Policy Search
Listed -
Emergence of Addictive Behaviors in Reinforcement Learning Agents
Listed -
Natural Environment Benchmarks for Reinforcement Learning
Listed -
Acknowledging Human Preference Types to Support Value Learning
Listed - Listed
- Listed
- Listed
-
Improving Generalization for Abstract Reasoning Tasks Using Disentangled Feature Representations
Listed -
Learning Latent Dynamics for Planning from Pixels
Listed -
Future directions for ambitious value learning
Listed - Listed
- Listed
-
Formal Limitations on the Measurement of Mutual Information
Listed -
Preface to the sequence on iterated amplification
Listed -
Specification gaming examples in AI
Listed -
A generic framework for privacy preserving deep learning
Listed -
Current AI Safety Roles for Software Engineers
Listed -
Model Mis-specification and Inverse Reinforcement Learning
Listed -
A Geometric Perspective on the Transferability of Adversarial Directions
Listed - Listed
- Listed
-
Intrinsic Geometric Vulnerability of High-Dimensional Artificial Intelligence
Listed -
Learning from Demonstration in the Wild
Listed -
Stovepiping and Malicious Software: A Critical Review of AGI Containment
Listed -
Integrative Biological Simulation, Neuropsychology, and AI Safety
Listed -
Latent Variables and Model Mis-Specification
Listed -
A Closer Look at Deep Policy Gradients
Listed -
A Model for General Intelligence
Listed - Listed
-
MixTrain: Scalable Training of Verifiably Robust Neural Networks
Listed -
Solomon's Code: Humanity in a World of Thinking Machines
Listed - Listed
- Listed
- Listed
-
Humans can be assigned any values whatsoever…
Listed -
Beliefs at different timescales
Listed - Listed
- Listed
- Listed
-
When does rationality-as-search have nontrivial implications?
Listed -
A Marauder's Map of Security and Privacy in Machine Learning
Listed -
The easy goal inference problem is still hard
Listed - Listed
- Listed
- Listed
-
Discussion on the machine learning approach to AI safety
Listed -
Discussion on the machine learning approach to AI safety
Listed - Listed
-
What is ambitious value learning?
Listed - Listed
- Listed
- Listed
-
On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models
Listed -
Preface to the sequence on value learning
Listed -
The Responsibility Quantification (ResQu) Model of Human Interaction with Automation
Listed - Listed
-
Announcing the new AI Alignment Forum
Listed -
Assessing Generalization in Deep Reinforcement Learning
Listed - Listed
- Listed
-
Introducing the AI Alignment Forum (FAQ)
Listed - Listed
-
Neural Modular Control for Embodied Question Answering
Listed -
Mimetic vs Anchored Value Alignment in Artificial Intelligence
Listed -
One-Shot Hierarchical Imitation Learning of Compound Visuomotor Tasks
Listed -
Inverse reinforcement learning for video games
Listed -
Toward an AI Physicist for Unsupervised Learning
Listed - Listed
- Listed
-
Applying Deep Learning To Airbnb Search
Listed -
Do Deep Generative Models Know What They Don't Know?
Listed - Listed
- Listed
-
Safe Reinforcement Learning with Model Uncertainty Estimates
Listed -
Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning
Listed -
BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning
Listed -
Establishing Appropriate Trust via Critical States
Listed - Listed
-
O2A: One-shot Observational learning with Action vectors
Listed -
Discriminator Rejection Sampling
Listed -
Finding Options that Minimize Planning Time
Listed -
On the Future: Prospects for Humanity
Listed - Listed
-
CURIOUS: Intrinsically Motivated Modular Multi-Goal Reinforcement Learning
Listed -
Deep Imitative Models for Flexible Inference, Planning, and Control
Listed -
Factorized Machine Self-Confidence for Decision-Making Agents
Listed -
Optimizing Agent Behavior over Long Time Scales by Transporting Value
Listed -
Successor Uncertainties: Exploration and Uncertainty in Temporal Difference Learning
Listed -
Hierarchical Game-Theoretic Planning for Autonomous Vehicles
Listed -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Listed - Listed
-
Learning under Misspecified Objective Spaces
Listed -
Batch Active Preference-Based Learning of Reward Functions
Listed -
Secure Deep Learning Engineering: A Software Quality Assurance Perspective
Listed -
Some cruxes on impactful alternatives to AI policy work
Listed -
Standard ML Oracles vs Counterfactual ones
Listed - Listed
-
Ethical Reflections on Artificial Intelligence
Listed -
Fast Context Adaptation via Meta-Learning
Listed - Listed
-
Sanity Checks for Saliency Maps
Listed -
The 30-Year Cycle In The AI Debate
Listed -
PPO-CMA: Proximal Policy Optimization with Covariance Matrix Adaptation
Listed -
A Rationality Condition for CDT Is That It Equal EDT (Part 1)
Listed -
Episodic Curiosity through Reachability
Listed - Listed
-
Unsupervised Learning via Meta-Learning
Listed - Listed
- Listed
-
Near-Optimal Representation Learning for Hierarchical Reinforcement Learning
Listed - Listed
-
Reinforcement Learning with Perturbed Rewards
Listed -
Bayesian Policy Optimization for Model Uncertainty
Listed