The catalog, page 35
Records 8,501 to 8,750 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
- Listed
-
SmartChoices: Hybridizing Programming and Machine Learning
Listed - Listed
-
EDT solves 5 and 10 with conditional oracles
Listed -
Few-Shot Goal Inference for Visuomotor Learning and Planning
Listed - Listed
- Listed
-
Stakeholders in Explainable AI
Listed -
Training Machine Learning Models by Regularizing their Explanations
Listed -
On the (in)applicability of corporate rights cases to digital minds
Listed -
Asymptotic Decision Theory (Improved Writeup)
Listed -
Building safe artificial intelligence: specification, robustness, and assurance
Listed -
Few-Shot Intent Inference via Meta-Inverse Reinforcement Learning
Listed -
Inferring Reward Functions from Demonstrators with Unknown Biases
Listed -
New DeepMind AI Safety Research Blog
Listed -
Uncovering Surprising Behaviors in Reinforcement Learning via Worst-case Analysis
Listed -
Adding Neural Network Controllers to Behavior Trees without Destroying Performance Guarantees
Listed - Listed
-
Wireheading as a potential problem with the new impact measure
Listed - Listed
- Listed
-
Reflective AIXI and Anthropics
Listed -
(Some?) Possible Multi-Agent Goodhart Interactions
Listed -
Interpretable Multi-Objective Reinforcement Learning through Policy Orchestration
Listed -
In Logical Time, All Games are Iterated Games
Listed -
Playing the Game of Universal Adversarial Perturbations
Listed -
Quantum theory cannot consistently describe the use of itself
Listed -
Bridging syntax and semantics, empirically
Listed -
Interpretable Reinforcement Learning with Ensemble Methods
Listed -
TStarBots: Defeating the Cheating Level Builtin AI in StarCraft II in the Full Game
Listed - Listed
-
Adversarial Imitation via Variational Inverse Reinforcement Learning
Listed - Listed
- Listed
-
Towards Better Interpretability in Deep Q-Networks
Listed -
Model-Based Reinforcement Learning via Meta-Policy Optimization
Listed -
CM3: Cooperative Multi-goal Multi-stage Multi-agent Reinforcement Learning
Listed -
Introducing the Unrestricted Adversarial Examples Challenge
Listed - Listed
- Listed
- Listed
- Listed
- Listed
-
Disagreement with Paul: alignment induction
Listed -
Expert-augmented actor-critic for ViZDoom and Montezumas Revenge
Listed - Listed
- Listed
-
Training for Faster Adversarial Robustness Verification via Inducing ReLU Stability
Listed -
Neural Guided Constraint Logic Programming for Program Synthesis
Listed -
Learning Invariances for Policy Generalization
Listed -
Challenges of Context and Time in Reinforcement Learning: Introducing Space Fortress as a Benchmark
Listed -
AI Governance: A Research Agenda
Listed -
Counterfactuals and reflective oracles
Listed -
Reinforcement Learning under Threats
Listed -
A Roadmap for Robust End-to-End Alignment
Listed -
Recurrent World Models Facilitate Policy Evolution
Listed - Listed
-
'The Hyperbolic Time Chamber & Brain Emulation'
Listed - Listed
- Listed
- Listed
- Listed
- Listed
- Listed
-
VOI is Only Nonnegative When Information is Uncorrelated With Future Action
Listed -
Do what we mean vs. do what we say
Listed -
History of the Development of Logical Induction
Listed - Listed
-
"Why Tool AIs Want to Be Agent AIs"
Listed - Listed
-
Corrigibility doesn't always have a good action to take
Listed -
Cycle-of-Learning for Autonomous Systems from Human Interaction
Listed - Listed
-
Reference Post: Formal vs. Effective Pre-Commitment
Listed -
Using expected utility for Good(hart)
Listed -
Why Self-Attention? A Targeted Evaluation of Neural Machine Translation Architectures
Listed -
The Social Cost of Strategic Classification
Listed -
Book Review: AI Safety and Security
Listed - Listed
-
Life-Long Disentangled Representation Learning with Cross-Domain Latent Homologies
Listed -
Reducing collective rationality to individual optimization in common-payoff games using MCMC
Listed - Listed
-
Multi-task Maximum Entropy Inverse Reinforcement Learning.
Listed -
Where Do You Think You’re Going?: Inferring Beliefs about Dynamics from Behavior.
Listed -
A Broader View on Bias in Automated Decision-Making: Reflecting on Epistemology and Dynamics.
Listed - Listed
-
Analyzing Inverse Problems with Invertible Neural Networks
Listed -
Confidence-aware motion prediction for real-time collision avoidance.
Listed -
Cost Functions for Robot Motion Style.
Listed - Listed
- Listed
- Listed
-
Evaluating the Stability of Non-Adaptive Trading in Continuous Double Auctions.
Listed -
Expert, Crowdsourced, and Machine Assessment of Suicide Risk via Online Postings.
Listed -
Expressing Robot Incapability.
Listed -
Learning from Physical Human Corrections, One Feature at a Time.
Listed -
Learning from Richer Human Guidance: Augmenting Comparison-Based Learning with Feature Queries.
Listed -
Learning Human Ergonomic Preferences for Handovers.
Listed -
Logical Counterfactuals & the Cooperation Game
Listed -
Many-Goals Reinforcement Learning.
Listed -
Minimax-regret querying on side effects for safe optimality in factored Markov decision processes.
Listed -
On Learning Intrinsic Rewards for Policy Gradient Methods.
Listed - Listed
- Listed
-
Rational metareasoning and the plasticity of cognitive control.
Listed -
Sensitivity to Shared Information in Social Learning.
Listed -
Social Cohesion in Autonomous Driving.
Listed -
SoK: Security and Privacy in Machine Learning.
Listed -
Solomon’s Code: Humanity in a World with Thinking Machines.
Listed -
Using Trusted Data to Train Deep Networks on Labels Corrupted by Severe Noise.
Listed -
Directed Policy Gradient for Safe Reinforcement Learning with Human Advice
Listed -
Risk-Sensitive Generative Adversarial Imitation Learning
Listed -
Building Safer AGI by introducing Artificial Stupidity
Listed -
A Note on the Existence of Ratifiable Acts.
Listed -
A Regression Approach for Modeling Games with Many Symmetric Players.
Listed -
Combining the Causal Judgments of Experts with Possibly Different Focus Areas.
Listed -
Empirical evidence for resource-rational anchoring and adjustment.
Listed -
Estimation with Incomplete Data: The Linear Case.
Listed - Listed
-
Is state-dependent valuation more adaptive than simpler rules?.
Listed -
Negotiable Reinforcement Learning for Pareto Optimal Sequential Decision-Making.
Listed -
On Handling Self-masking and Other Hard Missing Data Problems.
Listed -
Open Category Detection with PAC Guarantees.
Listed -
The anchoring bias reflects rational use of cognitive resources.
Listed - Listed
- Listed
-
Moral realism and AI alignment
Listed - Listed
-
Generalization Error in Deep Learning
Listed -
Learning Actionable Representations from Visual Observations
Listed - Listed
-
Sandboxing by Physical Simulation?
Listed -
Counterfactuals, thick and thin
Listed -
Safely and usefully spectating on AIs optimizing over toy worlds
Listed -
Security and Privacy Issues in Deep Learning
Listed -
Techniques for Interpretable Machine Learning
Listed - Listed
-
Reinforced Auto-Zoom Net: Towards Accurate and Fast Breast Cancer Segmentation in Whole-slide Images
Listed -
A Gym Gridworld Environment for the Treacherous Turn
Listed -
Decisions are not about changing the world, they are about learning what world you live in
Listed -
TensorFuzz: Debugging Neural Networks with Coverage-Guided Fuzzing
Listed - Listed
-
Evaluating and Understanding the Robustness of Adversarial Logit Pairing
Listed -
Multi-Agent Generative Adversarial Imitation Learning
Listed -
Variational Option Discovery Algorithms
Listed -
Differentiable Image Parameterizations
Listed - Listed
- Listed
- Listed
- Listed
-
Learning Plannable Representations with Causal InfoGAN
Listed -
Alignment Newsletter #16: 07/23/18
Listed -
Contrastive Explanations for Reinforcement Learning in terms of Expected Consequences
Listed -
Let's Discuss Functional Decision Theory
Listed -
EnsembleDAgger: A Bayesian Approach to Safe Imitation Learning
Listed -
Safe Option-Critic: Learning Safety in the Option-Critic Architecture
Listed -
Stable Pointers to Value III: Recursive Quantilization
Listed -
Can few-shot learning teach AI right from wrong?
Listed -
Knowledge Integration for Disease Characterization: A Breast Cancer Example
Listed -
Learning Heuristics for Quantified Boolean Formulas through Deep Reinforcement Learning
Listed -
Probability is Real, and Value is Complex
Listed - Listed
-
Backplay: "Man muss immer umkehren"
Listed -
Generative Adversarial Imitation from Observation
Listed -
Interpretable Latent Spaces for Learning from Demonstration
Listed -
Alignment Newsletter #15: 07/16/18
Listed -
Buridan's ass in coordination games
Listed - Listed
- Listed
-
Introducing Quantum-Like Influence Diagrams for Violations of the Sure Thing Principle
Listed -
Meta-Learning with Latent Embedding Optimization
Listed -
Safe Reinforcement Learning via Probabilistic Shields
Listed -
Announcement: AI alignment prize round 3 winners and next round
Listed - Listed
-
Exploring Hierarchy-Aware Inverse Reinforcement Learning
Listed -
Minimax-regret querying on side effects for safe optimality in factored Markov decision processes
Listed -
Model Reconstruction from Model Explanations
Listed -
An Agent is a Worldline in Tegmark V
Listed -
Historic trends in structure heights
Listed -
The Bottleneck Simulator: A Model-based Deep Reinforcement Learning Approach
Listed -
Visual Reinforcement Learning with Imagined Goals
Listed -
A comment on the IDA-AlphaGoZero metaphor; capabilities versus alignment
Listed -
Agents That Learn From Human Behavior Can't Learn Human Values That Humans Haven't Learned Yet
Listed -
An environment for studying counterfactuals
Listed -
Are pre-specified utility functions about the real world possible in principle?
Listed - Listed
-
Clarifying Consequentialists in the Solomonoff Prior
Listed -
Complete Class: Consequentialist Foundations
Listed -
Conceptual problems with utility functions
Listed -
Conditions under which misaligned subagents can (not) arise in classifiers
Listed -
Decision-theoretic problems and Theories; An (Incomplete) comparative list
Listed - Listed
-
Mechanistic Transparency for Machine Learning
Listed -
No, I won't go there, it feels like you're trying to Pascal-mug me
Listed -
On the Role of Counterfactuals in Learning
Listed -
A Game-Based Approximate Verification of Deep Neural Networks with Provable Guarantees
Listed -
A Simple Unified Framework for Detecting Out-of-Distribution Samples and Adversarial Attacks
Listed -
Announcing AlignmentForum.org Beta
Listed -
Bayesian Probability is for things that are Space-like Separated from You
Listed -
Conditioning, Counterfactuals, Exploration, and Gears
Listed -
Interpreting AI compute trends
Listed -
Logical Uncertainty and Functional Decision Theory
Listed -
Probability is fake, frequency is real
Listed -
Repeated (and improved) Sleeping Beauty problem
Listed -
Representation Learning with Contrastive Predictive Coding
Listed - Listed
- Listed
- Listed
-
Troubling Trends in Machine Learning Scholarship
Listed - Listed
-
Benchmarking Neural Network Robustness to Common Corruptions and Surface Variations
Listed -
Ranked Reward: Enabling Self-Play Reinforcement Learning for Combinatorial Optimization
Listed -
The Learning-Theoretic AI Alignment Research Agenda
Listed -
The Prediction Problem: A Variant on Newcomb's
Listed -
Intertheoretic utility comparison
Listed -
Alignment Newsletter #13: 07/02/18
Listed -
Accurate Uncertainties for Deep Learning Using Calibrated Regression
Listed -
Another take on agent foundations: formalizing zero-shot reasoning
Listed - Listed
- Listed
-
Machine learning 2.0 : Engineering Data Driven AI Products
Listed -
Minimax-Regret Querying on Side Effects for Safe Optimality in Factored Markov Decision Processes
Listed - Listed
-
The Facets of Artificial Intelligence: A Framework to Track the Evolution of AI
Listed -
Towards Mixed Optimization for Reinforcement Learning with Program Synthesis
Listed - Listed
-
Overcoming Clinginess in Impact Measures
Listed - Listed
-
“Cheating Death in Damascus” Solution to the Fermi Paradox
Listed -
A Benchmark for Interpretability Methods in Deep Neural Networks
Listed -
Adversarial Reprogramming of Neural Networks
Listed -
Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation
Listed -
New paper: “Forecasting using incomplete models”
Listed - Listed
-
Adversarial Active Exploration for Inverse Dynamics Model Learning
Listed -
Learning Existing Social Conventions via Observationally Augmented Self-Play
Listed -
Logical uncertainty and Mathematical uncertainty
Listed -
Multi-agent Inverse Reinforcement Learning for Certain General-sum Stochastic Games
Listed -
The Alignment Newsletter #12: 06/25/18
Listed -
DARTS: Differentiable Architecture Search
Listed -
UDT can learn anthropic probabilities
Listed - Listed
-
On Adversarial Examples for Character-Level Neural Machine Translation
Listed -
Human-Interactive Subgoal Supervision for Efficient Inverse Reinforcement Learning
Listed -
The Foundations of Deep Learning with a Path Towards General Intelligence
Listed -
Interpretable Discovery in Large Image Data Sets
Listed -
Interpretable to Whom? A Role-based Model for Analyzing Interpretable Machine Learning Systems
Listed -
RUDDER: Return Decomposition for Delayed Rewards
Listed -
A Survey of Inverse Reinforcement Learning: Challenges, Methods and Progress
Listed -
The Alignment Newsletter #11: 06/18/18
Listed