The catalog, page 37
Records 9,001 to 9,250 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
- Listed
-
A model I use when making plans to reduce AI x-risk
Listed -
Beware of black boxes in AI alignment research
Listed -
Symmetric Decomposition of Asymmetric Games
Listed -
Towards an Integrated Assessment of Global Catastrophic Risk
Listed -
Announcement: AI alignment prize winners and next round
Listed -
Counterfactual equivalence for POMDPs, and underlying deterministic environments
Listed - Listed
-
Spatially Transformed Adversarial Examples
Listed - Listed
-
Global online debate on the governance of AI
Listed - Listed
- Listed
-
A Rational Reinterpretation of Dual-Process Theories
Listed -
Adversarial Examples Are a Natural Consequence of Test Error in Noise
Listed - Listed
- Listed
-
Artificial General Intelligence: Coordination & Great Powers
Listed -
Artificial General Intelligence: Coordination and Great Powers
Listed -
Certified Defenses against Adversarial Examples
Listed -
Countering Superintelligence Misinformation
Listed - Listed
-
GLOBAL POLITICS AND THE GOVERNANCE OF ARTIFICIAL INTELLIGENCE
Listed -
Hacking the brain: dimensions of cognitive enhancement
Listed -
How rapidly are GPUs improving in price performance?
Listed - Listed
-
Insight-based AI timelines model
Listed - Listed
-
Motivating the Rules of the Game for Adversarial Example Research
Listed -
Negotiable Reinforcement Learning for Pareto Optimal Sequential Decision-Making
Listed -
Occam's razor is insufficient to infer the preferences of irrational agents
Listed -
On Calibration of Modern Neural Networks
Listed -
Predicting Human Deliberative Judgments with Machine Learning
Listed -
Public Policy and Superintelligent AI: A Vector Field Approach
Listed -
Reconciliation between factions focused on near-term and long-term artificial intelligence
Listed -
Superintelligence skepticism as a political tool
Listed - Listed
-
The new weapons of mass destruction?
Listed -
The State of Research in Existential Risk
Listed -
The vulnerable world hypothesis
Listed -
Toward A Working Theory of Mind
Listed -
Where Do You Think You're Going?: Inferring Beliefs about Dynamics from Behavior
Listed - Listed
- Listed
-
The Three Levels of Goodhart's Curse
Listed -
Effect of marginal hardware on artificial general intelligence
Listed -
Artificial Intelligence in Life Extension: from Deep Learning to Superintelligence
Listed -
Conceptual-Linguistic Superintelligence
Listed -
Superintelligence As a Cause or Cure For Risks of Astronomical Suffering
Listed -
2017 AI Safety Literature Review and Charity Comparison
Listed - Listed
- Listed
-
Indifference' methods for managing agent rewards
Listed - Listed
-
A Berkeley View of Systems Challenges for AI
Listed -
End-of-the-year matching challenge!
Listed -
Occam's razor is insufficient to infer the preferences of irrational agents
Listed -
Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
Listed -
Against the Linear Utility Hypothesis and the Leverage Penalty
Listed -
Three IQs of AI Systems and their Testing Methods
Listed - Listed
- Listed
-
A Low-Cost Ethics Shaping Approach for Designing Reinforcement Learning Agents
Listed - Listed
- Listed
-
Safety models and accident models
Listed - Listed
-
A reply to Francois Chollet on intelligence explosion
Listed - Listed
-
Using Artificial Intelligence to Augment Human Intelligence
Listed -
Implementation of Moral Uncertainty in Intelligent Machines
Listed - Listed
-
Policy Selection Solves Most Problems
Listed -
GoCAS talk on AI Impacts findings
Listed - Listed
-
Price performance Moore’s Law seems slow
Listed - Listed
-
Security Mindset and the Logistic Success Curve
Listed -
Security Mindset and Ordinary Paranoia
Listed - Listed
- Listed
- Listed
- Listed
-
Evaluating Robustness of Neural Networks with Mixed Integer Programming
Listed -
Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
Listed -
Announcing “Inadequate Equilibria”
Listed - Listed
-
Using KL-divergence to focus Deep Visual Explanation
Listed -
Good and safe uses of AI Oracles
Listed -
Fixing Weight Decay Regularization in Adam
Listed -
Rationalising humans: another mugging, but not Pascal's
Listed -
Military AI as a Convergent Goal of Self-Improving AI
Listed -
2017 trend in the cost of computing
Listed -
A Survey of Artificial General Intelligence Projects for Ethics, Risk, and Policy
Listed -
A major grant from the Open Philanthropy Project
Listed -
Price-performance trend in top supercomputers
Listed - Listed
- Listed
-
On the promotion of safe and socially beneficial artificial intelligence
Listed -
A Foundry of Human Activities and Infrastructures
Listed -
Mixed-Strategy Ratifiability Implies CDT=EDT
Listed -
Servant of Many Masters: Shifting priorities in Pareto-optimal sequential decision-making
Listed -
Learning Robust Rewards with Adversarial Inverse Reinforcement Learning
Listed - Listed
-
Logical Updatelessness as a Robust Delegation Problem
Listed -
Computing hardware performance data collections
Listed - Listed
-
Humans can be assigned any values whatsoever...
Listed -
Human-in-the-loop Artificial Intelligence
Listed -
A behaviorist approach to building phenomenological bridges
Listed -
New paper: “Functional Decision Theory”
Listed -
What Evidence Is AlphaGo Zero Re AGI Complexity?
Listed -
AlphaGo Zero and the Foom Debate
Listed -
AlphaGo Zero and the Foom Debate
Listed -
AlphaGo Zero and capability amplification
Listed -
Functional Decision Theory: A New Theory of Instrumental Rationality
Listed -
Decision Trees for Helpdesk Advisor Graphs
Listed - Listed
- Listed
- Listed
-
There's No Fire Alarm for Artificial General Intelligence
Listed -
There’s No Fire Alarm for Artificial General Intelligence
Listed -
Functional Decision Theory: A New Theory of Instrumental Rationality
Listed -
Robot Sex: Social and Ethical Implications
Listed -
There's No Fire Alarm for Artificial General Intelligence
Listed -
Consequence assessment: Estimating the impact of accident scenarios
Listed -
Toy model of the AI control problem: animated version
Listed -
Delegative Reinforcement Learning with a Merely Sane Advisor
Listed -
2016 ESPAI Narrow AI task forecast timeline
Listed - Listed
-
Global Catastrophes: The Most Extreme Risks
Listed -
An intervention to shape policy dialogue, communication, and AI research norms for AI safety
Listed - Listed
-
How feasible is the rapid development of artificial superintelligence?
Listed -
Deep TAMER: Interactive Agent Shaping in High-Dimensional State Spaces
Listed -
What do ML researchers think you are wrong about?
Listed -
When do ML Researchers Think Specific Tasks will be Automated?
Listed - Listed
- Listed
-
Autonomous Agents Modelling Other Agents: A Comprehensive Survey and Open Problems
Listed -
Naturalized induction – a challenge for evidential and causal decision theory
Listed -
A Voting-Based System for Ethical Decision Making
Listed -
Incorrigibility in the CIRL Framework
Listed -
DropoutDAgger: A Bayesian Approach to Safe Imitation Learning
Listed -
A Learning and Masking Approach to Secure Learning
Listed -
Automation of music production
Listed -
Aggregating incoherent agents who disagree
Listed -
Stuart Russell’s description of AI risk
Listed -
Knowledge Transfer Between Artificial Intelligence Systems
Listed -
New paper: “Incorrigibility in the CIRL Framework”
Listed - Listed
-
Why Does Deep and Cheap Learning Work So Well?
Listed -
The Doomsday argument in anthropic decision theory
Listed - Listed
-
Safe Reinforcement Learning via Shielding
Listed -
Using modal fixed points to formalize logical causality
Listed -
Life 3.0: Being Human in the Age of Artificial Intelligence (2017, Alfred A. Knopf)
Listed -
BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
Listed -
Logical Induction with incomputable sequences
Listed -
On Ensuring that Intelligent Machines Are Well-Behaved
Listed -
Stable Pointers to Value: An Agent Embedded in Its Own Utility Function
Listed - Listed
-
Portfolio approach to AI safety research
Listed - Listed
- Listed
-
Causality, Responsibility and Blame in Team Plans.
Listed -
Comparing Human-Centric and Robot-Centric Sampling for Robot Deep Learning from Demonstrations.
Listed -
Computational Extensive-Form Games.
Listed -
Do You Want Your Autonomous Car to Drive Like You?.
Listed -
Expressive Robot Motion Timing.
Listed -
Modeling Agents with Probabilistic Programs.
Listed -
Repeated Inverse Reinforcement Learning.
Listed -
Self-confirming price-prediction strategies for simultaneous one-shot auctions.
Listed -
The Computational Complexity of Structure-Based Causality.
Listed -
Toward a Rational and Mechanistic Account of Mental Effort.
Listed - Listed
-
Potential Risks from Advanced AI
Listed -
What does (and doesn't) AI mean for effective altruism?
Listed -
Daniel Dewey: The Open Philanthropy Project's work on potential risks from advanced AI
Listed -
Jan Leike, Helen Toner, Malo Bourgon, and Miles Brundage: Working in AI
Listed - Listed
- Listed
-
Owen Cotton-Barratt: What does (and doesn't) AI mean for effective altruism?
Listed -
Active Preference-Based Learning of Reward Functions.
Listed -
An automatic method for discovering rational heuristics for risky choice.
Listed -
Enhancing metacognitive reinforcement learning using reward structures and feedback.
Listed -
Multiverse-wide Cooperation via Correlated Decision Making
Listed - Listed
-
The evolution of cognitive mechanisms in response to cultural innovations.
Listed -
The Structure of Goal Systems Predicts Human Performance.
Listed -
When Does Bounded-Optimal Metareasoning Favor Few Cognitive Systems?.
Listed - Listed
- Listed
-
A Formal Approach to the Problem of Logical Non-Omniscience
Listed -
Together We Know How to Achieve: An Epistemic Logic of Know-How (Extended Abstract)
Listed -
The future of growth: near-zero growth rates
Listed -
Using Program Induction to Interpret Transition System Dynamics
Listed - Listed
-
Guidelines for Artificial Intelligence Containment
Listed -
Adversarial Examples for Evaluating Reading Comprehension Systems
Listed -
Pragmatic-Pedagogic Value Alignment
Listed -
RAIL: Risk-Averse Imitation Learning
Listed - Listed
-
Open Problems Regarding Counterfactuals: An Introduction For Beginners
Listed -
Trial without Error: Towards Safe Reinforcement Learning via Human Intervention
Listed -
Delegative Inverse Reinforcement Learning
Listed -
My current thoughts on MIRI's "highly reliable agent design" work
Listed -
Updates to the research team, and a major donation
Listed -
Approval-maximizing representations
Listed -
Artificial Intelligence and Global Security Initiative Research Agenda
Listed -
Teacher-Student Curriculum Learning
Listed - Listed
- Listed
-
A survey of polls on Newcomb’s problem
Listed -
Complications in evaluating neglectedness
Listed -
Expert and Non-Expert Opinion about Technological Unemployment
Listed -
Towards Deep Learning Models Resistant to Adversarial Attacks
Listed - Listed
- Listed
-
Media discussion of 2016 ESPAI
Listed -
Device Placement Optimization with Reinforcement Learning
Listed - Listed
- Listed
- Listed
-
SSC Journal Club: AI Timelines
Listed -
Cognitive Science/Psychology As a Neglected Approach to AI Safety
Listed -
Takeaways from self-tracking data
Listed -
Cooperative Oracles: Introduction
Listed -
Cooperative Oracles: Nonexploited Bargaining
Listed -
Cooperative Oracles: Stratified Pareto Optima and Almost Stratified Pareto Optima
Listed -
Acausal trade: different utilities, different trades
Listed -
Acausal trade: double decrease
Listed -
Acausal trade: universal utility, or selling non-existence insurance too late
Listed - Listed
-
Corrigibility thoughts I: caring about multiple things
Listed -
Counterfactually uninfluenceable agents
Listed - Listed
-
The AI revolution and international politics (Allan Dafoe)
Listed - Listed
- Listed
-
Why I am not currently working on the AAMLS agenda
Listed -
The Atari Grand Challenge Dataset
Listed - Listed
-
Constrained Policy Optimization
Listed - Listed
-
Low Impact Artificial Intelligences
Listed -
Universal Reinforcement Learning Algorithms: Survey and Experiments
Listed -
The Technological Singularity: Managing the Journey
Listed - Listed
-
Existential risk from AI without an intelligence explosion
Listed