The catalog, page 29
Records 7,001 to 7,250 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
- Listed
-
What Makes for Good Views for Contrastive Learning?
Listed -
Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text
Listed -
Learning and manipulating learning
Listed - Listed
-
Reward functions and updating assumptions can hide a multitude of sins
Listed -
The Mechanistic and Normative Structure of Agency
Listed -
Why you should minimax in two-player zero-sum games
Listed -
AI Governance Career Paths for Europeans
Listed - Listed
- Listed
-
How should AIs update a prior over human preferences?
Listed -
Language Conditioned Imitation Learning over Unstructured Data
Listed -
Overcoming Barriers to Cross-cultural Cooperation in AI Ethics and Governance
Listed -
[AN #99]: Doubling times for the efficiency of AI algorithms
Listed -
GovAI Webinars on the Governance and Economics of AI
Listed -
Planning to Explore via Self-Supervised World Models
Listed -
Simple Sensor Intentions for Exploration
Listed -
Book report: Theory of Games and Economic Behavior (von Neumann & Morgenstern)
Listed -
Critical Review of 'The Precipice': A Reassessment of the Risks of AI and Pandemics
Listed - Listed
-
Measuring the Algorithmic Efficiency of Neural Networks
Listed -
A Reinforcement Learning Potpourri
Listed -
Learning to Segment Actions from Observation and Narration
Listed -
[AN #98]: Understanding neural net training by seeing which gradients were helpful
Listed -
Maths writer/cowritter needed: how you can't distinguish early exponential from early sigmoid
Listed -
Modeling naturalized decision problems in linear logic
Listed -
Specification gaming: the flip side of AI ingenuity
Listed -
Specification gaming: the flip side of AI ingenuity
Listed -
A multi-component framework for the analysis and design of explainable artificial intelligence
Listed -
Competitive safety via gradated curricula
Listed -
Exploring Bayesian Optimization
Listed -
Writing Causal Models Like We Write Programs
Listed -
Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?
Listed - Listed
-
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Listed -
Open Loop In Natura Economic Planning
Listed - Listed
-
How does iterated amplification exceed human abilities?
Listed -
Stanford Encyclopedia of Philosophy on AI ethics and superintelligence
Listed - Listed
- Listed
-
Topological metaphysics: relating point-set topology and locale theory
Listed -
Optimising Society to Constrain Risk of War from an Artificial Superintelligence
Listed -
Reinforcement Learning with Augmented Data
Listed -
What is the alternative to intent alignment called?
Listed -
[AN #97]: Are there historical examples of large, robust discontinuities?
Listed -
Motivating Abstraction-First Decision Theory
Listed - Listed
-
Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels
Listed -
Pitfalls of learning a reward function online
Listed -
Is the Most Accurate AI the Best Teammate? Optimizing AI for Teamwork
Listed - Listed
-
Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
Listed - Listed
-
Fast Takeoff in Biological Intelligence
Listed -
What are the relative speeds of AI capabilities and AI safety?
Listed -
What makes counterfactuals comparable?
Listed -
DeepMind team on specification gaming
Listed -
Responsible AI and Its Stakeholders
Listed -
[AN #96]: Buck and I discuss/argue about AI Alignment
Listed -
A Neural Scaling Law from the Dimension of the Data Manifold
Listed -
Description vs simulated prediction
Listed - Listed
-
Problem relaxation as a tactic
Listed -
BERT-ATTACK: Adversarial Attack Against BERT Using BERT
Listed -
Databases of human behaviour and preferences?
Listed -
AI Services as a Research Paradigm
Listed -
Dark, Beyond Deep: A Paradigm Shift to Cognitive AI with Humanlike Common Sense
Listed -
Intuitions on Universal Behavior of Information at a Distance
Listed -
How do you talk about AI safety?
Listed -
Three Modern Roles for Logic in AI
Listed -
Discontinuous progress in history: an update
Listed - Listed
-
Improving Verifiability in AI Development
Listed -
Integrating Hidden Variables Improves Approximation
Listed -
Subjectifying Objectivity: Delineating Tastes in Theoretical Quantum Gravity Research
Listed -
[AN #95]: A framework for thinking about how to make AI go well
Listed - Listed
-
Aligning AI to Human Values means Picking the Right Metrics
Listed -
Database of existential risk estimates
Listed -
Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
Listed -
2019 recent trends in Geekbench score per CPU price
Listed -
Discontinuous progress in history: an update
Listed -
Precedents for economic n-year doubling before 4n-year doubling
Listed -
Resolutions of mathematical conjectures over time
Listed -
Surveys on fractional progress towards HLAI
Listed -
Trends in DRAM price per gigabyte
Listed -
"How conservative" should the partial maximisers be?
Listed -
Discontinuous progress in history: an update
Listed -
Certifiable Robustness to Adversarial State Uncertainty in Deep Reinforcement Learning
Listed -
Asymptotically Unambitious AGI
Listed -
[AN #94]: AI alignment as translation between humans and machines
Listed -
CURL: Contrastive Unsupervised Representations for Reinforcement Learning
Listed -
An Orthodox Case Against Utility Functions
Listed -
Takeaways from safety by default interviews
Listed -
TuringAdvice: A Generative and Dynamic Evaluation of Language Use
Listed -
Announcing Web-TAISU, May 13-17
Listed -
Preliminary survey of prescient actions
Listed -
Resources for AI Alignment Cartography
Listed -
Paul Christiano: Current work in AI alignment
Listed -
Robots Learning to Move like Animals
Listed -
Takeaways from safety by default interviews
Listed - Listed
- Listed
- Listed
- Listed
-
Equilibrium and prior selection problems in multipolar deployment
Listed -
Equilibrium and prior selection problems in multipolar deployment
Listed -
Interviews on plausibility of AI safety by default
Listed -
Three kinds of competitiveness
Listed -
What is the subjective experience of free will for agents?
Listed -
[AN #93]: The Precipice we’re standing at, and how we can back away from it
Listed -
AI Research with the Potential for Malicious Use: Publication Norms and Governance Considerations
Listed -
An Overview of Early Vision in InceptionV1
Listed - Listed
-
Counterfactual Multi-Agent Reinforcement Learning with Graph Convolution Communication
Listed -
How special are human brains among animal brains?
Listed -
How special are human brains among animal brains?
Listed - Listed
-
Meta-preferences two ways: generator vs. patch
Listed -
Two Alternatives to Logical Counterfactuals
Listed -
What achievements have people claimed will be warning signs for AGI?
Listed -
AI Services: Introduction v1.3
Listed -
Outperforming the human Atari benchmark
Listed -
Three Kinds of Competitiveness
Listed -
Three kinds of competitiveness
Listed -
Agent57: Outperforming the Atari Human Benchmark
Listed -
Book Review: 12 Rules For Life
Listed -
My current framework for thinking about AGI timelines
Listed -
Suphx: Mastering Mahjong with Deep Reinforcement Learning
Listed - Listed
- Listed
- Listed
- Listed
-
How important are MDPs for AGI (Safety)?
Listed -
What are the most plausible "AI Safety warning shot" scenarios?
Listed -
2019 recent trends in GPU price per FLOPS
Listed -
[AN #92]: Learning good representations with contrastive predictive coding
Listed -
An empirical investigation of the challenges of real-world reinforcement learning
Listed -
The Precipice: Existential Risk and the Future of Humanity
Listed -
Deconfusing Human Values Research Agenda v1
Listed -
[Meta] Do you want AIS Webinars?
Listed - Listed
- Listed
-
Rohin Shah_ WhatΓÇÖs been happening in AI alignment_-by EA Global Virtual 2020-date 20200321
Listed -
Abstraction = Information at a Distance
Listed - Listed
-
Thinking About Filtered Evidence Is (Very!) Hard
Listed -
[AN #91]: Concepts, implementations, problems, and a benchmark for impact measurement
Listed -
Proxy tasks and subjective measures can be misleading in evaluating explainable AI systems
Listed - Listed
-
AI Alignment Podcast: On Lethal Autonomous Weapons with Paul Scharre
Listed -
DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction
Listed -
Visualizing Neural Networks with the Grand Tour
Listed -
What are some exercises for building/generating intuitions about key disagreements in AI alignment?
Listed - Listed
-
Fast and Easy Infinitely Wide Networks with Neural Tangents
Listed -
The Conflict Between People's Urge to Punish AI and Legal Systems
Listed -
Sample Efficient Reinforcement Learning through Learning from Demonstrations in Minecraft
Listed -
[AN #90]: How search landscapes can contain self-reinforcing feedback loops
Listed - Listed
-
Visual Grounding in Video for Unsupervised Word Translation
Listed -
Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey
Listed -
Pruned Neural Networks are Surprisingly Modular
Listed -
Retrospective Analysis of the 2019 MineRL Competition on Sample Efficient Reinforcement Learning
Listed - Listed
-
Zoom In: An Introduction to Circuits
Listed -
Zoom In: An Introduction to Circuits
Listed -
Improved Baselines with Momentum Contrastive Learning
Listed -
"Other-Play" for Zero-Shot Coordination
Listed -
AutoML-Zero: Evolving Machine Learning Algorithms From Scratch
Listed -
Can ML predict the solution value for a difficult combinatorial problem?
Listed -
A critical agential account of free will, causation, and physics
Listed -
A critical agential account of free will, causation, and physics
Listed -
[AN #89]: A unifying formalism for preference learning algorithms
Listed - Listed
-
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
Listed -
Two Decades of AI4NETS-AI/ML for Data Networks: Challenges & Research Directions
Listed -
Anthropics over-simplified: it's about priors, not updates
Listed -
Cortés, Pizarro, and Afonso as Precedents for Takeover
Listed -
If I were a well-intentioned AI... IV: Mesa-optimising
Listed -
If I were a well-intentioned AI... IV: Mesa-optimising
Listed -
An Analytic Perspective on AI Alignment
Listed -
Cooperation, Conflict, and Transformative Artificial Intelligence - A Research Agenda
Listed -
Cortés, Pizarro, and Afonso as Precedents for Takeover
Listed -
Cortés, Pizarro, and Afonso as precedents for takeover
Listed -
My Updating Thoughts on AI policy
Listed -
Responsible AI—Two Frameworks for Ethical Design Practice
Listed -
Social choice ethics in artificial intelligence
Listed -
On Safety Assessment of Artificial Intelligence
Listed -
Conclusion to 'Reframing Impact'
Listed -
Efficiently Guiding Imitation Learning Agents with Human Gaze
Listed -
If I were a well-intentioned AI... III: Extremal Goodhart
Listed -
If I were a well-intentioned AI... III: Extremal Goodhart
Listed -
On Catastrophic Interference in Atari 2600 Games
Listed - Listed
-
[AN #88]: How the principal-agent literature relates to AI risk
Listed -
Attainable Utility Preservation: Scaling to Superhuman
Listed -
If I were a well-intentioned AI... II: Acting in a world
Listed -
If I were a well-intentioned AI... II: Acting in a world
Listed -
Reasons for Excitement about Impact of Impact Measure Research
Listed -
State-only Imitation with Transition Dynamics Mismatch
Listed -
Cautious Reinforcement Learning with Logical Constraints
Listed -
Generalized Hindsight for Reinforcement Learning
Listed -
If I were a well-intentioned AI... I: Image classifier
Listed -
If I were a well-intentioned AI... I: Image classifier
Listed -
Rethinking Bias-Variance Trade-off for Generalization of Neural Networks
Listed - Listed
-
Dividing the Ontology Alignment Task with Semantic Embeddings and Logic-based Modules
Listed -
How Low Should Fruit Hang Before We Pick It?
Listed - Listed
-
Other versions of "No free lunch in value learning"
Listed -
Rewriting History with Inverse RL: Hindsight Inference for Policy Improvement
Listed -
TanksWorld: A Multi-Agent Environment for AI Safety Research
Listed -
Subagents and impact measures, full and fully illustrated
Listed - Listed
-
Neuron Shapley: Discovering the Responsible Neurons
Listed -
Attainable Utility Preservation: Empirical Results
Listed -
The Pragmatic Turn in Explainable Artificial Intelligence (XAI)
Listed -
Unsupervised Question Decomposition for Question Answering
Listed - Listed
-
Safe Imitation Learning via Fast Bayesian Reward Inference from Preferences
Listed -
Will AI undergo discontinuous progress?
Listed -
A Road Map to Strong Intelligence
Listed -
Curiosity Killed the Cat and the Asymptotically Optimal Agent
Listed -
Goal-directed = Model-based RL?
Listed -
Tessellating Hills: a toy model for demons in imperfect search
Listed -
The Problem with Metrics is a Fundamental Problem for AI
Listed -
[AN #87]: What might happen as deep learning scales even further?
Listed -
Estimating Training Data Influence by Tracing Gradient Descent
Listed -
On unfixably unsafe AGI architectures
Listed -
Counterfactuals versus the laws of physics
Listed - Listed
-
Appendix: mathematics of indexical impact measures
Listed -
Attainable Utility Preservation: Concepts
Listed -
On the falsifiability of hypercomputation, part 2: finite input streams
Listed -
Does iterated amplification tackle the inner alignment problem?
Listed -
Reference Post: Trivial Decision Theory Problem
Listed -
The Archimedean trap: Why traditional reinforcement learning will probably not yield AGI
Listed -
Analyzing Differentiable Fuzzy Logic Operators
Listed -
Bayesian Evolving-to-Extinction
Listed -
Distinguishing definitions of takeoff
Listed -
The Catastrophic Convergence Conjecture
Listed -
The Next Decade in AI: Four Steps Towards Robust Artificial Intelligence
Listed -
The Reasonable Effectiveness of Mathematics or: AI vs sandwiches
Listed -
A Simple Framework for Contrastive Learning of Visual Representations
Listed - Listed
-
My personal cruxes for working on AI safety
Listed - Listed