The catalog, page 25
Records 6,001 to 6,250 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Preferences and biases, the information argument
Listed -
Replacing Rewards with Examples: Example-Based Policy Search via Recursive Classification
Listed -
Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of Pessimism
Listed -
Combining Reward Information from Multiple Sources
Listed -
Generalizing POWER to multi-agent games
Listed - Listed
-
Fisherian Runaway as a decision-theoretic problem
Listed -
Introducing The Nonlinear Fund: AI Safety research, incubation, and funding
Listed -
[AN #142]: The quest to understand a network well enough to reimplement it by hand
Listed - Listed
- Listed
- Listed
-
Comments on "The Singularity is Nowhere Near"
Listed -
Lyapunov Barrier Policy Optimization
Listed -
Using AI ethically to tackle covid-19
Listed -
AI x-risk reduction: why I chose academia over industry
Listed -
Success Weighted by Completion Time: A Dynamics-Aware Evaluation Criteria for Embodied Navigation
Listed - Listed
- Listed
-
Towards Risk Modeling for Collaborative AI
Listed -
Behavioral Sufficient Statistics for Goal-Directedness
Listed -
Four Motivations for Learning Normativity
Listed -
Resolutions to the Challenge of Resolving Forecasts
Listed -
Symbolic Reinforcement Learning for Safe RAN Control
Listed -
Systematic Mapping Study on the Machine Learning Lifecycle
Listed -
TASP Ep 3 - Optimal Policies Tend to Seek Power
Listed -
[AN #141]: The case for practicing alignment work on GPT-3 and other large models
Listed -
AXRP Episode 5 - Infra-Bayesianism with Vanessa Kosoy
Listed -
Designing Disaggregated Evaluations of AI Systems: Choices, Considerations, and Tradeoffs
Listed -
Extended Picture Theory or Models inside Models inside Models
Listed - Listed
-
[Link post] Coordination challenges for preventing AI conflict
Listed -
CLR's recent work on multi-agent systems
Listed -
Coordination challenges for preventing AI conflict
Listed -
Pretrained Transformers as Universal Computation Engines
Listed -
The AI Index 2021 Annual Report
Listed -
Towards a Mechanistic Understanding of Goal-Directedness
Listed -
A simple way to make GPT-3 follow instructions
Listed -
Epistemological Framing for AI Alignment Research
Listed -
Multi-agent learning in mixed-motive coordination problems
Listed -
What I'd change about different philosophy fields
Listed -
Collaborative game specification: arriving at common models in bargaining
Listed -
Causal Analysis of Agent Behavior for AI Safety
Listed -
MIRI comments on Cotra's "Case for Aligning Narrowly Superhuman Models"
Listed -
Multimodal Neurons in Artificial Neural Networks
Listed -
Rissanen Data Analysis: Examining Dataset Characteristics via Description Length
Listed -
The case for aligning narrowly superhuman models
Listed -
What mechanisms drive agent behaviour?
Listed -
[AN #140]: Theoretical models that predict scaling laws
Listed -
A non-logarithmic argument for Kelly
Listed -
A Semitechnical Introductory Dialogue on Solomonoff Induction
Listed -
Book review: "A Thousand Brains" by Jeff Hawkins
Listed -
Connecting the good regulator theorem with semantics and symbol grounding
Listed -
From-above vs Fine-grain diversity
Listed -
Growth Doesn't Care About Crises
Listed -
How does bee learning compare with machine learning?
Listed - Listed
- Listed
- Listed
-
Evaluating Robustness of Counterfactual Explanations
Listed -
The Importance of Artificial Sentience
Listed - Listed
-
AI, Governance Displacement, and the (De)Fragmentation of International Law
Listed - Listed
- Listed
-
International Control of Powerful Technology: Lessons from the Baruch Plan for Nuclear Weapons
Listed -
Key Concepts in AI Safety: An Overview
Listed -
Key Concepts in AI Safety: Interpretability in Machine Learning
Listed -
Key Concepts in AI Safety: Robustness and Adversarial Examples
Listed -
How might cryptocurrencies affect AGI timelines?
Listed - Listed
-
List sorting does not play well with few-shot
Listed -
Secure Evaluation of Knowledge Graph Merging Gain
Listed -
Bias-reduced Multi-step Hindsight Experience Replay for Efficient Multi-goal Reinforcement Learning
Listed - Listed
- Listed
-
[AN #139]: How the simplicity of reality explains the success of neural nets
Listed -
Beyond Fine-Tuning: Transferring Behavior in Reinforcement Learning
Listed -
Zero-Shot Text-to-Image Generation
Listed - Listed
-
A Citizen's Guide to Artificial Intelligence
Listed -
Software Architecture for Next-Generation AI Planning Systems
Listed -
A Game-Theoretic Approach for Hierarchical Epidemic Control
Listed - Listed
-
Google’s Ethical AI team and AI Safety
Listed - Listed
-
A maximum entropy model of bounded rational decision-making with prior beliefs and market feedback
Listed -
AXRP Episode 4 - Risks from Learned Optimization with Evan Hubinger
Listed -
Formal Solution to the Inner Alignment Problem
Listed -
Training a Resilient Q-Network against Observational Interference
Listed -
Utility Maximization = Description Length Minimization
Listed - Listed
-
[AN #138]: Why AI governance should find problems rather than just solving them
Listed -
Fully General Online Imitation Learning
Listed -
Graphical World Models, Counterfactuals, and Machine Learning Agents
Listed -
Safely controlling the AGI agent reward function
Listed -
Cartesian frames as generalised models
Listed -
Disentangling Corrigibility: 2015-2021
Listed -
Generalised models as a category
Listed -
Mathematical Models of Progress?
Listed -
Suggestions of posts on the AF to review
Listed -
Transferring Domain Knowledge with an Adviser in Continuous Tasks
Listed -
How RL Agents Behave When Their Actions Are Modified
Listed - Listed
-
On the Equilibrium Elicitation of Markov Games Through Information Design
Listed -
Interactive Learning from Activity Description
Listed -
Mitigating Negative Side Effects via Environment Shaping
Listed -
Modelling Cooperation in Network Games with Spatio-Temporal Complexity
Listed -
Weak identifiability and its consequences in strategic settings
Listed -
A Decentralized Approach towards Responsible AI in Social Ecosystems
Listed -
Discovery of Options via Meta-Learned Subgoals
Listed -
Explaining Neural Scaling Laws
Listed -
Mapping the Conceptual Territory in AI Existential Safety and Alignment
Listed -
Tournesol, YouTube and AI Risk
Listed -
Institute for Assured Autonomy (IAA) newsletter
Listed - Listed
-
Stuart Russell Human Compatible AI Roundtable with Allan Dafoe, Rob Reich, & Marietje Schaake
Listed -
[AN #137]: Quantifying the benefits of pretraining on downstream task performance
Listed -
Language models are 0-shot interpreters
Listed -
Some global catastrophic risk estimates
Listed -
Transfer Reinforcement Learning across Homotopy Classes
Listed - Listed
-
Equilibrium Refinements for Multi-Agent Influence Diagrams: Theory and Practice
Listed -
Fixing The Good Regulator Theorem
Listed -
Loom: interface to the multiverse
Listed -
13 Recent Publications on Existential Risk (Jan 2021 update)
Listed -
Alchemical marriage: GPT-3 x CLIP
Listed - Listed
-
Playing the Blame Game with Robots
Listed -
This Museum Does Not Exist: GPT-3 x CLIP
Listed - Listed
- Listed
-
Reflections on Artificial Intelligence for Humanity
Listed -
AI Can Stop Mass Shootings, and More
Listed -
Creating AGI Safety Interlocks
Listed -
Evolutions Building Evolutions: Layers of Generate and Test
Listed -
Learning Normativity: Language
Listed -
AI Development for the Public Interest: From Abstraction Traps to Sociotechnical Risks
Listed -
Exploring Beyond-Demonstrator via Meta Learning-Based Reward Extrapolation
Listed -
Feedback in Imitation Learning: The Three Regimes of Covariate Shift
Listed -
OpenAI: "Scaling Laws for Transfer", Hernandez et al.
Listed -
Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models
Listed - Listed
-
[AN #136]: How well will GPT-N perform on downstream tasks?
Listed - Listed
-
Counterfactual Planning in AGI Systems
Listed -
Distinguishing claims about training vs deployment
Listed - Listed
-
Agent Incentives: A Causal Perspective
Listed -
Data, Architecture, or Losses: What Contributes Most to Multimodal Transformer Success?
Listed - Listed
-
AI Verification: Mechanisms to Ensure AI Arms Control Compliance
Listed - Listed
- Listed
-
Institutionalizing ethics in AI through broader impact requirements
Listed - Listed
- Listed
-
Fairness through Social Welfare Optimization
Listed -
Limiting Causality by Complexity Class
Listed -
AMA on EA Forum: Ajeya Cotra, researcher at Open Phil
Listed -
Challenges for Using Impact Regularizers to Avoid Negative Side Effects
Listed -
Counterfactual Planning in AGI Systems
Listed -
AMA: Ajeya Cotra, researcher at Open Phil
Listed -
Extracting Money from Causal Decision Theorists
Listed -
Making Responsible AI the Norm rather than the Exception
Listed -
[AN #135]: Five properties of goal-directed systems
Listed - Listed
- Listed
-
Optimal play in human-judged Debate usually won't answer your question
Listed -
Muppet: Massive Multi-task Representations with Pre-Finetuning
Listed -
Accumulating Risk Capital Through Investing in Cooperation
Listed -
Language models are multiverse generators
Listed -
Measuring Intelligence and Growth Rate: Variations on Hibbard's Intelligence Measure
Listed -
What is a VNM stable set, really?
Listed -
FC final: Can Factored Cognition schemes scale?
Listed -
The Internet, mirrored by GPT-3
Listed -
[Podcast] Ajeya Cotra on worldview diversification and how big the future could be
Listed -
Baobao Zhang: How social science research can inform AI governance
Listed - Listed
-
Poll: Which variables are most strategically relevant?
Listed -
[AN #134]: Underspecification as a cause of fragility to distribution shift
Listed -
Counterfactual control incentives
Listed -
Singapore AI Policy Career Guide
Listed - Listed
-
Shielding Atari Games with Bounded Prescience
Listed -
UPDeT: Universal Multi-agent Reinforcement Learning via Policy Decoupling with Transformers
Listed -
Against the Backward Approach to Goal-Directedness
Listed -
Some thoughts on risks from narrow, non-agentic AI
Listed -
Some thoughts on risks from narrow, non-agentic AI
Listed - Listed
- Listed
-
Literature Review on Goal-Directedness
Listed - Listed
-
Adversarial Interaction Attack: Fooling AI to Misinterpret Human Intentions
Listed -
Excerpt from Arbital Solomonoff induction dialogue
Listed -
Teaming up with information agents
Listed -
The Challenge of Value Alignment: from Fairer Algorithms to AI Safety
Listed - Listed
-
[link] Centre for the Governance of AI 2020 Annual Report
Listed -
A canonical bit-encoding for ranged integers
Listed -
Evaluating the Robustness of Collaborative Agents
Listed -
Thoughts on Iason Gabriel’s Artificial Intelligence, Values, and Alignment
Listed -
[AN #133]: Building machines that can cooperate (with humans, institutions, or other machines)
Listed -
Some recent survey papers on (mostly near-term) AI safety, security, and assurance
Listed -
AGI Safety and Alignment with Robert Miles-by Machine Ethics-date 20210113
Listed -
AI and International Stability: Risks and Confidence-Building Measures
Listed -
How should we invest in "long-term short-termism" given the likelihood of transformative AI?
Listed - Listed
-
Review of 'Debate on Instrumental Convergence between LeCun, Russell, Bengio, Zador, and More'
Listed -
The Immigration Preferences of Top AI Researchers: New Survey Evidence | GovAI
Listed - Listed
-
Imitative Generalisation (AKA 'Learning the Prior')
Listed -
Prediction can be Outer Aligned at Optimum
Listed -
Review of Soft Takeoff Can Still Lead to DSA
Listed -
The Case for a Journal of AI Alignment
Listed -
What does it mean to become an expert in AI Hardware?
Listed -
Bridging In- and Out-of-distribution Samples for Their Better Discriminability
Listed -
Eight claims about multi-agent AGI safety
Listed - Listed
-
[AN #132]: Complex and subtly incorrect arguments as an obstacle to debate
Listed -
Legal Priorities Research: A Research Agenda
Listed -
Review of 'But exactly how complex and fragile?'
Listed -
The National Defense Authorization Act Contains AI Provisions
Listed -
The Pointers Problem: Clarifications/Variations
Listed -
Existential risks from a Thomist Christian perspective
Listed -
Multi-dimensional rewards for AGI interpretability and control
Listed - Listed
- Listed
-
Mental subagent implications for AI Safety
Listed -
A General Counterexample to Any Decision Theory and Some Responses
Listed - Listed
-
AI Alignment, Philosophical Pluralism, and the Relevance of Non-Western Philosophy
Listed -
AI and the Future of Cyber Competition
Listed -
AI CERTIFICATION: Advancing Ethical Practice by Reducing Information Asymmetries
Listed -
Artificial Canaries: Early Warning Signs for Anticipatory and Democratic Governance of AI
Listed -
Artificial Intelligence Governance Under Change: Foundations, Facets, Frameworks
Listed -
Emerging Technologies: More to explore
Listed -
Negative Side Effects and AI Agent Indicators: Experiments in SafeLife
Listed - Listed
-
Not a paper, but I find Chris Olah’s interview on the 80,000 Hours podcast super inspiring
Listed -
QNRs: Toward Language for Intelligent Machines
Listed -
Reflections on Larks’ 2020 AI alignment literature review
Listed -
Safe Pareto Improvements for Delegated Game Playing
Listed -
Safe Pareto Improvements for Delegated Game Playing
Listed -
Socially Responsible AI Algorithms: Issues, Purposes, and Challenges
Listed -
The case against economic values in the orbitofrontal cortex (or anywhere else in the brain)
Listed -
The Ethics of Sustainability for Artificial Intelligence
Listed -
Truthful AI: Developing and governing AI that does not lie
Listed -
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Listed -
WHAT IS THE UPPER LIMIT OF VALUE?
Listed