The catalog, page 30
Records 7,251 to 7,500 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
The Conditional Entropy Bottleneck
Listed -
[AN #86]: Improving debate and factored cognition through human experiments
Listed -
A Bounded Measure for Estimating the Benefit of Visualization
Listed -
AI Impacts: Historic trends in technological progress
Listed -
AI safety: state of the field through quantitative lens
Listed - Listed
-
Growing Neural Cellular Automata
Listed -
Leveraging Rationales to Improve Human Task Performance
Listed -
Short-Term AI Alignment as a Priority Cause
Listed -
Attainable Utility Landscape: How The World Is Changed
Listed -
Gricean communication and meta-preferences
Listed -
What can the principal-agent literature tell us about AI risk?
Listed -
Effect of AlexNet on historic trends in image recognition
Listed -
Historic trends in book production
Listed -
Historic trends in bridge span length
Listed - Listed
-
Historic trends in light intensity
Listed -
Historic trends in long-range military payload delivery
Listed -
Historic trends in slow light technology
Listed -
Historic trends in telecommunications performance
Listed -
Historic trends in the maximum superconducting temperature
Listed -
Historic trends in transatlantic message speed
Listed -
Incomplete case studies of discontinuous progress
Listed -
Penicillin and historic syphilis trends
Listed -
What can the principal-agent literature tell us about AI risk?
Listed -
Effect of Eli Whitney’s cotton gin on historic trends in cotton ginning
Listed -
Exploring AI Futures Through Role Play
Listed -
Historic trends in flight airspeed records
Listed -
On the falsifiability of hypercomputation
Listed -
Should Artificial Intelligence Governance be Centralised?: Design Lessons from History
Listed -
Student/Teacher Advising through Reward Augmentation
Listed -
The Windfall Clause: Distributing the Benefits of AI for the Common Good
Listed -
Plausibly, almost every powerful algorithm would be manipulative
Listed -
Quantifying Independently Reproducible Machine Learning
Listed - Listed
-
FHI Report: The Windfall Clause: Distributing the Benefits of AI for the Common Good
Listed -
Synthesizing amplification and debate
Listed -
Writeup: Progress on AI Safety via Debate
Listed - Listed
-
Pessimism About Unknown Unknowns Inspires Conservatism
Listed -
Philosophical self-ratification
Listed -
What are the challenges and problems with programming law-breaking constraints into AGI?
Listed - Listed
- Listed
-
Preventing Imitation Learning with Adversarial Policy Ensembles
Listed -
Brian Tse: Sino-Western cooperation in AI safety
Listed -
[AN #84] Reviewing AI alignment work in 2018-19
Listed -
A high-precision abundance analysis of the nuclear benchmark star HD 20
Listed -
Slide deck: Introduction to AI Safety
Listed - Listed
- Listed
- Listed
- Listed
- Listed
-
Appendix: how a subagent could get powerful
Listed -
Towards Learning Multi-agent Negotiations via Self-Play
Listed -
Using vector fields to visualise preferences and make them consistent
Listed -
Towards a Human-like Open-Domain Chatbot
Listed -
Silly rules improve the capacity of agents to learn stable enforcement and compliance behaviors
Listed -
The two-layer model of human values, and problems with synthesizing preferences
Listed -
Formulating Reductive Agency in Causal Models
Listed -
New paper: The Incentives that Shape Behaviour
Listed -
Scaling Laws for Neural Language Models
Listed -
What's a Good Prediction? Challenges in evaluating an agent's knowledge
Listed - Listed
-
[AN #83]: Sample-efficient deep learning with ReMixMatch
Listed -
Concerns Surrounding CEV: A case for human friendliness first
Listed -
Subjective Knowledge and Reasoning about Agents in Multi-Agent Systems
Listed -
Designing for the Long Tail of Machine Learning
Listed -
Explaining Data-Driven Decisions made by AI Systems: The Counterfactual Approach
Listed -
How Doomed are Large Organizations?
Listed -
Logical Representation of Causal Models
Listed -
Inner alignment requires making assumptions about human values
Listed -
The Incentives that Shape Behaviour
Listed -
Gradient Surgery for Multi-Task Learning
Listed -
Teaching Software Engineering for AI-Enabled Systems
Listed -
[Link] EAF Research agenda: "Cooperation, Conflict, and Transformative Artificial Intelligence"
Listed -
Activism by the AI Community: Analysing Recent Achievements and Future Prospects
Listed -
Engineering AI Systems: A Research Agenda
Listed -
Optimal by Design: Model-Driven Synthesis of Adaptation Strategies for Autonomous Systems
Listed -
[AN #82]: How OpenAI Five distributed their training computation
Listed -
ACDT: a hack-y acausal decision theory
Listed -
In Defense of the Arms Races… that End Arms Races
Listed - Listed
-
Predictors exist: CDT going bonkers... forever
Listed -
Social and Governance Implications of Improved Data Efficiency
Listed -
What are the most pressing issues in short-term AI policy?
Listed -
An Architectural Risk Analysis of Machine Learning Systems: Toward More Secure Machine Learning
Listed -
Artificial Intelligence, Values and Alignment
Listed -
Artificial Intelligence, Values and Alignment
Listed - Listed
-
Moral uncertainty: What kind of 'should' is involved?
Listed -
What is the relationship between Preference Learning and Value Learning?
Listed -
Malign generalization without internal search
Listed -
Update on Ought's experiments on factored evaluation of arguments
Listed -
Evaluating Arguments One Step at a Time
Listed -
I'm Cullen O'Keefe, a Policy Researcher at OpenAI, AMA
Listed -
Moral uncertainty vs related concepts
Listed - Listed
- Listed
-
Outer alignment and imitative amplification
Listed -
Visualizing the Impact of Feature Attribution Baselines
Listed - Listed
-
Preference synthesis illustrated: Star Wars
Listed -
The Logic of Strategic Assets: From Oil to Artificial Intelligence
Listed -
(Double-)Inverse Embedded Agency Problem
Listed -
[AN #81]: Universality as a potential solution to conceptual difficulties in intent alignment
Listed -
Algorithmic Fairness from a Non-ideal Perspective
Listed -
How to Throw Away Information in Causal DAGs
Listed -
Definitions of Causal Abstraction: Reviewing Beckers & Halpern
Listed - Listed
- Listed
-
Dissolving Confusion around Functional Decision Theory
Listed - Listed
-
[AN #80]: Why AI risk might be solved without additional intervention from longtermists
Listed -
Making decisions when both morally and empirically uncertain
Listed -
[AN #79]: Recursive reward modeling as an alignment technique integrated with deep RL
Listed -
A Framework for Democratizing AI
Listed -
AGI Safety From First Principles
Listed -
AI Paradigms and AI Safety: Mapping Artefacts and Techniques to Safety Issues
Listed -
Automating reasoning about the future at Ought
Listed -
Avoiding Negative Side Effects due to Incomplete Knowledge of AI Systems
Listed -
Avoiding Side Effects in Complex Environments
Listed -
Building Trust Through Testing
Listed -
Canaries in Technology Mines: Warning Signs of Transformative Progress in AI
Listed - Listed
-
Choice Set Misspecification in Reward Inference
Listed -
Classification of global catastrophic risks connected with artificial intelligence
Listed -
Decision Points in AI Governance
Listed -
Defence in Depth Against Human Extinction: Prevention, Response, Resilience, and Why They All Matter
Listed -
Fragmentation and the Future: Investigating Architectures for International AI Governance
Listed -
From the Standard Model of AI to Provably Beneficial Systems
Listed -
International evaluation of an AI system for breast cancer screening
Listed -
Learning to summarize with human feedback
Listed -
Open Problems in Cooperative AI
Listed - Listed
-
Pragmatic-Pedagogic Value Alignment
Listed -
Responsive safety in reinforcement learning by pid lagrangian methods
Listed -
Safer ML paradigms team: the story – AI Safety Research Program
Listed -
Since figuring out human values is hard, what about, say, monkey values?
Listed -
The MAGICAL Benchmark for Robust Imitation
Listed - Listed
-
Towards Cooperation in Learning Games
Listed -
human psycholinguists: a critical appraisal
Listed - Listed
- Listed
-
Uncertainty-Based Out-of-Distribution Classification in Deep Reinforcement Learning
Listed -
Making decisions under moral uncertainty
Listed -
Asking the Right Questions: Learning Interpretable Action Models Through Query Answering
Listed -
Safe exploration and corrigibility
Listed -
Conversation on AI risk with Adam Gleave
Listed -
Critiquing "What failure looks like"
Listed -
The Offense-Defense Balance of Scientific Knowledge: Does Publishing AI Research Reduce Misuse?
Listed - Listed
-
Brief summary of key disagreements in AI Risk
Listed -
New paper: (When) is Truth-telling Favored in AI debate?
Listed - Listed
-
Comparison of naturally evolved and engineered solutions
Listed - Listed
- Listed
-
Defining AI in Policy versus Practice
Listed -
Effects of breech loading rifles on historic trends in firearm progress
Listed - Listed
-
Humans Are Embedded Agents Too
Listed -
Might humans not be the most intelligent animals?
Listed -
Section 7: Foundations of Rational Agency
Listed -
Questions to Guide the Future of Artificial Intelligence Research
Listed -
The Counterfactual Prisoner's Dilemma
Listed -
Clarifying Power-Seeking and Instrumental Convergence
Listed -
Mastering Complex Control in MOBA Games with Deep Reinforcement Learning
Listed -
Retrospective on the specification gaming examples list
Listed -
Sections 5 & 6: Contemporary Architectures, Humans in the Loop
Listed -
2019 AI Alignment Literature Review and Charity Comparison
Listed - Listed
-
When Goodharting is optimal: linear vs diminishing returns, unlikely vs likely, and other factors
Listed -
[AN #77]: Double descent: a unification of statistical theory and modern ML practice
Listed -
Abstraction, Causality, and Embedded Maps: Here Be Monsters
Listed - Listed
-
Why we need an AI-resilient society
Listed -
A dilemma for prosaic AI alignment
Listed - Listed
-
Counterfactual Induction (Algorithm Sketch, Fixpoint proof)
Listed -
Counterfactual Induction (Lemma 4)
Listed -
Counterfactual Mugging: Why should you pay?
Listed - Listed
-
Is Causality in the Map or the Territory?
Listed -
Sections 1 & 2: Introduction, Strategy and Governance
Listed -
Sections 3 & 4: Credibility, Peaceful Bargaining Mechanisms
Listed -
More Data Can Hurt for Linear Regression: Sample-wise Double Descent
Listed -
Should Artificial Intelligence Governance be Centralised? Six Design Lessons from History
Listed - Listed
-
Is the term mesa optimizer too narrow?
Listed -
But exactly how complex and fragile?
Listed -
Dota 2 with Large Scale Deep Reinforcement Learning
Listed -
Preface to CLR's Research Agenda on Cooperation, Conflict, and TAI
Listed -
Examples of Causal Abstraction
Listed -
Causal Abstraction Toy Model: Medical Sensor
Listed -
Linear Mode Connectivity and the Lottery Ticket Hypothesis
Listed -
Regulatory Markets for AI Safety
Listed -
What Can Learned Intrinsic Rewards Capture?
Listed -
Deep Bayesian Reward Learning from Preferences
Listed -
Predictive coding = RL + SL + Bayes + MPC
Listed - Listed
-
Meta-Learning without Memorization
Listed -
Counterfactuals: Smoking Lesion vs. Newcomb's
Listed - Listed
-
Value-of-Information based Arbitration between Model-based and Model-free Control
Listed -
Are Humans 'Human Compatible'?
Listed -
Comment on Coherence arguments do not imply goal directed behavior
Listed -
Understanding “Deep Double Descent”
Listed - Listed
- Listed
-
Deep Ensembles: A Loss Landscape Perspective
Listed -
Historic trends in transatlantic passenger travel
Listed -
Learning Human Objectives by Evaluating Hypothetical Behavior
Listed -
Oracles: reject all deals - break superrationality, with superrationality
Listed -
Oracles: reject all deals - break superrationality, with superrationality
Listed -
Seeking Power is Often Convergently Instrumental in MDPs
Listed -
Values, Valence, and Alignment
Listed -
What are some non-purely-sampling ways to do deep RL?
Listed - Listed
- Listed
-
Learning Efficient Representation for Intrinsic Motivation
Listed -
Recent Progress in the Theory of Neural Networks
Listed -
Adaptive Online Planning for Continual Lifelong Learning
Listed -
Dream to Control: Learning Behaviors by Latent Imagination
Listed -
Measuring the intelligence of an idealized mechanical knowing agent
Listed -
SafeLife 1.0: Exploring Side Effects in Complex Environments
Listed -
A list of good heuristics that the case for AI x-risk fails
Listed -
Deep Learning for Symbolic Mathematics
Listed - Listed
- Listed
- Listed
-
Algorithmic Decision-Making and the Control Problem
Listed - Listed
- Listed
-
Interactive AI with a Theory of Mind
Listed -
Neural Networks and Deep Learning, Chapters 1-6 (more entry-level)
Listed -
Counterfactuals as a matter of Social Convention
Listed - Listed
-
What's been written about the nature of "son-of-CDT"?
Listed -
Induction of Subgoal Automata for Reinforcement Learning
Listed -
Anti-Alignments -- Measuring The Precision of Process Models and Event Logs
Listed - Listed
-
[AN #75]: Solving Atari and Go with learned game models, and thoughts from a MIRI employee
Listed -
The relationship between trust in AI and trustworthy machine learning technologies
Listed -
The Transformative Potential of Artificial Intelligence
Listed -
A test for symbol grounding methods: true zero-sum games
Listed -
Thoughts on implementing corrigible robust alignment
Listed -
Breaking Oracles: superrationality and acausal trade
Listed