The catalog, page 22
Records 5,251 to 5,500 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Some Existing Selection Theorems
Listed -
Takeoff Speeds and Discontinuities
Listed -
A brief review of the reasons multi-objective RL could be important in AI Safety Research
Listed -
Brain-inspired AGI and the "lifetime anchor"
Listed -
Self-Supervise, Refine, Repeat: Improving Unsupervised Anomaly Detection
Listed - Listed
-
Test Time Robustification of Deep Models via Adaptation and Augmentation
Listed - Listed
-
Untangling Braids with Multi-agent Q-Learning
Listed -
Argument for AI x-risk from large impacts
Listed -
Collection of arguments to expect (outer and inner) alignment failure?
Listed -
RAFT: A Real-World Few-Shot Text Classification Benchmark
Listed -
Selection Theorems: A Program For Understanding Agents
Listed -
Summary of history (empowerment and well-being lens)
Listed -
[Link post] How plausible are AI Takeover scenarios?
Listed -
AI takeoff story: a continuation of progress by other means
Listed -
AISC5 Retrospective: Mechanisms for Avoiding Tragedy of the Commons in Common Pool Resource Problems
Listed -
Beyond fire alarms: freeing the groupstruck
Listed - Listed
-
Transformative AI and Compute [Summary]
Listed -
AXRP Episode 11 - Attainable Utility and Power with Alex Turner
Listed -
Cognitive Biases in Large Language Models
Listed - Listed
-
Cartesian Frames and Factored Sets on ArXiv
Listed -
Cold Links: assorted fun basketball stuff
Listed -
Fanaticism in AI: SERI Project
Listed -
Seeking social science students / collaborators interested in AI existential risks
Listed -
The problem of artificial suffering
Listed -
Temporal Inference with Finite Factored Sets
Listed -
The Most Important Century (in a nutshell)
Listed -
What is Compute? - Transformative AI and Compute [1/4]
Listed -
[AN #165]: When large models are more likely to lie
Listed -
A sufficiently paranoid non-Friendly AGI might self-modify itself to become Friendly
Listed - Listed
-
Recursively Summarizing Books with Human Feedback
Listed -
UK's new 10-year "National AI Strategy," released today
Listed -
Announcing the Vitalik Buterin Fellowships in AI Existential Safety!
Listed - Listed
-
Redwood Research’s current project
Listed -
Why AI alignment could be hard with modern deep learning
Listed -
Why AI alignment could be hard with modern deep learning
Listed -
[Book Review] "The Alignment Problem" by Brian Christian
Listed -
Actionable Approaches to Promote Ethical AI in Libraries
Listed - Listed
-
Is working on AI safety as dangerous as ignoring it?
Listed - Listed
-
Testing The Natural Abstraction Hypothesis: Project Update
Listed - Listed
-
Towards Resilient Artificial Intelligence: Survey and Research Issues
Listed -
Asimov's Chronology of Science and Discovery
Listed - Listed
-
Immobile AI makes a move: anti-wireheading, ontology change, and model splintering
Listed -
Investigating AI Takeover Scenarios
Listed -
Is Curiosity All You Need? On the Utility of Emergent Behaviours from Curious Exploration
Listed - Listed
-
ThriftyDAgger: Budget-Aware Novelty and Risk Gating for Interactive Imitation Learning
Listed - Listed
- Listed
-
How truthful is GPT-3? A benchmark for language models
Listed -
I wanted to interview Eliezer Yudkowsky but he's busy so I simulated him instead
Listed -
Jitters No Evidence of Stupidity in RL
Listed -
The Metaethics and Normative Ethics of AGI Value Alignment: Many Questions, Some Implications
Listed -
[AN #164]: How well can language models write code?
Listed - Listed
-
Challenges in Detoxifying Language Models
Listed -
Challenges in Detoxifying Language Models
Listed -
Oracle predictions don't apply to non-existent worlds
Listed -
The Metaethics and Normative Ethics of AGI Value Alignment: Many Questions, Some Implications
Listed -
Von Neumann–Morgenstern utility theorem
Listed -
How to make the best of the most important century?
Listed -
How to make the best of the most important century?
Listed -
Augmenting Decision Making via Interactive What-If Analysis
Listed -
DeepMind is hiring Long-term Strategy & Governance researchers
Listed -
A Socially Aware Reinforcement Learning Agent for The Single Track Road Problem
Listed -
AI timelines and theoretical understanding of deep learning
Listed - Listed
-
Chris Olah on working at top AI labs without an undergrad degree
Listed -
Measurement, Optimization, and Take-off Speed
Listed -
Paths To High-Level Machine Intelligence
Listed -
The Blackwell order as a formalization of knowledge
Listed -
What is the EU AI Act and why should you care about it?
Listed -
A mesa-optimization perspective on AI valence and moral patienthood
Listed - Listed
- Listed
-
One Cold Link: “The Past and Future of Economic Growth: A Semi-Endogenous Perspective”
Listed -
The alignment problem in different capability regimes
Listed -
User Tampering in Reinforcement Learning Recommender Systems
Listed -
[AN #163]: Using finite factored sets for causal and temporal inference
Listed -
Distinguishing AI takeover scenarios
Listed -
Gradient descent is not just more efficient genetic algorithms
Listed - Listed
-
AI Timelines: Where the Arguments, and the "Experts," Stand
Listed -
AI Timelines: Where the Arguments, and the "Experts," Stand
Listed -
Alignment via manually implementing the utility function
Listed -
It takes 5 layers and 1000 artificial neurons to simulate a single biological neuron [Link]
Listed - Listed
-
List of AI safety courses and resources
Listed - Listed
- Listed
-
How to get more academics enthusiastic about doing AI Safety research?
Listed -
Robust fine-tuning of zero-shot models
Listed -
Finetuned Language Models Are Zero-Shot Learners
Listed - Listed
-
A Gentle Introduction to Graph Neural Networks
Listed - Listed
- Listed
-
Formalizing Objections against Surrogate Goals
Listed - Listed
-
Understanding Convolutions on Graphs
Listed -
AI Education in China and the United States
Listed - Listed
-
NIST AI Risk Management Framework request for information (RFI)
Listed -
Problem Learning: Towards the Free Will of Machines
Listed - Listed
- Listed
-
The DOD’s Hidden Artificial Intelligence Workforce
Listed - Listed
-
Call for research on evaluating alignment (funding + advice available)
Listed -
Finite Factored Sets: Applications
Listed -
Finite Factored Sets: Inferring Time
Listed -
Forecasting transformative AI: the "biological anchors" method in a nutshell
Listed -
Grokking the Intentional Stance
Listed -
Reward splintering as reverse of interpretability
Listed -
The Telephone Theorem: Information At A Distance Is Mediated By Deterministic Constraints
Listed -
What are biases, anyway? Multiple type signatures
Listed -
"Epistemic maps" for AI Debates? (or for other issues)
Listed -
A short introduction to machine learning
Listed -
Alignment Research = Conceptual Alignment Research + Applied Alignment Research
Listed - Listed
- Listed
-
The Governance Problem and the "Pretty Good" X-Risk
Listed -
Brain-Computer Interfaces and AI Alignment
Listed -
What are good alignment conference papers?
Listed -
Why and How Governments Should Monitor AI Development
Listed -
[AN #162]: Foundation models: a paradigm shift within AI
Listed - Listed
-
More on “multiple world-size economies per atom”
Listed -
Introduction to Reducing Goodhart
Listed -
The gloves are off, the pants are on
Listed -
(apologies for Alignment Forum server outage last night)
Listed -
MIRI/OP exchange about decision theory
Listed -
Reasoning about Counterfactuals and Explanations: Problems, Results and Directions
Listed -
What are the top priorities in a slow-takeoff, multipolar world?
Listed -
Are we "trending toward" transformative AI? (How would we know?)
Listed -
Extraction of human preferences 👨→🤖
Listed -
Extraction of human preferences 👨→🤖
Listed - Listed
- Listed
- Listed
- Listed
-
AI Risk for Epistemic Minimalists
Listed - Listed
-
AI Safety Papers: An App for the TAI Safety Database
Listed -
Designing a Combinatorial Financial Options Market
Listed -
From Optimizing Engagement to Measuring Value
Listed -
Learning Causal Models of Autonomous Agents using Interventions
Listed - Listed
-
Analogies and General Priors on Intelligence
Listed -
Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions
Listed -
How DeepMind's Generally Capable Agents Were Trained
Listed -
Provide feedback on Open Philanthropy’s AI alignment RFP
Listed -
Safe Transformative AI via a Windfall Clause
Listed -
Cold Links: heartwarming sports stuff
Listed - Listed
-
Updates and Lessons from AI Forecasting
Listed -
Finite Factored Sets: Polynomials and Probability
Listed -
Forecasting transformative AI: what's the burden of proof?
Listed -
1h-volunteers needed for a small AI Safety-related research project
Listed -
1h-volunteers needed for a small AI Safety-related research project
Listed -
Downstream Evaluations of Rotary Position Embeddings
Listed -
Is it worth making a database for moral predictions?
Listed -
Modelling Transformative AI Risks (MTAIR) Project: Introduction
Listed -
On the Opportunities and Risks of Foundation Models
Listed -
Program Synthesis with Large Language Models
Listed -
Decision Transformer: Reinforcement Learning via Sequence Modeling.
Listed -
Designing Recommender Systems to Depolarize.
Listed -
Feature Expansive Reward Learning: Rethinking Human Input.
Listed -
kolmogorov complexity objectivity and languagespace
Listed -
Measuring Coding Challenge Competence With APPS.
Listed -
Unsolved Problems in ML Safety.
Listed -
A neural circuit for flexible control of persistent behavioral states.
Listed -
A Robust Control Framework for Human Motion Prediction.
Listed -
Agent-aware state estimation for autonomous vehicles.
Listed -
Agnostic Learning with Unknown Utilities.
Listed -
AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control.
Listed -
Analyzing Human Models that Adapt Online.
Listed -
Approaches to gradient hacking
Listed -
APS: Active Pretraining with Successor Features.
Listed -
Behavior From the Void: Unsupervised Active Pre-Training.
Listed -
Beyond Engagement: Aligning Algorithmic Recommendations With Prosocial Goals.
Listed -
book recommendation: Greg Egan's
Listed -
Building efficient, reliable, and ethical autonomous systems.
Listed -
Causal Inference Struggles with Agency on Online Platforms.
Listed -
Clusterability in Neural Networks.
Listed -
Contrastive Code Representation Learning.
Listed -
CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review.
Listed -
Decoupling Representation Learning from Reinforcement Learning.
Listed -
Dynamically Switching Human Prediction Models for Efficient Planning.
Listed -
Efficient Dynamics Estimation With Adaptive Model Sets.
Listed -
Estimating and Penalizing Preference Shift in Recommender Systems.
Listed -
Ethically compliant planning within moral communities.
Listed -
Ethically compliant sequential decision making.
Listed -
Evaluating Strategy Exploration in Empirical Game-Theoretic Analysis.
Listed -
Evolution Strategies for Approximate Solution of Bayesian Games.
Listed - Listed
-
Explore and Control with Adversarial Surprise.
Listed -
Improving Competence via Iterative State Space Refinement.
Listed -
Improving Computational Efficiency in Visual Reinforcement Learning via Stored Embeddings.
Listed -
Iterative Empirical Game Solving via Single Policy Best Response.
Listed -
Learning State Representations from Random Deep Action-Conditional Predictions.
Listed -
Mapping the Political Economy of Reinforcement Learning Systems: The Case of Autonomous Vehicles.
Listed -
Mastering Atari Games with Limited Data.
Listed -
Measuring mathematical problem solving with the math dataset.
Listed - Listed
-
Offline-to-Online Reinforcement Learning via Balanced Replay and Pessimistic Q-Ensemble.
Listed -
On complementing end-to-end human behavior predictors with planning.
Listed -
On the benefits of randomly adjusting anytime weighted A*.
Listed -
On the Expressivity of Markov Reward.
Listed -
Optimal Cost Design for Model Predictive Control.
Listed -
Passive Attention in Artificial Neural Networks Predicts Human Visual Selectivity.
Listed -
Perceptual Adversarial Robustness: Defense Against Unseen Threat Models.
Listed -
Physical interaction as communication: Learning robot objectives online from human corrections.
Listed -
Policy Gradient Bayesian Robust Optimization for Imitation Learning.
Listed -
Pragmatic Image Compression for Human-in-the-Loop Decision-Making.
Listed - Listed
-
Putting NeRF on a Diet: Semantically Consistent Few-Shot View Synthesis.
Listed -
Reinforcement Learning of Implicit and Explicit Control Flow Instructions.
Listed -
Reinforcement Learning with Latent Flow.
Listed -
Reward is Enough for Convex MDPs.
Listed -
Scalable Online Planning via Reinforcement Learning Fine-Tuning.
Listed -
Show me the algorithm: Transparency in recommendation systems.
Listed -
Situational Confidence Assistance for Lifelong Shared Autonomy.
Listed -
Skill Preferences: Learning to Extract and Execute Robotic Skills from Human Feedback.
Listed -
State Entropy Maximization with Random Encoders for Efficient Exploration.
Listed -
Unsupervised Learning of Visual 3D Keypoints for Control.
Listed -
URLB: Unsupervised Reinforcement Learning Benchmark.
Listed -
Using metareasoning to maintain and restore safety for reliable autonomy.
Listed -
[AN #160]: Building AIs that learn and think like people
Listed -
A review of "Agents and Devices"
Listed -
A few quick links re: COVID-19/Delta
Listed -
Power-seeking for successive choices
Listed -
Some criteria for sandwiching projects
Listed -
Automating Auditing: An ambitious concrete technical research proposal
Listed -
Beyond Fairness Metrics: Roadblocks and Challenges for Ethical AI in Practice
Listed - Listed
-
A Qualitative and Intuitive Explanation of Expected Value
Listed -
A Rational Account of Anchor Effects in Hindsight Bias.
Listed -
A rational model of people’s inferences about others’ preferences based on response times.
Listed -
A Strategic Analysis of Portfolio Compression.
Listed -
An Agent-Based Model of Strategic Adoption of Real-Time Payments.
Listed