The catalog, page 27
Records 6,501 to 6,750 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Avoiding Side Effects By Considering Future Tasks
Listed -
Do's and Don'ts for Human and Digital Worker Integration
Listed -
The case for taking AI seriously as a threat to humanity (Kelsey Piper)
Listed -
Time for AI to cross the human performance range in chess
Listed -
[AN #121]: Forecasting transformative AI timelines using biological anchors
Listed -
The Colliding Exponentials of AI
Listed -
The Solomonoff Prior is Malign
Listed -
Knowledge, manipulation, and free will
Listed -
Longtermist reasons to work for innovative governments
Listed -
The Achilles Heel Hypothesis for AI
Listed -
Toy Problem: Detective Story Alignment
Listed - Listed
-
Safe Reinforcement Learning with Natural Language Constraints
Listed -
Logical Foundations of Government Policy
Listed -
If GPT-6 is human-level AGI but costs $200 per page of output, what would happen?
Listed -
Parameterized Reinforcement Learning for Optical System Optimization
Listed -
[Link] How understanding valence could help make future AIs safer
Listed -
Information-Driven Adaptive Sensing Based on Deep Reinforcement Learning
Listed -
[AN #120]: Tracing the intellectual roots of AI and AI alignment
Listed -
A framework for predicting, interpreting, and improving Learning Outcomes
Listed -
Providing Actionable Feedback in Hiring Marketplaces using Generative Adversarial Networks
Listed -
Safety Aware Reinforcement Learning (SARL)
Listed -
The Alignment Problem: Machine Learning and Human Values
Listed -
The Alignment Problem: Machine Learning and Human Values
Listed -
Learning to Generalize for Sequential Decision Making
Listed -
AGI safety from first principles: Conclusion
Listed -
AI race considerations in a report by the U.S. House Committee on Armed Services
Listed -
Feedback Request on EA Philippines' Career Advice Research for Technical AI Safety
Listed - Listed
- Listed
-
Socialism as a conspiracy theory
Listed - Listed
- Listed
-
AGI safety from first principles: Control
Listed -
AGI safety from first principles: Alignment
Listed -
Economic growth under transformative AI
Listed -
Emergent Social Learning via Multi-agent Reinforcement Learning
Listed -
Hiring engineers and researchers to help align GPT-3
Listed -
Hiring engineers and researchers to help align GPT-3
Listed -
Mediating Artificial Intelligence Developments through Negative and Positive Incentives
Listed -
Quantifying the probability of existential catastrophe: A reply to Beard et al.
Listed - Listed
-
[AN #119]: AI safety when agents are shaped by environments, not rewards
Listed - Listed
-
Learning Rewards from Linguistic Feedback
Listed -
“Unsupervised” translation as an (intent) alignment problem
Listed -
“Unsupervised” translation as an (intent) alignment problem
Listed -
AGI safety from first principles: Goals and Agency
Listed -
Learning to Play Against Any Mixture of Opponents
Listed -
Trust-Region Method with Deep Reinforcement Learning in Analog Design Space Exploration
Listed -
AGI safety from first principles: Introduction
Listed -
AGI safety from first principles: Superintelligence
Listed -
Benefits of Assistance over Reward Learning
Listed -
The EMPATHIC Framework for Task Learning from Implicit Human Feedback
Listed -
The Grey Hoodie Project: Big Tobacco, Big Tech, and the threat on academic integrity
Listed -
What Decision Theory is Implied By Predictive Processing?
Listed - Listed
-
What to do with imitation humans, other than asking them what the right thing to do is?
Listed -
Inverse Rational Control with Partially Observable Continuous Nonlinear Dynamics
Listed -
Neurosymbolic Reinforcement Learning with Formally Verified Exploration
Listed -
Examples of self-governance to reduce technology risk?
Listed -
[AN #118]: Risks, solutions, and prioritization in a world with many AI systems
Listed - Listed
- Listed
-
Anthropomorphisation vs value learning: type 1 vs type 2 errors
Listed -
AMA: Markus Anderljung (PM at GovAI, FHI)
Listed - Listed
-
Clarifying “What failure looks like”
Listed -
Hidden Incentives for Auto-Induced Distributional Shift
Listed -
Humans learn too: Better Human-AI Interaction using Optimized Human Inputs
Listed -
Why GPT wants to mesa-optimize & how we might change this
Listed - Listed
-
Efficient Reinforcement Learning Development with RLzoo
Listed -
Enterprise AI Canvas -- Integrating Artificial Intelligence into Business
Listed -
The "Backchaining to Local Search" Technique in AI Alignment
Listed -
AI Governance: Opportunity and Theory of Impact
Listed -
Alignment for Advanced Machine Learning Systems
Listed -
Distributional Generalization: A New Kind of Generalization
Listed -
Learnable Strategies for Bilateral Agent Negotiation over Multiple Issues
Listed -
[AN #117]: How neural nets would fare under the TEVV framework
Listed -
Applying the Counterfactual Prisoner's Dilemma to Logical Uncertainty
Listed -
Are social media algorithms an existential risk?
Listed -
New report on how much computational power it takes to match the human brain (Open Philanthropy)
Listed - Listed
-
My computational framework for the brain
Listed -
Decision Theory is multifaceted
Listed - Listed
-
Towards the Quantification of Safety Risks in Deep Neural Networks
Listed -
How Much Computational Power Does It Take to Match the Human Brain?
Listed - Listed
-
Communicating with Interactive Articles
Listed -
How Much Computational Power Does It Take to Match the Human Brain?
Listed - Listed
-
The AIQ Meta-Testbed: Pragmatically Bridging Academic AI Testing and Industrial Q Needs
Listed -
DeepSpeed: Extreme-scale model training for everyone
Listed -
Do mesa-optimizer risk arguments rely on the train-test paradigm?
Listed -
Importance Weighted Policy Learning and Adaptation
Listed -
Measurement in AI Policy: Opportunities and Challenges
Listed -
Safety via selection for obedience
Listed -
[AN #116]: How to make explanations of neurons compositional
Listed -
Beneficial and Harmful Explanatory Machine Learning
Listed -
Safer sandboxing via collective separation
Listed -
Determining core values & existential self-determination
Listed -
(The Cartoon Guide to) Lob’s Theorem
Listed - Listed
-
A Technical Explanation of Technical Explanation
Listed -
An Intuitive Explanation of Bayes’ Theorem
Listed - Listed
-
Artificial Intelligence as a Positive and Negative Factor in Global Risk
Listed -
Cognitive Biases Potentially Affecting Judgment of Global Risks
Listed - Listed
- Listed
- Listed
- Listed
- Listed
- Listed
- Listed
- Listed
-
Three Major Singularity Schools
Listed -
Transhumanism as Simplified Humanism
Listed - Listed
-
Using GPT-N to Solve Interpretability of Neural Networks: A Research Agenda
Listed -
[AN #115]: AI safety research problems in the AI-GA framework
Listed - Listed
-
Are we living at the hinge of history
Listed - Listed
-
(Humor) AI Alignment Critical Failure Table
Listed -
A course for the general public on AI
Listed -
interpreting GPT: the logit lens
Listed - Listed
-
Updates and additions to "Embedded Agency"
Listed -
A Framework for Improving Scholarly Neural Network Diagrams
Listed - Listed
-
Belief Functions And Decision Theory
Listed -
Model splintering: moving from one imperfect model to another
Listed -
Preface to the sequence on economic growth
Listed -
Proofs Section 1.1 (Initial results to LF-duality)
Listed -
Proofs Section 1.2 (Mixtures, Updates, Pushforwards)
Listed -
Proofs Section 2.1 (Theorem 1, Lemmas)
Listed -
Proofs Section 2.2 (Isomorphism to Expectations)
Listed -
Proofs Section 2.3 (Updates, Decision Theory)
Listed - Listed
-
Technical model refinement formalism
Listed -
Thread: Differentiable Self-organizing Systems
Listed -
[AN #114]: Theory-inspired safety solutions for powerful Bayesian RL agents
Listed -
Asya Bergal: Reasons you might think human-level AI is unlikely to happen soon
Listed -
Introduction To The Infra-Bayesianism Sequence
Listed -
nostalgebraist: Recursive Goodhart's Law
Listed -
Singapore’s Technical AI Alignment Research Career Guide
Listed -
What is the interpretation of the do() operator?
Listed -
Learning human preferences: black-box, white-box, and structured white-box access
Listed -
Forecasting Thread: AI Timelines
Listed -
A Composable Specification Language for Reinforcement Learning Tasks
Listed -
Thoughts on the Feasibility of Prosaic AGI Alignment?
Listed -
Understanding View Selection for Contrastive Learning
Listed - Listed
-
What's a Decomposable Alignment Topic?
Listed -
Animal Rights, The Singularity, and Astronomical Suffering
Listed -
[AN #113]: Checking the ethical intuitions of large language models
Listed -
AI safety as featherless bipeds *with broad flat nails*
Listed -
Alex Irpan: "My AI Timelines Have Sped Up"
Listed -
Looking for adversarial collaborators to test our Debate protocol
Listed -
Deploying Lifelong Open-Domain Dialogue Learning
Listed -
Learning human preferences: optimistic and pessimistic scenarios
Listed - Listed
- Listed
-
A way to beat superrational/EDT agents?
Listed -
Forward and inverse reinforcement learning sharing network weights and hyperparameters
Listed -
Runtime-Safety-Guided Policy Repair
Listed -
Goal-Directedness: What Success Looks Like
Listed - Listed
-
Adversarial Policies: Attacking Deep Reinforcement Learning.
Listed -
Conservative agency via attainable utility preservation..
Listed -
Incomplete Contracting and AI Alignment.
Listed -
LESS is More: Rethinking Probabilistic Models of Human Behavior.
Listed - Listed
-
My Understanding of Paul Christiano's Iterated Amplification AI Safety Research Agenda
Listed -
My Understanding of Paul Christiano's Iterated Amplification AI Safety Research Agenda
Listed -
On the Geometry of Adversarial Examples.
Listed - Listed
-
SQIL: Imitation Learning via Regularized Behavioral Cloning..
Listed -
What are you optimizing for? Aligning Recommender Systems with Human Values.
Listed -
A rational model of sequential self-assessment.
Listed -
A Rational Reinterpretation of Dual-Process Theories.
Listed -
Adaptive Autonomous Secure Cyber Systems.
Listed -
Advancing rational analysis to the algorithmic level.
Listed -
AI Research Considerations for Human Existential Safety (ARCHES).
Listed -
Aligning AI With Shared Human Values.
Listed -
Aligning with Heterogeneous Preferences for Kidney Exchange.
Listed -
Artificial Intelligence: A Modern Approach (Textbook, 4th Edition).
Listed -
Assessing Mathematics Misunderstandings via Bayesian Inverse Planning.
Listed -
AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty.
Listed -
Choice Set Misspecification in Reward Inference.
Listed -
Combining experts’ causal judgments.
Listed - Listed
-
Decentralized Reinforcement Learning: Global Decision-Making via Local Economic Transactions.
Listed -
DERAIL: Diagnostic Environments for Reward And Imitation Learning.
Listed -
Downloading Culture.zip: Social learning by program induction.
Listed - Listed
- Listed
-
Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design.
Listed -
Extracting low-dimensional psychological representations from convolutional neural networks.
Listed - Listed
-
Hidden Community Detection on Two-layer Stochastic Models: a Theoretical Perspective.
Listed -
How Should an Agent Practice?.
Listed -
Interpretable and Pedagogical Examples.
Listed -
Measuring Massive Multitask Language Understanding.
Listed -
misc raw responses to a tract of Critical Rationalism
Listed -
Multi-Principal Assistance Games.
Listed - Listed
-
People Do Not Just Plan,They Plan to Plan.
Listed -
Pretrained Transformers Improve Out-of-Distribution Robustness.
Listed -
Reconciling novelty and complexity through a rational analysis of curiosity.
Listed -
Resource-rational Task Decomposition to Minimize Planning Costs.
Listed -
Scaling up psychology via Scientific Regret Minimization.
Listed -
SLIP: Learning to predict in unknown dynamical systems with long-term memory.
Listed -
Solving hard AI planning instances using curriculum-driven deep reinforcement learning.
Listed -
Sparse Graphical Memory for Robust Planning.
Listed -
Structure Learning for Approximate Solution of Many-Player Games.
Listed -
The Efficiency of Human Cognition Reflects Planned Information Processing.
Listed -
The MAGICAL Benchmark for Robust Imitation.
Listed -
The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization.
Listed -
The method of loci is an optimal policy for memory search.
Listed -
Translucent players: Explaining cooperative behavior in social dilemmas.
Listed -
Understanding Learned Reward Functions.
Listed -
Value-laden Disciplinary Shifts in Machine Learning.
Listed -
[AN #112]: Engineering a Safer World
Listed - Listed
-
Considerations, Good Practices, Risks and Pitfalls in Developing AI Solutions Against COVID-19
Listed - Listed
-
Blog post: A tale of two research communities
Listed -
Matt Botvinick on the spontaneous emergence of learning algorithms
Listed -
Strong implication of preference uncertainty
Listed -
Book review: Architects of Intelligence by Martin Ford (2018)
Listed -
Will OpenAI's work unintentionally increase existential risks related to AI?
Listed -
Adapting a kidney exchange algorithm to align with human values.
Listed -
Approximate Causal Abstractions.
Listed -
ASNets: Deep Learning for Generalised Planning.
Listed -
AvE: Assistance via Empowerment.
Listed -
Bounded Rationality in Las Vegas: Probabilistic Finite Automata Play Multi-Armed Bandits.
Listed -
Cognitive prostheses for goal achievement.
Listed -
Exploring AI Safety in Degrees: Generality, Capability and Control
Listed -
Forecasting AI Progress: A Research Agenda
Listed -
How to Be Helpful to Multiple People at Once.
Listed -
Inconsistency evaluation in pairwise comparison using norm-based distances.
Listed -
Learning Rewards from Linguistic Feedback.
Listed -
Market Manipulation: An Adversarial Learning Framework for Detection and Evasion.
Listed -
Predicting responsibility judgments from dispositional inferences and causal attributions.
Listed -
Preference learning along multiple criteria: A game-theoretic perspective.
Listed -
Rational use of episodic and working memory: A normative account of prospective memory.
Listed