The catalog, page 23
Records 5,501 to 5,750 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Asking the Right Questions: Learning Interpretable Action Models Through Query Answering.
Listed -
Evaluating models of robust word recognition with serial reproduction.
Listed -
Extending rational models of communication from beliefs to actions.
Listed -
Fixation patterns in simple choice reflect optimal information sampling.
Listed -
Forecasting Transformative AI, Part 1: What Kind of AI?
Listed -
Forecasting Transformative AI: What Kind of AI?
Listed -
From convolutional neural networks to models of higher level cognition (and back again).
Listed -
Hindsight Task Relabelling: Experience Replay for Sparse Reward Meta-RL.
Listed -
Human biases limit cumulative innovation.
Listed -
Human-Compatible Artificial Intelligence.
Listed -
Improving Transferability of Representations via Augmentation-Aware Self-Supervision.
Listed -
Intuitions about magic track the development of intuitive physics.
Listed -
Learning What To Do by Simulating the Past.
Listed -
Making Algorithms Work for Reporting.
Listed -
Meta-Learning of Structured Task Distributions in Humans and Machines.
Listed -
Quantifying Differences in Reward Functions.
Listed -
Replay-Guided Adversarial Environment Design.
Listed - Listed
-
Serial reproduction reveals the geometry of visuospatial representations.
Listed -
Spoofing the Limit Order Book: A Strategic Agent-Based Analysis.
Listed -
Stability Effects of Arbitrage in Exchange Traded Funds: An Agent-Based Model.
Listed -
Teachable Reinforcement Learning via Advice Distillation.
Listed -
The Dynamics of Exemplar and Prototype Representations Depend on Environmental Statistics.
Listed - Listed
-
Transforming Worlds: Automated Involutive MCMC for Open-Universe Probabilistic Models.
Listed -
Tuning the hyperparameters of anytime planning: A deep reinforcement learning approach.
Listed -
Unifying Principles and Metrics for Safe and Assistive AI.
Listed -
X2T: Training an X-to-Text Typing Interface with Online Learning from User Feedback.
Listed -
Goal-Directedness and Behavior, Redux
Listed -
When Most VNM-Coherent Preference Orderings Have Convergent Instrumental Incentives
Listed -
Applications for Deconfusing Goal-Directedness
Listed -
Seeking Power is Convergently Instrumental in a Broad Class of Environments
Listed -
DySR: A Dynamic Representation Learning and Aligning based Model for Service Bundle Recommendation
Listed - Listed
- Listed
-
What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
Listed -
Evaluating CLIP: Towards Characterization of Broader Capabilities and Downstream Implications
Listed -
Sharing the World with Digital Minds
Listed -
Traps of Formalization in Deconfusion
Listed -
Why talk about 10,000 years from now?
Listed -
[AN #159]: Building agents that know how to experiment, by training on procedurally generated games
Listed -
Chris Olah on what the hell is going on inside neural networks
Listed -
Garrabrant and Shah on human modeling in AGI
Listed - Listed
-
The Great Depression, Recession and Stagnation in Full Historical Context
Listed -
Value loading in the human brain: a worked example
Listed -
With One Voice: Composing a Travel Voice Assistant from Re-purposed Models
Listed -
How Do AI Timelines Affect Giving Now vs. Later?
Listed -
How should my timelines influence my career choice?
Listed -
LCDT, A Myopic Decision Theory
Listed - Listed
- Listed
-
What does GPT-3 understand? Symbol grounding and Chinese rooms
Listed -
Bridging the gap: the case for an ‘Incompletely Theorized Agreement’ on AI policy
Listed - Listed
-
Indonesia’s AI Promise in Perspective
Listed -
Military AI Cooperation Toolbox
Listed -
Reputations for Resolve and Higher-Order Beliefs in Crisis Bargaining
Listed -
Responsible and Ethical Military AI
Listed -
Soft Calibration Objectives for Neural Networks
Listed - Listed
-
[AN #158]: Should we be optimistic about generalization?
Listed -
An Ethical Framework for Guiding the Development of Affectively-Aware Artificial Intelligence
Listed -
Did they or didn't they learn tool use?
Listed -
How much compute was used to train DeepMind's generally capable agents?
Listed -
Imagining yourself as a digital person (two sketches)
Listed -
A Reflection on Learning from Data: Epistemology Issues and Limitations
Listed -
Discovering User-Interpretable Capabilities of Black-Box Planning Agents
Listed -
Does X cause Y? An in-depth evidence review
Listed - Listed
-
DeepMind: Generally capable agents emerge from open-ended play
Listed -
DeepMind: Generally capable agents emerge from open-ended play
Listed - Listed
-
Digital People Would Be An Even Bigger Deal
Listed -
Human-Level Reinforcement Learning through Theory-Based Modeling, Exploration, and Planning
Listed -
Open-Ended Learning Leads to Generally Capable Agents
Listed -
Towards Industrial Private AI: A two-tier framework for data and model security
Listed -
AMA: The new Open Philanthropy Technology Policy Fellowship
Listed -
Refactoring Alignment (attempt #2)
Listed - Listed
-
[AN #157]: Measuring misalignment in the technology underlying Copilot
Listed -
AXRP Episode 10 - AI’s Future and Impacts with Katja Grace
Listed - Listed
- Listed
-
Enabling high-accuracy protein structure prediction at the proteome scale
Listed - Listed
-
What are you optimizing for? Aligning Recommender Systems with Human Values
Listed -
Reward splintering for AI design
Listed -
Track records for those who have made lots of predictions
Listed -
Apply to the new Open Philanthropy Technology Policy Fellowship!
Listed - Listed
-
Entropic boundary conditions towards safe artificial superintelligence
Listed - Listed
-
The Duplicator: Instant Cloning Would Make the World Economy Explode
Listed -
A Modulation Layer to Increase Neural Network Robustness Against Data Quality Issues
Listed -
Is the argument that AI is an xrisk valid?
Listed - Listed
-
A model of decision-making in the brain (the short version)
Listed -
Books and lecture series relevant to AI governance?
Listed -
botched alignment and alignment awareness
Listed - Listed
- Listed
-
[AN #156]: The scaling hypothesis: a plan for building AGI
Listed -
A personal take on longtermist AI governance
Listed -
AI alignment and wolfram physics
Listed -
Bayesianism versus conservatism versus Goodhart
Listed -
Underlying model of an imperfect morphism
Listed -
A closer look at chess scalings (into the past)
Listed - Listed
-
Fractional progress estimates for AI timelines and implied resource requirements
Listed -
Generalizing Koopman-Pitman-Darmois
Listed -
Highly accurate protein structure prediction with AlphaFold | Nature
Listed -
Phil Birnbaum's "bad regression" puzzles
Listed - Listed
-
Conservative Objective Models for Effective Offline Model-Based Optimization
Listed -
Deep Adaptive Multi-Intention Inverse Reinforcement Learning
Listed - Listed
-
Melting Pot: an evaluation suite for multi-agent reinforcement learning
Listed -
Model-based RL, Desires, Brains, Wireheading
Listed -
Scalable Evaluation of Multi-Agent Reinforcement Learning with Melting Pot
Listed - Listed
-
All Possible Views About Humanity's Future Are Wild
Listed - Listed
- Listed
- Listed
- Listed
-
What will the twenties look like if AGI is 30 years away?
Listed - Listed
-
Anthropic decision theory for self-locating beliefs
Listed -
The inescapability of knowledge
Listed -
The More Power At Stake, The Stronger Instrumental Convergence Gets For Optimal Policies
Listed - Listed
-
The accumulation of knowledge: literature review
Listed -
A Simple Model of AGI Deployment Risk
Listed -
Aligning an optical interferometer with beam divergence control and continuous action space
Listed -
estimating the amount of populated intelligence explosion timelines
Listed -
Finite Factored Sets: Conditional Orthogonality
Listed -
Generalised models: imperfect morphisms and informational entropy
Listed - Listed
-
The Centre for the Governance of AI is becoming a nonprofit
Listed -
[AN #155]: A Minecraft benchmark for algorithms that learn without reward functions
Listed -
A world in which the alignment problem seems lower-stakes
Listed -
Anthropics and Fermi: grabby, visible, zoo-keeping, and early aliens
Listed -
Anthropics in infinite universes
Listed -
BASALT: A Benchmark for Learning from Human Feedback
Listed -
Intermittent Distillations #4: Semiconductors, Economics, Intelligence, and Technological Progress.
Listed -
Intermittent Distillations #4: Semiconductors, Economics, Intelligence, and Technological Progress.
Listed - Listed
- Listed
-
The SIA population update can be surprisingly small
Listed - Listed
-
A second example of conditional orthogonality in finite factored sets
Listed -
Agency and the unreliable autonomous car
Listed -
Evaluating Large Language Models Trained on Code
Listed -
How much chess engine progress is about adapting to bigger computers?
Listed -
Not Quite 'Ask a Librarian': AI on the Nature, Value, and Future of LIS
Listed - Listed
-
What A Long, Strange Trip It's Been: EleutherAI One Year Retrospective
Listed -
A simple example of conditional orthogonality in finite factored sets
Listed - Listed
-
Getting started independently in AI Safety
Listed -
Is keeping AI "in the box" during training enough?
Listed - Listed
-
Anthropic Effects in Estimating Evolution Difficulty
Listed -
Corporate Governance of Artificial Intelligence in the Public Interest
Listed -
Logic Locking at the Frontiers of Machine Learning: A Survey on Developments and Opportunities
Listed -
The MineRL BASALT Competition on Learning from Human Feedback
Listed -
Towards solving the 7-in-a-row game
Listed -
Evolution as Backstop for Reinforcement Learning
Listed -
Mauhn Releases AI Safety Documentation
Listed -
Confusions re: Higher-Level Game Theory
Listed - Listed
- Listed
-
AI Accidents: An Emerging Threat
Listed -
Experimentally evaluating whether honesty generalizes
Listed - Listed
-
[AN #154]: What economic growth theory has to say about transformative AI
Listed -
How to get technological knowledge on AI/ML (for non-tech people)
Listed -
Musings on general systems alignment
Listed -
Progress on Causal Influence Diagrams
Listed -
Progress on Causal Influence Diagrams
Listed -
The Threat of Offensive AI to Organizations
Listed -
Thoughts on safety in predictive learning
Listed - Listed
- Listed
- Listed
-
How teams went about their research at AI Safety Camp edition 5
Listed -
Brute force searching for alignment
Listed -
Finite Factored Sets: LW transcript with running commentary
Listed -
[AN #153]: Experiments that demonstrate failures of objective robustness
Listed -
Anthropics and Embedded Agency
Listed -
aiSTROM -- A roadmap for developing a successful AI strategy
Listed -
The positive case for a focus on achieving safe AI?
Listed -
AXRP Episode 9 - Finite Factored Sets with Scott Garrabrant
Listed -
classifying computational frameworks
Listed -
degrees of runtime metaprogrammability
Listed -
Modeling the Mistakes of Boundedly Rational Agents Within a Bayesian Theory of Mind
Listed -
Shallow evaluations of longtermist organizations
Listed - Listed
-
Alex Turner's Research, Comprehensive Information Gathering
Listed -
Discussion: Objective Robustness and Inner Alignment Terminology
Listed -
Empirical Observations of Objective Robustness Failures
Listed -
Frequent arguments about alignment
Listed -
How Well do Feature Visualizations Support Causal Understanding of CNN Activations?
Listed -
IQ-Learn: Inverse soft-Q Learning for Imitation
Listed - Listed
- Listed
-
Environmental Structure Can Cause Instrumental Convergence
Listed - Listed
- Listed
- Listed
-
Parameter counts in Machine Learning
Listed -
Uncertain Decisions Facilitate Better Preference Learning
Listed -
Conditional offers and low priors: the problem with 1-boxing Newcomb's dilemma
Listed -
Knowledge is not just precipitation of action
Listed -
MADE: Exploration via Maximizing Deviation from Explored Regions
Listed -
Non-poisonous cake: anthropic updates are normal
Listed -
categories of knowledge representation
Listed -
Poisoning and Backdooring Contrastive Learning
Listed -
Pros and cons of working on near-term technical AI safety and assurance
Listed -
Thoughts on a "Sequences Inspired" PhD Topic
Listed -
[AN #152]: How we’ve overestimated few-shot learning capabilities
Listed -
Aligning AI Regulation to Sociotechnical Change
Listed -
Developing a Fidelity Evaluation Approach for Interpretable Machine Learning
Listed - Listed
- Listed
-
Open problem: how can we quantify player alignment in 2x2 normal-form games?
Listed - Listed
-
Futureproof: Artificial Intelligence Chapter | GovAI
Listed -
Knowledge is not just digital abstraction layers
Listed -
my answer to the fermi paradox
Listed -
refusing to answer ≠ giving a negative answer
Listed -
Revisiting the Calibration of Modern Neural Networks
Listed -
the many faces of chaos magick
Listed -
the persistent data structure argument against linear consciousness
Listed -
the systematic absence of libertarian thought
Listed - Listed
-
Vignettes Workshop (AI Impacts)
Listed -
Vignettes Workshop (AI Impacts)
Listed -
The case for strong longtermism
Listed -
What is an example of recent, tangible progress in AI safety research?
Listed -
Answering questions honestly given world-model mismatches
Listed -
Avoiding the instrumental policy by hiding information about humans
Listed - Listed
-
A New Formalism, Method and Open Issues for Zero-Shot Coordination
Listed -
A naive alignment strategy and optimism about generalization
Listed -
Finite Factored Sets: Orthogonality and Time
Listed -
Hard Choices in Artificial Intelligence
Listed -
Knowledge is not just mutual information
Listed -
Synthesising Reinforcement Learning Policies through Set-Valued Inductive Rule Learning
Listed