The catalog, page 26
Records 6,251 to 6,500 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
- Listed
-
[AN #131]: Formalizing the argument of ignored attributes in a utility function
Listed - Listed
- Listed
-
A canonical and efficient byte-encoding for ints
Listed -
Against GDP as a metric for timelines and takeoff speeds
Listed -
Against GDP as a metric for timelines and takeoff speeds
Listed -
AXRP Episode 1 - Adversarial Policies with Adam Gleave
Listed -
AXRP Episode 2 - Learning Human Biases with Rohin Shah
Listed -
AXRP Episode 3 - Negotiable Reinforcement Learning with Andrew Critch
Listed - Listed
-
Debate Minus Factored Cognition
Listed -
Multi-Principal Assistance Games: Definition and Collegial Mechanisms
Listed -
Why Neural Networks Generalise, and Why They Are (Kind of) Bayesian
Listed -
You are your information system
Listed -
[AN #130]: A new AI x-risk podcast, and reviews of the field
Listed - Listed
- Listed
-
Operationalizing compatibility with strategy-stealing
Listed -
2019 Review Rewrite: Seeking Power is Often Robustly Instrumental in MDPs
Listed -
Announcing AXRP, the AI X-risk Research Podcast
Listed - Listed
-
Augmenting Policy Learning with Routines Discovered from a Single Demonstration
Listed -
Debate update: Obfuscated arguments problem
Listed - Listed
- Listed
-
CFP for the Largest Annual Meeting of Political Science: Get Help With Your Research Submission
Listed - Listed
-
TAI Safety Bibliographic Database
Listed -
TAI Safety Bibliographic Database
Listed -
2020 AI Alignment Literature Review and Charity Comparison
Listed - Listed
- Listed
-
Evaluating Agents without Rewards
Listed -
How Lesswrong helped me make $25K: A rational pricing strategy
Listed -
Taking Principles Seriously: A Hybrid Approach to Value Alignment
Listed - Listed
-
Hierarchical planning: context agents
Listed -
Probabilistic Dependency Graphs
Listed -
Exploring Fluent Query Reformulations with Text-to-Text Transformers and Reinforcement Learning
Listed -
Extrapolating GPT-N performance
Listed -
AI Impacts key questions of interest
Listed -
Can Transformers Reason About Effects of Actions?
Listed -
Extracting and Using Preference Information from the State of the World
Listed -
How long till Inverse AlphaFold?
Listed -
Mapping the Conceptual Territory in AI Existential Safety and Alignment
Listed -
Mitigating x-risk through modularity
Listed -
Open Philanthropy's AI governance grantmaking (so far)
Listed -
[AN #129]: Explaining double descent by measuring bias and variance
Listed -
Homogeneity vs. heterogeneity in AI takeoff scenarios
Listed -
Less Basic Inframeasure Theory
Listed -
Our AI governance grantmaking so far
Listed -
Challenges of Aligning Artificial Intelligence with Human Values
Listed - Listed
- Listed
-
Some AI research areas and their relevance to existential safety
Listed -
Efficient Querying for Cooperative Probabilistic Commitments
Listed -
Extracting Training Data from Large Language Models
Listed -
What are the best precedents for industries failing to invest in valuable AI research?
Listed -
Wilds: A Benchmark of in-the-Wild Distribution Shifts
Listed - Listed
-
Avoiding Side Effects in Complex Environments
Listed -
Imitating Interactive Intelligence
Listed -
Interdisciplinary Approaches to Understanding Artificial Intelligence's Impact on Society
Listed -
How energy efficient are human-engineered flight designs relative to natural ones?
Listed -
Imitating Interactive Intelligence
Listed -
Learning to Resolve Conflicts for Multi-Agent Path Finding with Conflict-Based Search
Listed -
Neurosymbolic AI: The 3rd Wave
Listed -
What technologies could cause world GDP doubling times to be <8 years?
Listed - Listed
-
What the AI Community Can Learn From Sneezing Ferrets and a Mutant Virus Debate
Listed -
Conservatism in neocortex-like AGIs
Listed -
Naturally Occurring Equivariance in Neural Networks
Listed -
Idea: an AI governance group colocated with every AI research group!
Listed -
Launching the Forecasting AI Progress Tournament
Listed - Listed
-
Minimal Maps, Semi-Decisions, and Neural Representations
Listed -
AI Problems Shared by Non-AI Systems
Listed -
AI Winter Is Coming - How to profit from it?
Listed - Listed
- Listed
-
Values Form a Shifting Landscape (and why you might care)
Listed -
An overview of 11 proposals for building safe advanced AI
Listed -
Learning in two-player games between transparent opponents
Listed -
LessWrong is now a book, available for pre-order!
Listed -
Long-Term Future Fund: Ask Us Anything!
Listed -
Understanding meta-trained algorithms through a Bayesian lens
Listed - Listed
-
Centre for the Study of Existential Risk Four Month Report June - September 2020
Listed - Listed
- Listed
-
Aligning AI Optimization to Community Well-Being
Listed -
Fast reinforcement learning with generalized policy updates
Listed - Listed
- Listed
-
Reinforcement Learning in Newcomblike Environments
Listed - Listed
-
Sharing the World with Digital Minds
Listed - Listed
- Listed
-
Preface to the Sequence on Factored Cognition
Listed -
Is this a good way to bet on short timelines?
Listed -
Is this a good way to bet on short timelines?
Listed -
[AN #126]: Avoiding wireheading by decoupling action feedback from action effects
Listed - Listed
- Listed
-
Energy efficiency of monarch butterfly flight
Listed -
Energy efficiency of wandering albatross flight
Listed -
European Strategy on AI: Are we truly fostering social good?
Listed -
Contract Scheduling With Predictions
Listed -
Critiques of the Agent Foundations agenda?
Listed -
Energy efficiency of paramotors
Listed -
The next AI winter will be due to energy costs
Listed -
Commentary on AGI Safety from First Principles
Listed -
Continuing the takeoffs debate
Listed -
Syntax, semantics, and symbol grounding, simplified
Listed -
Transforming Worlds: Automated Involutive MCMC for Open-Universe Probabilistic Models
Listed - Listed
-
BARS: Joint Search of Cell Topology and Layout for Accurate and Efficient Binary ARchitectures
Listed -
Emergent Road Rules In Multi-Agent Driving Environments
Listed -
Jaan Tallinn: Fireside chat (2020)
Listed - Listed
-
Non-Obstruction: A Simple Concept Motivating Corrigibility
Listed -
Tan Zhi Xuan: AI alignment, philosophical pluralism, and the relevance of non-Western philosophy
Listed -
UDT might not pay a Counterfactual Mugger
Listed -
Assessing Generalization in Reward Learning: Intro and Background
Listed - Listed
-
Persuasion Tools: AI takeover without AGI or agency?
Listed -
Persuasion Tools: AI takeover without AGI or agency?
Listed - Listed
-
Inner Alignment in Salt-Starved Rats
Listed -
Misalignment and misuse: whose values are manifest?
Listed - Listed
-
Some AI research areas and their relevance to existential safety
Listed -
A Prototypeness Hierarchy of Realities
Listed -
Energy efficiency of The Spirit of Butt’s Farm
Listed - Listed
-
Should we postpone AGI until we reach safety?
Listed -
The ethics of AI for the Routledge Encyclopedia of Philosophy
Listed -
The Pointers Problem: Human Values Are A Function Of Humans' Latent Variables
Listed -
Using Unity to Help Solve Intelligence
Listed -
AI transparency: a matter of reconciling design with critique
Listed -
Avoiding Tampering Incentives in Deep RL via Decoupled Approval
Listed - Listed
-
Preventing Repeated Real World AI Failures by Cataloging Incidents: The AI Incident Database
Listed -
REALab: An Embedded Perspective on Tampering
Listed - Listed
-
Was the industrial revolution a drastic departure from historic trends?
Listed -
Donating against Short Term AI risks
Listed -
Extortion beats brinksmanship, but the audience matters
Listed -
How Roodman's GWP model translates to TAI timelines
Listed -
How Roodman's GWP model translates to TAI timelines
Listed -
A guide to Iterated Amplification & Debate
Listed - Listed
-
Early Thoughts on Ontology/Grounding Problems
Listed -
A Self-Embedded Probabilistic Model
Listed -
Active Reinforcement Learning: Observing Rewards at a Cost
Listed -
Misalignment and misuse: whose values are manifest?
Listed -
Communication Prior as Alignment Strategy
Listed - Listed
-
Learning Latent Representations to Influence Multi-Agent Interaction
Listed -
Performance of Bounded-Rational Agents With the Ability to Self-Modify
Listed -
[AN #125]: Neural network scaling laws across multiple modalities
Listed -
A Correspondence Theorem in the Maximum Entropy Framework
Listed - Listed
-
Fooling the primate brain with minimal, targeted image manipulation
Listed -
I Know What You Meant: Learning Human Objectives by (Under)estimating Their Choice Set
Listed -
Learning Normativity: A Research Agenda
Listed - Listed
-
Eight Definitions of Observability
Listed -
Energy efficiency of MacCready Gossamer Albatross
Listed -
It Takes a Village: The Shared Responsibility of 'Raising' an Autonomous Weapon
Listed -
Natural Language Inference in Context -- Investigating Contextual Reasoning over Long Texts
Listed -
What Did You Think Would Happen? Explaining Agent Behaviour Through Intended Outcomes
Listed -
A Theory of Universal Learning
Listed -
Clarifying inner alignment terminology
Listed -
Committing, Assuming, Externalizing, and Internalizing
Listed -
Risk Assessment for Machine Learning Models
Listed -
Why You Should Care About Goal-Directedness
Listed - Listed
-
When Hindsight Isn't 20/20: Incentive Design With Imperfect Credit Allocation
Listed -
How can I bet on short timelines?
Listed -
How can I bet on short timelines?
Listed -
the scaling “inconsistency”: openAI’s new insight
Listed -
Additive and Multiplicative Subagents
Listed -
Does SGD Produce Deceptive Alignment?
Listed -
Energy efficiency of Airbus A320
Listed -
Energy efficiency of Boeing 747-400
Listed -
Energy efficiency of North American P-51 Mustang
Listed -
Consider paying me to do AI safety research work
Listed -
Defining capability and alignment in gradient descent
Listed -
Energy efficiency of Vickers Vimy plane
Listed -
Energy efficiency of Wright model B
Listed - Listed
-
What considerations influence whether I have more influence over short or long timelines?
Listed -
What considerations influence whether I have more influence over short or long timelines?
Listed -
[AN #124]: Provably safe exploration through shielding
Listed -
Energy efficiency of Wright Flyer
Listed -
Multiplicative Operations on Cartesian Frames
Listed - Listed
- Listed
-
Ask Your Humans: Using Human Instructions to Improve Generalization in Reinforcement Learning
Listed -
Automated intelligence is not AI
Listed -
"Inner Alignment Failures" Which Are Actually Outer Alignment Failures
Listed -
Containing the AI... Inside a Simulated Reality
Listed -
Amplifying GPT-3 on closed-ended questions
Listed - Listed
-
Responses to Christiano on takeoff speeds?
Listed - Listed
- Listed
-
Controllables and Observables, Revisited
Listed -
Recovery RL: Safe Reinforcement Learning with Learned Recovery Zones
Listed -
[AN #123]: Inferring what is valuable in order to align recommender systems
Listed - Listed
-
Draft papers for REALab and Decoupled Approval on tampering
Listed -
Scaling Laws for Autoregressive Generative Modeling
Listed -
Dutch-Booking CDT: Revised Argument
Listed -
Generative Temporal Difference Learning for Infinite-Horizon Prediction
Listed -
Learning to be Safe: Deep RL with a Safety Critic
Listed - Listed
-
Security Mindset and Takeoff Speeds
Listed - Listed
-
Additive Operations on Cartesian Frames
Listed -
Supervised learning of outputs in the brain
Listed -
Time for AI to cross the human range in English draughts
Listed -
4 Years Later: President Trump and Global Catastrophic Risk
Listed -
Artificial intelligence career stories
Listed -
Buck Shlegeris: How I think students should orient to AI safety
Listed -
How to build a safe advanced AI (Evan Hubinger) | What's up in AI safety? (Asya Bergal)
Listed -
Reply to Jebari and Lundborg on Artificial Superintelligence
Listed -
Exemplary natural images explain CNN activations better than feature visualizations
Listed -
Humans are stunningly rational and stunningly irrational
Listed - Listed
-
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Listed - Listed
-
Exploring the Nuances of Designing (with/for) Artificial Intelligence
Listed -
Introduction to Cartesian Frames
Listed -
The date of AI Takeover is not the day the AI takes over
Listed -
[AN #122]: Arguing for AGI-driven existential risk from first principles
Listed -
AGI safety from first principles
Listed - Listed
-
Problems Involving Abstraction?
Listed -
Robust Imitation Learning from Noisy Demonstrations
Listed -
Time for AI to cross the human range in StarCraft
Listed -
Chance-Constrained Control with Lexicographic Deep Reinforcement Learning
Listed - Listed
-
RobustBench: a standardized adversarial robustness benchmark
Listed -
Time for AI to cross the human performance range in ImageNet image classification
Listed - Listed
-
Time for AI to cross the human performance range in Go
Listed