The catalog, page 20
Records 4,751 to 5,000 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Identifying Adversarial Attacks on Text Classifiers
Listed - Listed
-
[linkpost] Sharing powerful AI models: the emerging paradigm of structured access
Listed -
Action: Help expand funding for AI Safety by coordinating on NSF response
Listed -
Book non-review: The Dawn of Everything
Listed -
Estimating training compute of Deep Learning models
Listed -
Priors, Hierarchy, and Information Asymmetry for Skill Transfer in Reinforcement Learning
Listed -
Safe Deep RL in 3D Environments using Human Feedback
Listed -
Safety-Aware Multi-Agent Apprenticeship Learning
Listed -
What's Up With Confusingly Pervasive Goal Directedness?
Listed -
AI acceleration from a safety perspective: Trade-offs and considerations
Listed -
European Union AI Development and Governance Partnerships
Listed -
Improving Behavioural Cloning with Human-Driven Dynamic Dataset Augmentation
Listed -
My plan for a “Most Important Century” reading group
Listed - Listed
-
Clarifications about structural risk from AI
Listed -
Empowerment and Stakeholder Management
Listed -
Solving Dynamic Principal-Agent Problems with a Rationally Inattentive Principal
Listed -
Spurious normativity enhances learning of compliance and enforcement behavior in artificial agents
Listed -
The longtermist AI governance landscape: a basic overview
Listed -
Challenges with Breaking into MIRI-Style Research
Listed -
Different way classifiers can be diverse
Listed -
FLI launches Worldbuilding Contest with $100,000 in prizes
Listed - Listed
-
PIBBSS Fellowship: Bounty for Referrals & Deadline Extension
Listed -
Planning Not to Talk: Multiagent Systems that are Robust to Communication Loss
Listed -
Scalar reward is not enough for aligned AGI
Listed -
Truthful LMs as a warm-up for aligned AGI
Listed - Listed
- Listed
-
The Greedy Doctor Problem... turns out to be relevant to the ELK problem?
Listed -
Tools and Practices for Responsible AI Engineering
Listed -
EU's importance for AI governance is conditional on AI trajectories - a case study
Listed -
ULTRA: A Data-driven Approach for Recommending Team Formation in Response to Proposal Calls
Listed -
New year, new research agenda post
Listed -
Revelation of Task Difficulty in AI-aided Education
Listed -
The Concept of Criticality in AI Safety
Listed -
Value extrapolation partially resolves symbol grounding
Listed -
An Open Philanthropy grant proposal: Causal representation learning of human preferences
Listed -
Danijar Hafner - Gaming our way to AGI-by Towards Data Science-video_id Bgz9eMcE5Do-date 20220112
Listed -
Future ML Systems Will Be Qualitatively Different
Listed - Listed
-
The Turing Trap: The Promise & Peril of Human-Like Artificial Intelligence
Listed -
Understanding the two-head strategy for teaching ML to answer questions honestly
Listed -
uploading people for alignment purposes
Listed -
What is the role of Bayesian ML for AI alignment/safety?
Listed -
Why it matters if "ideas get harder to find"
Listed -
Critiquing Scasper's Definition of Subjunctive Dependence
Listed -
The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Listed - Listed
-
Modeling Human-AI Team Decision Making
Listed -
How artistic ideas could get harder to find
Listed -
questions about the cosmos and rich computations
Listed - Listed
-
Brain Efficiency: Much More than You Wanted to Know
Listed -
Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
Listed -
Importance of foresight evaluations within ELK
Listed - Listed
-
Signaling isn't about signaling, it's about Goodhart
Listed -
A Reaction to Wolfgang Schwarz's "On Functional Decision Theory"
Listed - Listed
-
brittle physics and the nature of X-risks
Listed -
Consider trying the ELK contest (I am)
Listed - Listed
-
Robust Self-Supervised Audio-Visual Speech Recognition
Listed - Listed
-
China’s New AI Governance Initiatives Shouldn’t Be Ignored
Listed - Listed
-
Promising posts on AF that have fallen through the cracks
Listed - Listed
-
Apply for research internships at ARC!
Listed -
Execute Order 66: Targeted Data Poisoning for Reinforcement Learning
Listed -
Have I done enough planning or should I plan more?
Listed -
How an alien theory of mind might be unlearnable
Listed - Listed
- Listed
-
$1000 USD prize - Circular Dependency of Counterfactuals
Listed - Listed
-
Locating and Editing Factual Associations in GPT
Listed -
Will the EU regulations on AI matter to the rest of the world?
Listed -
Counterexamples to some ELK proposals
Listed -
Eliciting Latent Knowledge Via Hypothetical Sensors
Listed -
We need a theory of anthropic measure binding
Listed - Listed
-
Reverse-engineering using interpretability
Listed -
Gradient Hacking via Schelling Goals
Listed - Listed
-
13 Very Different Stances on AGI
Listed -
AGI alignment results from a series of aligned actions
Listed -
less quantum immortality? • carado.moe
Listed -
Why don't governments seem to mind that companies are explicitly trying to make AGIs?
Listed -
database transactions: you guessed it, it's WASM again
Listed -
My Overview of the AI Alignment Landscape: Threat Models
Listed -
non-scarce compute means moral patients might not get optimized out
Listed -
thinking about psi: as a more general json
Listed - Listed
- Listed
-
Should the EA community have a DL engineering fellowship?
Listed -
2021 AI Alignment Literature Review and Charity Comparison
Listed -
Free Guy, a rom-com on the moral patienthood of digital sentience
Listed -
Reply to Eliezer on Biological Anchors
Listed - Listed
-
Why don't governments seem to mind that companies are explicitly trying to make AGIs?
Listed -
Worst-case thinking in AI alignment
Listed -
A Mathematical Framework for Transformer Circuits
Listed - Listed
-
Potential gears level explanations of smooth progress
Listed - Listed
-
Worldbuilding exercise: The Highwayverse.
Listed - Listed
-
DB-BERT: a Database Tuning Tool that "Reads the Manual"
Listed -
Demanding and Designing Aligned Cognitive Architectures
Listed -
Introducing a New Course on the Economics of AI
Listed -
Researcher incentives cause smoother progress on benchmarks
Listed -
Demanding and Designing Aligned Cognitive Architectures
Listed -
Don't Influence the Influencers!
Listed -
[Extended Deadline: Jan 23rd] Announcing the PIBBSS Summer Research Fellowship
Listed -
Disentangling Perspectives On Strategy-Stealing in AI Safety
Listed -
Exploring Decision Theories With Counterfactuals and Dynamic Agent Self-Pointers
Listed - Listed
- Listed
-
WebGPT: Browser-assisted question-answering with human feedback
Listed -
Elicitation for Modeling Transformative AI Risks
Listed -
Evidence Sets: Towards Inductive-Biases based Analysis of Prosaic AGI
Listed -
Housing Markets, Satisficers, and One-Track Goodhart
Listed -
Motivations, Natural Selection, and Curriculum Engineering
Listed -
Reviews of "Is power-seeking AI an existential risk?"
Listed -
Reviews of “Is power-seeking AI an existential risk?”
Listed - Listed
-
AI Safety: Applying to Graduate Studies
Listed -
Framing approaches to alignment and the hard problem of AI cognition
Listed -
My Overview of the AI Alignment Landscape: A Bird's Eye View
Listed -
My Overview of the AI Alignment Landscape: A Bird’s Eye View
Listed -
Ngo’s view on alignment difficulty
Listed - Listed
-
ARC's first technical report: Eliciting Latent Knowledge
Listed -
Consequentialism & corrigibility
Listed -
Decision Theory Breakdown—Personal Attempt at a Review
Listed -
Filling gaps in trustworthy development of AI
Listed -
Interlude: Agents as Automobiles
Listed -
Ngo's view on alignment difficulty
Listed -
Programmatic Reward Design by Example
Listed -
Should we rely on the speed prior for safety?
Listed -
The Natural Abstraction Hypothesis: Implications and Evidence
Listed - Listed
-
Hard-Coding Neural Computation
Listed -
Language Model Alignment Research Internships
Listed - Listed
- Listed
-
Summary of the Acausal Attack Issue for AIXI
Listed -
Understanding and controlling auto-induced distributional shift
Listed -
What’s the backward-forward FLOP ratio for Neural Networks?
Listed -
Redwood's Technique-Focused Epistemic Strategy
Listed -
Some abstract, non-technical reasons to be non-maximally-pessimistic about AI alignment
Listed -
Teaser: Hard-coding Transformer Models
Listed -
Moore's Law, AI, and the pace of progress
Listed - Listed
-
What role should evolutionary analogies play in understanding AI takeoff speeds?
Listed -
What role should evolutionary analogies play in understanding AI takeoff speeds?
Listed - Listed
- Listed
-
The Promise and Peril of Finite Sets
Listed -
There is essentially one best-validated theory of cognition.
Listed -
TV shows I wish I could watch: Intergalactic Immigration Wars
Listed -
Understanding Gradient Hacking
Listed -
[MLSN #2]: Adversarial Training
Listed -
Conversation on technology forecasting and gradualism
Listed -
Conversation on technology forecasting and gradualism
Listed -
emotionally appreciating grand political visions
Listed -
EU AI Act now has a section on general purpose AI systems
Listed -
freedom and diversity in Albion's Seed
Listed -
Introduction to inaccessible information
Listed - Listed
-
non-interfering superintelligence and remaining philosophical progress: a deterministic utopia
Listed -
PixMix: Dreamlike Pictures Comprehensively Improve Safety Measures
Listed - Listed
-
Supervised learning and self-modeling: What's "superhuman?"
Listed -
unoptimal superintelligence doesn't lose
Listed -
[AN #170]: Analyzing the argument for risk from power-seeking AI
Listed -
Creating Interactive Agents with Imitation Learning
Listed -
Finding the multiple ground truths of CoinRun and image classification
Listed -
Improving language models by retrieving from trillions of tokens
Listed -
Some thoughts on why adversarial training might be useful
Listed -
Considerations on interaction between AI and expected value of the future
Listed -
Exterminating humans might be on the to-do list of a Friendly AI
Listed -
HIRING: Inform and shape a new project on AI safety at Partnership on AI
Listed -
MESA: Offline Meta-RL for Safe Adaptation and Fault Tolerance
Listed -
More Christiano, Cotra, and Yudkowsky on AI progress
Listed -
Theoretical Neuroscience For Alignment Theory
Listed -
Why Describing Utopia Goes Badly
Listed -
A Framework to Explain Bayesian Models
Listed -
A Possible Resolution To Spurious Counterfactuals
Listed -
Are there alternative to solving value transfer and extrapolation?
Listed -
Candidate for “highest-stakes question of the next several months” (rare hot take)
Listed -
Contribute by facilitating the AGI Safety Fundamentals Programme
Listed -
Declustering, reclustering, and filling in thingspace
Listed -
Do neural networks learn human concepts?
Listed -
Information bottleneck for counterfactual corrigibility
Listed -
ML Alignment Theory Program under Evan Hubinger
Listed -
Modeling Failure Modes of High-Level Machine Intelligence
Listed -
More Christiano, Cotra, and Yudkowsky on AI progress
Listed -
Retrospective on the Summer 2021 AGI Safety Fundamentals
Listed -
Are limited-horizon agents a good heuristic for the off-switch problem?
Listed -
Behavior Cloning is Miscalibrated
Listed -
Interpreting Yudkowsky on Deep vs Shallow Knowledge
Listed - Listed
-
Agency: What it is and why it matters
Listed - Listed
-
Misc. questions about EfficientZero
Listed -
Shulman and Yudkowsky on AI progress
Listed -
Shulman and Yudkowsky on AI progress
Listed - Listed
- Listed
-
$100/$50 rewards for good references
Listed -
[Linkpost] A General Language Assistant as a Laboratory for Alignment
Listed -
Biology-Inspired AGI Timelines: The Trick That Never Works
Listed - Listed
-
Formalizing Policy-Modification Corrigibility
Listed -
Shulman and Yudkowsky on AI progress
Listed - Listed
-
AXRP Episode 12 - AI Existential Risk with Paul Christiano
Listed - Listed
- Listed
- Listed
- Listed
-
A General Language Assistant as a Laboratory for Alignment
Listed -
Biology-Inspired AGI Timelines: The Trick That Never Works
Listed - Listed
- Listed
-
Hypotheses about Finding Knowledge and One-Shot Causal Entanglements
Listed -
On the Expressivity of Markov Reward
Listed -
AI Governance Fundamentals - Curriculum and Application
Listed -
Did life get better during the pre-industrial era? (Ehhhh)
Listed -
From AI for People to AI for the World and the Universe
Listed -
Infra-Bayesian physicalism: a formal theory of naturalized induction
Listed -
Infra-Bayesian physicalism: proofs part I
Listed -
Infra-Bayesian physicalism: proofs part II
Listed -
My take on higher-order game theory
Listed -
Pyramid Adversarial Training Improves ViT Performance
Listed -
Visible Thoughts Project and Bounty Announcement
Listed -
Visible Thoughts Project and Bounty Announcement
Listed -
AI Governance Course - Curriculum and Application
Listed -
Comments on Allan Dafoe on AI Governance
Listed -
How to measure FLOP/s for Neural Networks empirically?
Listed - Listed
-
Question/Issue with the 5/10 Problem
Listed -
Redwood Research is hiring for several roles
Listed -
Soares, Tallinn, and Yudkowsky discuss AGI cognition
Listed -
Soares, Tallinn, and Yudkowsky discuss AGI cognition
Listed -
Soares, Tallinn, and Yudkowsky discuss AGI cognition
Listed