The catalog, page 16
Records 3,751 to 4,000 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
2022 Expert Survey on Progress in AI
Listed -
Announcing the SPT Model Web App for AI Governance
Listed -
Convergence Towards World-Models: A Gears-Level Model
Listed -
Does China have AI alignment resources/institutions? How can we prioritize creating more?
Listed -
Surprised by ELK report's counterexample to Debate, IDA
Listed - Listed
- Listed
-
What do ML researchers think about AI in 2022?
Listed -
What do ML researchers think about AI in 2022?
Listed -
Why we need a new agency to regulate advanced artificial intelligence
Listed -
Would "Manhattan Project" style be beneficial or deleterious for AI Alignment?
Listed -
Ajeya's TAI timeline shortened from 2050 to 2040
Listed - Listed
-
Externalized reasoning oversight: a research direction for language model alignment
Listed -
Precursor checking for deceptive alignment
Listed - Listed
- Listed
-
tiling the cosmos might be unavoidable
Listed -
What if AI development goes well?
Listed -
Exploratory survey on psychology of AI risk perception
Listed -
Information in risky technology races
Listed -
isn't it weird that we have a chance at all?
Listed -
Law-Following AI 4: Don't Rely on Vicarious Liability
Listed -
Two-year update on my personal AI timelines
Listed -
Announcing the GovAI Policy Team
Listed -
Few-shot Adaptation Works with UnpredicTable Data
Listed -
Preventing an AI-related catastrophe
Listed -
chinchilla's wild implications
Listed -
Wanted: Notation for credal resilience
Listed -
AI timelines by bio anchors: the debate in one place
Listed -
How transparency changed over time
Listed - Listed
-
Abstracting The Hardness of Alignment: Unbounded Atomic Optimization
Listed -
Closing the Feedback Loop on AI Safety Research.
Listed -
Comparing Four Approaches to Inner Alignment
Listed -
Conjecture: Internal Infohazard Policy
Listed - Listed
-
AI Alignment is intractable (and we humans should stop working on it)
Listed - Listed
-
Efficient training of language models to fill in the middle
Listed -
Latent Properties of Lifelong Learning Systems
Listed -
Safety without oppression: an AI governance problem
Listed - Listed
-
AGI ruin scenarios are likely (and disjunctive)
Listed -
FLI is hiring a new Director of US Policy
Listed - Listed
- Listed
-
Moral strategies at different capability levels
Listed -
Principles of Privacy for Alignment Research
Listed -
Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks
Listed -
Unifying Bargaining Notions (2/2)
Listed -
Active Inference as a formalisation of instrumental convergence
Listed - Listed
-
How much should you optimize for the short-timelines scenario?
Listed -
Humanity’s vast future and its implications for cause prioritization
Listed -
Neartermists should consider AGI timelines in their spending decisions
Listed -
NeurIPS ML Safety Workshop 2022
Listed - Listed
-
«Boundaries», Part 1: a key missing concept from utility theory
Listed -
A hazard analysis framework for code synthesis large language models
Listed -
AGI Safety Needs People With All Skillsets!
Listed -
Does agent foundations cover all future ML systems?
Listed -
How much should we worry about mesa-optimization challenges?
Listed -
Reward is not the optimization target
Listed -
Unifying Bargaining Notions (1/2)
Listed -
Brainstorm of things that could force an AI team to burn their lead
Listed -
Finding Skeletons on Rashomon Ridge
Listed -
We Did AGISF’s 8-week Course in 3 Days. Here’s How it Went
Listed -
Robustness to Scaling Down: More Important Than I Thought
Listed -
Which singularity schools plus the no singularity school was right?
Listed -
Connor Leahy on Conjecture and Dying with Dignity
Listed -
Maybe AI risk shouldn't affect your life plan all that much
Listed -
Reasons I’ve been hesitant about high levels of near-ish AI risk
Listed -
[AN #173] Recent language model results from DeepMind
Listed -
A "Solipsistic" Repugnant Conclusion
Listed -
Conditioning Generative Models with Restrictions
Listed -
How much to optimize for the short-timelines scenario?
Listed -
Our Existing Solutions to AGI Alignment (semi-safe)
Listed -
UK AI Policy Report: Content, Summary, and its Impact on EA Cause Areas
Listed -
Discriminator-Weighted Offline Imitation Learning from Suboptimal Demonstrations
Listed -
How to Diversify Conceptual AI Alignment: the Model Behind Refine
Listed -
How to Diversify Conceptual Alignment: the Model Behind Refine
Listed - Listed
-
The Need for a Meta-Architecture for Robot Autonomy
Listed -
A Critique of AI Alignment Pessimism
Listed -
Abram Demski's ELK thoughts and proposal - distillation
Listed -
Bounded complexity of solving ELK and its implications
Listed -
Help ARC evaluate capabilities of current language models (still need people)
Listed - Listed
-
A distillation of Evan Hubinger's training stories (for SERI MATS)
Listed -
A Survey of the Potential Long-term Impacts of AI
Listed -
Boolean Decision Rules for Reinforcement Learning Policy Summarisation
Listed -
Conditioning Generative Models for Alignment
Listed -
Deception?! I ain’t got time for that!
Listed -
Forecasting ML Benchmarks in 2023
Listed -
GPT-2 as step toward general intelligence (Alexander, 2019)
Listed -
How Interpretability can be Impactful
Listed -
Quantilizers and Generative Models
Listed -
Training goals for large language models
Listed -
What’s so dangerous about AI anyway? – Or: What it means to be a superintelligence
Listed -
Why EAs are skeptical about AI Safety
Listed -
Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover
Listed -
Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover
Listed -
Do EA folks think that a path to zero AGI development is feasible or worthwhile for safety from AI?
Listed -
Examples of AI Increasing AI Progress
Listed - Listed
-
Four questions I ask AI safety researchers
Listed - Listed
-
Why you might expect homogeneous take-off: evidence from ML research
Listed - Listed
-
All AGI safety questions welcome (especially basic ones) [July 2022]
Listed - Listed
-
Does the idea of AGI that benevolently control us appeal to EA folks?
Listed -
How would a language model become goal-directed?
Listed -
Perceiver AR: general-purpose, long-context autoregressive generation
Listed -
A note about differential technological development
Listed -
More to explore on 'Risks from Artificial Intelligence'
Listed - Listed
-
Safety Implications of LeCun's path to machine intelligence
Listed -
What if we don't need a "Hard Left Turn" to reach AGI?
Listed -
Circumventing interpretability: How to defeat mind-readers
Listed -
Humans provide an untapped wealth of evidence about alignment
Listed -
It's OK not to go into AI (for students)
Listed -
Resilience Via Fragmented Power
Listed -
Why policymakers should beware claims of new "arms races" (Bulletin of the Atomic Scientists)
Listed -
Artificial Sandwiching: When can we test scalable alignment protocols without humans?
Listed -
Deep learning curriculum for large language model alignment
Listed -
Goal Alignment Is Robust To the Sharp Left Turn
Listed -
Making decisions using multiple worldviews
Listed -
MIRI Conversations: Technology Forecasting & Gradualism (Distillation)
Listed -
Pile of Law and Law-Following AI
Listed -
Searle vs Bostrom: crucial considerations for EA AI work?
Listed -
Slowing down AI progress is an underexplored alignment strategy
Listed -
Acceptability Verification: A Research Agenda
Listed -
AI ethics: the case for including animals (my first published paper, Peter Singer's first on AI)
Listed -
Mosaic and Palimpsests: Two Shapes of Research
Listed -
On how various plans miss the hard bits of the alignment challenge
Listed -
Recommendations for non-technical books on AI?
Listed -
Response to Blake Richards: AGI, generality, alignment, & loss functions
Listed -
What is wrong with this approach to corrigibility?
Listed - Listed
-
Intuitive physics learning in a deep-learning model inspired by developmental psychology
Listed -
Language Models (Mostly) Know What They Know
Listed - Listed
-
Immanuel Kant and the Decision Theory App Store
Listed -
Grouped Loss may disfavor discontinuous capabilities
Listed -
Making it harder for an AGI to "trick" us, with STVs
Listed -
Report from a civilizational observer on Earth
Listed -
Train first VS prune first in neural networks.
Listed -
Visualizing Neural networks, how to blame the bias
Listed -
Reinforcement Learner Wireheading
Listed -
Research Notes: What are we aligning for?
Listed -
Human values & biases are inaccessible to the genome
Listed -
Principles for Alignment/Agency Projects
Listed - Listed
-
Safety considerations for online generative modeling
Listed - Listed
-
How humanity would respond to slow takeoff, with takeaways from the entire COVID-19 pandemic
Listed -
Inferring and Conveying Intentionality: Beyond Numerical Rewards to Logical Intentions
Listed -
Introducing the Fund for Alignment Research (We're Hiring!)
Listed -
Introducing the Fund for Alignment Research (We're Hiring!)
Listed -
Outer vs inner misalignment: three framings
Listed -
The History of AI Rights Research
Listed -
What work has been done on the post-AGI distribution of wealth?
Listed -
[AN #172] Sorry for the long hiatus!
Listed -
A central AI alignment problem: capabilities generalization, and the sharp left turn
Listed -
Facilitator Help Wanted for Columbia EA AI Safety Groups
Listed -
The curious case of Pretty Good human inner/outer alignment
Listed -
A compressed take on recent disagreements
Listed - Listed
-
Benchmark for successful concept extrapolation/avoiding goal misgeneralization
Listed -
Future Matters #3: digital sentience, AGI ruin, and forecasting track records
Listed -
Human-centred mechanism design with Democratic AI
Listed -
Is General Intelligence "Compact"?
Listed -
New US Senate Bill on X-Risk Mitigation [Linkpost]
Listed -
Please help us communicate AI xrisk. It could save the world.
Listed -
Remaking EfficientZero (as best I can)
Listed -
Decision theory and dynamic inconsistency
Listed -
Why AGI Timeline Research/Discourse Might Be Overrated
Listed -
[Linkpost] Existential Risk Analysis in Empirical Research Papers
Listed -
Components of Strategic Clarity [Strategic Perspectives on Long-term AI Governance, #2]
Listed -
Follow along with Columbia EA's Advanced AI Safety Fellowship!
Listed -
Naive Hypotheses on AI Alignment
Listed -
Research + Reality Graphing to Support AI Policy (and more): Summary of a Frozen Project
Listed -
Strategic Perspectives on Transformative AI Governance: Introduction
Listed -
The Linguistic Blind Spot of Value-Aligned Agency, Natural and Artificial
Listed -
The Tree of Life: Stanford AI Alignment Theory of Change
Listed -
When 2/3rds of the world goes against you
Listed -
AI safety university groups: a promising opportunity to reduce existential risk
Listed -
Artificial Intelligence, Morality, and Sentience (AIMS) Survey: 2021
Listed -
AXRP Episode 16 - Preparing for Debate AI with Geoffrey Irving
Listed -
generalized values: testing for patterns in computation
Listed - Listed
- Listed
-
Trends in GPU price-performance
Listed -
What Is The True Name of Modularity?
Listed -
$500 bounty for alignment contest ideas
Listed -
(Even) More Early-Career EAs Should Try AI Safety Technical Research
Listed -
Announcing the Harvard AI Safety Team
Listed -
Forecasting Future World Events with Neural Networks
Listed -
Formal Philosophy and Alignment Possible Projects
Listed -
Quick survey on AI alignment resources
Listed -
The Track Record of Futurists Seems ... Fine
Listed -
What are some current, already present challenges from AI?
Listed -
Gradient hacking: definitions and examples
Listed - Listed
-
The inordinately slow spread of good AGI conversations in ML
Listed -
Will Capabilities Generalise More?
Listed -
Doom doubts - is inner alignment a likely problem?
Listed -
Four reasons I find AI safety emotionally compelling
Listed -
Some alternative AI safety research projects
Listed - Listed
- Listed
-
Announcing Epoch: A research organization investigating the road to Transformative AI
Listed -
Announcing Epoch: A research organization investigating the road to Transformative AI
Listed -
Announcing the Inverse Scaling Prize ($250k Prize Pool)
Listed -
Auditing Visualizations: Transparency Methods Struggle to Detect Anomalous Behavior
Listed -
Deliberation Everywhere: Simple Examples
Listed - Listed
-
Exploring Mild Behaviour in Embedded Agents
Listed -
Military Artificial Intelligence as Contributor to Global Catastrophic Risk
Listed -
Parametrically Retargetable Decision-Makers Tend To Seek Power
Listed - Listed
-
Training Trace Priors and Speed Priors
Listed -
[LQ] Some Thoughts on Messaging Around AI Risk
Listed -
AI-Written Critiques Help Humans Notice Flaws
Listed -
Conditioning Generative Models
Listed -
7 essays on Building a Better Future
Listed -
Multi-Modal and Multi-Factor Branching Time Active Inference
Listed -
Raphaël Millière on the Limits of Deep Learning and AI x-risk skepticism
Listed -
Updated Deference is not a strong argument against the utility uncertainty approach to alignment
Listed -
20 Critiques of AI Safety That I Found on Twitter
Listed -
Formalizing the Problem of Side Effect Regularization
Listed -
Half-baked ideas thread (EA / AI Safety)
Listed -
Learning to play Minecraft with Video PreTraining
Listed - Listed
- Listed
-
On Avoiding Power-Seeking by Artificial Intelligence
Listed -
Confusion about neuroscience/cognitive science as a danger for AI Alignment
Listed -
Google's new text-to-image model - Parti, a demonstration of scaling benefits
Listed -
Reflection Mechanisms as an Alignment target: A survey
Listed -
A Quick List of Some Problems in AI Alignment As A Field
Listed -
Getting from an unaligned AGI to an aligned AGI?
Listed -
Getting from an unaligned AGI to an aligned AGI?
Listed -
Love and AI: Relational Brain/Mind Dynamics in AI Development
Listed -
Technical AI safety in the United Arab Emirates
Listed -
The inordinately slow spread of good AGI conversations in ML
Listed -
Uncertainty Quantification for Competency Assessment of Autonomous Agents
Listed -
A Toy Model of Gradient Hacking
Listed -
BYOL-Explore: Exploration with Bootstrapped Prediction
Listed