The catalog, page 9
Records 2,001 to 2,250 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
New Artificial Intelligence quiz: can you beat ChatGPT?
Listed -
Robin Hanson’s latest AI risk position statement
Listed -
Situational awareness in Large Language Models
Listed -
state of my alignment research, and what needs work
Listed -
state of my alignment research, and what needs work
Listed -
The Waluigi Effect (mega-post)
Listed -
Why are counterfactuals elusive?
Listed -
Call to demand answers from Anthropic about joining the AI race
Listed - Listed
-
Game theory work on AI alignment with diverse AI systems, human individuals, & human groups?
Listed -
Joscha Bach on Synthetic Intelligence [annotated]
Listed -
Payor's Lemma in Natural Language
Listed -
Scoring forecasts from the 2016 “Expert Survey on Progress in AI”
Listed -
The View from 30,000 Feet: Preface to the Second EleutherAI Retrospective
Listed -
What are some sources related to big-picture AI strategy?
Listed -
Call for Cruxes by Rhyme, a Longtermist History Consultancy
Listed -
Call for Cruxes by Rhyme, a Longtermist History Consultancy
Listed -
Existential Risk from Power-Seeking AI
Listed -
Extreme GDP growth is a bad operating definition of "slow takeoff"
Listed -
Implied "utilities" of simulators are broad, dense, and shallow
Listed -
Inside the mind of a superhuman Go model: How does Leela Zero read ladders?
Listed -
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small
Listed -
on strong/general coherent agents
Listed -
Predictions for shard theory mechanistic interpretability results
Listed -
Problems of people new to AI safety and my project ideas to mitigate them
Listed -
Scoring forecasts from the 2016 “Expert Survey on Progress in AI”
Listed -
Some Variants of Sleeping Beauty
Listed - Listed
-
$20 Million in NSF Grants for Safety Research
Listed -
A mostly critical review of infra-Bayesianism
Listed - Listed
-
Heuristics on bias to action versus status quo?
Listed -
Performance guarantees in classical learning theory and infra-Bayesianism
Listed -
Power-seeking can be probable and predictive for trained agents
Listed -
Scarce Channels and Abstraction Coupling
Listed -
Some Things I Heard about AI Governance at EAG
Listed -
Transcript: Testing ChatGPT's Performance in Engineering
Listed -
What does Bing Chat tell us about AI risk?
Listed -
What does Bing Chat tell us about AI risk?
Listed -
[Simulators seminar sequence] #2 Semiotic physics - revamped
Listed -
Counting-down vs. counting-up coherence
Listed -
Seeking input on a list of AI books for broader audience
Listed -
Some thoughts pointing to slower AI take-off
Listed -
The idea of an "aligned superintelligence" seems misguided
Listed -
Why I think it's important to work on AI forecasting
Listed -
[Link Post] Cyber Digital Authoritarianism (National Intelligence Council Report)
Listed -
A library for safety research in conditioning on RLHF tasks
Listed -
A mechanistic explanation for SolidGoldMagikarp-like tokens in GPT2
Listed -
An economics of AI gov - best resources for
Listed -
How to ‘troll for good’: Leveraging IP for AI governance
Listed -
Incentives and Selection: A Missing Frame From AI Threat Discussions?
Listed -
some thoughts about terminal alignment
Listed -
The Preference Fulfillment Hypothesis
Listed - Listed
-
clarifying formal alignment implementation
Listed -
Cognitive Emulation: A Naive AI Safety Proposal
Listed -
Which is more important for reducing s-risks, researching on AI sentience or animal welfare?
Listed -
Would more model evals teams be good?
Listed -
2023 Stanford Existential Risks Conference
Listed -
Agents vs. Predictors: Concrete differentiating factors
Listed -
Christiano (ARC) and GA (Conjecture) Discuss Alignment Cruxes
Listed -
How major governments can help with the most important century
Listed -
How major governments can help with the most important century
Listed -
How popular is ChatGPT? Part 1: more popular than Taylor Swift
Listed -
Meta "open sources" LMs competitive with Chinchilla, PaLM, and code-davinci-002 (Paper)
Listed -
Retrospective on the 2022 Conjecture AI Discussions
Listed -
Sam Altman: "Planning for AGI and beyond"
Listed -
Training for corrigability: obvious problems?
Listed -
AI that shouldn't work, yet kind of does
Listed -
Automated Sandwiching & Quantifying Human-LLM Cooperation: ScaleOversight hackathon results
Listed - Listed
- Listed
-
Full Transcript: Eliezer Yudkowsky on the Bankless podcast
Listed - Listed
-
Searching for a model's concepts by their shape – a theoretical framework
Listed -
Taking a leave of absence from Open Philanthropy to work on AI safety
Listed - Listed
-
Cyborg Periods: There will be multiple AI transitions
Listed - Listed
-
Intervening in the Residual Stream
Listed - Listed
-
Power-Seeking = Minimising free energy
Listed - Listed
-
The shallow reality of 'deep learning theory'
Listed -
Video/animation: Neel Nanda explains what mechanistic interpretability is
Listed -
[Preprint] Pretraining Language Models with Human Preferences
Listed -
A proof of inner Löb's theorem
Listed -
Basic facts about language models during training
Listed -
Breaking the Optimizer’s Curse, and Consequences for Existential Risks and Value Learning
Listed -
Deceptive Alignment is <1% Likely by Default
Listed -
Does most of your impact come from what you do soon?
Listed -
EIS X: Continual Learning, Modularity, Compression, and Biological Brains
Listed -
Instrumentality makes agents agenty
Listed -
Pretraining Language Models with Human Preferences
Listed -
What is it like doing AI safety work?
Listed -
You're not a simulation, 'cause you're hallucinating
Listed - Listed
- Listed
-
A circuit for Python docstrings in a 4-layer attention-only transformer
Listed -
AGI doesn't need understanding, intention, or consciousness in order to kill us, only intelligence
Listed -
Behavioral and mechanistic definitions (often confuse AI alignment discussions)
Listed -
Bing finding ways to bypass Microsoft's filters without being asked. Is it reproducible?
Listed - Listed
-
EIS IX: Interpretability and Adversaries
Listed -
Emergent Deception and Emergent Optimization
Listed - Listed
-
There are no coherence theorems
Listed -
There are no coherence theorems
Listed -
Validator models: A simple approach to detecting goodharting
Listed -
What AI companies can do today to help with the most important century
Listed -
What AI companies can do today to help with the most important century
Listed -
What to think when a language model tells you it's sentient
Listed -
A Neural Network undergoing Gradient-based Training as a Complex System
Listed - Listed
-
Does novel understanding imply novel agency / values?
Listed -
EIS VIII: An Engineer’s Understanding of Deceptive Alignment
Listed -
AGI in sight: our look at the game board
Listed -
EIS VII: A Challenge for Mechanists
Listed -
Interview with Roman Yampolskiy about AGI on The Reality Check
Listed -
Parametrically retargetable decision-makers tend to seek power
Listed -
Should ChatGPT make us downweight our belief in the consciousness of non-human animals?
Listed -
AI Safety Info Distillation Fellowship
Listed - Listed
-
EIS VI: Critiques of Mechanistic Interpretability Work in AI Safety
Listed -
How good/bad is the new Bing AI for the world?
Listed -
How should AI systems behave, and who should decide? [OpenAI blog]
Listed -
I Am Scared of Posting Negative Takes About Bing's AI
Listed -
One-layer transformers aren’t equivalent to a set of skip-trigrams
Listed -
Powerful mesa-optimisation is already here
Listed -
The public supports regulating AI for safety
Listed -
Two problems with ‘Simulators’ as a frame
Listed -
don't censor yourself, silly !
Listed -
EIS V: Blind Spots In AI Safety Interpretability Research
Listed -
Non-Unitary Quantum Logic -- SERI MATS Research Sprint
Listed -
Paper: The Capacity for Moral Self-Correction in Large Language Models (Anthropic)
Listed -
Pretraining Language Models with Human Preferences
Listed -
a narrative explanation of the QACI alignment plan
Listed -
AI alignment researchers may have a comparative advantage in reducing s-risks
Listed -
Don't accelerate problems you're trying to solve
Listed -
EIS IV: A Spotlight on Feature Attribution/Saliency
Listed -
EIS IV: A Spotlight on Feature Attribution/Saliency
Listed -
Huh. Bing thing got me real anxious about AI. Resources to help with that please?
Listed -
Order Matters for Deceptive Alignment
Listed -
The Capacity for Moral Self-Correction in Large Language Models
Listed -
EIS III: Broad Critiques of Interpretability Research
Listed - Listed
-
Explaining SolidGoldMagikarp by looking at it from random directions
Listed -
Qualities that alignment mentors value in junior researchers
Listed -
SolidGoldMagikarp III: Glitch token archaeology
Listed -
The Cave Allegory Revisited: Understanding GPT's Worldview
Listed -
The Linguistic Blind Spot of Value-Aligned Agency, Natural and Artificial
Listed -
The Linguistic Blind Spot of Value-Aligned Agency, Natural and Artificial
Listed -
Whole Bird Emulation requires Quantum Mechanics
Listed -
4 ways to think about democratizing AI [GovAI Linkpost]
Listed -
is intelligence program inversion?
Listed -
LLM Basics: Embedding Spaces - Transformer Token Vectors Are Not Points in Space
Listed -
Morphological intelligence, superhuman empathy, and ethical arbitration
Listed -
fuzzies & utils: check that you're getting either
Listed -
High impact job opportunity at ARIA (UK)
Listed -
Jobs that can help with the most important century
Listed -
The conceptual Doppelgänger problem
Listed -
Why almost every RL agent does learned optimization
Listed - Listed
-
GPT is dangerous because it is useful at all
Listed -
my takeoff speeds? depends how you define that
Listed -
Shortening Timelines: There's No Buffer Anymore
Listed -
The Importance of AI Alignment, explained in 5 points
Listed - Listed
- Listed
-
A proposed method for forecasting transformative AI
Listed -
Conditioning Predictive Models: Open problems, Conclusion, and Appendix
Listed - Listed
-
FLI Podcast: Connor Leahy on AI Progress, Chimps, Memes, and Markets (Part 1/3)
Listed -
Jobs that can help with the most important century
Listed -
Many important technologies start out as science fiction before becoming real
Listed -
Mechanism Design for AI Safety - Agenda Creation Retreat
Listed -
Why I’m not working on {debate, RRM, ELK, natural abstractions}
Listed -
Anomalous tokens reveal the original identities of Instruct models
Listed -
Anomalous tokens reveal the original identities of Instruct models
Listed -
Apply to the Cambridge ML for Alignment Bootcamp (CaMLAB) [26 March - 8 April]
Listed - Listed
-
Conditioning Predictive Models: Deployment strategy
Listed -
Do the Safety Properties of Powerful AI Systems Need to be Adversarially Robust? Why?
Listed -
EIS II: What is “Interpretability”?
Listed -
EIS II: What is “Interpretability”?
Listed -
Notes on the Mathematics of LLM Architectures
Listed -
On Developing a Mathematical Theory of Interpretability
Listed -
Security Mindset - Fire Alarms and Trigger Signatures
Listed - Listed
-
Technological developments that could increase risks from nuclear weapons: A shallow review
Listed -
Technology is Power: Raising Awareness Of Technological Risks
Listed -
The Engineer’s Interpretability Sequence (EIS) I: Intro
Listed -
Using PICT against PastaGPT Jailbreaking
Listed -
A (EtA: quick) note on terminology: AI Alignment != AI x-safety
Listed -
A (EtA: quick) note on terminology: AI Alignment != AI x-safety
Listed -
A multi-disciplinary view on AI safety research
Listed -
Conditioning Predictive Models: Interactions with other approaches
Listed -
Dear Anthropic people, please don't release Claude
Listed -
[ASoT] Policy Trajectory Visualization
Listed - Listed
-
Conditioning Predictive Models: Making inner alignment as easy as possible
Listed - Listed
-
How evals might (or might not) prevent catastrophic risks from AI
Listed -
OpenAI/Microsoft announce "next generation language model" integrated into Bing/Edge
Listed -
Review of AI Alignment Progress
Listed -
so you think you're not qualified to do technical alignment research?
Listed -
so you think you're not qualified to do technical alignment research?
Listed - Listed
- Listed
-
A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations
Listed -
Addendum: More Efficient FFNs via Attention
Listed -
Conditioning Predictive Models: The case for competitiveness
Listed -
Decision Transformer Interpretability
Listed -
Donation recommendations for xrisk + ai safety
Listed -
Early situational awareness and its implications, a story
Listed - Listed
-
Gradient surfing: the hidden role of regularization
Listed -
Launching The Collective Intelligence Project: Whitepaper and Pilots
Listed -
SolidGoldMagikarp II: technical details and more recent findings
Listed -
Are short timelines actually bad?
Listed -
Call for submissions: AI Safety Special Session at the Conference on Artificial Life (ALIFE 2023)
Listed -
Evaluations (of new AI Safety researchers) can be noisy
Listed -
Modal Fixpoint Cooperation without Löb's Theorem
Listed -
Questions about AI that bother me
Listed -
SolidGoldMagikarp (plus, prompt generation)
Listed -
A discussion with ChatGPT on value-based models vs. large language models, etc..
Listed -
Attribution Patching: Activation Patching At Industrial Scale
Listed -
AXRP Episode 19 - Mechanistic Interpretability with Neel Nanda
Listed -
Criticism Thread: What things should OpenPhil improve on?
Listed -
Empathy as a natural consequence of learnt reward models
Listed -
Mech Interp Project Advising Call: Memorisation in GPT-2 Small
Listed -
Some miscellaneous thoughts on ChatGPT, stories, and mechanical interpretability
Listed -
An audio version of the alignment problem from a deep learning perspective by Richard Ngo Et Al
Listed -
Assessing China's importance as an AI superpower
Listed -
ChatGPT: Tantalizing afterthoughts in search of story trajectories [induction heads]
Listed -
Focus on the places where you feel shocked everyone’s dropping the ball
Listed -
Google invests $300mn in artificial intelligence start-up Anthropic | FT
Listed -
Many AI governance proposals have a tradeoff between usefulness and feasibility
Listed -
What I mean by “alignment is in large part about making cognition aimable at all”
Listed -
40,000 reasons to worry about AI safety
Listed -
A Brief Overview of AI Safety/Alignment Orgs, Fields, Researchers, and Resources for ML Researchers
Listed -
Conditioning Predictive Models: Large language models as predictors
Listed -
Conditioning Predictive Models: Outer alignment via careful conditioning
Listed -
Interviews with 97 AI Researchers: Quantitative Analysis
Listed -
More findings on maximal data dimension
Listed -
Normative vs Descriptive Models of Agency
Listed -
Predicting researcher interest in AI alignment
Listed -
Research agenda: Formalizing abstractions of computations
Listed -
Retrospective on the AI Safety Field Building Hub
Listed -
Temporally Layered Architecture for Adaptive, Distributed and Continuous Control
Listed