The catalog, page 11
Records 2,501 to 2,750 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
What AI Safety Materials Do ML Researchers Find Compelling?
Listed -
Can we efficiently distinguish different mechanisms?
Listed -
How to Catch a ChatGPT Cheat: 7 Practical Tips
Listed -
I have thousands of copies of HPMOR in Russian. How to use them with the most impact?
Listed -
Institutions Cannot Restrain Dark-Triad AI Exploitation
Listed -
My Reservations about Discovering Latent Knowledge (Burns, Ye, et al)
Listed -
Reflections on my 5-month alignment upskilling grant
Listed -
The AIA and its Brussels Effect
Listed -
Why The Focus on Expected Utility Maximisers?
Listed -
Air-gapping evaluation and support
Listed -
An overview of some promising work by junior alignment researchers
Listed -
Analogies between Software Reverse Engineering and Mechanistic Interpretability
Listed -
Announcing: The Independent AI Safety Registry
Listed -
Avoiding perpetual risk from TAI
Listed -
Coherent extrapolated dreaming
Listed -
How long till Brussels?: A light investigation into the Brussels Gap
Listed - Listed
-
Slightly against aligning with neo-luddites
Listed -
[Hebbian Natural Abstractions] Mathematical Foundations
Listed -
Accurate Models of AI Risk Are Hyperexistential Exfohazards
Listed -
Concrete Steps to Get Started in Transformer Mechanistic Interpretability
Listed -
I've updated towards AI boxing being surprisingly easy
Listed -
Oracle AGI - How can it escape, other than security issues? (Steganography?)
Listed -
Take 14: Corrigibility isn't that great.
Listed -
Is Eric Schmidt funding AI capabilities research by the US government?
Listed - Listed
-
Löb's Lemma: an easier approach to Löb's Theorem
Listed -
Practical AI risk I: Watching large compute
Listed - Listed
-
Katja Grace: Let's think about slowing down AI
Listed - Listed
-
Article Review: Discovering Latent Knowledge (Burns, Ye, et al)
Listed -
being only polynomial capabilities away from alignment: what a great problem to have that would be!
Listed -
December 2022 updates and fundraising
Listed -
Let’s think about slowing down AI
Listed -
Let’s think about slowing down AI
Listed -
one-shot AI, delegating embedded agency and decision theory, and one-shot QACI
Listed -
Racing through a minefield: the AI deployment problem
Listed -
Response to Holden’s alignment plan
Listed -
Some Notes on the mathematics of Toy Autoencoding Problems
Listed -
Take 13: RLHF bad, conditioning good.
Listed - Listed
-
A Comprehensive Mechanistic Interpretability Explainer & Glossary
Listed -
Applications Open: GovAI Summer Fellowship 2023
Listed -
Background for "Understanding the diffusion of large language models"
Listed - Listed
-
Conclusion and Bibliography for "Understanding the diffusion of large language models"
Listed -
Decisions: Ontologically Shifting to Determinism
Listed -
Drivers of large language model diffusion: incremental research, publicity, and cascades
Listed -
GPT-3-like models are now much easier to access and deploy than to develop
Listed -
Implications of large language model diffusion for AI governance
Listed -
New AI risk intro from Vox [link post]
Listed -
Publication decisions for large language models, and their impacts
Listed -
Questions for further investigation of AI diffusion
Listed -
The replication and emulation of GPT-3
Listed -
the scarcity of moral patient involvement
Listed -
Understanding the diffusion of large language models: summary
Listed -
An Open Agency Architecture for Safe Transformative AI
Listed -
Discovering Language Model Behaviors with Model-Written Evaluations
Listed -
High-level hopes for AI alignment
Listed -
Note on algorithms with multiple trained components
Listed - Listed
-
Take 12: RLHF's use is evidence that orgs will jam RL at real-world problems.
Listed -
The "Minimal Latents" Approach to Natural Abstractions
Listed -
AGI Timelines in Governance: Different Strategies for Different Timeframes
Listed -
Conditions for Superrationality-motivated Cooperation in a one-shot Prisoner's Dilemma
Listed -
Discovering Language Model Behaviors with Model-Written Evaluations
Listed -
Event [Berkeley]: Alignment Collaborator Speed-Meeting
Listed - Listed
-
Results from a survey on tool use and workflows in alignment research
Listed -
Shard Theory in Nine Theses: a Distillation and Critical Appraisal
Listed -
The ‘Old AI’: Lessons for AI governance from early electricity regulation
Listed - Listed
-
Why I think that teaching philosophy is high impact
Listed -
Why I think that teaching philosophy is high impact
Listed -
Will research in AI risk jinx it? Consequences of training AI on AI risk arguments
Listed -
Take 11: "Aligning language models" should be weirder.
Listed -
Looking for an alignment tutor
Listed -
Positive values seem more robust and lasting than prohibitions
Listed -
There have been 3 planes (billionaire donors) and 2 have crashed
Listed -
There have been 3 planes (billionaire donors) and 2 have crashed
Listed - Listed
-
AI overhangs depend on whether algorithms, compute and data are substitutes or complements
Listed -
Can we efficiently explain model behaviors?
Listed -
Concrete actionable policies relevant to AI safety (written 2019)
Listed -
How important are accurate AI timelines for the optimal spending schedule on AI risk interventions?
Listed -
How would you estimate the value of delaying AGI by 1 day, in marginal donations to GiveWell?
Listed -
Paper: Constitutional AI: Harmlessness from AI Feedback (Anthropic)
Listed -
Paper: Transformers learn in-context by gradient descent
Listed -
Point-E: A system for generating 3D point clouds from complex prompts
Listed -
Proper scoring rules don’t guarantee predicting fixed points
Listed -
We should say more than “x-risk is high”
Listed -
Who will be in charge once alignment is achieved?
Listed -
AI Neorealism: a threat model & success criterion for existential safety
Listed - Listed
-
High-level hopes for AI alignment
Listed -
High-level hopes for AI alignment
Listed - Listed
- Listed
-
How is ARC planning to use ELK?
Listed - Listed
-
The next decades might be wild
Listed -
all claw, no world — and other thoughts on the universal distribution
Listed -
Discovering Latent Knowledge in Language Models Without Supervision
Listed - Listed
-
Extracting and Evaluating Causal Direction in LLMs' Activations
Listed -
Is the AI timeline too short to have children?
Listed -
My AGI safety research—2022 review, ’23 plans
Listed - Listed
-
Seeking participants for study of AI safety researchers
Listed -
Trying to disambiguate different questions about whether RLHF is “good”
Listed -
«Boundaries», Part 3b: Alignment problems in terms of boundaries
Listed -
[Interim research report] Taking features out of superposition with sparse autoencoders
Listed -
AI alignment is distinct from its near-term applications
Listed -
Alignment with argument-networks and assessment-predictions
Listed -
An exploration of GPT-2's embedding weights
Listed -
Applications open for AGI Safety Fundamentals: Alignment Course
Listed -
Applications open for AGI Safety Fundamentals: Alignment Course
Listed -
Are lawsuits against AGI companies extending AGI timelines?
Listed -
Best introductory overviews of AGI safety?
Listed -
Existential AI Safety is NOT separate from near-term applications
Listed - Listed
-
Take 10: Fine-tuning with RLHF is aesthetically unsatisfying.
Listed - Listed
-
Concept extrapolation for hypothesis generation
Listed -
Join the AI Testing Hackathon this Friday
Listed -
Side-channels: input versus output
Listed -
Take 9: No, RLHF/IDA/debate doesn't solve outer alignment.
Listed -
a rough sketch of formal aligned AI using QACI
Listed -
AI Safety Seems Hard to Measure
Listed -
An appraisal of the Future of Life Institute AI existential risk program
Listed -
Benchmarks for Comparing Human and AI Intelligence
Listed -
Finite Factored Sets in Pictures
Listed -
Please provide feedback on AI-safety grant proposal, thanks!
Listed -
Reflections on the PIBBSS Fellowship 2022
Listed -
Reflections on the PIBBSS Fellowship 2022
Listed - Listed
-
[ASoT] Natural abstractions and AlphaZero
Listed - Listed
-
Cooperation, Avoidance, and Indifference: Alternate Futures for Misaligned AGI
Listed -
How promising are legal avenues to restrict AI training data?
Listed -
My thoughts on OpenAI's Alignment plan
Listed - Listed
-
Fear mitigated the nuclear threat, can it do the same to AGI risks?
Listed -
ML Safety at NeurIPS & Paradigmatic AI Safety? MLAISU W49
Listed -
Prosaic misalignment from the Solomonoff Predictor
Listed -
Take 8: Queer the inner/outer alignment dichotomy.
Listed -
Working towards AI alignment is better
Listed -
You can still fetch the coffee today if you're dead tomorrow
Listed -
AI Safety Seems Hard to Measure
Listed -
AI Safety Seems Hard to Measure
Listed -
I Believe we are in a Hardware Overhang
Listed -
If Wentworth is right about natural abstractions, it would be bad for alignment
Listed -
Main paths to impact in EU AI Policy
Listed -
Notes on OpenAI’s alignment plan
Listed - Listed
-
Take 7: You should talk about "the human's utility function" less.
Listed - Listed
-
Discovering Latent Knowledge in Language Models Without Supervision
Listed -
Promoting compassionate longtermism
Listed -
Simple Way to Prevent Power-Seeking AI
Listed -
Something to make myself fascinated with computing science and AI.
Listed -
Take 6: CAIS is actually Orwellian.
Listed -
Thoughts on AGI organizations and capabilities work
Listed -
Thoughts on AGI organizations and capabilities work
Listed -
AI for the board game Diplomacy
Listed -
AI Safety in a Vulnerable World: Requesting Feedback on Preliminary Thoughts
Listed -
In defense of probably wrong mechanistic models
Listed - Listed
-
Take 5: Another problem for natural abstractions is laziness.
Listed -
Using GPT-Eliezer against ChatGPT Jailbreaking
Listed -
Verification Is Not Easier Than Generation In General
Listed -
[Link] Why I’m optimistic about OpenAI’s alignment approach
Listed -
A Tentative Timeline of The Near Future (2022-2025) for Self-Accountability
Listed -
AI Safety Pitches post ChatGPT
Listed -
Aligned Behavior is not Evidence of Alignment Past a Certain Level of Intelligence
Listed -
Analysis of AI Safety surveys for field-building insights
Listed -
ChatGPT on Spielberg’s A.I. and AI Alignment
Listed -
Foresight for AGI Safety Strategy: Mitigating Risks and Identifying Golden Opportunities
Listed -
Have your timelines changed as a result of ChatGPT?
Listed -
Is the "Valley of Confused Abstractions" real?
Listed -
Probably good projects for the AI safety ecosystem
Listed -
Share your requests for ChatGPT
Listed -
Steering Behaviour: Testing for (Non-)Myopia in Language Models
Listed -
Take 4: One problem with natural abstractions is there's too many of them.
Listed - Listed
- Listed
-
AI can exploit safety plans posted on the Internet
Listed -
Race to the Top: Benchmarks for AI Safety
Listed -
Race to the Top: Benchmarks for AI Safety
Listed -
Take 3: No indescribable heavenworlds.
Listed -
Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research]
Listed - Listed
-
Causal scrubbing: results on a paren balance checker
Listed -
Causal scrubbing: results on induction heads
Listed -
Logical induction for software engineers
Listed -
Take 2: Building tools to help build FAI is a legitimate strategy, but it's dual-use.
Listed -
Will the first AGI agent have been designed as an agent (in addition to an AGI)?
Listed -
[ASoT] Finetuning, RL, and GPT's world prior
Listed -
Announcing the Cambridge Boston Alignment Initiative [Hiring!]
Listed -
Apply for the ML Winter Camp in Cambridge, UK [2-10 Jan]
Listed -
Deconfusing Direct vs Amortised Optimization
Listed -
Inner and outer alignment decompose one hard problem into two extremely hard problems
Listed -
Jailbreaking ChatGPT on Release Day
Listed -
Subsets and quotients in interpretability
Listed -
Takeoff speeds, the chimps analogy, and the Cultural Intelligence Hypothesis
Listed - Listed
-
A challenge for AGI organizations, and a challenge for readers
Listed -
Concrete actions to improve AI governance: the behaviour science approach
Listed -
Distillation of "How Likely is Deceptive Alignment?"
Listed -
Finding gliders in the game of life
Listed - Listed
-
Research request (alignment strategy): Deep dive on "making AI solve alignment for us"
Listed -
Take 1: We're not going to reverse-engineer the AI.
Listed - Listed
-
Theories of impact for Science of Deep Learning
Listed -
AI takeover tabletop RPG: "The Treacherous Turn"
Listed -
Biological Anchors external review by Jennifer Lin (linkpost)
Listed -
Compute Accounting Principles Can Help Reduce AI Risks
Listed -
Multi-Component Learning and S-Curves
Listed -
"The Physicists": A play about extinction and the responsibility of scientists
Listed -
Alignment allows "nonrobust" decision-influences and doesn't require robust grading
Listed -
Distinguishing test from training
Listed -
When to diversify? Breaking down mission-correlated investing
Listed - Listed
-
Why Would AI "Aim" To Defeat Humanity?
Listed -
Why Would AI "Aim" To Defeat Humanity?
Listed -
Why Would AI "Aim" To Defeat Humanity?
Listed -
Future Bowl Forecasting Tournament
Listed -
My take on Jacob Cannell’s take on AGI safety
Listed - Listed
-
The Singular Value Decompositions of Transformer Weight Matrices are Highly Interpretable
Listed -
Good Futures Initiative: Winter Project Internship
Listed -
More Academic Diversity in Alignment?
Listed -
Don't align agents to evaluations of plans
Listed -
Three Alignment Schemas & Their Problems
Listed -
Fair Collective Efficient Altruism
Listed -
Mechanistic anomaly detection and ELK
Listed - Listed
-
Planes are still decades away from displacing most bird jobs
Listed - Listed
-
Refining the Sharp Left Turn threat model
Listed -
Refining the Sharp Left Turn threat model, part 2: applying alignment techniques
Listed -
Rethink Priorities’ 2022 Impact, 2023 Strategy, and Funding Gaps
Listed -
Semi-conductor / AI stocks discussion.
Listed - Listed
-
Clarifying wireheading terminology
Listed -
Corrigibility Via Thought-Process Deference
Listed - Listed
-
Open technical problem: A Quinean proof of Löb's theorem, for an easier cartoon guide
Listed