The catalog, page 10
Records 2,251 to 2,500 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
You are probably not a good alignment researcher, and other blatant lies
Listed -
“AI Risk Discussions” website: Exploring interviews from 97 AI Researchers
Listed -
AI Safety Arguments: An Interactive Guide
Listed -
Eli Lifland on Navigating the AI Alignment Landscape
Listed -
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small
Listed -
Language Models can be Utility-Maximising Agents
Listed -
More findings on Memorization and double descent
Listed -
Product safety is a poor model for AI governance
Listed -
Product safety is a poor model for AI governance
Listed -
The effect of horizon length on scaling laws
Listed -
Trends in the dollar training cost of machine learning systems
Listed -
[Linkpost] Human-narrated audio version of "Is Power-Seeking AI an Existential Risk?"
Listed -
Apply to HAIST/MAIA’s AI Governance Workshop in DC (Feb 17-20)
Listed -
Criticism of the main framework in AI alignment
Listed -
How to hedge investment portfolio against AI risk?
Listed -
How to use AI speech transcription and analysis to accelerate social science research
Listed -
Inner Misalignment in "Simulator" LLMs
Listed -
Mechanistic Interpretability Quickstart Guide
Listed -
On value in humans, other animals, and AI
Listed -
Questions about AI that bother me
Listed -
What Are The Biggest Threats To Humanity? (A Happier World video)
Listed -
Against Boltzmann mesaoptimizers
Listed -
Call for submissions: “(In)human Values and Artificial Agency”, ALIFE 2023
Listed - Listed
-
Model-driven feedback could amplify alignment failures
Listed -
Time-stamping: An urgent, neglected AI safety measure
Listed -
What I mean by "alignment is in large part about making cognition aimable at all"
Listed -
Why I hate the "accident vs. misuse" AI x-risk dichotomy (quick thoughts on "structural risk")
Listed -
a guess at my intrinsic values
Listed -
communicating with successful alignment timelines
Listed -
Compendium of problems with RLHF
Listed -
formal alignment: what it is, and some proposals
Listed -
formal alignment: what it is, and some proposals
Listed -
Structure, creativity, and novelty
Listed - Listed
-
Optimality is the tiger, and annoying the user is its teeth
Listed -
Spooky action at a distance in the loss landscape
Listed -
Stop-gradients lead to fixed point predictions
Listed -
Assigning Praise and Blame: Decoupling Epistemology and Decision Theory
Listed -
Literature review of TAI timelines
Listed -
The role of Bayesian ML in AI safety - an overview
Listed -
to me, it's instrumentality that is alienating
Listed -
WaPo: "Big Tech was moving cautiously on AI. Then came ChatGPT."
Listed -
"How to Escape from the Simulation" - Seeds of Science call for reviewers
Listed -
AI Risk Management Framework | NIST
Listed -
All AGI Safety questions welcome (especially basic ones) [~monthly thread]
Listed -
Excerpts from "Doing EA Better" on x-risk methodology
Listed - Listed
-
AGI will have learnt utility functions
Listed -
Quick thoughts on "scalable oversight" / "super-human feedback" research
Listed -
Spreading messages to help with the most important century
Listed -
Spreading messages to help with the most important century
Listed -
Thoughts on the impact of RLHF research
Listed -
Alexander and Yudkowsky on AGI goals
Listed - Listed
-
Gradient hacking is extremely difficult
Listed -
How-to Transformer Mechanistic Interpretability—in 50 lines of code or less!
Listed -
Inverse Scaling Prize: Second Round Winners
Listed -
Large Language Models as Fiduciaries to Humans
Listed -
Parameter Scaling Comes for RL, Maybe
Listed -
Some of my disagreements with List of Lethalities
Listed -
Thoughts on hardware / compute requirements for AGI
Listed -
Update to Samotsvety AGI timelines
Listed -
Why people want to work on AI safety (but don’t)
Listed - Listed
- Listed
- Listed
-
My highly personal skepticism braindump on existential risk from artificial intelligence.
Listed -
There should be a public adversarial collaboration on AI x-risk
Listed -
What a compute-centric framework says about AI takeoff speeds
Listed -
What a compute-centric framework says about AI takeoff speeds
Listed -
Emotional attachment to AIs opens doors to problems
Listed - Listed
-
Large language models learn to represent the world
Listed -
NYT: Google will ‘recalibrate’ the risk of releasing AI due to competition with OpenAI
Listed -
AI Safety "Textbook". Test chapter. Orthogonality Thesis, Goodhart Law and Instrumental Convergency
Listed - Listed
- Listed
-
[TIME magazine] DeepMind’s CEO Helped Take AI Mainstream. Now He’s Urging Caution (Perrigo, 2023)
Listed -
Critique of some recent philosophy of LLMs’ minds
Listed -
Shard theory alignment has important, often-overlooked free parameters.
Listed -
Transcript of Sam Altman's interview touching on AI safety
Listed -
What’s going on with ‘crunch time’?
Listed -
"Heretical Thoughts on AI" by Eli Dourado
Listed -
200 COP in MI: Studying Learned Features in Language Models
Listed -
6-paragraph AI risk intro for MAISI
Listed -
AGI safety field building projects I’d like to see
Listed - Listed
- Listed
-
List of technical AI safety exercises and projects
Listed -
nostalgia: a value pointing home
Listed -
Thoughts on refusing harmful requests to large language models
Listed -
Any Philosophy PhD recommendations for students interested in Alignment Efforts?
Listed -
Approfondimenti sui rischi dell’IA (materiali in inglese)
Listed -
Emerging Paradigms: The Case of Artificial Intelligence Safety
Listed - Listed
-
Help me to understand AI alignment!
Listed -
Neural networks generalize because of this one weird trick
Listed -
Vitalik on science, his philanthropy and effective altruism.
Listed -
AGISF adaptation for in-person groups
Listed -
Collin Burns on Alignment Research And Discovering Latent Knowledge Without Supervision
Listed -
How many people are working (directly) on reducing existential risk from AI?
Listed -
Il panorama della governance lungoterminista delle intelligenze artificiali
Listed -
Le Tempistiche delle IA: il dibattito e il punto di vista degli “esperti”
Listed -
Lessons learned and review of the AI Safety Nudge Competition
Listed -
Löbian emotional processing of emergent cooperation: an example
Listed -
L’importanza delle IA come possibile minaccia per l’umanità
Listed -
Perché il deep learning moderno potrebbe rendere difficile l’allineamento delle IA
Listed -
Preparing for AI-assisted alignment research: we need data!
Listed -
Prevenire una catastrofe legata all'intelligenza artificiale
Listed -
Ricerca sulla sicurezza delle IA: panoramica delle carriere
Listed -
Should AI writers be prohibited in education?
Listed -
What can thought-experiments do?
Listed -
Aligning the Aligners: Ensuring Aligned AI acts for the common good of all mankind
Listed -
Can GPT-3 produce new ideas? Partially automating Robin Hanson and others
Listed -
Consequentialists: One-Way Pattern Traps
Listed - Listed
-
Experiment Idea: RL Agents Evading Learned Shutdownability
Listed -
How we could stumble into AI catastrophe
Listed -
Import AI - coming soon to Substack
Listed -
Reflections on Trusting Trust & AI
Listed -
Should AI writers be prohibited in education?
Listed -
Showing versus doing: Teaching by demonstration
Listed -
Deceptive failures short of full catastrophe.
Listed -
Non-directed conceptual founding
Listed -
Speculation on Path-Dependance in Large Language Models.
Listed -
Underspecification of Oracle AI
Listed -
Concrete Reasons for Hope about AI
Listed -
World-Model Interpretability Is All We Need
Listed -
[ASoT] Simulators show us behavioural properties by default
Listed -
AGISF adaptation for in-person groups
Listed - Listed
-
Concerns about AI safety career change
Listed -
Disentangling Shard Theory into Atomic Claims
Listed -
How does GPT-3 spend its 175B parameters?
Listed -
How we could stumble into AI catastrophe
Listed -
Some Arguments Against Strong Scaling
Listed -
The AI Control Problem in a wider intellectual context
Listed -
Tracr: Compiled Transformers as a Laboratory for Interpretability | DeepMind
Listed -
Tracr: Compiled Transformers as a Laboratory for Interpretability | DeepMind
Listed -
[Linkpost] Scaling Laws for Generative Mixed-Modal Language Models
Listed - Listed
-
Announcing the 2023 PIBBSS Summer Research Fellowship
Listed -
Announcing the 2023 PIBBSS Summer Research Fellowship
Listed -
Categorical-measure-theoretic approach to optimal policies tending to seek power
Listed -
ChatGPT struggles to respond to the real world
Listed -
ea.domains - Domains Free to a Good Home
Listed -
How it feels to have your mind hacked by an AI
Listed -
Microsoft Plans to Invest $10B in OpenAI; $3B Invested to Date | Fortune
Listed -
ML Summer Bootcamp Reflection: Aalto EA Finland
Listed -
Reward is not Necessary: How to Create a Compositional Self-Preserving Agent for Life-Long Learning
Listed - Listed
-
Victoria Krakovna on AGI Ruin, The Sharp Left Turn and Paradigms of AI Alignment
Listed -
Forecasting potential misuses of language models for disinformation campaigns and how to reduce risk
Listed -
200 COP in MI: Interpreting Reinforcement Learning
Listed -
Against using stock prices to forecast AI timelines
Listed -
Against using stock prices to forecast AI timelines
Listed -
AGI and the EMH: markets are not expecting aligned or unaligned AI in the next 30 years
Listed -
Review AI Alignment posts to help figure out how to make a proper AI Alignment review
Listed -
The Alignment Problem from a Deep Learning Perspective (major rewrite)
Listed - Listed
-
What AI Take-Over Movies or Books Will Scare Me Into Taking AI Seriously?
Listed -
[MLSN #7]: an example of an emergent internal optimizer
Listed - Listed
-
Is anyone else also getting more worried about hard takeoff AGI scenarios?
Listed - Listed
-
Nearcast-based “deployment problem” analysis (Karnofsky, 2022)
Listed -
Trying to isolate objectives: approaches toward high-level interpretability
Listed -
Wentworth and Larsen on buying time
Listed -
You're Not One "You" - How Decision Theories Are Talking Past Each Other
Listed -
200 COP in MI: Image Model Interpretability
Listed -
Is this community over-emphasizing AI alignment?
Listed -
Learning as much Deep Learning math as I could in 24 hours
Listed -
Research ideas (AI Interpretability & Neurosciences) for a 2-months project
Listed - Listed
-
David Krueger on AI Alignment in Academia and Coordination
Listed -
How to create curriculum for self-study towards AI alignment work?
Listed -
Looking for Spanish AI Alignment Researchers
Listed -
Protectionism will Slow the Deployment of AI
Listed -
200 COP in MI: Techniques, Tooling and Automation
Listed - Listed
-
AI Safety Camp, Virtual Edition 2023
Listed -
AI Safety Camp: Machine Learning for Scientific Discovery
Listed -
AI security might be helpful for AI alignment
Listed -
Categorizing failures as “outer” or “inner” misalignment is often confused
Listed -
Definitions of “objective” should be Probable and Predictive
Listed -
Machine Learning for Scientific Discovery - AI Safety Camp
Listed -
Metaculus Year in Review: 2022
Listed -
Transformative AI issues (not just misalignment): an overview
Listed -
Illusion of truth effect and Ambiguity effect: Bias in Evaluating AGI X-Risks
Listed -
Paper: Superposition, Memorization, and Double Descent (Anthropic)
Listed -
Skill up in ML for AI safety with the Intro to ML Safety course (Spring 2023)
Listed -
Superposition, Memorization, and Double Descent
Listed -
Transformative AI issues (not just misalignment): an overview
Listed - Listed
- Listed
-
200 COP in MI: Analysing Training Dynamics
Listed -
2022 was the year AGI arrived (Just don't call it that)
Listed -
Announcing Insights for Impact
Listed -
Basic Facts about Language Model Internals
Listed -
Causal representation learning as a technique to prevent goal misgeneralization
Listed -
ChatGPT understands, but largely does not generate Spanglish (and other code-mixed) text
Listed - Listed
-
Large Language Models as Corporate Lobbyists, and Implications for Societal-AI Alignment
Listed -
List of links for getting into AI safety
Listed -
Normalcy bias and Base rate neglect: Bias in Evaluating AGI X-Risks
Listed -
200 COP in MI: Exploring Polysemanticity and Superposition
Listed -
Holden Karnofsky Interview about Most Important Century & Transformative AI
Listed -
How have shorter AI timelines been affecting you, and how have you been responding to them?
Listed -
I have thousands of copies of HPMOR in Russian. How to use them with the most impact?
Listed -
Is recursive self-alignment possible?
Listed -
Touch reality as soon as possible (when doing machine learning research)
Listed - Listed
-
[Simulators seminar sequence] #1 Background & shared assumptions
Listed -
AI Safety Doesn't Have to be Weird
Listed -
Alignment, Anger, and Love: Preparing for the Emergence of Superintelligent AI
Listed -
Large language models can provide "normative assumptions" for learning human preferences
Listed - Listed
-
On the Importance of Open Sourcing Reward Models
Listed -
Results from the AI testing hackathon
Listed -
Soft optimization makes the value target bigger
Listed -
A Löbian argument pattern for implicit reasoning in natural language: Löbian party invitations
Listed -
Summary of 80k's AI problem profile
Listed - Listed
- Listed
-
Visualizing what ConvNets learn
Listed -
Would it be good or bad for the US military to get involved in AI risk?
Listed -
'simulator' framing and confusions about LLMs
Listed -
200 COP in MI: Interpreting Algorithmic Problems
Listed -
Are Mixture-of-Experts Transformers More Interpretable Than Dense Transformers?
Listed - Listed
-
Racing through a minefield: the AI deployment problem
Listed -
Self-Limiting AI in AI Alignment
Listed -
Should AI systems have to identify themselves?
Listed -
Beyond Rewards and Values: A Non-dualistic Approach to Universal Intelligence
Listed -
But is it really in Rome? An investigation of the ROME model editing technique
Listed -
Future Matters #6: FTX collapse, value lock-in, and counterarguments to AI x-risk
Listed - Listed
-
My thoughts on OpenAI's alignment plan
Listed -
200 COP in MI: Looking for Circuits in the Wild
Listed -
CFP for Rebellion and Disobedience in AI workshop
Listed -
Internal Interfaces Are a High-Priority Interpretability Target
Listed -
The commercial incentive to intentionally train AI to deceive us
Listed -
200 Concrete Open Problems in Mechanistic Interpretability: Introduction
Listed -
200 COP in MI: The Case for Analysing Toy Language Models
Listed -
Book recommendations for the history of ML?
Listed -
Getting up to Speed on the Speed Prior in 2022
Listed - Listed
-
making decisions as our approximately simulated selves
Listed -
Reflections on my 5-month AI alignment upskilling grant
Listed