The catalog, page 3
Records 501 to 750 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
We are not alone: many communities want to stop Big Tech from scaling unsafe AI
Listed -
AI is centralizing by default; let's not make it worse
Listed -
Is there much need for frontend engineers in AI alignment?
Listed -
Sparse Autoencoders Find Highly Interpretable Directions in Language Models
Listed -
Sparse Autoencoders Find Highly Interpretable Directions in Language Models
Listed -
Sparse Autoencoders: Future Work
Listed - Listed
-
There should be more AI safety orgs
Listed -
Careless talk on US-China AI competition? (and criticism of CAIS coverage)
Listed -
Existential Cybersecurity Risks & AI (A Research Agenda)
Listed -
Interpretability Externalities Case Study - Hungry Hungry Hippos
Listed -
The Case for AI Safety Advocacy to the Public
Listed -
[Link post] Michael Nielsen's "Notes on Existential Risk from Artificial Superintelligence"
Listed -
AISN #22: The Landscape of US AI Legislation - Hearings, Frameworks, Bills, and Laws
Listed -
Anthropic's Responsible Scaling Policy & Long-Term Benefit Trust
Listed -
Anthropic's Responsible Scaling Policy & Long-Term Benefit Trust
Listed -
Formalizing «Boundaries» with Markov blankets + Criticism of this approach
Listed -
Protest against Meta's irreversible proliferation (Sept 29, San Francisco)
Listed -
The possibility of an indefinite AI pause
Listed -
Comments on Manheim's "What's in a Pause?"
Listed -
Knowledge Database 1: The structure and the method of building
Listed -
Knowledge Database 2: Shopping advisor and other uses of knowledge base about products
Listed -
Relationship between EA Community and AI safety
Listed -
Stuart J. Russell on "should we press pause on AI?"
Listed -
Technical AI Safety Research Landscape [Slides]
Listed -
The omnizoid - Heighn FDT Debate #5
Listed -
US public opinion on AI, September 2023
Listed -
Where might I direct promising-to-me researchers to apply for alignment jobs/grants?
Listed - Listed
-
How to talk about reasons why AGI might not be near?
Listed - Listed
-
Microdooms averted by working on AI Safety
Listed -
Microdooms averted by working on AI Safety
Listed -
Reflexive decision theory is an unsolved problem
Listed -
Telopheme, telophore, and telotect
Listed - Listed
-
Policy ideas for mitigating AI risk
Listed - Listed
-
A conversation with Pi, a conversational AI.
Listed - Listed
-
Destroying the fabric of the universe as an instrumental goal.
Listed -
Instrumental Convergence Bounty
Listed -
The state of AI in different countries — an overview
Listed -
Uncovering Latent Human Wellbeing in LLM Embeddings
Listed -
AI-Risk in the State of the European Union Address
Listed -
Applications for EU Tech Policy Fellowship 2024 now open
Listed -
Apply to lead a project during the next virtual AI Safety Camp
Listed -
Is AI Safety dropping the ball on privacy?
Listed - Listed
-
UDT shows that decision theory is more puzzling than ever
Listed -
Who should we interview for The 80,000 Hours Podcast?
Listed -
Automatically finding feature vectors in the OV circuits of Transformers without using probing
Listed - Listed
-
Theory: “WAW might be of higher impact than x-risk prevention based on utilitarianism”
Listed -
Focus on the Hardest Part First
Listed -
How should technical AI researchers best transition into AI governance and policy?
Listed -
How teams went about their research at AI Safety Camp edition 8
Listed -
Panel discussion on AI consciousness with Rob Long and Jeff Sebo
Listed -
Possible Divergence in AGI Risk Tolerance between Selfish and Altruistic agents
Listed -
US presidents discuss AI alignment agendas
Listed -
A case study of regulation done well? Canadian biorisk regulations
Listed -
Debate series: should we push for a pause on the development of AI?
Listed -
Explained Simply: Quantilizers
Listed -
Explaining grokking through circuit efficiency
Listed - Listed
-
How long will reaching a Risk Awareness Moment and CHARTS agreement take?
Listed -
What I would do if I wasn’t at ARC Evals
Listed -
What term to use for AI in different policy contexts?
Listed -
What's in your list of important technical projects/experiments to run for AI alignment?
Listed - Listed
- Listed
-
Benchmarks for Detecting Measurement Tampering [Redwood Research]
Listed -
Decision theory is not policy theory is not agent theory
Listed -
Strongest real-world examples supporting AI risk claims?
Listed -
The Evolutionary Pathway from Biological to Digital Intelligence: A Cosmic Perspective
Listed -
What I would do if I wasn’t at ARC Evals
Listed -
[CFP] NeurIPS workshop: AI meets Moral Philosophy and Moral Psychology
Listed - Listed
-
Data Poisoning for Dummies (No Code, No Math)
Listed -
Getting Washington and Silicon Valley to tame AI (Mustafa Suleyman on the 80,000 Hours Podcast)
Listed -
Hertford, Sourbut (rationality lessons from University Challenge)
Listed -
No. Impending AGI doesn't make everything else unimportant.
Listed -
Notes on nukes, IR, and AI from "Arsenals of Folly" (and other books)
Listed -
Paper: On measuring situational awareness in LLMs
Listed -
Transformative AI and Compute - Reading List
Listed -
[Linkpost] Beware the Squirrel by Verity Harding
Listed -
Fundamental question: What determines a mind's effects?
Listed -
Series of absurd upgrades in nature's great search
Listed - Listed
- Listed
-
Rational Agents Cooperate in the Prisoner's Dilemma
Listed - Listed
-
Meta Questions about Metaphilosophy
Listed -
[Linkpost] Michael Nielsen remarks on 'Oppenheimer'
Listed -
Alignment & Capabilities: What's the difference?
Listed - Listed
- Listed
-
Agency Foundations Challenge: September 8th-24th, $10k Prizes
Listed -
An adversarial example for Direct Logit Attribution: memory management in gelu-4l
Listed -
Invulnerable Incomplete Preferences: A Formal Statement
Listed -
Report on Frontier Model Training
Listed -
Responses to apparent rationalist confusions about game / decision theory
Listed -
Updates from Campaign for AI Safety
Listed -
AI Deception: A Survey of Examples, Risks, and Potential Solutions
Listed -
AISN #20: LLM Proliferation, AI Deception, and Continuing Drivers of AI Capabilities
Listed -
An Interpretability Illusion for Activation Patching of Arbitrary Subspaces
Listed -
An OV-Coherent Toy Model of Attention Head Superposition
Listed -
Anyone want to debate publicly about FDT?
Listed -
Apply to a small iteration of MLAB to be run in Oxford
Listed -
Barriers to Mechanistic Interpretability for AGI Safety
Listed - Listed
-
Impact Academy is hiring an AI Governance Lead - more information, upcoming Q&A and $500 bounty
Listed -
Incentives affecting alignment-researcher encouragement
Listed - Listed
- Listed
-
OpenAI API base models are not sycophantic, at any size
Listed -
Paper Walkthrough: Automated Circuit Discovery with Arthur Conmy
Listed -
AI Deception: A Survey of Examples, Risks, and Potential Solutions
Listed -
Information warfare historically revolved around human conduits
Listed -
Introducing the Center for AI Policy (& we're hiring!)
Listed -
Navigating the Future: A Guide on How to Stay Safe with AI | Emmanuel Katto Uganda
Listed -
Paradigms and Theory Choice in AI: Adaptivity, Economy and Control
Listed - Listed
- Listed
-
Apply to a small iteration of MLAB to be run in Oxford
Listed - Listed
-
A list of core AI safety problems and how I hope to solve them
Listed -
EA is underestimating intelligence agencies and this is dangerous
Listed -
Mesa-Optimization: Explain it like I'm 10 Edition
Listed -
Ramble on STUFF: intelligence, simulation, AI, doom, default mode, the usual
Listed -
Red-teaming language models via activation engineering
Listed -
A Model-based Approach to AI Existential Risk
Listed -
A model-based approach to AI Existential Risk
Listed - Listed
-
What AI Posts Do You Want Distilled?
Listed -
[Crosspost] AI Regulation May Be More Important Than AI Alignment For Existential Safety
Listed -
AI Regulation May Be More Important Than AI Alignment For Existential Safety
Listed - Listed
-
Assessment of intelligence agency functionality is difficult yet important
Listed -
Enhancing Corrigibility in AI Systems through Robust Feedback Loops
Listed -
Health, morality, and goal alignment of systems, agents, and organs
Listed -
Would it be useful to collect the contexts, where various LLMs think the same?
Listed - Listed
-
Implications of evidential cooperation in large worlds
Listed -
The Ethical Basilisk Thought Experiment
Listed -
Why Is No One Trying To Align Profit Incentives With Alignment Research?
Listed -
An argument for accelerating international AI governance research (part 2)
Listed -
Why does an AI have to have specified goals?
Listed -
Causality and a Cost Semantics for Neural Networks
Listed -
Ideas for improving epistemics in AI safety outreach
Listed - Listed
-
Large Language Models will be Great for Censorship
Listed - Listed
-
Call for Papers on Global AI Governance from the UN
Listed -
Jan Kulveit's Corrigibility Thoughts Distilled
Listed -
Longtermism Fund: August 2023 Grants Report
Listed -
Memetic Judo #3: The Intelligence of Stochastic Parrots v.2
Listed -
XPT forecasts on (some) Direct Approach model inputs
Listed -
“Dirty concepts” in AI alignment discourses, and some guesses for how to deal with them
Listed - Listed
-
Clarifying how misalignment can arise from scaling LLMs
Listed -
Supervised Program for Alignment Research (SPAR) at UC Berkeley: Spring 2023 summary
Listed -
Supervised Program for Alignment Research (SPAR) at UC Berkeley: Spring 2023 summary
Listed -
We can do better than DoWhatIMean
Listed -
Will AI kill everyone? Here's what the godfathers of AI have to say [RA video]
Listed -
6 non-obvious mental health issues specific to AI safety
Listed - Listed
-
An Overview of Catastrophic AI Risks: Summary
Listed -
Managing risks of our own work
Listed -
AIのタイムライン ─ 提案されている論証と「専門家」の立ち位置
Listed -
Autonomous replication and adaptation: an attempt at a concrete danger threshold
Listed -
Corporate campaigns work: a key learning for AI Safety
Listed - Listed
-
Looking for judges for critiques of Alignment Plans
Listed -
Making EA more inclusive, representative, and impactful in Africa
Listed -
The State of AI Governance in Africa: Musings from the Global South
Listed -
A Proof of Löb's Theorem using Computability Theory
Listed -
An argument for accelerating international AI governance research (part 1)
Listed -
One example of how LLM propaganda attacks can hack the brain
Listed -
Stampy's AI Safety Info - New Distillations #4 [July 2023]
Listed -
Understanding and visualizing sycophancy datasets
Listed -
A bill to prevent AI from hiring people instead of human enployers in NY
Listed -
Am I taking crazy pills? Why aren't EAs advocating for a pause on AI capabilities?
Listed -
An Overview of Catastrophic AI Risks
Listed -
Bio-x-AI policy: call for ideas from the Federation of American Scientists
Listed -
Credo AI is hiring for AI Gov Researcher & more!
Listed -
Principles of Cyber-Physical Systems, Chapters 1-7,9
Listed -
Why some people disagree with the CAIS statement on AI
Listed -
$1,000 bounty for an AI Programme Lead recommendation
Listed -
A short calculation about a Twitter poll
Listed -
Decomposing independent generalizations in neural networks via Hessian analysis
Listed -
Import AI 336: Financialized AI; public and elite AI opinion; one million insects.
Listed - Listed
- Listed
-
Summary of “The Precipice” (2 of 4): We are a danger to ourselves
Listed -
We Should Prepare for a Larger Representation of Academia in AI Safety
Listed -
What do we know about Mustafa Suleyman's position on AI Safety?
Listed -
Biological Anchors: The Trick that Might or Might Not Work
Listed -
AI Safety Concepts Writeup: WebGPT
Listed -
When discussing AI risks, talk about capabilities, not intelligence
Listed -
A selection of some writings and considerations on the cause of artificial sentience
Listed -
Could We Automate AI Alignment Research?
Listed -
Ilya Sutskever's thoughts on AI safety (July 2023): a transcript with my comments
Listed -
Seeking Input to AI Safety Book for non-technical audience
Listed -
The positional embedding matrix and previous-token heads: how do they actually work?
Listed -
UN Public Call for Nominations For High-level Advisory Body on Artificial Intelligence
Listed -
Update on cause area focus working group
Listed - Listed
-
4 types of AGI selection, and how to constrain them
Listed -
Acausal Now: We could totally acausally bargain with aliens at our current tech level if desired
Listed -
Modulating sycophancy in an RLHF model via activation steering
Listed -
Modulating sycophancy in an RLHF model via activation steering
Listed -
When discussing AI risks, talk about capabilities, not intelligence
Listed - Listed
-
Beginner's question about RLHF
Listed - Listed
-
Fundamentals of Global Priorities Research in Economics Syllabus
Listed -
Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research
Listed - Listed
-
Podcast (+transcript): Nathan Barnard on how US financial regulation can inform AI governance
Listed -
An interactive introduction to grokking and mechanistic interpretability
Listed -
Optimisation Measures: Desiderata, Impossibility, Proposals
Listed -
Studying Large Language Model Generalization with Influence Functions
Listed -
Updates from Campaign for AI Safety
Listed -
Rebooting AI Governance: An AI-Driven Approach to AI Governance
Listed -
Safety-First Agents/Architectures Are a Promising Path to Safe AGI
Listed -
Yann LeCun on AGI and AI Safety
Listed -
An appeal to people who are smarter than me: please help me clarify my thinking about AI
Listed -
Ground-Truth Label Imbalance Impairs Contrast-Consistent Search Performance
Listed - Listed
-
Join AISafety.info's Writing & Editing Hackathon (Aug 25-28) (Prizes to be won!)
Listed -
[Linkpost] Multimodal Neurons in Pretrained Text-Only Transformers
Listed -
Apollo Research is hiring evals and interpretability engineers & scientists
Listed -
Apollo Research is hiring evals and interpretability engineers & scientists
Listed -
AI #23: Fundamental Problems with RLHF
Listed -
Password-locked models: a stress case for capabilities evaluation
Listed -
Training for Good is hiring (and why you should join us): AI Programme Lead and Operations Associate
Listed -
3 levels of threat obfuscation
Listed -
3 levels of threat obfuscation
Listed -
Alignment Grantmaking is Funding-Limited Right Now [crosspost]
Listed -
Four part playbook for dealing with AI (Holden Karnofsky on the 80,000 Hours Podcast)
Listed -
How many people are neartermist and have high P(doom)?
Listed -
AI romantic partners will harm society if they go unregulated
Listed - Listed
-
ARC Evals new report: Evaluating Language-Model Agents on Realistic Autonomous Tasks
Listed -
Artificially sentient beings: Moral, political, and legal issues
Listed -
Confidence-Building Measures for Artificial Intelligence: Workshop proceedings
Listed -
Evaluating Superhuman Models with Consistency Checks
Listed -
Riesgos Catastróficos Globales needs funding
Listed -
What is autonomy, and how does it lead to greater risk from AI?
Listed