The catalog, page 2
Records 251 to 500 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
AI Alignment: A Comprehensive Survey
Listed - Listed
-
Forecasting Questions: What do you want to predict on AI?
Listed -
My thoughts on the social response to AI risk
Listed - Listed
-
Reactions to the Executive Order
Listed -
Robustness of Contrast-Consistent Search to Adversarial Prompting
Listed -
Singular learning theory and bridging from ML to brain emulations
Listed -
Snapshot of narratives and frames against regulating AI
Listed -
The Bletchley Declaration on AI Safety
Listed -
Agent Foundations track in MATS
Listed -
AI Safety 101 - Chapter 5.1 - Debate
Listed - Listed
- Listed
-
Preventing Language Models from hiding their reasoning
Listed -
The UK AI Safety Summit tomorrow
Listed -
Thoughts on the AI Safety Summit company policy requests and responses
Listed -
Urging an International AI Treaty: An Open Letter
Listed -
5 Reasons Why Governments/Militaries Already Want AI for Information Warfare
Listed -
[Linkpost] Two major announcements in AI governance today
Listed -
Charbel-Raphaël and Lucius discuss Interpretability
Listed - Listed
-
President Biden Issues Executive Order on Safe, Secure, and Trustworthy Artificial Intelligence
Listed - Listed
-
Will releasing the weights of large language models grant widespread access to pandemic agents?
Listed -
Would it make sense to bring a civil lawsuit against Meta for recklessly open sourcing models?
Listed -
Clarifying the free energy principle (with quotes)
Listed -
The AI Boom Mainly Benefits Big Firms, but long-term, markets will concentrate
Listed -
Regrant up to $600,000 to AI safety projects with GiveWiki
Listed -
Summary: Existential risk from power-seeking AI by Joseph Carlsmith
Listed -
AI safety field-building survey: Talent needs, infrastructure needs, and relationship to EA
Listed -
Efficacy of AI Activism: Have We Ever Said No?
Listed -
Linkpost: Rishi Sunak's Speech on AI (26th October)
Listed -
New report on the state of AI safety in China
Listed -
Value systematization: how values become coherent (and misaligned)
Listed -
We're Not Ready: thoughts on "pausing" and responsible scaling policies
Listed -
Wireheading and misalignment by composition on NetHack
Listed -
1. Premise one: Values are malleable
Listed -
2. Premise two: Some cases of value change are (il)legitimate
Listed - Listed
-
4. Risks from causing illegitimate value change (performative predictors)
Listed -
5. Risks from preventing legitimate value change (value collapse)
Listed -
AI #35: Responsible Scaling Policies
Listed -
Apply to the Constellation Visiting Researcher Program and Astra Fellowship, in Berkeley this Winter
Listed -
Apply to the Constellation Visiting Researcher Program and Astra Fellowship, in Berkeley this Winter
Listed -
CHAI internship applications are open (due Nov 13)
Listed -
Codebook Features: Sparse and Discrete Interpretability for Neural Networks
Listed -
Disagreements over the prioritization of existential risk from AI
Listed -
OpenAI’s new Preparedness team is hiring
Listed -
UK Government publishes "Frontier AI: capabilities and risks" Discussion Paper
Listed -
UK Prime Minister Rishi Sunak's Speech on AI
Listed -
What we learned from running an Australian AI Safety Unconference
Listed -
AI as a science, and three obstacles to alignment strategies
Listed -
Announcing Epoch's newly expanded Parameters, Compute and Data Trends in Machine Learning database
Listed - Listed
-
Compositional preference models for aligning LMs
Listed -
Compositional preference models for aligning LMs
Listed -
Responsible Scaling Policies Are Risk Management Done Wrong
Listed -
Successif: Join our AI program to help mitigate the catastrophic risks of AI
Listed - Listed
-
[Interview w/ Quintin Pope] Evolution, values, and AI Safety
Listed -
Announcing #AISummitTalks featuring Professor Stuart Russell and many others
Listed -
Announcing #AISummitTalks featuring Professor Stuart Russell and many others
Listed -
Go Mobilize? Lessons from GM Protests for Pausing AI
Listed - Listed
-
Largest AI model in 2 years from $10B
Listed -
Lying is Cowardice, not Strategy
Listed -
The Self in Artificial Consciousness: A Buddhist Investigation into Advanced AI
Listed -
Thoughts on responsible scaling policies and regulation
Listed -
Thoughts on responsible scaling policies and regulation
Listed -
Towards Understanding Sycophancy in Language Models
Listed -
Towards Understanding Sycophancy in Language Models
Listed -
Who is Harry Potter? Some predictions.
Listed -
Fundamental Challenges in AI Governance
Listed -
Help us design the interface for aisafety.com
Listed -
Machine Unlearning Evaluations as Interpretability Benchmarks
Listed -
Machine Unlearning Evaluations as Interpretability Benchmarks
Listed -
Open Source Replication & Commentary on Anthropic's Dictionary Learning Paper
Listed -
Pausing AI might be good policy, but it's bad politics
Listed -
Programmatic backdoors: DNNs can use SGD to run arbitrary stateful computation
Listed -
The Shutdown Problem: Three Theorems
Listed -
VLM-RM: Specifying Rewards with Natural Language
Listed -
VLM-RM: Specifying Rewards with Natural Language
Listed - Listed
- Listed
- Listed
-
AI Safety is Dropping the Ball on Clown Attacks, and Mind Control in General
Listed -
Alignment Implications of LLM Successes: a Debate in One Act
Listed -
Apply for MATS Winter 2023-24!
Listed -
Apply for MATS Winter 2023-24!
Listed -
How toy models of ontology changes can be misleading
Listed -
Thoughts On (Solving) Deep Deception
Listed -
Announcing new round of "Key Phenomena in AI Risk" Reading Group
Listed -
I Would Have Solved Alignment, But I Was Worried That Would Advance Timelines
Listed -
Internal Target Information for AI Oversight
Listed -
Revealing Intentionality In Language Models Through AdaVAE Guided Sampling
Listed -
Specific versus General Principles for Constitutional AI
Listed -
TOMORROW: the largest AI Safety protest ever!
Listed -
Towards Understanding Sycophancy in Language Models
Listed - Listed
-
New roles on my team: come build Open Phil's technical AI safety program with me!
Listed -
Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Listed -
(Non-deceptive) Suboptimality Alignment
Listed - Listed
- Listed
-
Alignment 101 - Ch.2 - Reward Misspecification
Listed -
Metaculus Launches Conditional Cup to Explore Linked Forecasts
Listed -
On Interpretability's Robustness
Listed -
The (partial) fallacy of dumb superintelligence
Listed -
Beginner’s guide to reducing s-risks [link-post]
Listed -
Investigating the learning coefficient of modular addition: hackathon project
Listed -
The theoretical computational limit of the Solar System is 1.47x10^49 bits per second.
Listed -
Goodhart's Law in Reinforcement Learning
Listed -
Knowledge Base 4: General applications
Listed - Listed
-
UNGA General Debate speeches on AI
Listed - Listed
-
Mapping ChatGPT’s ontological landscape, gradients and choices [interpretability]
Listed -
Politico article on Open Phil, Horizon Fellowship, and EA
Listed -
Assessing the Dangerousness of Malevolent Actors in AGI Governance: A Preliminary Exploration
Listed -
ChatGPT tells 20 versions of its prototypical story, with a short note on method
Listed -
Natural Abstraction: Convergent Preferences Over Information Structures
Listed - Listed
- Listed
-
[Paper] All's Fair In Love And Love: Copy Suppression in GPT-2 Small
Listed -
At Our World in Data we're hiring our first Communications & Outreach Manager
Listed -
FLI podcast series, "Imagine A World", about aspirational futures with AGI
Listed -
How can I best use my career to pass impactful AI and Biosecurity policy.
Listed -
Paper: Understanding and Controlling a Maze-Solving Policy Network
Listed -
To open-source or to not open-source, that is (an oversimplification of) the question.
Listed -
What he’s learned as an AI policy insider (Tantum Collins on the 80,000 Hours Podcast)
Listed - Listed
- Listed
-
LoRA Fine-tuning Efficiently Undoes Safety Training from Llama 2-Chat 70B
Listed -
Opportunities for Impact Beyond the EU AI Act
Listed -
Relevance of 'Harmful Intelligence' Data in Training Datasets (WebText vs. Pile)
Listed -
Resources & opportunities for careers in European AI Policy
Listed -
The International PauseAI Protest: Activism under uncertainty
Listed - Listed
-
unRLHF - Efficiently undoing LLM safeguards
Listed -
Attributing to interactions with GCPD and GWPD
Listed -
Attributing to interactions with GCPD and GWPD
Listed - Listed
-
Update on the UK AI Taskforce & AI Safety Summit
Listed -
You’re Measuring Model Complexity Wrong
Listed -
A New Model for Compute Center Verification
Listed -
AI+bio cannot be half of AI catastrophe risk, right?
Listed -
Become a PIBBSS Research Affiliate
Listed -
Documenting Journey Into AI Safety
Listed -
Non-superintelligent paperclip maximizers are normal
Listed -
Pause For Thought: The AI Pause Debate
Listed - Listed
-
The Bostrom Buckle: Visualising the Vulnerable World Hypothesis
Listed -
We don't understand what happened with culture enough
Listed -
We don't understand what happened with culture enough
Listed -
Perspective Based Reasoning Could Absolve CDT
Listed -
Silicon Valley’s Rabbit Hole Problem
Listed -
Time is homogeneous sequentially-composable determination
Listed -
Comparing Anthropic's Dictionary Learning to Ours
Listed -
Don't Dismiss Simple Alignment Approaches
Listed -
Fixing Insider Threats in the AI Supply Chain
Listed -
Risk-averse Batch Active Inverse Reward Design
Listed -
Utilitarianism is irrational or self-undermining
Listed -
A personal explanation of ELK concept and task.
Listed -
What AI could mean for animals
Listed -
Best project management software for research projects and labs?
Listed -
Evaluating the historical value misspecification argument
Listed -
Ideation and Trajectory Modelling in Language Models
Listed -
Pause For Thought: The AI Pause Debate (Astral Codex Ten)
Listed -
Stampy's AI Safety Info soft launch
Listed -
Stampy's AI Safety Info soft launch
Listed -
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
Listed -
AISN #23: New OpenAI Models, News from Anthropic, and Representation Engineering
Listed -
Apply to Spring 2024 policy internships (we can help)
Listed -
Entanglement and intuition about words and meaning
Listed -
Ethical Considerations in regard to Outsourcing Labour Needs to the Global South
Listed -
Fiscal sponsorship, ops support, or incubation?
Listed -
Graphical tensor notation for interpretability
Listed -
How Rethink Priorities’ Research could inform your grantmaking
Listed -
How to solve deception and still fail.
Listed -
I don’t find the lie detection results that surprising (by an author of the paper)
Listed -
What are some examples of AIs instantiating the 'nearest unblocked strategy problem'?
Listed -
Why isn't there a Charity Entrepreneurship program for AI Safety?
Listed -
AXRP Episode 25 - Cooperative AI with Caspar Oesterheld
Listed -
De Dicto and De Se Reference Matters for Alignment
Listed -
Early Experiments in Reward Model Interpretation Using Sparse Autoencoders
Listed - Listed
-
What would it mean to understand how a large language model (LLM) works? Some quick notes.
Listed -
Why We Use Money? - A Walrasian View
Listed -
Announcing FAR Labs, an AI safety coworking space
Listed -
Automated Parliaments — A Solution to Decision Uncertainty and Misalignment in Language Models
Listed - Listed
-
Expectations for Gemini: hopefully not a big deal
Listed -
Modelling large-scale cyber attacks from advanced AI systems with Advanced Persistent Threats
Listed -
Observations on the funding landscape of EA and AI safety
Listed -
Representation Engineering: A Top-Down Approach to AI Transparency
Listed -
AI Safety Impact Markets: Your Charity Evaluator for AI Safety
Listed -
Join AISafety.info's Distillation Hackathon (Oct 6-9th)
Listed -
New Tool: the Residual Stream Viewer
Listed -
Announcing the Winners of the 2023 Open Philanthropy AI Worldviews Contest
Listed -
Focusing your impact on short vs long TAI timelines
Listed -
How model editing could help with the alignment problem
Listed -
Introducing Future Matters – a strategy consultancy
Listed -
"Diamondoid bacteria" nanobots: deadly threat or dead-end? A nanotech investigation
Listed -
Anki deck for learning the main AI safety orgs, projects, and programs
Listed -
Steering subsystems: capabilities, agency, and alignment
Listed -
The Retroactive Funding Landscape: Innovations for Donors and Grantmakers
Listed -
Alibaba Group releases Qwen, 14B parameter LLM
Listed - Listed
-
ARC Evals: Responsible Scaling Policies
Listed -
Culture and Programming Retrospective: ERA Fellowship 2023
Listed -
Different views of alignment have different consequences for imperfect methods
Listed -
High-level interpretability: detecting an AI's objectives
Listed -
How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions
Listed -
Tarbell Fellowship 2024 - Applications Open (AI Journalism)
Listed -
Projects I would like to see (possibly at AI Safety Camp)
Listed -
[Linkpost] Prospect Magazine - How to save humanity from extinction
Listed -
Announcing the CNN Interpretability Competition
Listed - Listed
- Listed
- Listed
-
International AI Institutions: a literature review of models, examples, and proposals
Listed -
It Is Powerful, It Can't Be Aimed
Listed -
Let's think about...lowering the burden of proof for liability for harms associated with AI.
Listed -
News: Spanish AI image outcry + US AI workforce "regulation"
Listed - Listed
-
Amazon to invest up to $4 billion in Anthropic
Listed -
How to pursue a career in AI governance and coordination
Listed -
Impact stories for model internals: an exercise for interpretability researchers
Listed -
Public Opinion on AI Safety: AIMS 2023 and 2021 Summary
Listed -
Public Opinion on AI Safety: AIMS 2023 and 2021 Summary
Listed -
Understanding strategic deception and deceptive alignment
Listed -
Understanding strategic deception and deceptive alignment
Listed -
Welcome to Apply: The 2024 Vitalik Buterin Fellowships in AI Existential Safety by FLI!
Listed -
What causes a decision theory to be used?
Listed -
What is wrong with this "utility switch button problem" approach?
Listed -
“X distracts from Y” as a thinly-disguised fight over group status / politics
Listed -
Five neglected work areas that could reduce AI risk
Listed -
Five neglected work areas that could reduce AI risk
Listed - Listed
-
"We can Prevent AI Disaster Like We Prevented Nuclear Catastrophe"
Listed -
I designed an AI safety course (for a philosophy department)
Listed -
I designed an AI safety course (for a philosophy department)
Listed -
It’s not obvious that getting dangerous AI later is better
Listed - Listed
-
Evidence to prioritize or working on AI as the most impactful thing?
Listed - Listed
-
Intro to AI risk for AI grad students?
Listed - Listed
-
Let's talk about Impostor syndrome in AI safety
Listed