The catalog, page 8
Records 1,751 to 2,000 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
ChatGPT banned in Italy over privacy concerns
Listed -
Critiques of prominent AI safety labs: Redwood Research
Listed - Listed
-
Human Values and AGI Risk | William James
Listed -
Imagine a world where Microsoft employees used Bing
Listed -
It Can't Be Mesa-Optimizers All The Way Down (Or Else It Can't Be Long-Term Supercoherence?)
Listed - Listed
- Listed
-
Maze-solving agents: Add a top-right vector, make the agent go to the top-right
Listed -
Seeking advice on impactful career paths given my unique capabilities and interests
Listed -
What are the biggest obstacles on AI safety research career?
Listed -
Widening Overton Window - Open Thread
Listed -
[Event] Join Metaculus Tomorrow, March 31st, for Forecast Friday!
Listed - Listed
-
ChatGPT is capable of cognitive empathy!
Listed -
Deference on AI timelines: survey results
Listed -
How is AI governed and regulated, around the world?
Listed -
Imitation Learning from Language Feedback
Listed -
Longtermism and shorttermism can disagree on nuclear war to stop advanced AI
Listed -
Nuclear brinksmanship is not a good AI x-risk strategy
Listed -
Recruit the World’s best for AGI Alignment
Listed -
Role Architectures: Applying LLMs to consequential tasks
Listed - Listed
-
You Can’t Predict a Game of Pinball
Listed -
"Sorcerer's Apprentice" from Fantasia as an analogy for alignment
Listed -
Actually, Othello-GPT Has A Linear Emergent World Representation
Listed - Listed
-
Nobody’s on the ball on AGI alignment
Listed -
Othello-GPT: Future Work I Am Excited About
Listed -
Othello-GPT: Reflections on the Research Process
Listed -
Pausing AI Developments Isn't Enough. We Need to Shut it All Down by Eliezer Yudkowsky
Listed - Listed
-
Want to win the AGI race? Solve alignment.
Listed - Listed
- Listed
-
100 Dinners And A Workshop: Information Preservation And Goals
Listed -
A rough and incomplete review of some of John Wentworth's research
Listed -
Corrigibility, Self-Deletion, and Identical Strawberries
Listed -
Explorers in a virtual country: Navigating the knowledge landscape of large language models
Listed - Listed
-
I had a chat with GPT-4 on the future of AI and AI safety
Listed -
Improving Code Generation by Training with Natural Language Feedback
Listed -
It Looks Like You’re Trying To Take Over The World
Listed -
Natural Selection Favors AIs over Humans
Listed -
Some of My Current Impressions Entering AI Safety
Listed -
Training Language Models with Language Feedback at Scale
Listed -
What longtermist projects would you like to see implemented?
Listed -
When Will We Spend Enough to Train Transformative AI
Listed -
Why does advanced AI want not to be shut down?
Listed -
Are there cause priortizations estimates for s-risks supporters?
Listed -
Best resources to learn philosophy of mind and AI?
Listed -
CAIS-inspired approach towards safer and more interpretable AGIs
Listed -
ChatGPT bug leaked users' conversation histories
Listed - Listed
-
GPT-4 is bad at strategic thinking
Listed - Listed
-
Lessons from Convergent Evolution for AI Alignment
Listed -
New blog: Planned Obsolescence
Listed -
Nobody knows how to reliably test for AI safety
Listed - Listed
-
Practical Pitfalls of Causal Scrubbing
Listed - Listed
-
Descriptive vs. specifiable values
Listed -
EleutherAI Second Retrospective: The long version
Listed -
LLM Modularity: The Separability of Capabilities in Large Language Models
Listed -
The alignment stability problem
Listed -
What happens with logical induction when...
Listed -
What would a compute monitoring plan look like? [Linkpost]
Listed -
$500 Bounty/Contest: Explain Infra-Bayes In The Language Of Game Theory
Listed -
A stylized dialogue on John Wentworth's claims about markets and optimization
Listed -
Aligned AI as a wrapper around an LLM
Listed -
Can independent researchers get a sponsored visa for the US or UK?
Listed - Listed
-
Mitigating existential risks associated with human nature and AI: Thoughts on serious measures.
Listed -
My attempt at explaining the case for AI risk in a straightforward way
Listed -
opinions on the consequences of AI
Listed -
Are extrapolation-based AIs alignable?
Listed -
Does GPT-4 exhibit agency when summarizing articles?
Listed -
Exploring Metaculus’ community predictions
Listed -
GPT-2005: A conversation with ChatGPT (featuring semi-functional Wolfram Alpha plugin!)
Listed -
Grinding slimes in the dungeon of AI alignment research
Listed -
Metaculus Predicts Weak AGI in 2 Years and AGI in 10
Listed -
More experiments in GPT-4 agency: writing memos
Listed -
The Concept of Boundary Layer in Language Games and Its Implications for AI
Listed -
Wittgenstein and ML — parameters vs architecture
Listed -
continue working on hard alignment! don't give up!
Listed - Listed
- Listed
-
Join the AI governance and interpretability hackathons!
Listed -
The Overton Window widens: Examples of AI risk in the media
Listed -
The Quantization Model of Neural Scaling
Listed -
Truth and Advantage: Response to a draft of “AI safety seems hard to measure”
Listed -
[Linkpost] Shorter version of report on existential risk from power-seeking AI
Listed -
[Linkpost] Shorter version of report on existential risk from power-seeking AI
Listed -
Announcing the European Network for AI Safety (ENAIS)
Listed -
Empirical risk minimization is fundamentally confused
Listed -
Is Bill Gates overly optomistic about AI?
Listed -
Key Questions for Digital Minds
Listed - Listed
-
The space of systems and the space of maps
Listed -
The space of systems and the space of maps
Listed -
Truth and Advantage: Response to a draft of "AI safety seems hard to measure"
Listed -
Whether you should do a PhD doesn't depend much on timelines.
Listed - Listed
-
"Aligned with who?" Results of surveying 1,000 US participants on AI values
Listed -
Capabilities Denial: The Danger of Underestimating AI
Listed - Listed
- Listed
- Listed
- Listed
-
Future Matters #8: Bing Chat, AI labs on safety, and pausing Future Matters
Listed -
My Objections to "We’re All Gonna Die with Eliezer Yudkowsky"
Listed -
New 'South Park' episode on AI & Chat GPT
Listed -
Some constructions for proof-based cooperation without Löb
Listed -
the QACI alignment plan: table of contents
Listed -
Where I'm at with AI risk: convinced of danger but not (yet) of doom
Listed - Listed
-
RLHF does not appear to differentially cause mode-collapse
Listed -
Sentience in Machines - How Do We Test for This Objectively?
Listed - Listed
-
the QACI alignment plan: table of contents
Listed -
The Wizard of Oz Problem: How incentives and narratives can skew our perception of AI developments
Listed -
"Wide" vs "Tall" superintelligence
Listed -
A tension between two prosaic alignment subgoals
Listed -
How AI could workaround goals if rated by people
Listed -
How much should governments pay to prevent catastrophes? Longtermism’s limited role
Listed -
More information about the dangerous capability evaluations we did with GPT-4 and Claude.
Listed - Listed
-
QACI blob location: an issue with firstness
Listed - Listed
-
you can't simulate the universe from the beginning?
Listed - Listed
-
An Appeal to AI Superintelligence: Reasons to Preserve Humanity
Listed -
Potential employees have a unique lever to influence the behaviors of AI labs
Listed -
Pros and Cons of boycotting paid Chat GPT
Listed -
Would you pursue software engineering as a career today?
Listed -
[Linkpost] Alpaca 7B release | Budget ChatGPT for everybody?
Listed -
Are nested jailbreaks inevitable?
Listed -
GPTs are GPTs: An early look at the labor market impact potential of large language models
Listed -
Survey on intermediate goals in AI governance
Listed - Listed
-
Unjournal: Evaluations of "Artificial Intelligence and Economic Growth", and new hosting space
Listed -
[Appendix] Natural Abstractions: Key Claims, Theorems, and Critiques
Listed -
[ASoT] Some thoughts on human abstractions
Listed -
Are AI developers playing with fire?
Listed -
Attribution Patching: Activation Patching At Industrial Scale
Listed -
ChatGPT getting out of the box
Listed -
Conceding a short timelines bet early
Listed -
Donation offsets for ChatGPT Plus subscriptions
Listed - Listed
- Listed
-
Natural Abstractions: Key claims, Theorems, and Critiques
Listed -
Privileged Bases in the Transformer Residual Stream
Listed - Listed
-
We are fighting a shared battle (a call for a different approach to AI Strategy)
Listed -
What organizations other than Conjecture have (esp. public) info-hazard policies?
Listed -
AI Safety - 7 months of discussion in 17 minutes
Listed -
AI safety and consciousness research: A brainstorm
Listed -
ARC tests to see if GPT-4 can escape human control; GPT-4 failed to do so
Listed -
Good depictions of speed mismatches between advanced AI systems and humans?
Listed -
How well did Manifold predict GPT-4?
Listed -
Shutting Down the Lightcone Offices
Listed -
Towards understanding-based safety evaluations
Listed -
2023 Open Philanthropy AI Worldviews Contest: Odds of Artificial General Intelligence by 2043
Listed -
A better analogy and example for teaching AI takeover: the ML Inferno
Listed -
Comments on OpenAI’s "Planning for AGI and beyond"
Listed -
Eliciting Latent Predictions from Transformers with the Tuned Lens
Listed -
Fixed points in mortal population games
Listed -
GPT can write Quines now (GPT-4)
Listed -
GPT-4 is out: thread (& links)
Listed -
Human preferences as RL critic values - implications for alignment
Listed -
Storytelling Makes GPT-3.5 Deontologist: Unexpected Effects of Context on LLM Behavior
Listed -
What is a definition, how can it be extrapolated?
Listed -
Yudkowsky on AGI risk on the Bankless podcast
Listed -
"Can We Survive Technology?" by John von Neumann
Listed -
Could Roko's basilisk acausally bargain with a paperclip maximizer?
Listed -
Discussion with Nate Soares on a key alignment difficulty
Listed -
Foundations for a Longtermist Foreign Policy
Listed - Listed
- Listed
- Listed
-
Plan for mediocre alignment of brain-like [model-based RL] AGI
Listed -
What Discovering Latent Knowledge Did and Did Not Find
Listed -
your terminal values are complex and not objective
Listed -
Yudkowsky on AGI risk on the Bankless podcast
Listed -
An AI risk argument that resonates with NYTimes readers
Listed - Listed
-
Paper Replication Walkthrough: Reverse-Engineering Modular Addition
Listed -
the quantum amplitude argument against ethics deduplication
Listed -
Thoughts on self-inspecting neural networks.
Listed -
[Linkpost] Scott Alexander reacts to OpenAI's latest post
Listed -
Compositional language for hypotheses about computations
Listed - Listed
-
Understanding and controlling a maze-solving policy network
Listed -
Announcing the Open Philanthropy AI Worldviews Contest
Listed -
Everything's normal until it's not
Listed -
Everything's normal until it's not
Listed - Listed
- Listed
- Listed
-
Reflections On The Feasibility Of Scalable-Oversight
Listed -
Stop calling it "jailbreaking" ChatGPT
Listed - Listed
-
A Roundtable for Safe AI (RSAI)?
Listed -
A Windfall Clause for CEO could worsen AI race dynamics
Listed -
Anthropic's Core Views on AI Safety
Listed -
Anthropic: Core Views on AI Safety: When, Why, What, and How
Listed -
Challenge: construct a Gradient Hacker
Listed -
How bad a future do ML researchers expect?
Listed -
Near-term motivation for AI alignment
Listed - Listed
-
QACI blobs and interval illustrated
Listed -
The Translucent Thoughts Hypotheses and Their Implications
Listed -
Utility uncertainty vs. expected information gain
Listed -
Why Not Just Outsource Alignment Research To An AI?
Listed -
[Crosspost] Why Uncontrollable AI Looks More Likely Than Ever
Listed - Listed
-
AI Safety in a World of Vulnerable Machine Learning Systems
Listed -
QACI blob location: no causality & answer signature
Listed -
Squeezing foundations research assistance out of formal logic narrow AI.
Listed -
Why Uncontrollable AI Looks More Likely Than Ever
Listed -
[Linkpost] Some high-level thoughts on the DeepMind alignment team's strategy
Listed - Listed
-
Introducing AI Alignment Inc., a California public benefit corporation...
Listed -
Should people get neuroscience phD to work in AI safety field?
Listed -
What‘s in your list of unsolved problems in AI alignment?
Listed -
before the sharp left turn: what wins first?
Listed - Listed
-
Introducing Leap Labs, an AI interpretability startup
Listed -
Model-Based Policy Analysis under Deep Uncertainty
Listed -
A concerning observation from media coverage of AI industry dynamics
Listed -
Do humans derive values from fictitious imputed coherence?
Listed -
EA Infosec: skill up in or make a transition to infosec via this book club
Listed -
QACI: the problem of blob location, causality, and counterfactuals
Listed -
Research proposal: Leveraging Jungian archetypes to create values-based models
Listed -
Who Aligns the Alignment Researchers?
Listed -
Why Not Just... Build Weak AI Tools For AI Alignment Research?
Listed - Listed
-
How to navigate potential infohazards
Listed -
Misalignment Museum opens in San Francisco: ‘Sorry for killing most of humanity’
Listed -
More money with less risk: sell services instead of model access
Listed -
The Benefits of Distillation in Research
Listed -
A reply to Byrnes on the Free Energy Principle
Listed - Listed
- Listed
-
AI Governance & Strategy: Priorities, talent gaps, & opportunities
Listed -
Aspiring AI safety researchers should ~argmax over AGI timelines
Listed -
ChatGPT tells stories, and a note about reverse engineering: A Working Paper
Listed -
How popular is ChatGPT? Part 2: slower growth than Pokémon GO
Listed -
Introducing the new Riesgos Catastróficos Globales team
Listed