The catalog, page 7
Records 1,501 to 1,750 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
A very non-technical explanation of the basics of infra-Bayesianism
Listed - Listed
- Listed
-
How Many Bits Of Optimization Can One Bit Of Observation Unlock?
Listed -
I was Wrong, Simulator Theory is Real
Listed -
Infra-Bayesianism naturally leads to the monotonicity principle, and I think this is a problem
Listed -
Is there EA discussion on non-x-risk transformative AI?
Listed -
Join a ‘learning by writing' group
Listed -
LM Situational Awareness, Evaluation Proposal: Violating Imitation
Listed -
AI Safety Newsletter #3: AI policy proposals and a new challenger approaches
Listed -
Briefly how I've updated since ChatGPT
Listed - Listed
-
Making Nanobots isn't a one-shot process, even for an artificial superintelligance
Listed -
My Assessment of the Chinese AI Safety Community
Listed -
Notes on Potential Future AI Tax Policy
Listed - Listed
- Listed
-
UK Government announces £100 million in funding for Foundation Model Taskforce.
Listed -
A concise sum-up of the basic argument for AI doom
Listed -
Consequentialism is in the Stars not Ourselves
Listed -
For alignment, we should simultaneously use multiple theories of cognition and value
Listed -
FT: We must slow down the race to God-like AI
Listed - Listed
-
No, the EMH does not imply that markets have long AGI timelines
Listed -
Student competition for drafting a treaty on moratorium of large-scale AI capabilities R&D
Listed -
The Case For Civil Disobedience For The AI Movement
Listed -
Value Learning – Towards Resolving Confusion
Listed - Listed
-
A great talk for AI noobs (according to an AI noob)
Listed -
A great talk for AI noobs (according to an AI noob)
Listed -
Endo-, Dia-, Para-, and Ecto-systemic novelty
Listed -
Preventing AI Misuse: State of the Art Research and its Flaws
Listed -
Why do we care about agency for alignment?
Listed -
PhD Position: AI Interpretability in Berlin, Germany
Listed -
The Security Mindset, S-Risk and Publishing Prosaic Alignment Research
Listed -
"Who Will You Be After ChatGPT Takes Your Job?"
Listed - Listed
- Listed
- Listed
-
Should we publish mechanistic interpretability research?
Listed -
The basic reasons I expect AGI ruin
Listed -
the multiverse argument argument against automated alignment
Listed -
Thinking about maximization and corrigibility
Listed - Listed
-
An open letter to SERI MATS program organisers
Listed -
Behavioural statistics for a maze-solving agent
Listed - Listed
-
Japan AI Alignment Conference Postmortem
Listed -
Language Models are a Potentially Safe Path to Human-Level AGI
Listed -
Merger of DeepMind and Google Brain
Listed -
OpenAI could help X-risk by wagering itself
Listed -
Proposal: Using Monte Carlo tree search instead of RLHF for alignment research
Listed - Listed
-
Responsible Deployment in 20XX
Listed -
Stability AI releases StableLM, an open-source ChatGPT counterpart
Listed -
The Economist feature articles on LLMs
Listed -
'AI Emergency Eject Criteria' Survey
Listed -
12 tentative ideas for US AI policy (Luke Muehlhauser)
Listed - Listed
-
Approximation is expensive, but the lunch is cheap
Listed -
Artificial Intelligence: Challenges and Opportunities for the Department of Defense
Listed -
Davidad's Bold Plan for Alignment: An In-Depth Explanation
Listed -
Is there any literature on using socialization for AI alignment?
Listed -
Organizing a debate with experts and MPs to raise AI xrisk awareness: a possible blueprint
Listed -
Orthogonal: A new agent foundations alignment organization
Listed - Listed
-
The Learning-Theoretic Agenda: Status 2023
Listed -
[Linkpost] AI Alignment, Explained in 5 Points (updated)
Listed -
AI Safety Newsletter #2: ChaosGPT, Natural Selection, and AI Safety in the Media
Listed -
AI Safety Newsletter #2: ChaosGPT, Natural Selection, and AI Safety in the Media
Listed -
Capabilities and alignment of LLM cognitive architectures
Listed - Listed
- Listed
- Listed
- Listed
- Listed
-
AI Alignment Research Engineer Accelerator (ARENA): call for applicants
Listed -
AI Impacts Quarterly Newsletter, Jan-Mar 2023
Listed -
AI Impacts Quarterly Newsletter, Jan-Mar 2023
Listed - Listed
-
An alternative of PPO towards alignment
Listed - Listed
- Listed
-
Import AI 325: Automated mad science; AI vs democracy; and a 12B parameter language model
Listed -
Prediction: any uncontrollable AI will turn earth into a giant computer
Listed - Listed
-
What is your timelines for ADI (artificial disempowering intelligence)?
Listed -
[Link/crosspost] [US] NTIA: AI Accountability Policy Request for Comment
Listed -
Mechanistically interpreting time in GPT-2 small
Listed - Listed
-
Summary: The Case for Halting AI Development - Max Tegmark on the Lex Fridman Podcast
Listed -
An example elevator pitch for AI doom
Listed -
Brain-computer interfaces and brain organoids in AI alignment?
Listed - Listed
-
Concrete, existing examples of high-impact risks from AI?
Listed -
FLI report: Policymaking in the Pause
Listed -
Open-source LLMs may prove Bostrom's vulnerable world hypothesis
Listed -
SmartyHeaderCode: anomalous tokens for GPT3.5 and GPT-4
Listed -
Who is testing AI Safety public outreach messaging?
Listed -
"Do X because decision theory" ~= "Do X because bayes theorem"
Listed -
"Risk Awareness Moments" (Rams): A concept for thinking about AI governance interventions
Listed -
[linkpost] "What Are Reasonable AI Fears?" by Robin Hanson, 2023-04-23
Listed -
[Linkpost] The A.I. Dilemma - March 9, 2023, with Tristan Harris and Aza Raskin
Listed -
A freshman year during the AI midgame: my approach to the next year
Listed -
AI Safety Europe Retreat 2023 Retrospective
Listed -
Anti-'FOOM' (stop trying to make your cute pet name the thing)
Listed -
GPT-4 is easily controlled/exploited with tricky decision theoretic dilemmas.
Listed -
List of requests for an AI slowdown/halt.
Listed -
Prospects for AI safety agreements between countries
Listed -
Research Report: Incorrectness Cascades
Listed -
Shapley Value Attribution in Chain of Thought
Listed - Listed
-
What we’ve learned so far from our technological temptations project
Listed -
"Aligned" foundation models don't imply aligned systems
Listed -
[US] NTIA: AI Accountability Policy Request for Comment
Listed -
AGI - alignment - paperclip maximizer - pause - defection - incentives
Listed -
Announcing Epoch’s dashboard of key trends and figures in Machine Learning
Listed -
Financial Times: We must slow down the race to God-like AI
Listed -
Identifying semantic neurons, mechanistic circuits & interpretability web apps
Listed -
Intro to Ontogenetic Curriculum
Listed -
Navigating the Open-Source AI Landscape: Data, Funding, and Safety
Listed -
What is the best source to explain short AI timelines to a skeptical person?
Listed -
[Link] Sarah Constantin: "Why I am Not An AI Doomer"
Listed -
[linkpost] AI NOW Institute's 2023 Annual Report & Roadmap
Listed -
AGI goal space is big, but narrowing might not be as hard as it seems.
Listed -
AI x-risk, approximately ordered by embarrassment
Listed - Listed
- Listed
-
Apply to >50 AI safety funders in one application with the Nonlinear Network [Round Closed]
Listed -
Artificial Intelligence as exit strategy from the age of acute existential risk
Listed -
AXRP Episode 20 - ‘Reform’ AI Alignment with Scott Aaronson
Listed -
Boundaries-based security and AI safety approaches
Listed -
Gradient Descent in Activation Space: a Tale of Two Papers
Listed -
Localizing Model Behavior With Path Patching
Listed - Listed
-
Navigating the Open-Source AI Landscape: Data, Funding, and Safety
Listed -
Towards a solution to the alignment problem via objective detection and evaluation
Listed -
[Linkpost] 538 Politics Podcast on AI risk & politics
Listed - Listed
- Listed
-
AI Risk US Presidental Candidate
Listed - Listed
-
Data Taxation: A Proposal for Slowing Down AGI Progress
Listed -
Evolution provides no evidence for the sharp left turn
Listed -
Existential risk x Crypto: An unconference at Zuzalu
Listed -
FLI And Eliezer Should Reach Consensus
Listed -
Four mindset disagreements behind existential risk disagreements in ML
Listed - Listed
-
Introducing the Mental Health Roadmap Series
Listed -
Measuring artificial intelligence on human benchmarks is naive
Listed -
Metaculus’ predictions are much better than low-information priors
Listed - Listed
- Listed
-
NTIA - AI Accountability Announcement
Listed -
Paleontological study of extinctions supports AI as a existential threat to humanity
Listed -
Preliminary investigations on if STEM and EA communities could benefit from more overlap
Listed -
Request to AGI organizations: Share your views on pausing AI progress
Listed -
Some Intuitions Around Short AI Timelines Based on Recent Progress
Listed - Listed
-
AI Safety Newsletter #1 [CAIS Linkpost]
Listed -
An AI Realist Manifesto: Neither Doomer nor Foomer, but a third more reasonable thing
Listed -
Current UK government levers on AI development
Listed -
Humans are not prepared to operate outside their moral training distribution
Listed -
Misgeneralization as a misnomer
Listed -
Which stocks or ETFs should you invest in to take advantage of a possible AGI explosion, and why?
Listed -
Why I'm not worried about imminent doom
Listed -
Why Simulator AIs want to be Active Inference AIs
Listed -
Agentized LLMs will change the alignment landscape
Listed -
Expanding the domain of discourse reveals structure already there but hidden
Listed -
Foom seems unlikely in the current LLM training paradigm
Listed - Listed
-
All AGI Safety questions welcome (especially basic ones) [April 2023]
Listed -
All images from the WaitButWhy sequence on AI
Listed -
Can we evaluate the "tool versus agent" AGI prediction?
Listed -
GPTs are Predictors, not Imitators
Listed -
How does a company like Instadeep fit into the current AI landscape?
Listed -
Pausing AI Developments Isn't Enough. We Need to Shut it All Down
Listed -
Pausing AI Developments Isn’t Enough. We Need to Shut it All Down
Listed -
SERI MATS - Summer 2023 Cohort
Listed -
An 'AGI Emergency Eject Criteria' consensus could be really useful.
Listed -
Beren's "Deconfusing Direct vs Amortised Optimisation"
Listed -
Environments for Measuring Deception, Resource Acquisition, and Ethical Violations
Listed -
Goal alignment without alignment on epistemology, ethics, and science is futile
Listed -
How much should states invest in contingency plans for widespread internet outage?
Listed -
If Alignment is Hard, then so is Self-Improvement
Listed -
Imagine AGI killed us all in three years. What would have been our biggest mistakes?
Listed -
n=3 AI Risk Quick Math and Reasoning
Listed -
Risks from GPT-4 Byproduct of Recursively Optimizing AIs
Listed -
Select Agent Specifications as Natural Abstractions
Listed -
Should we publish arguments for the preservation of humanity?
Listed -
Stampy's AI Safety Info - New Distillations #1 [March 2023] (Expansive interactive FAQ)
Listed -
Superintelligence Is Not Omniscience
Listed -
AISafety.world is a map of the AIS ecosystem
Listed -
AISafety.world is a map of the AIS ecosystem
Listed -
Daisy-chaining epsilon-step verifiers
Listed -
Debates on reducing long-term s-risks?
Listed - Listed
- Listed
-
Is "Recursive Self-Improvement" Relevant in the Deep Learning Paradigm?
Listed - Listed
-
Misgeneralization as a misnomer
Listed -
Some Preliminary Opinions on AI Safety Problems
Listed -
The Computational Anatomy of Human Values
Listed - Listed
-
Yoshua Bengio: "Slowing down development of AI systems passing the Turing test"
Listed -
"Corrigibility at some small length" by dath ilan
Listed -
Best arguments against instrumental convergence?
Listed -
Empathy bandaid for immediate AI catastrophe
Listed -
OpenAI: Our approach to AI safety
Listed -
The Orthogonality Thesis is Not Obviously True
Listed -
Universality and Hidden Information in Concept Bottleneck Models
Listed -
What to suggest companies & entrepreneurs do to use AI safely?
Listed - Listed
-
Excessive AI growth-rate yields little socio-economic benefit.
Listed -
Giant (In)scrutable Matrices: (Maybe) the Best of All Possible Worlds
Listed -
Keep Chasing AI Safety Press Coverage
Listed -
Penalize Model Complexity Via Self-Distillation
Listed -
Why might AI be a x-risk? Succinct explanations please
Listed - Listed
-
Apply to the Cavendish Labs Fellowship (by 4/15)
Listed -
Exploratory Analysis of RLHF Transformers with TransformerLens
Listed -
If interpretability research goes well, it may get dangerous
Listed -
Import AI 323: AI researcher warns about AI; BloombergGPT; and an open source Flamingo
Listed -
Mati's introduction to pausing giant AI experiments
Listed -
Platform for Project Spitballing? (e.g., for AI field building)
Listed -
Reducing profit motivations in AI development
Listed -
Repeated Play of Imperfect Newcomb's Paradox in Infra-Bayesian Physicalism
Listed -
Towards empathy in RL agents and beyond: Insights from cognitive science for AI Alignment
Listed - Listed
-
AISC 2023, Progress Report for March: Team Interpretable Architectures
Listed -
Exploratory Analysis of TRLX RLHF Transformers with TransformerLens
Listed - Listed
- Listed
-
Predictions for future AI governance?
Listed -
Research Summary: Forecasting with Large Language Models
Listed -
Transparency for Generalizing Alignment from Toy Models
Listed -
Ultimate ends may be easily hidable behind convergent subgoals
Listed - Listed
-
A policy guaranteed to increase AI timelines
Listed -
AI community building: EliezerKart
Listed - Listed
-
Campaign for AI Safety: Please join me
Listed -
How to persuade a non-CS background person to believe AGI is 50% possible in 2040?
Listed - Listed
-
Policy discussions follow strong contextualizing norms
Listed -
Singularities against the Singularity: Announcing Workshop on Singular Learning Theory and Alignment
Listed -
Vael Gates: Risks from Highly-Capable AI (March 2023)
Listed -
Γαμινγκ the Algorithms: Large Language Models as Mirrors
Listed -
AI, Cybersecurity, and Malware: A Shallow Report [General]
Listed -
AI, Cybersecurity, and Malware: A Shallow Report [Technical]
Listed