The catalog, page 18
Records 4,251 to 4,500 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
A Bird's Eye View of the ML Field [Pragmatic AI Safety #2]
Listed -
AI Alignment YouTube Playlists
Listed - Listed
-
Aligned with Whom? Direct and social goals for AI systems
Listed -
Conditions for mathematical equivalence of Stochastic Gradient Descent and Natural Selection
Listed -
Introduction to Pragmatic AI Safety [Pragmatic AI Safety #1]
Listed -
Introduction to Pragmatic AI Safety [Pragmatic AI Safety #1]
Listed -
Jobs: Help scale up LM alignment research at NYU
Listed -
Student project for engaging with AI alignment
Listed -
Transcripts of interviews with AI researchers
Listed - Listed
-
When is AI safety research harmful?
Listed -
A Survey on AI Sustainability: Emerging Trends on Learning Algorithms and Research Challenges
Listed -
Algorithmic formalization of FDT?
Listed - Listed
-
README-by Vael Gates-date 20220509
Listed -
Video and Transcript of Presentation on Existential Risk from Power-Seeking AI
Listed -
Video and Transcript of Presentation on Existential Risk from Power-Seeking AI
Listed -
What are the coolest topics in AI safety, to a hopelessly pure mathematician?
Listed -
What does Functional Decision Theory say to do in imperfect Newcomb situations?
Listed -
Active offline policy selection
Listed - Listed
-
Apply to the second ML for Alignment Bootcamp (MLAB 2) in Berkeley [Aug 15 - Fri Sept 2]
Listed -
Open Problems in Negative Side Effect Minimization
Listed -
The case for becoming a black-box investigator of language models
Listed -
A Deep Reinforcement Learning Framework for Rapid Diagnosis of Whole Slide Pathological Images
Listed -
High-stakes alignment via adversarial training [Redwood Research report]
Listed - Listed
- Listed
- Listed
-
Ethan Caballero-by The Inside View-date 20220505
Listed -
Introducing the ML Safety Scholars Program
Listed -
Introducing the ML Safety Scholars Program
Listed -
Adversarial Training for High-Stakes Reliability
Listed -
Is evolutionary influence the mesa objective that we're interested in?
Listed -
Information security considerations for AI and the long term future
Listed -
My thoughts on nanotechnology strategy research as an EA cause area
Listed -
The AI Index 2022 Annual Report
Listed -
What are the best journals to publish AI governance papers in?
Listed -
A tale of 2.5 orthogonality theses
Listed - Listed
-
Note-Taking without Hidden Messages
Listed -
Quick Thoughts on A.I. Governance
Listed - Listed
-
Do FDT (or similar) recommend reparations?
Listed - Listed
-
Prize for Alignment Research Tasks
Listed - Listed
-
Training Language Models with Language Feedback
Listed -
Slides: Potential Risks From Advanced AI
Listed -
[Intro to brain-like-AGI safety] 13. Symbol grounding & human social instincts
Listed - Listed
- Listed
-
If you’re very optimistic about ELK then you should be optimistic about outer alignment
Listed -
Law-Following AI 1: Sequence Introduction and Structure
Listed -
Law-Following AI 2: Intent Alignment + Superintelligence → Lawless AI (By Default)
Listed -
Law-Following AI 2: Intent Alignment + Superintelligence → Lawless AI (By Default)
Listed -
Law-Following AI 3: Lawless AI Agents Undermine Stabilizing Agreements
Listed -
Law-Following AI 3: Lawless AI Agents Undermine Stabilizing Agreements
Listed -
SERI ML Alignment Theory Scholars Program 2022
Listed -
SERI ML Alignment Theory Scholars Program 2022
Listed -
The Speed + Simplicity Prior is probably anti-deceptive
Listed -
[$20K in Prizes] AI Safety Arguments Competition
Listed -
[$20K In Prizes] AI Safety Arguments Competition
Listed -
Framings of Deceptive Alignment
Listed -
How to engage with AI 4 Social Justice actors
Listed -
Why Copilot Accelerates Timelines
Listed - Listed
-
Intuitions about solving hard problems
Listed -
Key questions about artificial sentience: an opinionated guide
Listed -
Make a neural network in ~10 minutes
Listed -
Towards Evaluating Adaptivity of Model-Based Reinforcement Learning Methods
Listed -
What is being improved in recursive self improvement?
Listed -
Which Post Idea Is Most Effective?
Listed -
Examining Evolution as an Upper Bound for AGI Timelines
Listed -
Skilling-up in ML Engineering for Alignment: request for comments
Listed -
[ASoT] Consequentialist models as a superset of mesaoptimizers
Listed -
Calling for Student Submissions: AI Safety Distillation Contest
Listed - Listed
- Listed
-
Instrumental Convergence To Offer Hope?
Listed -
Choice := Anthropics uncertainty? And potential implications for agency
Listed -
For every choice of AGI difficulty, conditioning on gradual take-off implies shorter timelines.
Listed -
Path-Specific Objectives for Safer Agent Incentives
Listed -
The Risks of Machine Learning Systems
Listed - Listed
-
[Intro to brain-like-AGI safety] 12. Two paths forward: “Controlled AGI” and “Social-instinct AGI”
Listed -
GPT-3 and concept extrapolation
Listed -
Why No *Interesting* Unaligned Singularity?
Listed -
[Closed] Hiring a mathematician to work on the learning-theoretic AI alignment agenda
Listed -
Another argument that you will let the AI out of the box
Listed -
Chaining Retroactive Funders to Borrow Against Unlikely Utopias
Listed -
Concept extrapolation: key posts
Listed -
Deceptive Agents are a Good Way to Do Things
Listed -
“Pivotal Act” Intentions: Negative Consequences and Fallacious Arguments
Listed -
“Pivotal Act” Intentions: Negative Consequences and Fallacious Arguments
Listed -
Hierarchical Optimal Transport for Comparing Histopathology Datasets
Listed -
How will the world respond to "AI x-risk warning shots" according to reference class forecasting?
Listed -
How I failed to form views on AI safety
Listed -
What is causality to an evidential decision theorist?
Listed -
Why not offer a multi-million / billion dollar prize for solving the Alignment Problem?
Listed -
A grand strategy to recruit AI capabilities researchers into AI safety research
Listed -
Begging, Pleading AI Orgs to Comment on NIST AI Risk Management Framework
Listed - Listed
-
Everything I Need To Know About Takeoff Speeds I Learned From Air Conditioner Ratings On Amazon
Listed -
Please Share Your Perspectives on the Degree of Societal Impact from Transformative AI Outcomes
Listed -
Refine: An Incubator for Conceptual Alignment Research Bets
Listed -
Some reasons why a predictor wants to be a consequentialist
Listed - Listed
-
How dath ilan coordinates around solving AI alignment
Listed -
Methodical Advice Collection and Reuse in Deep Reinforcement Learning
Listed -
Redwood Research is hiring for several roles (Operations and Technical)
Listed -
Another list of theories of impact for interpretability
Listed -
Flexible Multiple-Objective Reinforcement Learning for Chip Placement
Listed -
Hierarchical text-conditional image generation with CLIP latents
Listed -
Takeoff speeds have a huge effect on what it means to work on AI x-risk
Listed -
What more compute does for brain-like models: response to Rohin
Listed -
What to include in a guest lecture on existential risks from AI?
Listed - Listed
-
6 Year Decrease of Metaculus AGI Prediction
Listed -
A broad basin of attraction around human values?
Listed -
A primer & some reflections on recent CSER work (EAB talk)
Listed -
A Small Negative Result on Debate
Listed -
AdaTest:Reinforcement Learning and Adaptive Sampling for On-chip Hardware Trojan Detection
Listed -
AI governance student hackathon on Saturday, April 23: register now!
Listed -
An empirical analysis of compute-optimal large language model training
Listed -
finding earth in the universal program
Listed -
Help us find pain points in AI safety
Listed -
How to become an AI safety researcher
Listed -
Is technical AI alignment research a net positive?
Listed - Listed
- Listed
-
Reward model hacking as a challenge for reward learning
Listed -
Three questions about mesa-optimizers
Listed -
Tips for conducting worldview investigations
Listed -
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Listed -
Useful Vices for Wicked Problems
Listed - Listed
-
Credo AI is hiring for several roles
Listed -
Goodhart's Law Causal Diagrams
Listed -
Linguistic communication as (inverse) reward design
Listed -
Metaethical Perspectives on 'Benchmarking' AI Ethics
Listed - Listed
-
The Regulatory Option: A response to near 0% survival odds
Listed -
A visualization of some orgs in the AI Safety Pipeline
Listed -
Crucial considerations in the field of Wild Animal Welfare (WAW)
Listed -
Enhancing the Robustness, Efficiency, and Diversity of Differentiable Architecture Search
Listed -
A concrete bet offer to those with short AGI timelines
Listed - Listed
-
AMA Conjecture, A New Alignment Startup
Listed -
bracing for the alignment tunnel
Listed -
Elicit: Language Models as Research Assistants
Listed - Listed
-
The right to protection from catastrophic AI risk
Listed -
[RETRACTED] It's time for EA leadership to pull the short-timelines fire alarm.
Listed -
AIs should learn human preferences, not biases
Listed -
Different perspectives on concept extrapolation
Listed - Listed
-
Language Model Tools for Alignment Research
Listed -
We Are Conjecture, A New Alignment Research Startup
Listed -
[ASoT] Some thoughts about imperfect world modeling
Listed - Listed
-
Ideal governance (for companies, countries and more)
Listed -
Is GPT3 a Good Rationalist? - InstructGPT3 [2/2]
Listed -
New Sequence - Towards a worldwide, watertight Windfall Clause
Listed -
Productive Mistakes, Not Perfect Answers
Listed -
Robust Event-Driven Interactions in Cooperative Multi-Agent Learning
Listed -
Truthfulness, standards and credibility
Listed -
What Should We Optimize - A Conversation
Listed -
[Cross-post] Change my mind: we should define and measure the effectiveness of advanced AI
Listed -
[Intro to brain-like-AGI safety] 11. Safety ≠ alignment (but they’re close!)
Listed -
[Link] A minimal viable product for alignment
Listed -
[Link] Why I’m excited about AI-assisted human feedback
Listed -
A Cognitive Framework for Delegation Between Error-Prone AI and Human Agents
Listed -
PaLM in "Extrapolating GPT-N performance"
Listed - Listed
-
AXRP Episode 14 - Infra-Bayesian Physicalism with Vanessa Kosoy
Listed -
Ideal governance (for companies, countries and more)
Listed -
Supervise Process, not Outcomes
Listed -
The case for Doing Something Else (if Alignment is doomed)
Listed -
What an actually pessimistic containment strategy looks like
Listed -
Yudkowsky and Christiano on AI Takeoff Speeds [LINKPOST]
Listed - Listed
-
Project Intro: Selection Theorems for Modularity
Listed -
Theories of Modularity in the Biological Literature
Listed -
AI Governance across Slow/Fast Takeoff and Easy/Hard Alignment spectra
Listed -
Is it valuable to the field of AI Safety to have a neuroscience background?
Listed -
On Agent Incentives to Manipulate Human Feedback in Multi-Agent Reward Learning Scenarios
Listed -
Optimality is the tiger, and agents are its teeth
Listed -
What are the best ideas of how to regulate AI from the US executive branch?
Listed - Listed
- Listed
- Listed
-
New Scaling Laws for Large Language Models
Listed -
Questions about ''formalizing instrumental goals"
Listed -
Replacing Karma with Good Heart Tokens (Worth $1!)
Listed -
[Link] Training Compute-Optimal Large Language Models
Listed -
AXRP Episode 13 - First Principles of AGI Safety with Richard Ngo
Listed -
[ASoT] Some thoughts about LM monologue limitations and ELK
Listed -
[Intro to brain-like-AGI safety] 10. The alignment problem
Listed -
Announcing the EU Tech Policy Fellowship
Listed -
ELK Computational Complexity: Three Levels of Difficulty
Listed - Listed
-
No, EDT Did Not Get It Right All Along: Why the Coin Flip Creation Problem Is Irrelevant
Listed -
Pitching AI Safety in 3 sentences
Listed -
Procedurally evaluating factual accuracy: a request for research
Listed -
Request for Assistance - Research on Scenario Development for Advanced AI Risk
Listed -
8 possible high-level goals for work on nuclear risk
Listed -
A dataset for AI/superintelligence stories and other media?
Listed - Listed
-
Can we simulate human evolution to create a somewhat aligned AGI?
Listed -
Debating myself on whether “extra lives lived” are as good as “deaths prevented”
Listed -
Gears-Level Mental Models of Transformer Interpretability
Listed - Listed
-
should we implement free will?
Listed -
Towards a better circuit prior: Improving on ELK state-of-the-art
Listed -
What would make you confident that AGI has been achieved?
Listed -
[ASoT] Some thoughts about deceptive mesaoptimization
Listed -
A Primer on God, Liberalism and the End of History
Listed - Listed
-
Seeking Survey Responses - Attitudes Towards AI risks
Listed -
The role of academia in AI Safety.
Listed - Listed
-
[ASoT] Searching for consequentialist structure
Listed -
[ASoT] Some ways ELK could still be solvable in practice
Listed -
Practical everyday human strategizing
Listed -
Scenario Mapping Advanced AI Risk: Request for Participation with Data Collection
Listed - Listed
-
Compute Governance: The Role of Commodity Hardware
Listed -
When people ask for your P(doom), do you give them your inside view or your betting odds?
Listed -
I'm interviewing Nova Das Sarma about AI safety and information security. What shouId I ask her?
Listed -
What's the best machine learning newsletter? How do you keep up to date?
Listed -
Why Agent Foundations? An Overly Abstract Explanation
Listed -
A Rationale-Centric Framework for Human-in-the-loop Machine Learning
Listed -
AI Safety Overview: CERI Summer Research Fellowship
Listed -
Data Publication for the 2021 Artificial Intelligence, Morality, and Sentience (AIMS) Survey
Listed - Listed
- Listed
-
On expected utility, part 4: Dutch books, Cox, and Complete Class
Listed -
Your Policy Regulariser is Secretly an Adversary
Listed -
[Intro to brain-like-AGI safety] 9. Takeaways from neuro 2/2: On AGI motivation
Listed -
A survey of tool use and workflows in alignment research
Listed -
Meditations on careers in AI Safety
Listed -
NeurIPSorICML_bj9ne-by Vael Gates-date 20220324
Listed -
NeurIPSorICML_cvgig-by Vael Gates-date 20220324
Listed -
goals for emergency unaligned AI
Listed - Listed
-
are there finitely many moral patients?
Listed - Listed
- Listed