The catalog, page 19
Records 4,501 to 4,750 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
- Listed
-
making the UD and UDASSA less broken: identifying time steps
Listed -
NeurIPSorICML_lgu5f-by Vael Gates-date 20220322
Listed -
the word "syntax" in programming, linguistics and LISP
Listed -
values system as test-driven development
Listed -
ViM: Out-Of-Distribution with Virtual-logit Matching
Listed -
How might a herd of interns help with AI or biosecurity research tasks/questions?
Listed -
individuallyselected_92iem-by Vael Gates-date 20220321
Listed - Listed
- Listed
-
Robust Action Gap Increasing with Clipped Advantage Learning
Listed -
What EAG sessions would you like on AI?
Listed -
Can you be Not Even Wrong in AI Alignment?
Listed -
Exploring Finite Factored Sets with some toy examples
Listed -
NeurIPSorICML_7oalk-by Vael Gates-date 20220320
Listed - Listed
-
Career Advice: Philosophy + Programming -> AI Safety
Listed -
Cold Takes reader survey - let me know what you want more and less of!
Listed - Listed
- Listed
- Listed
- Listed
- Listed
- Listed
-
AI Risk Management Framework: Initial Draft
Listed -
individuallyselected_7ujun-by Vael Gates-date 20220318
Listed -
individuallyselected_84py7-by Vael Gates-date 20220318
Listed -
individuallyselected_w5cb5-by Vael Gates-date 20220318
Listed -
individuallyselected_zlzai-by Vael Gates-date 20220318
Listed -
NeurIPSorICML_a0nfw-by Vael Gates-date 20220318
Listed -
NeurIPSorICML_q243b-by Vael Gates-date 20220318
Listed -
[Intro to brain-like-AGI safety] 8. Takeaways from neuro 1/2: On AGI development
Listed -
Building AI Innovation Labs together with Companies
Listed -
Danger(s) of theorem-proving AI?
Listed -
GopherCite: Teaching language models to support answers with verified quotes
Listed -
Mediocre AI safety as existential risk
Listed -
Resilient Neural Forecasting Systems
Listed -
Teaching language models to support answers with verified quotes
Listed -
Dual use of artificial-intelligence-powered drug discovery
Listed -
Early-warning Forecasting Center: What it is, and why it'd be cool
Listed -
ELK contest submission: route understanding through the human ontology
Listed -
There should be an AI safety project board
Listed -
Twitter-length responses to 24 AI alignment arguments
Listed -
Algebraic Learning: Towards Interpretable Information Modeling
Listed -
CMKD: CNN/Transformer-Based Cross-Model Knowledge Distillation for Audio Classification
Listed -
Label-only Model Inversion Attack: The Attack that Requires the Least Information
Listed -
New GPT3 Impressive Capabilities - InstructGPT3 [1/2]
Listed -
Compute Trends — Comparison to OpenAI’s AI and Compute
Listed - Listed
-
A Longlist of Theories of Impact for Interpretability
Listed -
[Intro to brain-like-AGI safety] 7. From hardcoded drives to foresighted plans: A worked example
Listed -
A Rephrasing Of and Footnote To An Embedded Agency Proposal
Listed -
Ask AI companies about what they are doing for AI safety?
Listed - Listed
-
ELK Sub - Note-taking in internal rollouts
Listed -
It Looks Like You're Trying To Take Over The World
Listed -
On presenting the case for AI risk
Listed - Listed
-
Towards a Roadmap on Software Engineering for Responsible AI
Listed -
“Intro to brain-like-AGI safety” series—halfway point!
Listed -
[MLSN #3]: NeurIPS Safety Paper Roundup
Listed -
AI Risk is like Terminator; Stop Saying it's Not
Listed -
In-context Learning and Induction Heads
Listed - Listed
- Listed
-
Value extrapolation, concept extrapolation, model splintering
Listed -
An Intuitive Introduction to Causal Decision Theory
Listed -
An Intuitive Introduction to Evidential Decision Theory
Listed -
An Intuitive Introduction to Functional Decision Theory
Listed -
Basic Concepts in Decision Theory
Listed -
Projecting compute trends in Machine Learning
Listed -
Enabling Automated Machine Learning for Model-Driven AI Engineering
Listed -
experience/moral patient deduplication and ethics
Listed -
Preserving and continuing alignment research through a severe global catastrophe
Listed - Listed
-
Is transformative AI the biggest existential risk? Why or why not?
Listed -
A Typology for Exploring the Mitigation of Shortcut Behavior
Listed -
AutoDIME: Automatic Design of Interesting Multi-Agent Environments
Listed - Listed
-
A research agenda for assessing the economic impacts of code generation models
Listed - Listed
-
Graph Neural Networks for Multimodal Single-Cell Data Integration
Listed -
Learning Robust Real-Time Cultural Transmission without Human Data
Listed -
Reasoning about Counterfactuals to Improve Human Inverse Reinforcement Learning
Listed -
What will be some of the most impactful applications of advanced AI in the near term?
Listed -
3D Common Corruptions and Data Augmentation
Listed -
[Intro to brain-like-AGI safety] 6. Big picture of motivation, decision-making, and RL
Listed -
do not hold on to your believed intrinsic values — follow your heart!
Listed - Listed
-
Ngo and Yudkowsky on scientific reasoning and pivotal acts
Listed -
Ordinary and unordinary decision theory
Listed -
Responsible-AI-by-Design: a Pattern Collection for Designing Responsible AI Systems
Listed -
Shah and Yudkowsky on alignment failures
Listed - Listed
-
Would (myopic) general public good producers significantly accelerate the development of AGI?
Listed - Listed
-
AGI x-risk timelines: 10% chance (by year X) estimates should be the headline, not 50%.
Listed - Listed
-
AI Value Alignment Speaker Series Presented By EA Berkeley
Listed -
AI views and disagreements AMA: Christiano, Ngo, Shah, Soares, Yudkowsky
Listed -
Christiano and Yudkowsky on AI predictions and human intelligence
Listed - Listed
-
Being an individual alignment grantmaker
Listed - Listed
-
Late 2021 MIRI Conversations: AMA / Discussion
Listed -
Shah and Yudkowsky on alignment failures
Listed -
Shah and Yudkowsky on alignment failures
Listed -
The dangers in algorithms learning humans' values and irrationalities
Listed -
How I Formed My Own Views About AI Safety
Listed -
How do new models from OpenAI, DeepMind and Anthropic perform on TruthfulQA?
Listed -
IMO challenge bet with Eliezer
Listed -
New Speaker Series on AI Alignment Starting March 3
Listed -
The Quest for a Common Model of the Intelligent Decision Maker
Listed -
University community building seems like the wrong model for AI safety
Listed -
Composing Complex and Hybrid AI Solutions
Listed -
OCR-IDL: OCR Annotations for Industry Document Library Dataset
Listed -
Re: Some thoughts on vegetarianism and veganism
Listed -
The Big Picture Of Alignment (Talk Part 2)
Listed -
The “Slicing Problem” for Computational Theories of Consciousness
Listed - Listed
-
A comment on Ajeya Cotra's draft report on AI timelines
Listed -
All You Need Is Supervised Learning: From Imitation Learning to Meta-RL With Upside Down RL
Listed -
Important, actionable research questions for the most important century
Listed -
Transformer inductive biases & RASP
Listed -
[Intro to brain-like-AGI safety] 5. The “long-term predictor”, and TD learning
Listed -
Christiano and Yudkowsky on AI predictions and human intelligence
Listed -
Christiano and Yudkowsky on AI predictions and human intelligence
Listed - Listed
-
More GPT-3 and symbol grounding
Listed - Listed
-
Probing Image-Language Transformers for Verb Understanding
Listed -
ELK Proposal: Thinking Via A Human Imitator
Listed - Listed
-
Retrieval Augmented Classification for Long-Tail Visual Recognition
Listed - Listed
-
Favorite / most obscure research on understanding DNNs?
Listed -
HCMD-zero: Learning Value Aligned Mechanisms from Data
Listed -
Inferring Lexicographically-Ordered Rewards from Preferences
Listed -
Investigations of Performance and Bias in Human-AI Teamwork in Hiring
Listed -
Ngo and Yudkowsky on scientific reasoning and pivotal acts
Listed -
Ngo and Yudkowsky on scientific reasoning and pivotal acts
Listed -
The Big Picture Of Alignment (Talk Part 1)
Listed - Listed
-
Deconstructing Distributions: A Pointwise Framework of Learning
Listed -
Alignment researchers, how useful is extra compute for you?
Listed -
Analogy of AI Alignment as Raising a Child?
Listed - Listed
-
Thoughts on Dangerous Learned Optimization
Listed -
Critical Checkpoints for Evaluating Defence Models Against Adversarial Attack and Robustness
Listed -
Implications of automated ontology identification
Listed - Listed
- Listed
-
[Intro to brain-like-AGI safety] 4. The “short-term predictor”
Listed -
[Intro to brain-like-AGI safety] 4. The “short-term predictor”
Listed -
Compute Trends Across Three eras of Machine Learning
Listed -
Defending One-Dimensional Ethics
Listed -
How harmful are improvements in AI? + Poll
Listed -
Is ELK enough? Diamond, Matrix and Child AI
Listed -
Predictability and Surprise in Large Generative Models
Listed -
REPL's: a type signature for agents
Listed -
Safe Reinforcement Learning by Imagining the Near Future
Listed - Listed
-
What Does The Natural Abstraction Framework Say About ELK?
Listed -
Zero-Shot Assistance in Sequential Decision Problems
Listed -
[Linkpost] How To Get Into Independent Research On Alignment/Agency
Listed -
A Map to Navigate AI Governance
Listed -
Question 5: The timeline hyperparameter
Listed -
A Simplified Variant of Gödel's Ontological Argument
Listed -
Abstractions as Redundant Information
Listed -
Is a career in making AI systems more secure a meaningful way to mitigate the X-risk posed by AGI?
Listed -
Question 4: Implementing the control proposals
Listed -
Defending against Adversarial Policies in Reinforcement Learning with Alternating Training
Listed -
Question 3: Control proposals for minimizing bad outcomes
Listed -
Uncalibrated Models Can Improve Human-AI Collaboration
Listed - Listed
-
Predicting Out-of-Distribution Error with the Projection Norm
Listed -
Question 2: Predicted bad outcomes of AGI learning architecture
Listed -
To Match the Greats, Don’t Follow In Their Footsteps
Listed -
A summary of aligning narrowly superhuman models
Listed -
Inferring utility functions from locally non-transitive preferences
Listed -
Interpretable pipelines with evolutionarily optimized modules for RL tasks with visual inputs
Listed -
Locating and Editing Factual Associations in GPT
Listed -
Proceedings of the Robust Artificial Intelligence System Assurance (RAISA) Workshop 2022
Listed -
Question 1: Predicted architecture of AGI learning algorithm(s)
Listed - Listed
-
[Intro to brain-like-AGI safety] 3. Two subsystems: Learning & Steering
Listed - Listed
-
"Moral progress" vs. the simple passage of time
Listed -
Defending Functional Decision Theory
Listed -
How complex are myopic imitators?
Listed -
Hypothesis: gradient descent prefers general circuits
Listed -
Local Explanations for Reinforcement Learning
Listed -
Machine Explanations and Human Understanding
Listed -
Metaculus launches contest for essays with quantitative predictions about AI
Listed -
Paradigm-building: Introduction
Listed -
Software engineering - Career review
Listed -
Red Teaming Language Models with Language Models
Listed -
Red Teaming Language Models with Language Models
Listed -
forking bitrate and entropy control
Listed -
Human rights, democracy, and the rule of law assurance framework for AI systems: A proposal
Listed -
Science Facing Interoperability as a Necessary Condition of Success and Evil
Listed - Listed
- Listed
- Listed
-
Do mesa-optimization problems correlate with low-slack?
Listed -
Knowledge-Integrated Informed AI for National Security
Listed - Listed
-
The 6-Ds of Creating AI-Enabled Systems
Listed -
Certifying Out-of-Domain Generalization for Blackbox Functions
Listed - Listed
-
Investigating musical genius by listening to the Beach Boys a lot
Listed -
Observed patterns around major technological advancements
Listed -
QNR prospects are important for AI alignment research
Listed -
Reward is not enough: can we liberate AI from the reinforcement learning paradigm?
Listed -
Technology Ethics in Action: Critical and Interdisciplinary Perspectives
Listed -
The Met Dataset: Instance-level Recognition for Artworks
Listed -
[Intro to brain-like-AGI safety] 2. “Learning from scratch” in the brain
Listed - Listed
- Listed
- Listed
-
Impossibility results for unbounded utilities
Listed -
OpenAI Solves (Some) Formal Math Olympiad Problems
Listed -
Thoughts on AGI safety from the top
Listed -
VOS: Learning What You Don't Know by Virtual Outlier Synthesis
Listed -
CIC: Contrastive Intrinsic Control for Unsupervised Skill Discovery
Listed -
Interactive configurator with FO(.) and IDP-Z3
Listed - Listed
-
AMA: Future of Life Institute's EU Team
Listed -
Argument Against Impact: EU Is Not an AI Superpower
Listed -
Should you work in the European Union to do AGI governance?
Listed -
Explaining Reinforcement Learning Policies through Counterfactual Trajectories
Listed -
Certifying Model Accuracy under Distribution Shifts
Listed -
Towards Safe Reinforcement Learning with a Safety Editor Policy
Listed -
Arguments about Highly Reliable Agent Designs as a Useful Path to Artificial Intelligence Safety
Listed -
Causality, Transformative AI and alignment - part I
Listed -
Cost disease and civilizational decline
Listed -
Human-centered mechanism design with Democratic AI
Listed - Listed
-
[Intro to brain-like-AGI safety] 1. What's the problem & Why work on it now?
Listed -
Cybertrust: From Explainable to Actionable and Interpretable AI (AI2)
Listed -
ELK First Round Contest Winners
Listed - Listed
-
Reader reactions and update on "Where's Today's Beethoven"
Listed -
Safe AI -- How is this Possible?
Listed -
CSER is hiring for a senior research associate on longterm AI risk and governance
Listed -
Scaling Up Knowledge Graph Creation to Large and Heterogeneous Data Sources
Listed -
Alignment Problems All the Way Down
Listed -
Instrumental Convergence For Realistic Agent Objectives
Listed -
[AN #171]: Disagreements between alignment "optimists" and "pessimists"
Listed -
Avoiding Unsafe States in 3D Environments using Human Feedback
Listed