The catalog, page 28
Records 6,751 to 7,000 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Sparse Skill Coding: Learning Behavioral Hierarchies with Sparse Codes.
Listed -
Special issue on autonomous agents modelling other agents: Guest editorial.
Listed -
SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards..
Listed -
What Can Learned Intrinsic Rewards Capture?.
Listed -
What the Baldwin Effect affects depends on the nature of plasticity.
Listed -
10/50/90% chance of GPT-N Transformative AI?
Listed -
Non-Adversarial Imitation Learning and its Connections to Adversarial Methods
Listed -
The Fusion Power Generator Scenario
Listed -
Towards a Formalisation of Logical Counterfactuals
Listed -
Impact of meta-roles on the evolution of organisational institutions
Listed -
Analyzing the Problem GPT-3 is Trying to Solve
Listed -
Decoupling Exploration and Exploitation for Meta-Reinforcement Learning without Sacrifices
Listed - Listed
-
[AN #111]: The Circuits hypotheses for deep learning
Listed - Listed
-
Collecting the Public Perception of AI and Robot Rights
Listed -
Forecasting AI Progress: A Research Agenda
Listed -
Infinite Data/Compute Arguments in Alignment
Listed -
Interpretability in ML: A Broad Overview
Listed -
(Ir)rationality of Pascal's wager
Listed -
AI Risk: Increasing Persuasion Power
Listed -
Is GPT-3 the death of the paperclip maximizer?
Listed -
Three mental images from thinking about AGI debate & corrigibility
Listed -
What are the most important papers/post/resources to read to understand more of GPT-3?
Listed -
What do we do if AI doesn't take over the world, but still causes a significant global problem?
Listed -
Inner Alignment: Explain like I'm 12 Edition
Listed -
Power as Easily Exploitable Opportunities
Listed -
Testing the Automation Revolution Hypothesis
Listed -
"Go west, young man!" - Preferences in (imperfect) maps
Listed -
On Single Point Forecasts for Fat-Tailed Variables
Listed -
Is the work on AI alignment relevant to GPT?
Listed -
The academic contribution to AI safety seems large
Listed -
What if memes are common in highly capable minds?
Listed -
[AN #110]: Learning features from human feedback to enable reward learning
Listed -
Engaging Seriously with Short Timelines
Listed -
Learning the prior and generalization
Listed -
Rohin Shah: What’s been happening in AI alignment?
Listed -
The "best predictor is malicious optimiser" problem
Listed -
What Failure Looks Like: Distilling the Discussion
Listed -
Does the lottery ticket hypothesis suggest the scaling hypothesis?
Listed - Listed
-
Probability that other architectures will scale as well as Transformers?
Listed -
To what extent are the scaling properties of Transformer networks exceptional?
Listed - Listed
-
What specific dangers arise when asking GPT-N to write an Alignment Forum post?
Listed - Listed
- Listed
-
Combining Deep Reinforcement Learning and Search for Imperfect-Information Games
Listed -
Generalizing the Power-Seeking Theorems
Listed - Listed
-
Automated Database Indexing using Model-free Reinforcement Learning
Listed -
Constraints from naturalized ethics.
Listed -
Markus Anderljung and Ben Garfinkel: Fireside chat on AI governance
Listed -
Bridging the Imitation Gap by Adaptive Insubordination
Listed -
Can you get AGI from a Transformer?
Listed -
Improving Competence for Reliable Autonomy
Listed - Listed
-
Scrutinizing AI Risk (80K, #81) - v. quick summary
Listed -
Toward Campus Mail Delivery Using BDI
Listed -
Why the Orthogonality Thesis's veracity is not the point:
Listed -
[AN #109]: Teaching neural nets to generalize the way humans would
Listed -
Intellectual Diversity in AI Safety
Listed - Listed
-
$1000 bounty for OpenAI to show whether GPT3 was "deliberately" pretending to be stupider than it is
Listed -
[Preprint] The Computational Limits of Deep Learning
Listed -
AI Benefits Post 5: Outstanding Questions on Governing Benefits
Listed -
Alignment As A Bottleneck To Usefulness Of GPT-3
Listed -
Competition: Amplify Rohin’s Prediction on AGI researchers & Safety Concerns
Listed -
How strong is the evidence of unaligned AI systems causing harm?
Listed -
Artificial Intelligence is stupid and causal reasoning won't fix it
Listed - Listed
-
Parallels Between AI Safety by Debate and Evidence Law
Listed -
Parallels Between AI Safety by Debate and Evidence Law
Listed -
To what extent is GPT-3 capable of reasoning?
Listed -
What Would I Do? Self-prediction in Simple Algorithms
Listed - Listed
- Listed
-
Modulation of viability signals for self-regulatory control
Listed -
Why is pseudo-alignment "worse" than other ways ML can fail to generalize?
Listed -
Environments as a bottleneck in AGI development
Listed - Listed
-
Technologies for Trustworthy Machine Learning: A Survey in a Socio-Technical Context
Listed -
[AN #107]: The convergent instrumental subgoals of goal-directed agents
Listed -
[AN #108]: Why we should scrutinize arguments for AI risk
Listed -
A list of good heuristics that the case for AI X-risk fails
Listed -
Alignment proposals and complexity classes
Listed -
Artificial Interdisciplinarity: Artificial Intelligence for Research on Complex Societal Problems
Listed -
LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning
Listed -
Failures of Contingent Thinking
Listed -
How should AI debate be judged?
Listed -
New paper: AGI Agent Safety by Iteratively Improving the Utility Function
Listed -
AI Benefits Post 4: Outstanding Questions on Selecting Benefits
Listed -
The Goldbach conjecture is probably correct; so was Fermat's last theorem
Listed -
What are the mostly likely ways AGI will emerge?
Listed -
3-P Group optimal for discussion?
Listed -
AMA or discuss my 80K podcast episode: Ben Garfinkel, FHI researcher
Listed - Listed
-
Does generality pay? GPT-3 can provide preliminary evidence.
Listed - Listed
-
Meta Programming GPT: A route to Superintelligence?
Listed -
A space of proposals for building safe advanced AI
Listed -
Machine Learning Explainability for External Stakeholders
Listed -
Mesa-Optimizers vs “Steered Optimizers”
Listed -
Talk: Key Issues In Near-Term AI Safety Research
Listed -
AI Research Considerations for Human Existential Safety (ARCHES)
Listed -
Arguments against myopic training
Listed -
Why is the impact penalty time-inconsistent?
Listed -
Decolonial AI: Decolonial Theory as Sociotechnical Foresight in Artificial Intelligence
Listed - Listed
- Listed
-
Mahendra Prasad: Rational group decision-making
Listed -
Sunday July 12 — talks by Scott Garrabrant, Alexflint, alexei, Stuart_Armstrong
Listed -
What does it mean to apply decision theory?
Listed -
Antitrust-Compliant AI Industry Self-Regulation
Listed -
Dynamic inconsistency of the inaction and initial state baseline
Listed - Listed
- Listed
-
Reducing long-term risks from malevolent actors
Listed -
Robust Learning with Frequency Domain Regularization
Listed -
AI Benefits Post 3: Direct and Indirect Approaches to AI Benefits
Listed -
Better priors as a safety problem
Listed -
Better priors as a safety problem
Listed -
Decentralized Reinforcement Learning: Global Decision-Making via Local Economic Transactions
Listed - Listed
- Listed
-
Tradeoff between desirable properties for baseline choices in impact measures
Listed -
Customized Handling of Unintended Interface Operation in Assistive Robots
Listed -
Tradeoff between desirable properties for baseline choices in impact measures
Listed -
AI Unsafety via Non-Zero-Sum Debate
Listed -
Research ideas to study humans with AI Safety in mind
Listed -
Splitting Debate up into Two Subsystems
Listed - Listed
- Listed
-
Verifiably Safe Exploration for End-to-End Reinforcement Learning
Listed -
[AN #106]: Evaluating generalization ability of learned reward models
Listed -
Artificial intelligence in a crisis needs ethics with urgency
Listed - Listed
-
Evan Hubinger on Inner Alignment, Outer Alignment, and Proposals for Building Safe Advanced AI
Listed - Listed
-
Unifying Model Explainability and Robustness via Machine-Checkable Concepts
Listed -
Comparing AI Alignment Approaches to Minimize False Positive Risk
Listed -
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Listed - Listed
-
AI Benefits Post 2: How AI Benefits Differs from AI Alignment & AI for Good
Listed -
How do takeoff speeds affect the probability of bad outcomes from AGI?
Listed -
Gary Marcus vs Cortical Uniformity
Listed -
Have general decomposers been formalized?
Listed -
Song Pairs that can be listened to together
Listed - Listed
-
AvE: Assistance via Empowerment
Listed -
Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance
Listed -
Is SGD a Bayesian sampler? Well, almost
Listed -
Radical Probabilism [Transcript]
Listed -
Some promising career ideas beyond 80,000 Hours' priority paths
Listed - Listed
-
“Explaining” machine learning reveals policy challenges
Listed -
AI Governance Reading Group Guide
Listed - Listed
-
[AN #105]: The economic trajectory of humanity, and what we might mean by optimization
Listed -
Abstraction, Evolution and Gears
Listed -
Compositional Explanations of Neurons
Listed -
Models, myths, dreams, and Cheshire cat grins
Listed -
Quantifying Differences in Reward Functions
Listed -
RL Unplugged: Benchmarks for Offline Reinforcement Learning
Listed - Listed
-
Adversarial Soft Advantage Fitting: Imitation Learning without Policy Optimization
Listed - Listed
-
AI Benefits Post 1: Introducing “AI Benefits”
Listed - Listed
-
Plausible cases for HRAD work, and locating the crux in the "realism about rationality" debate
Listed -
Safe Reinforcement Learning via Curriculum Induction
Listed - Listed
-
The flaws that make today's AI architecture unsafe and a new approach that could fix it
Listed - Listed
-
Will AGI cause mass technological unemployment?
Listed -
Relevant pre-AGI possibilities
Listed -
Relevant pre-AGI possibilities
Listed - Listed
-
Relevant pre-AGI possibilities
Listed - Listed
-
IReEn: Reverse-Engineering of Black-Box Functions via Iterative Neural Program Synthesis
Listed -
Big Self-Supervised Models are Strong Semi-Supervised Learners
Listed - Listed
-
Our take on CHAI’s research agenda in under 1500 words
Listed -
Results of $1,000 Oracle contest!
Listed -
Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
Listed -
Relating HCH and Logical Induction
Listed -
What are the high-level approaches to AI alignment?
Listed -
Causality Adds Up to Normality
Listed -
dm_control: Software and Tasks for Continuous Control
Listed -
Formal Verification of End-to-End Learning in Cyber-Physical Systems: Progress and Challenges
Listed - Listed
-
Pessimism About Unknown Unknowns Inspires Conservatism
Listed - Listed
- Listed
- Listed
-
Ethical Considerations for AI Researchers
Listed -
Online Bayesian Goal Inference for Boundedly-Rational Planning Agents
Listed -
Preparing for "The Talk" with AI projects
Listed -
Open Questions in Creating Safe Open-ended AI: Tensions Between Control and Creativity
Listed -
SAMBA: Safe Model-Based & Active Reinforcement Learning
Listed -
Cartesian Boundary as Abstraction Boundary
Listed -
Multi-Agent Informational Learning Processes
Listed -
[AN #103]: ARCHES: an agenda for existential safety, and combining natural language with deep RL
Listed -
What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
Listed - Listed
- Listed
-
More on disambiguating "discontinuity"
Listed -
Public Static: What is Abstraction?
Listed -
Goal-directedness is behavioral, not structural
Listed -
Learning to Play No-Press Diplomacy with Best Response Policy Iteration
Listed -
Reinforcement Learning Under Moral Uncertainty
Listed -
Curiosity Killed or Incapacitated the Cat and the Asymptotically Optimal Agent
Listed -
Reply to Paul Christiano on Inaccessible Information
Listed -
[AN #102]: Meta learning by GPT-3, and a list of full proposals for AI alignment
Listed -
Focus: you are allowed to be bad at accomplishing your goals
Listed - Listed
- Listed
-
MultiXNet: Multiclass Multistage Multimodal Motion Prediction
Listed -
AI Definitions Affect Policymaking
Listed -
Aligning Superhuman AI with Human Behavior: Chess as a Model System
Listed -
Building brain-inspired AGI is infinitely easier than understanding the brain
Listed -
Acme: A new framework for distributed reinforcement learning
Listed -
Assessing the Risks Posed by the Convergence of Artificial Intelligence and Biotechnology
Listed -
How to Be Helpful to Multiple People at Once
Listed -
Medium-Term Artificial Intelligence and Society
Listed -
Recordings from AI Safety Discussion Days
Listed -
Shaping the Terrain of AI Competition
Listed -
Sparsity and interpretability?
Listed - Listed
-
Possible takeaways from the coronavirus pandemic for slow AI takeoff
Listed -
Possible takeaways from the coronavirus pandemic for slow AI takeoff
Listed -
AI Research Considerations for Human Existential Safety (ARCHES)
Listed - Listed
-
An overview of 11 proposals for building safe advanced AI
Listed - Listed
- Listed
-
Language Models are Few-Shot Learners
Listed -
[AN #101]: Why we should rigorously measure and forecast AI progress
Listed -
AI Forensics: Did the Artificial Intelligence System Do It? Why?
Listed - Listed
- Listed
-
How can Interpretability help Alignment?
Listed - Listed
-
Danny Hernandez on forecasting and the drivers of AI progress
Listed -
From ImageNet to Image Classification: Contextualizing Progress on Benchmarks
Listed -
AI Research Considerations for Human Existential Safety (ARCHES)
Listed -
Comparing reward learning/reward tampering formalisms
Listed -
[AN #100]: What might go wrong if you learn a reward function while acting
Listed -
Probabilities, weights, sums: pretty much the same for reward functions
Listed