The catalog, page 14
Records 3,251 to 3,500 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
What Should AI Owe To Us? Accountable and Aligned AI Systems via Contractualist AI Alignment
Listed -
What Should AI Owe To Us? Accountable and Aligned AI Systems via Contractualist AI Alignment
Listed - Listed
- Listed
-
AI-assisted list of ten concrete alignment things to do right now
Listed -
Can "Reward Economics" solve AI Alignment?
Listed - Listed
- Listed
-
A New York Times article on AI risk
Listed - Listed
-
Alex Lawsen On Forecasting AI Progress
Listed -
Community Building for Graduate Students: A Targeted Approach
Listed -
ethics juice and anthropic juice
Listed - Listed
-
How can we secure more research positions at our universities for x-risk researchers?
Listed -
How Josiah became an AI safety researcher
Listed -
In conversation with AI: building better language models
Listed - Listed
-
A Game About AI Alignment (& Meta-Ethics): What Are the Must Haves?
Listed -
AI Governance Needs Technical Work
Listed -
An entire category of risks is undervalued by EA [Summary of previous forum post]
Listed - Listed
-
Do AI companies make their safety researchers sign a non-disparagement clause?
Listed -
Three scenarios of pseudo-alignment
Listed -
Breaking Newcomb's Problem with Non-Halting states
Listed -
Help me find a good Hackathon subject
Listed -
How To Know What the AI Knows - An ELK Distillation
Listed - Listed
-
The shard theory of human values
Listed -
An Update on Academia vs. Industry (one year into my faculty job)
Listed -
AXRP Episode 18 - Concept Extrapolation with Stuart Armstrong
Listed -
Behaviour Manifolds and the Hessian of the Total Loss - Notes and Criticism
Listed - Listed
-
Three scenarios of pseudo-alignment
Listed -
We may be able to see sharp left turns coming
Listed -
Levelling Up in AI Safety Research Engineering
Listed - Listed
- Listed
- Listed
- Listed
-
Sticky goals: a concrete experiment for understanding deceptive alignment
Listed -
Systemic Cascading Risks: Relevance in Longtermism & Value Lock-In
Listed -
We Can’t Do Long Term Utilitarian Calculations Until We Know if AIs Can Be Conscious or Not
Listed -
A Technique to Create Weaker Abstract Board Game Agents via Reinforcement Learning
Listed -
AI coordination needs clear wins
Listed -
AI Safety and Neighboring Communities: A Quick-Start Guide, as of Summer 2022
Listed -
Alignment is hard. Communicating that, might be harder
Listed - Listed
-
Gradient Hacker Design Principles From Biology
Listed -
I Tripped and Became GPT! (And How This Updated My Timelines)
Listed - Listed
-
My take on What We Owe the Future
Listed -
Reasons for my negative feelings towards the AI risk discussion
Listed -
Strategy For Conditioning Generative Models
Listed -
Values lock-in is already happening (without AGI)
Listed -
A Critique of AI Takeover Scenarios
Listed -
AI Box Experiment: Are people still interested?
Listed -
From motor control to embodied intelligence
Listed -
Survey of NLP Researchers: NLP is contributing to AGI progress; major catastrophe plausible
Listed -
The great energy descent (short version) - An important thing EA might have missed
Listed -
The great energy descent - Part 2: Limits to growth and why we probably won’t reach the stars
Listed -
Chaining the evil genie: why "outer" AI safety is probably easy
Listed -
Correct-by-Construction Runtime Enforcement in AI -- A Survey
Listed -
How likely is deceptive alignment?
Listed -
Inner Alignment via Superpowers
Listed - Listed
-
The Happiness Maximizer: Why EA is an x-risk
Listed -
Worlds Where Iterative Design Fails
Listed -
(My understanding of) What Everyone in Technical Alignment is Doing and Why
Listed -
*New* Canada AI Safety & Governance community
Listed -
Are Generative World Models a Mesa-Optimization Risk?
Listed -
How Do AI Timelines Affect Existential Risk?
Listed -
How might we align transformative AI if it’s developed very soon?
Listed -
How might we align transformative AI if it’s developed very soon?
Listed -
Preventing an AI-related catastrophe - Problem profile
Listed -
Reinforcement Learning for Hardware Security: Opportunities, Developments, and Challenges
Listed -
Breaking down the training/deployment dichotomy
Listed -
Who ordered alignment's apple?
Listed - Listed
-
Basin broadness depends on the size and number of orthogonal features
Listed -
Help Understanding Preferences And Evil
Listed -
Solving Alignment by "solving" semantics
Listed -
The History of AI Rights Research
Listed -
AGI Safety Fundamentals programme is contracting a low-code engineer
Listed -
AI Risk in Terms of Unstable Nuclear Software
Listed - Listed
- Listed
-
ARIA is looking for topics for roundtables
Listed -
DETERRENT: Detecting Trojans using Reinforcement Learning
Listed -
Seeking Student Submissions: Edit Your Source Code Contest
Listed -
Taking the parameters which seem to matter and rotating them until they don't
Listed -
A Test for Language Model Consciousness
Listed - Listed
-
Common misconceptions about OpenAI
Listed -
Discovering when an agent is present in a system
Listed -
Some conceptual alignment research projects
Listed -
The Shard Theory Alignment Scheme
Listed -
Who would you have on your dream team for solving AGI Alignment?
Listed - Listed
-
AI Safety For Dummies (Like Me)
Listed -
Beliefs and Disagreements about Automating Alignment Research
Listed -
Ethan Perez on the Inverse Scaling Prize, Language Feedback and Red Teaming
Listed -
Google AI integrates PaLM with robotics: SayCan update [Linkpost]
Listed -
Interspecies diplomacy as a potentially productive lens on AGI alignment
Listed -
Should I force myself to work on AGI alignment?
Listed - Listed
- Listed
-
What Makes A Good Measurement Device?
Listed -
AGI Timelines Are Mostly Not Strategically Relevant To Alignment
Listed -
AI alignment as “navigating the space of intelligent behaviour”
Listed -
First call for EA Data Science/ML/AI
Listed -
Philanthropists Probably Shouldn't Mission-Hedge AI Progress
Listed -
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Listed -
The Brussels Effect and Artificial Intelligence: How EU regulation will impact the global AI market
Listed -
Artificial Intelligence: A Modern Approach, Chapters 1-17
Listed -
Finding Goals in the World Model
Listed - Listed
-
What if we solve AI Safety but no one cares
Listed -
AXRP Episode 17 - Training for Very High Reliability with Daniel Ziegler
Listed -
My Plan to Build Aligned Superintelligence
Listed -
Pivotal acts using an unaligned AGI?
Listed -
Reward Reports for Reinforcement Learning
Listed -
Benchmarking Proposals on Risk Scenarios
Listed - Listed
- Listed
-
Less Threat-Dependent Bargaining Solutions?? (3/2)
Listed -
No One-Size-Fit-All Epistemic Strategy
Listed -
PreDCA: vanessa kosoy's alignment protocol
Listed -
Reducing Goodhart: Announcement, Executive Summary
Listed - Listed
-
What if we approach AI safety like a technical engineering safety problem
Listed -
BERI, Epoch, and FAR will explain their work & current job openings online this Sunday
Listed - Listed
- Listed
-
Epistemic Artefacts of (conceptual) AI alignment research
Listed -
How to do theoretical research, a personal perspective
Listed -
PreDCA: vanessa kosoy's alignment protocol
Listed - Listed
-
Announcing Encultured AI: Building a Video Game
Listed - Listed
-
Discovering when an agent is present in a system
Listed -
Intellectual Property Evaluation Utilizing Machine Learning
Listed -
An Exercise in Speed-Reading: The National Security Commission on AI (NSCAI) Final Report
Listed -
Autonomy as taking responsibility for reference maintenance
Listed -
Concrete Advice for Forming Inside Views on AI Safety
Listed -
Conditioning, Prompts, and Fine-Tuning
Listed -
Could realistic depictions of catastrophic AI risks effectively reduce said risks?
Listed - Listed
-
Human Mimicry Mainly Works When We’re Already Close
Listed -
Interpretability Tools Are an Attack Channel
Listed - Listed
-
Mesa-optimization for goals defined only within a training environment is dangerous
Listed -
The Core of the Alignment Problem is...
Listed - Listed
-
A concern about the “evolutionary anchor” of Ajeya Cotra’s report on AI timelines.
Listed -
A Review of the Convergence of 5G/6G Architecture and Deep Learning
Listed -
alignment research is very weird
Listed -
alignment researchspace is potentially malign
Listed - Listed
-
Deception as the optimal: mesa-optimizers and inner alignment
Listed -
Deception as the optimal: mesa-optimizers and inner alignment
Listed -
guiding your brain: go with your gut!
Listed -
Supplement to "The Brussels Effect and AI: How EU AI regulation will impact the global AI market"
Listed -
The Credibility of Apocalyptic Claims: A Critique of Techno-Futurism within Existential Risk
Listed -
What Makes an Idea Understandable? On Architecturally and Culturally Natural Ideas.
Listed -
A Mechanistic Interpretability Analysis of Grokking
Listed -
Seeking Interns/RAs for Mechanistic Interpretability Projects
Listed -
The Parable of the Boy Who Cried 5% Chance of Wolf
Listed -
What's General-Purpose Search, And Why Might We Expect To See It In Trained ML Systems?
Listed -
A brief note on Simplicity Bias
Listed -
A general framework for reward function distances.
Listed -
A neural network ensemble with feature engineering for improved credit card fraud detection.
Listed -
A Penalty Default Approach to Preemptive Harm Disclosure and Mitigation for AI Systems.
Listed -
A Primer on Maximum Causal Entropy Inverse Reinforcement Learning.
Listed - Listed
- Listed
-
Active uncertainty learning for human-robot interaction: An implicit dual control approach.
Listed -
Active uncertainty reduction for human-robot interaction: An implicit dual control approach.
Listed -
AdaCat: Adaptive Categorical Discretization for Autoregressive Models.
Listed -
Adversarial Motion Priors Make Good Substitutes for Complex Reward Functions.
Listed -
Adversarial Policies Beat Professional-Level Go AIs.
Listed -
AI experts are increasingly afraid of what they’re creating.
Listed -
All the posts I will never write
Listed -
An empirical investigation of representation learning for imitation.
Listed -
An interpretable machine learning approach for hepatitis b diagnosis.
Listed -
APReL: A Library for Active Preference-based Reward Learning Algorithms.
Listed -
Are we living in an AGI World?.
Listed -
ASHA: Assistive Teleoperation via Human-in-the-Loop Reinforcement Learning. .
Listed -
Assistive Teaching of Motor Control Tasks to Humans.
Listed -
Automatic Correction of Human Translations.
Listed -
Autoregressive Latent Video Prediction with High-Fidelity Image Generator.
Listed -
Autoregressive Uncertainty Modeling for 3D Bounding Box Prediction.
Listed -
Back to the Future: Efficient, Time-Consistent Solutions in Reach-Avoid Games.
Listed -
Balancing Efficiency and Comfort in Robot-Assisted Bite Transfer.
Listed -
Banning Lethal Autonomous Weapons: An Education.
Listed - Listed
-
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.
Listed -
Brain computation as fast spiking neural Monte Carlo inference in probabilistic programs.
Listed -
Brain-like AGI project "aintelope"
Listed -
Building Human Values into Recommender Systems: An Interdisciplinary Synthesis.
Listed -
Calculus on MDPs: Potential Shaping as a Gradient.
Listed -
Chain of Thought Imitation with Procedure Cloning.
Listed -
Choices, Risks, and Reward Reports: Charting Public Policy for Reinforcement Learning Systems.
Listed -
COLA: Consistent Learning with Opponent-Learning Awareness.
Listed -
Conditional Imitation Learning for Multi-Agent Games.
Listed - Listed
-
Cooperative Multi-Agent Fairness and Equivariant Policies.
Listed -
COVID-19 diagnosis: a review of rapid antigen, RT-PCR and artificial intelligence methods.
Listed -
Cross-Domain Imitation Learning via Optimal Transport.
Listed -
DayDreamer: World Models for Physical Robot Learning.
Listed -
Deciding to be authentic: Intuition is favored over deliberation when authenticity matters.
Listed -
Democratic Control of Recommender Systems.
Listed -
Demography of Machine Learning Education Within the K12.
Listed -
Diagnostics for Deep Neural Networks with Automated Copy/Paste Attacks.
Listed -
Differential Assessment of Black-Box AI Agents..
Listed -
Director: Deep Hierarchical Planning from Pixels.
Listed -
Discovered policy optimisation.
Listed -
Disentangling Abstraction from Statistical Pattern Matching in Human and Machine Learning..
Listed -
Eliciting Compatible Demonstrations for Multi-Human Imitation Learning.
Listed -
Enhanced prediction of chronic kidney disease using feature selection and boosted classifiers.
Listed -
essential inequality vs functional inequivalence
Listed -
Estimating and Penalizing Induced Preference Shifts in Recommender Systems.
Listed - Listed
-
Evaluations of Causal Claims Reflect a Trade-Off Between Informativeness and Compression.
Listed -
Experiments on causal exclusion.
Listed -
Explaining Reinforcement Learning Policies through Counterfactual Trajectories.
Listed - Listed
-
Exploiting Extensive-Form Structure in Empirical Game-Theoretic Analysis.
Listed -
Few-Shot Preference Learning for Human-in-the-Loop RL.
Listed -
First Contact: Unsupervised Human-Machine Co-Adaptation via Mutual Information Maximization. .
Listed -
Fleet-DAgger: Interactive Robot Fleet Learning with Scalable Human Supervision.
Listed -
For learning in symmetric teams, local optima are global nash equilibria.
Listed -
Get It in Writing: Formal Contracts Mitigate Social Dilemmas in Multi-Agent RL.
Listed -
Goal Misgeneralization: Why Correct Specifications Aren’t Enough For Correct Goals.
Listed - Listed
-
Guided imitation of task and motion planning.
Listed -
Haptic perception using optoelectronic robotic flesh for embodied artificially intelligent agents.
Listed -
Heterogeneous-agent mirror learning: A continuum of solutions to cooperative marl.
Listed -
Hidden Gold for IT Professionals, Educators, and Students: Insights From Stack Overflow Survey.
Listed -
Hierarchical Few-Shot Imitation with Skill Transition Models.
Listed -
How do people incorporate advice from artificial agents when making physical judgments?.
Listed -
How to talk so AI will learn: Instructions, descriptions, and autonomy.
Listed -
How to talk so your robot will learn: Instructions, descriptions, and pragmatics.
Listed -
How Would The Viewer Feel? Estimating Wellbeing From Video Scenarios.
Listed -
How “is” shapes “ought” for folk-biological concepts.
Listed -
Human-Centered Evaluation of Explanations.
Listed - Listed
-
Imitation Learning by Estimating Expertise of Demonstrators.
Listed -
imitation: Clean Imitation Learning Implementations.
Listed -
Inferring Rewards from Language in Context.
Listed