The catalog, page 21
Records 5,001 to 5,250 of 10,618. Explained and Verified records first, then newest. Every row carries the label that says how far the checking went.
-
Weighing the Milky Way and Andromeda with Artificial Intelligence
Listed -
Compute Research Questions and Metrics - Transformative AI and Compute [4/4]
Listed - Listed
-
The Terminology of Artificial Sentience
Listed -
Learning from learning machines: a new generation of AI technology to meet the needs of science
Listed -
Normative Disagreement as a Challenge for Cooperative AI
Listed -
AI and the Everything in the Whole Wide World Benchmark
Listed - Listed
-
larger language models may disappoint you [or, an eternally unfinished draft]
Listed -
Machines & Influence: An Information Systems Lens
Listed -
Sentience Institute 2021 End of Year Summary
Listed -
Christiano, Cotra, and Yudkowsky on AI progress
Listed -
Christiano, Cotra, and Yudkowsky on AI progress
Listed -
Christiano, Cotra, and Yudkowsky on AI progress
Listed -
[AN #169]: Collaborating with humans without human data
Listed -
Artificial Intelligence and the Problem of Control
Listed -
HIRING: Inform and shape a new project on AI safety at Partnership on AI
Listed -
HIRING: Inform and shape a new project on AI safety at Partnership on AI
Listed -
ReAct: Out-of-distribution Detection With Rectified Activations
Listed -
[linkpost] Acquisition of Chess Knowledge in AlphaZero
Listed -
AI Safety Needs Great Engineers
Listed -
AI Safety researcher career review
Listed -
AI Tracker: monitoring current and near-future risks from superscale models
Listed -
Integrating Three Models of (Human) Cognition
Listed - Listed
-
Slightly advanced decision theory 102: Four reasons not to be a (naive) utility maximizer
Listed -
What is most confusing to you about AI stuff?
Listed -
Branching Time Active Inference: empirical study and complexity class analysis
Listed -
Morally underdefined situations can be deadly
Listed -
Potential Alignment mental tool: Keeping track of the types
Listed -
Some real examples of gradient hacking
Listed -
Yudkowsky and Christiano discuss "Takeoff Speeds"
Listed -
Yudkowsky and Christiano discuss "Takeoff Speeds"
Listed -
Yudkowsky and Christiano discuss “Takeoff Speeds”
Listed -
From language to ethics by automated reasoning
Listed -
Genuineness, Existential Selfdetermination, Satisfaction: pick 2
Listed - Listed
-
A Certain Formalization of Corrigibility Is VNM-Incoherent
Listed -
Discrete Representations Strengthen Vision Transformer Robustness
Listed - Listed
-
More detailed proposal for measuring alignment of current models
Listed - Listed
-
rust & wasm, without wasm-pack
Listed -
unoptimal superintelligence loses
Listed - Listed
-
How To Get Into Independent Research On Alignment/Agency
Listed -
Ngo and Yudkowsky on AI capability gains
Listed -
Ngo and Yudkowsky on AI capability gains
Listed - Listed
-
Finding Useful Predictions by Meta-gradient Descent to Improve Decision-making
Listed -
Ngo and Yudkowsky on AI capability gains
Listed -
Satisficers Tend To Seek Power: Instrumental Convergence Via Retargetability
Listed -
Software Engineering for Responsible AI: An Empirical Study and Operationalised Patterns
Listed -
“Biological anchors” is about bounding, not pinpointing, AI timelines
Listed -
Acquisition of Chess Knowledge in AlphaZero
Listed -
Applications for AI Safety Camp 2022 Now Open!
Listed -
A positive case for how we might succeed at prosaic AI alignment
Listed -
Artificial Intelligence Needs Environmental Ethics | Global Catastrophic Risk Institute
Listed -
Falling everyday violence, bigger wars and atrocities: how do they net out?
Listed -
Improving Learning from Demonstrations by Learning from Experience
Listed -
Ngo and Yudkowsky on alignment difficulty
Listed -
Quantilizer ≡ Optimizer with a Bounded Amount of Output
Listed -
Solving Probability and Statistics Problems by Program Synthesis
Listed - Listed
-
Attempted Gears Analysis of AGI Intervention Discussion With Eliezer
Listed -
My understanding of the alignment problem
Listed -
Ngo and Yudkowsky on alignment difficulty
Listed -
Ngo and Yudkowsky on alignment difficulty
Listed -
"Slower tech development" can be about ordering, gradualness, or distance from now
Listed -
Artificial Intelligence Needs Environmental Ethics
Listed -
What would we do if alignment were futile?
Listed -
A FLI postdoctoral grant application: AI alignment via causal analysis and design of agents
Listed -
Comments on Carlsmith's “Is power-seeking AI an existential risk?”
Listed -
Is Functional Decision Theory still an active area of research?
Listed -
What’s the likelihood of only sub exponential growth for AGI?
Listed -
A Defense of Functional Decision Theory
Listed -
Human irrationality: both bad and good for reward inference
Listed - Listed
-
Why I'm excited about Redwood Research's current project
Listed -
Discussion with Eliezer Yudkowsky on AGI interventions
Listed -
Discussion with Eliezer Yudkowsky on AGI interventions
Listed -
Discussion with Eliezer Yudkowsky on AGI interventions
Listed -
Explainable AI (XAI): A Systematic Meta-Survey of Current Challenges and Future Opportunities
Listed -
Model-Free Risk-Sensitive Reinforcement Learning
Listed -
Reflections on the first year of parenting
Listed -
Weak point in “most important century”: lock-in
Listed -
BERI is hiring an ML Software Engineer
Listed -
What exactly is GPT-3's base objective?
Listed -
Building an AI-ready RSE Workforce
Listed -
Data Augmentation Can Improve Robustness
Listed -
Long-term AI policy strategy research and implementation
Listed -
Lymph Node Detection in T2 MRI with Transformers
Listed -
Possible research directions to improve the mechanistic explanation of neural networks
Listed -
Rowing, Steering, Anchoring, Equity, Mutiny
Listed - Listed
- Listed
-
Efficient estimates of optimal transport via low-dimensional embeddings
Listed -
How do we become confident in the safety of a machine learning system?
Listed -
psi: a universal format for structured information
Listed -
What are red flags for Neural Network suffering?
Listed -
A Word on Machine Ethics: A Response to Jiang et al. (2021)
Listed -
Using Brain-Computer Interfaces to get more data for AI alignment
Listed -
[Discussion] Best intuition pumps for AI safety
Listed - Listed
-
Linguistic Cues of Deception in a Multilingual April Fools' Day Context
Listed - Listed
-
Comments on OpenPhil's Interpretability RFP
Listed -
Drug addicts and deceptively aligned agents - a comparative analysis
Listed -
Modeling the impact of safety agendas
Listed -
Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models
Listed -
B-Pref: Benchmarking Preference-Based Reinforcement Learning
Listed - Listed
-
Apply to the ML for Alignment Bootcamp (MLAB) in Berkeley [Jan 3 - Jan 22]
Listed -
Apply to the ML for Alignment Bootcamp (MLAB) in Berkeley [Jan 3 - Jan 22]
Listed -
How to Improve China-Western Coordination on EA Issues?
Listed -
The case for long-term corporate governance of AI
Listed -
AI Ethics Statements -- Analysis and lessons learnt from NeurIPS Broader Impact Statements
Listed -
EfficientZero: human ALE sample-efficiency w/MuZero+self-supervised
Listed - Listed
-
Unraveling the evidence about violence among very early humans
Listed -
Artificial intelligence, systemic risks, and sustainability
Listed - Listed
- Listed
- Listed
-
Moral consideration of nonhumans in the ethics of artificial intelligence
Listed -
saving the server-side of the internet: just WASM,
Listed -
Apply to be a Stanford HAI Junior Fellow (Assistant Professor- Research) by Nov. 15, 2021
Listed -
Nate Soares on the Ultimate Newcomb's Problem
Listed -
A very crude deception eval is already passed
Listed - Listed
- Listed
- Listed
-
Measuring and forecasting risks
Listed -
Request for proposals for projects in AI alignment that work with deep learning systems
Listed -
Stuart Russell and Melanie Mitchell on Munk Debates
Listed -
Techniques for enhancing human feedback
Listed - Listed
-
[AN #168]: Four technical topics for which Open Phil is soliciting grant proposals
Listed -
Forecasting progress in language models
Listed -
Selfishness, preference falsification, and AI alignment
Listed -
Weak point in “most important century”: full automation
Listed -
Toward a Theory of Justice for Artificial Intelligence
Listed -
Understanding Interlocking Dynamics of Cooperative Rationalization
Listed -
Was life better in hunter-gatherer times?
Listed -
A Preliminary Exploration into Factored Cognition with Language Models
Listed -
QuantifyML: How Good is my Machine Learning Model?
Listed -
What Would Jiminy Cricket Do? Towards Agents That Behave Morally
Listed - Listed
-
Phil Trammell on Economic Growth Under Transformative AI
Listed - Listed
-
Towards Deconfusing Gradient Hacking
Listed -
Inference cost limits the impact of ever larger models
Listed -
alignment is an optimization processes problem
Listed -
AMA on Truthful AI: Owen Cotton-Barratt, Owain Evans & co-authors
Listed -
Epistemic Strategies of Safety-Capabilities Tradeoffs
Listed -
General alignment plus human values, or alignment via human values?
Listed -
Emergent modularity and safety
Listed -
Ethical Norms for New Generation Artificial Intelligence Released
Listed -
Podcast: Krister Bykvist on moral uncertainty, rationality, metaethics, AI and future populations
Listed -
to wasm and back again: the essence of portable programs
Listed -
[AN #167]: Concrete ML safety problems and their relevance to x-risk
Listed -
AGI Safety Fundamentals curriculum and application
Listed -
AGI Safety Fundamentals curriculum and application
Listed -
Reading books vs. engaging with them
Listed -
Shaking the foundations: delusions in sequence models for interaction and control
Listed - Listed
-
Pre-agriculture gender relations seem bad
Listed -
Risks of AI Foundation Models in Education
Listed -
[Creative Writing Contest] An AI Safety Limerick
Listed -
[MLSN #1]: ICLR Safety Paper Roundup
Listed -
An ML safety insurance company - shower thoughts
Listed -
Beyond the human training distribution: would the AI CEO create almost-illegal teddies?
Listed -
Epistemic Strategies of Selection Theorems
Listed -
MEMO: Test Time Robustness via Adaptation and Augmentation
Listed - Listed
-
New Working Paper Series of the Legal Priorities Project
Listed -
On The Risks of Emergent Behavior in Foundation Models
Listed -
Truthful AI: Developing and governing AI that does not lie
Listed -
Value alignment: a formal approach
Listed -
Improving End-To-End Modeling for Mispronunciation Detection with Effective Augmentation Mechanisms
Listed -
[Creative Writing Contest] Metal or Mortal
Listed -
Analyzing Dynamic Adversarial Training Data in the Limit
Listed -
Memetic hazards of AGI architecture posts
Listed -
Optimization Concepts in the Game of Life
Listed - Listed
-
Cold Links: assorted sports longreads
Listed -
Collaborating with Humans without Human Data
Listed - Listed
-
General vs specific arguments for the longtermist importance of shaping AI development
Listed -
NLP Position Paper: When Combatting Hype, Proceed with Caution
Listed -
Robustness of different loss functions and their impact on networks learning capability
Listed -
Can Machines Learn Morality? The Delphi Experiment
Listed -
Classical symbol grounding and causal graphs
Listed -
Compute Governance and Conclusions - Transformative AI and Compute [3/4]
Listed - Listed
-
[Creative Writing Contest] The Puppy Problem
Listed -
[Proposal] Method of locating useful subnets in large models
Listed - Listed
- Listed
-
Is it crunch time yet? If so, who can help?
Listed - Listed
-
Quantifying Local Specialization in Deep Neural Networks
Listed - Listed
-
EDT with updating double counts
Listed - Listed
-
Has life gotten better?: the post-industrial era
Listed -
Modeling Risks From Learned Optimization
Listed -
Certified Patch Robustness via Smoothed Vision Transformers
Listed -
Multiple Choice Normalization in LM Evaluation
Listed -
NVIDIA and Microsoft releases 530B parameter transformer model, Megatron-Turing NLG
Listed -
On Solving Problems Before They Appear: The Weird Epistemologies of Alignment
Listed - Listed
-
The evaluation function of an AI is not its aim
Listed - Listed
-
Why aren't you freaking out about OpenAI? At what point would you start?
Listed - Listed
-
Steelman arguments against the idea that AGI is inevitable and will arrive soon
Listed -
[AN #166]: Is it crazy to claim we're in the most important century?
Listed - Listed
-
Fingerprinting Multi-exit Deep Neural Network Models via Inference Time
Listed -
Safety-capabilities tradeoff dials are inevitable in AGI
Listed -
“Technological unemployment” AI vs. “most important century” AI: how far apart?
Listed -
[Job ad] Research important longtermist topics at Rethink Priorities!
Listed -
Automated Fact Checking: A Look at the Field
Listed -
Preferences from (real and hypothetical) psychology papers
Listed -
We're Redwood Research, we do applied alignment research, AMA
Listed -
Force neural nets to use models, then detect these
Listed - Listed
-
Procedure Planning in Instructional Videos via Contextual Modeling and Model-based Policy Learning
Listed -
Thinking Fast and Slow in AI: the Role of Metacognition
Listed -
Learning to Assist Agents by Observing Them
Listed -
Nuclear Espionage and AI Governance
Listed -
Nuclear Espionage and AI Governance
Listed -
A Framework of Prediction Technologies
Listed -
Occam's Razor and the Universal Prior
Listed -
The Dark Side of Cognition Hypothesis
Listed -
Why does (any particular) AI safety work reduce s-risks more than it increases them?
Listed -
A collection of AI Governance-related Podcasts, Newsletters, Blogs, and more
Listed -
Forecasting Compute - Transformative AI and Compute [2/4]
Listed - Listed
- Listed
-
Mapping the AI Investment Activities of Top Global Defense Companies
Listed -
Meta learning to gradient hack
Listed -
No Permits, No Fabs: The Importance of Regulatory Reform for Semiconductor Manufacturing
Listed -
Proposal: Scaling laws for RL generalization
Listed -
The Simulation Hypothesis Undercuts the SIA/Great Filter Doomsday Argument
Listed -
U.S. AI Workforce: Policy Recommendations
Listed -
What Selection Theorems Do We Expect/Want?
Listed -
AI learns betrayal and how to avoid it
Listed -
My take on Vanessa Kosoy's take on AGI safety
Listed