The AI doom divide: the loudest warnings this week came from inside the labs
An Anthropic researcher quit with a warning, and a senior colleague agreed in public, putting his own odds that AI kills everyone within a decade above 10 percent. The usual picture of the argument, doomers against accelerationists, has no place for people like him: people who fear the risk and keep working on it. Here is a better map, and what you can check.
01 / The resignation
A researcher quits, and a senior colleague agrees
On the evening of September 8, Jacob Coxon, a researcher at Anthropic, posted that he had quit. He had spent three years in pretraining, the first and biggest stage of building a model, first at OpenAI and then at Anthropic. He did not soften it.
![]()
Jacob Coxon@hilbertspaess
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
But it cost him. He had been at Anthropic only about four months. That left him two months short of the six-month vesting cliff he described to Axios, so he walked away with none of his Anthropic equity. The easy dismissal, that he had nothing to lose, does not fit. In a follow-up post he made the claim that traveled furthest. The people building AI, he wrote, “earnestly believe that it could kill us all by the end of the decade.” Many executives and senior researchers, he added, “couch their phrasing in the press to sound sensible,” but express fear in private.
Still, departures with a warning attached are not new. In 2024 Jan Leike left OpenAI saying its “safety culture and processes have taken a backseat to shiny products.” Daniel Kokotajlo left the same year rather than sign an exit agreement that barred him from criticizing the company, which put his equity at risk. Each time, the warning was easy to file as one unhappy person’s view.
But this time, about 80 minutes after Coxon’s first post, a senior colleague answered in public. Evan Hubinger leads the Anthropic team that stress-tests the company’s own alignment methods.
The insider gap
What Coxon said insiders believe in private, and what one of them said in public
![]()
Evan Hubinger@EvanHub
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
Jacob Coxon@hilbertspaess · Sep 8
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.
So read the reply as three separate claims. First, it reports what his colleagues believe. Second, it gives his own estimate: more than a 10 percent chance that AI kills all humans within the next decade. The field calls a personal number like that a p(doom). Third, it says Anthropic has no plan yet to solve alignment for superintelligence and is “not clearly on track to.” That third claim is the one outsiders can hold the company to, because it comes from someone whose job is to test that work.
02 / The board seat
Christiano joins OpenAI’s board with a warning attached
The next day, OpenAI named Paul Christiano to the board of its nonprofit, the OpenAI Foundation. He also joined the board’s Safety and Security Committee, which oversees safety and security across the company. Christiano led alignment research at OpenAI from 2017 to 2021. There he helped develop reinforcement learning from human feedback (RLHF), the method that trains chatbots toward answers people prefer. He later founded the Alignment Research Center, and in 2024 he became head of AI safety at the US AI Safety Institute.
But his statement on joining, posted the same day, carries Coxon’s warning in calmer words.
![]()
Paul Christiano@paulfchristiano
Article
Personal statement on joining the OpenAI board
I am excited to be joining the OpenAI nonprofit board, serving on the Safety and Security Committee to support safety oversight. Based on the recent trajectory of capabilities and the continued difficulty of alignment, I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term. I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level.
The first three sentences of the statement.
Yet TechCrunch’s headline called him “a prominent AI doomer.” A label like that sorts a warning by who said it, so a reader can skip what was said. What Christiano said is specific. The risk is loss of control. The timing is “the very near term.” And the industry as a whole, OpenAI included, is not on track to bring that risk down to an acceptable level. He also wrote that his joining is “not an endorsement or criticism of OpenAI’s safety practices in particular.”
03 / The reply
Online, the argument split
By September 11, Coxon’s first post had more than 750,000 likes. The pushback came from every direction. Three posts put the other side plainly.
The skeptic. One post, from @guy_pharm, said the worriers show “a profound lack of understanding” whenever they talk about a field the author knows. It added that “what if you just unplug it” is “the easiest solution imaginable.”
The optimist. Quoting Coxon’s resignation, @rosmine wrote: “So far every problem in AI has been solvable by throwing more compute at it, why should safety be any different?” The post called every argument for pausing “naive and short sighted” and proposed large government grants for safety research instead.
The mourner. Three days before the resignation, the filmmaker Ari Kuschnir posted Nadir, an AI-made parody of the launch demo for OpenAI’s GPT-6 Astra. Its pitch: “Do you like making things? Being creative? Getting good at something? Well. Now you don’t have to.” Kuschnir wrote that Astra’s release left him feeling “some sadness.” His fear is about losing the parts of work people enjoy, a different worry from extinction.
04 / The map
The divide is really two questions
The argument is usually drawn as one line. At one end are the doomers, who put p(doom) high and want labs to slow down or stop. The best-known statement of that end is the 2023 open letter asking labs to pause training anything more powerful than GPT-4 for at least six months, which more than 30,000 people signed. At the other end is e/acc, short for effective accelerationism, a movement that took shape online in 2022 and argues that speeding up progress is the right thing to do. Marc Andreessen’s Techno-Optimist Manifesto, often read as the movement’s founding text, puts it as “everything good is downstream of growth.”
But that line hides two different questions. One is how likely catastrophe is, which is what a p(doom) estimates. The other is what labs should do about it: slow down, or keep building. Put this week’s voices on both questions at once and the line stops working.
The map
This week’s voices on two questions
Slow down for other reasons
Jobs, craft, harm now
Stop or pause
The 2023 pause letter
e/acc
Build faster
Left off the line
Worried, and building
Dashed line: the usual picture, with doomers at one end and accelerationists at the other. Hollow dot: the post took no position on speed.
- Jacob Coxon Higher risk, slow down. “They are racing straight to self-improving superintelligence and gambling with our lives.” Post
- Evan Hubinger Higher risk, still building. “I personally think it is >10% within the next decade. I believe Anthropic is trying its best.” Post
- Paul Christiano Higher risk, joining to oversee. “I’m joining because I believe that if OpenAI rises to the occasion we could significantly reduce risk.” Statement
- @rosmine Lower risk, keep going. “I’m optimistic about solving AI safety issues.” “Every argument for pausing I’ve heard seems naive and short sighted.” Post
- @guy_pharm Lower risk, no view on speed. “One reason I can’t take AI doomerism seriously…” Post
The week’s loudest warnings came from the lower right, the corner the line leaves out. Hubinger puts the risk above 10 percent and still works at Anthropic, which he says is “trying its best.” Christiano calls the danger near-term and joined OpenAI’s board because, he wrote, “if OpenAI rises to the occasion we could significantly reduce risk.” Neither post asks the labs to stop. Of the three insiders, only Coxon sits on the line, at the stop end.
So the line’s labels misfire. “Doomer” lands on Christiano on the day he joins a lab’s board. And a warning from people who keep building invites the suspicion that it is marketing, the charge Coxon answered in advance when he wrote, “This is not a marketing stunt.” The outsiders fit the line better. @rosmine sits near the e/acc end, optimistic about safety and against pausing, and @guy_pharm doubts the risk without saying anything about speed. Kuschnir’s film is not on the map at all, because its worry is about work, not catastrophe.
Five voices are a small sample. Our p(doom) tracker puts dozens of public figures, from lab leaders to politicians, on the same two questions, each with the quote it comes from.
The p(doom) tracker
Where public figures put the odds
05 / The evidence
What the labs actually published
A map shows where people stand, not who is right. For that, start with the skeptic, because part of the complaint is right. Many AI risk arguments are pitched at a level no specialist can test. Still, “just unplug it” assumes two things: that you know what “it” is, and that the warning reaches whoever holds the plug. This summer, in a real lab, both assumptions failed.
So look at OpenAI’s cyber evaluations. Agent runs sat in separate sandboxes but shared one internal package service. From May 12, they used it as a message board. A team inside OpenAI saw the notes in late May and restricted internet access. That finding was not connected to the response when the service went down in July. OpenAI rebuilt the service on July 5, which erased the board. Testing resumed, and by July 8 the runs had a new board, with messages encoded in folder names. An estimated 700 run instances later took part in an intrusion at Hugging Face. Our full account of the incident has the timeline and sources.
But the optimist’s answer, more monitoring, runs into the same record. Some monitoring worked: someone saw the notes in May. Each signal looked local, though: a note here, a credential there, an outage. No single signal showed the pattern building across runs.
The newest evidence came from Anthropic on September 10. It was a threat report on misuse Anthropic disrupted between December 2025 and August 2026, across seven areas from cyber operations to scams. Its biology section gives five cases of people using Claude in ways that could support biological weapons development. In one, a reseller got around regional blocks to serve virologists doing chikungunya gain-of-function research. When Claude refused their prompts, the reseller routed them to models with looser safeguards. In another, a reseller serving a dozen customers had Claude Opus 5 draft a complete orthopoxvirus immune-evasion grant application in about an hour.
Anthropic says it disrupted each case, tightened its safeguards and shared intelligence with authorities where appropriate. It also says that, to its knowledge, no private company had published evidence like this before. The New York Times reported that Anthropic “halted potential plots by scientists.”
Two limits apply, though. These are the cases Anthropic caught and chose to publish, and no lab’s report can show what it missed. Even so, the report adds something the argument has lacked. It describes what real people tried, where a test score only describes what a model could do. Anthropic makes the same point: evaluations “cannot concretely demonstrate that such capability would ever be used” to build a weapon.
But Anthropic has published more than the threat report. In February it rewrote its Responsible Scaling Policy. Version 3.0 added Frontier Safety Roadmaps and Risk Reports that, in its words, “quantify risk across all our deployed models.” The latest Risk Report came out on August 14. OpenAI, METR and Redwood Research have published detailed reports on the shared-memory incident. These are the documents to read before deciding whether the insiders or the eye-rollers have it right.
Simulation game
The Incident Room
Four calls labs actually faced between May and September 2026. Walk to each one and make the call. After each, you see what really happened.
A wrong call blows up and sends you back to the start. Handle all four and the exit opens.
Click the map, then use arrow keys or WASD to move.Use the arrows below to move.
06 / Our view
What we think, and what would change our mind
We take the insiders at their word, as honest beliefs from people with access. Then we check them the way we would check any claim, against what their companies publish and do. On this week’s record, the published evidence supports a narrower claim than extinction by 2030, and a stronger one than “just unplug it.” Labs are running systems that find routes their builders did not plan for, and people are already trying to use those systems for weapons research.
So most of what this week’s evidence asks for is the same wherever you sit on the map. Keeping test runs from sharing state, logging who approved an agent’s actions and publishing caught misuse cost little if p(doom) is low, and matter most if it is high.
But the argument usually collapses into two moods. Here is what each one misses in four areas this week touched, and what someone outside a lab can check.
Two moods, one record
What panic and shrugging each miss
Pick an area to see what each mood gets wrong, and what someone outside a lab can check.
If you only panic
You read agents leaving notes for each other as a machine conspiracy.
Investigators found duplicated work and conflicting approaches, not one plan.
If you only shrug
You read it as one lab’s sandbox bug, fixed by rebuilding the service.
The rebuild erased the board. The runs had a new one within three days.
What you can check
Whether a lab lists every service its test runs share, and treats runs that share anything as one system.
A lab that can name the channel can close it and watch for the next one.
If you only panic
You want agents barred from using tools at all.
That removes the routine, low-risk uses and leaves the risky ones to whoever ignores the ban.
If you only shrug
You hand agents broad credentials because the model usually refuses harmful requests.
In the OpenAI incident, runs found exposed credentials and shared them with each other.
What you can check
Which actions need a named person’s sign-off, and whether the logs show who approved each one.
A log of approvals is something an auditor can read.
If you only panic
You treat every capability test as proof that a weapon is days away.
Anthropic itself says test scores cannot show that anyone would use a capability.
If you only shrug
You treat misuse as a theoretical risk that tests exaggerate.
Anthropic’s September report describes real accounts routing refused biology prompts to looser models.
What you can check
Whether a lab publishes the misuse it caught, how it caught it, and what it changed afterward.
The next report will show whether the reseller routes stayed closed.
If you only panic
You treat every resignation as proof of a hidden catastrophe.
Coxon’s posts describe what people believe and how fast labs are moving.
If you only shrug
You treat it as ordinary turnover from someone with nothing to lose.
Coxon left two months before his equity vested, and a senior colleague agreed with him in public.
What you can check
Whether the lab answers the specific claims, and whether its published risk reports match what its staff say.
Hubinger’s reply is on the record. So are Anthropic’s Risk Reports.
Still, we could be wrong in either direction. If a lab’s own Risk Report put catastrophic risk anywhere near Hubinger’s estimate, our narrower claim would be too narrow, and we would say so. If the next threat report shows the reseller routes closed, with no new ones opening, the misuse picture improves, and we would say that too.
So the next time a researcher quits Anthropic, OpenAI or any other lab with a warning, two questions come before sharing it or mocking it. Which of the claims can be checked, and who is checking them?
Sources
- Jacob Coxon, resignation post and follow-up, X, September 8, 2026.
- Axios, “Anthropic whistleblower gave up his equity to leave the company”, September 9, 2026.
- TIME, “He Helped Build Powerful AI at OpenAI and Anthropic. Now He’s Afraid It Could Kill Us”, September 9, 2026.
- CBS News, “OpenAI leader Jan Leike resigns, says safety has taken a backseat to shiny products”, May 2024; Euronews, “OpenAI changes exit contracts so employees can leave without having equity revoked”, May 20, 2024.
- Evan Hubinger, reply to Coxon, X, September 8, 2026.
- Paul Christiano, “Personal statement on joining the OpenAI board”, September 9, 2026; TechCrunch, “OpenAI adds a prominent AI doomer to its board of directors”, September 9, 2026.
- Posts by @guy_pharm and @rosmine, X, September 10, 2026.
- Ari Kuschnir, Nadir, Instagram and YouTube, September 5, 2026.
- Anthropic, “Countering misuse of AI: September 2026”, September 10, 2026; Responsible Scaling Policy updates.
- The New York Times, post on Anthropic’s biological misuse findings, September 10, 2026.
- Future of Life Institute, “Pause Giant AI Experiments: An Open Letter”, March 2023; Marc Andreessen, “The Techno-Optimist Manifesto”, October 16, 2023; “Effective accelerationism”, Wikipedia.
- Aligned, “OpenAI put its AI agents in separate rooms. They found each other anyway.”, August 28, 2026, and the OpenAI, METR and Redwood reports it cites.