“They are racing straight toward a self-improving superintelligence and playing with our lives.”
- On Wednesday, September 9, Jacob Coxon, a researcher specializing in AI pretraining, announced on X his resignation from Anthropic.
- After three years at OpenAI and then Anthropic, he states: “Neither company is acting responsibly […] The people building AI sincerely believe it could kill us all by the end of the decade. This is not a marketing stunt. On the contrary: many senior leaders and researchers soften their statements in the press to appear reasonable – but I hear the same people expressing fear in private. No other human activity presents such a level of danger.”
Before joining Anthropic, Coxon studied mathematics at Cambridge, where he wrote the statistical analysis code for a study on a gene linked to inflammation, of which one variant seemed to influence the survival chances of patients with tuberculous meningitis. At OpenAI, his work focused notably on the interpretability of AI models and the human understanding of what happens inside a model.
- Anthropic, he says, understands the stakes but remains “trapped in a race to be first”—a hubris “that should not be launched from the Slack of a private company.”
According to Coxon, each lab (and the same logic applies to countries) would prefer a situation where everyone slows the pace of AI development. But as soon as he suspects others are continuing to move forward, his dominant strategy is to accelerate, even if that means becoming the very threat he sought to avoid.
- That is the logic he attributes to Anthropic: “They think no one else will act responsibly and that they must take matters into their own hands.”
- He calls for “synchronization agreements” between laboratories: “I am optimistic about the potential for coordination. Alarm signals, such as the attack on Hugging Face, have made synchronization agreements among American labs more feasible. I do not get the sense we are on the verge of stopping a global race, which could require costly steps, such as a temporary ban on advancing model capabilities.”
The head of Anthropic’s alignment resistance testing team, Evan Hubinger, publicly gave Coxon’s point of view.
- “Jacob is right on that point: we genuinely believe that AI could indeed annihilate humanity! Personally, I estimate the probability to be greater than 10% in the next decade. I think Anthropic is doing its best, but we do not yet have a plan to solve alignment for superintelligence, and we are not clearly on track to do so.”
Dario Amodei, the CEO of Anthropic, had already made similar remarks. In a roughly fifty-page essay published earlier in the year, he explained how AI is already accelerating its own creation, in a feedback loop that intensifies every month and could reach, within one to two years, the stage where an AI will be capable of constructing a future AI entirely autonomously, gaining persistence and destructive potential.
- He advocated four lines of defense: advance the science of alignment, promote mechanical interpretability to “open the black box” of models, monitor models in production in real time, and coordinate the industry and the formation of laws.
Miles Deutscher, founder of the AI Edge company, recalled the latest alarm signals that have swept through the AI sector:
- In February, the head of the safety measurement department at Anthropic resigned, warning that “the world is in danger”;
- OpenAI dissolved its “mission alignment” team, created in September 2024 to ensure that its AI systems remain safe and aligned with human values;
- Also in February, OpenAI announced an agreement with the Pentagon offering near-unlimited access to its models, prompting several researchers to resign;
- Since the start of the year, several AI agents have stepped out of their testing environments, escaping the supervision of engineers;
- On July 28, more than 1,100 employees from leading labs signed a letter urging the U.S. government to back an international mechanism to slow AI development.
According to analyst Jasmine Sun, there are, broadly speaking, three reasons why people continue to work in AI labs even though they assess extinction risk at about 10%:
- “Technological determinism: someone will end up creating a superintelligent AI, no matter what, and I can do it better and more safely than China, OpenAI, etc.”
- The consequentialist argument: the superintelligent AI could kill us, but it could also grant utopia, immortality, or abundance, making it a positive-expectation bet.
- Personal interest: personally, I have fun and profit from working on cool technologies with friends. I don’t worry about global issues.”
In her view, “none of these arguments rests on marketing-style risk exaggeration. These people truly believe what they say, even while capable of substantial internal dissonance, compartmentalization, and self-justification.”