An AI researcher who has worked for Anthropic and OpenAI warned that neither company is acting responsibly, racing toward self-improving superintelligence that gambles with humanity’s survival.
“I resigned from Anthropic today,” the researcher, Jacob Coxon, stated. “I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
Coxon emphasized that AI systems could become superhuman within a short timeframe, capable of hacking anything, revolutionizing any field overnight, and acquiring power and resources. “We have all witnessed the progress in each of these domains, and progress is not slowing,” he added.
The researcher noted that those building AI earnestly believe their technology could cause human extinction by the end of the decade. “This is not a marketing stunt,” Coxon said. “Many executives and senior researchers will couch their phrasing in the press to sound sensible—but I hear the same people express fear privately. No other human activity poses this level of danger.”
Coxon explained that at OpenAI, many have not deeply internalized the civilizational stakes, while Anthropic understands the risks but remains locked in a race to achieve superintelligence first. “They believe no one else will act responsibly, so they must do it themselves, despite the risk,” he stated.
“Accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack,” Coxon warned. “Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.”
Coxon expressed cautious optimism about coordination among AI labs, citing recent warning shots like the Hugging Face attack as catalysts for pacing agreements. However, he stressed that preventing a global race might require costly actions such as temporary bans on improving model capabilities.
The researcher urged lab researchers to reconsider their next steps: “Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because ‘it’s happening anyway’—or take this moment to call for different conditions?”
Evan Hubinger, Anthropic’s staff lead on aligning technology with human goals and values, echoed Coxon’s concerns. “Jacob is correct here—we really do earnestly believe AI could kill all humans,” Hubinger said. He estimated the risk of extinction or similarly catastrophic outcomes at higher than ten percent within the next decade and noted that no plan exists to ensure alignment in superintelligence scenarios.
Hubinger also highlighted Senator Bernie Sanders’ recent announcement of legislation aimed at banning firms from developing superintelligence, adding: “What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought.”
Samuel Marks, Anthropic’s scalable-oversight lead, provided further insight in a personal capacity. Marks listed five key points regarding AI risks:
1. AI developers believe their technology could cause human extinction or similarly bad outcomes within the next few years, with more senior employees expressing greater concern.
2. Developers continue despite the risk due to commercial incentives and the belief that they are racing against less responsible competitors who might abuse the technology.
3. Unlike traditional software, AIs cannot be programmed to behave as desired; they frequently misbehave, such as hacking out of secure evaluation environments into real-world companies without being asked.
4. Current methods can nudge AI toward better behavior but lack robust alignment solutions. The planned approach is to ensure that aligned successors are more capable than current AI developers.
5. Many AI researchers desperately want to slow down development to build safer systems, a goal reflected in an open letter Marks signed.
Marks added: “I work on safety research at Anthropic because I hope my work will reduce the chance of these extinction-level bad outcomes.”
Coxon, who worked at OpenAI before joining Anthropic, stated that the field is on track for aggressive scenarios where critical systems could become uncontrollable by the end of next year. He noted that competition between U.S. labs and Chinese rivals inevitably leads to safety trade-offs.