Jacob Coxon's resignation from Anthropic wasn't an isolated warning — the company's alignment lead has separately said AI could kill humanity with over 10% probability.
Jacob Coxon's resignation from Anthropic wasn't a lone voice. According to The Wall Street Journal and Hindustan Times, Coxon — who spent three years in pretraining research across OpenAI and Anthropic — left saying neither company is acting responsibly, calling the race toward self-improving superintelligence "a gamble with human lives."
But the more striking part of this story is what's happening inside Anthropic at the same time. Forbes reports that Evan Hubinger, identified as Anthropic's alignment science lead, said the company does not yet have a plan to solve alignment for superintelligent systems — and separately put his own estimate of AI eventually killing all humans within the next decade at greater than 10%.
That's an important distinction to sit with: this is Hubinger's personal probability estimate, not a stated Anthropic corporate position, and Anthropic hasn't publicly responded to either his comments or Coxon's allegations. But two people connected to the company — one departing, one still leading its alignment research — are independently pointing at the same unresolved question: can researchers reliably control systems significantly more capable than today's models?
It's worth being precise about what's actually being claimed here. Coxon's warnings — that advanced AI could compromise digital systems, accelerate work across fields, and acquire real-world resources — are forecasts about future capability, not documented behavior in currently deployed models. There's no evidence cited that existing Anthropic or OpenAI systems can self-improve, evade control, or act autonomously today.
Bottom line: This isn't proof of imminent catastrophe, and it isn't a claim that today's models are dangerous. It's something narrower but still serious: a former frontier researcher believes the industry is racing ahead of its safety case, and Anthropic's own alignment leadership has publicly admitted it doesn't have a complete plan for the problem it was founded to solve.
Jacob Coxon's resignation from Anthropic wasn't a lone voice.
It's worth being precise about what's actually being claimed here.
There's no evidence cited that existing Anthropic or OpenAI systems can self improve, evade control, or act autonomously today.
Continue reading