ai governance
a 10 percent chance of what, exactly
the head of alignment at anthropic put a number on the end of the world, in a reply, with no analysis attached. the number cannot be checked, and it has the same shape as the benchmark score in the vendor deck on your desk. 4 questions work on both.
on september 9, 2026, a researcher named jacob coxon quit anthropic after 3 years of pre-training research there and at openai. he said neither company is acting responsibly, and that they are "racing straight to self-improving superintelligence and gambling with our lives"1.
a researcher saying that on the way out is a story, and not an unusual one. the part worth your time is who agreed with him.
evan hubinger leads alignment science at anthropic and still works there. he replied in public: "we really do earnestly believe ai could kill all humans! i personally think it is >10 per cent within the next decade." he added that anthropic is trying its best, but does not yet have a plan to solve alignment for superintelligence and is "not clearly on track to"1.
so the 10 percent is hubinger's own number, put there while employed. he marked it as personal. the title under his name is anthropic's.
the part that should bother a buyer
set aside whether the number is right. nobody can check it, and that is the part to look at, because you have a number like it on your desk.
the ai vendor's deck has one: a benchmark score, a share of a workflow automated, hours saved per analyst per week. that number and hubinger's share a shape. each comes from someone positioned to know. each makes the technology sound more powerful than you had assumed. and nothing that happens next quarter can prove either one wrong, because nobody outside the company can see the analysis behind it. a number with that shape does sales work whether or not anyone means it to. hubinger, who is telling the public his own field has no plan, plainly believes his. sincerity does not make a number testable.
so you are being asked to believe 2 things at once about the same technology: that it may end the world, and that it is ready for you to restructure a department around. both arrive with confidence. neither arrives with a way to check.
show the work
a researcher posted on his way out. an executive replied. a reporter wrote up the exchange. no document accompanied any of it, and anthropic's published safety frameworks and model cards do not contain this estimate or the reasoning behind it.
ed elson, a business podcast host, said on september 9 that he does not believe the number, and that "that's not a good enough reason not to interrogate them"2. we agree. his ask is public testimony and the documentation behind the claim. nobody can settle in a hearing whether the technology is dangerous, so the ask worth making is narrower: 4 artifacts, each of which either exists or does not.
1. the analysis behind the number, with its assumptions and the name of whoever signed it.
2. what changed at the company after it reached that conclusion.
3. whether the same evidence went to the investors in anthropic's coming ipo3. a claim material enough to say in public is material enough to appear in a filing.
4. what result would move the number. a forecast nothing could change is a belief, not a risk estimate.
anthropic has not refused any of this. it was never published alongside the claim.
the version you can use next week
the same 4 questions work on the vendor deck, and you do not have to wait for washington to ask them. what analysis is behind the number. who ran it. what would make it come out differently. can we see it. a vendor who can answer those has earned the score. a vendor who cannot has told you what the score is worth.
sources
- 1abc news, "anthropic researcher quits saying ai developers believe 'it could kill us all'," september 9, 2026. jacob coxon's role, tenure and quoted words; evan hubinger's title and his quotes, including ">10 per cent within the next decade" and "not clearly on track to."
- 2prof g markets, episode of september 9, 2026 (prof g media). ed elson's closing segment: his stated disbelief in the number, his call for congressional testimony and subpoenas, and his argument that the statement was material to investors. quotes transcribed from the episode captions. we agree with the call for the documents and take no position on the securities question.
- 3anthropic, "anthropic confidentially submits draft s-1 to the sec," june 1, 2026. the company's own announcement of a confidential draft registration statement toward a proposed initial public offering.