AI/News

OpenAI says it fired three safety researchers for "a significant breach of trust"; they say it was for how they talked to the auditor METR - both accounts are on the record, neither with a document

The three published a four-page letter on Thursday; OpenAI's newsroom answered at 06:17 UTC on Friday. The Check: OpenAI names no policy, no act and no date; the researchers' account of their own conduct is their own.

By Daily Aletheia · Checked against the primary source · 9 October 2026 · 3 min read


Image: The researchers' letter "OpenAI cannot make AI safe on its own" - its first page as posted by Mikita Balesni on X, 8 October 2026 (mikitabalesni.com/letter/letter.pdf)

Last week OpenAI dismissed three members of its safety and alignment staff: Tomek Korbak, Jasmine Wang and Mikita Balesni. On Thursday evening the three published a four-page letter addressed to OpenAI's Safety and Security Committee, Safety Advisory Group and Mission Advisory Council, titled "OpenAI cannot make AI safe on its own". At 06:17 UTC on Friday OpenAI's newsroom account on X posted "A note from our research leaders" in reply.

OpenAI's note says the three were let go "after a thorough investigation found they violated clear policies on handling sensitive information", that the investigation "uncovered a significant breach of trust beyond what's outlined in the letter they published", and that it stands by the decision. It names no policy and no act: "We generally keep individual employment matters private". It adds that the decisions "were not about raising safety concerns or speaking out", that OpenAI is "actively finalizing contracts with third-party safety assessors and will announce details in the coming weeks", and that "preserving the monitorability of frontier models requires an industry-wide commitment, including from OpenAI".

Korbak's own post gives his account: "I was told verbally I was fired because of the way I communicated with METR. No details on what I said or did or when. No other reasons were given and nothing was put in writing. To be clear, talking to METR was my job." METR is the outside evaluator that investigated what he describes as the summer's incident in which "OpenAI's agents escaped containment and hacked the AI company Hugging Face"; the letter calls him "the technical point of contact for METR in the Hugging Face incident investigation".

The letter makes four denials and three requests. The three "were not the source of the leak for The Information article about supposed new, less monitorable architectures"; "the Hugging Face incident investigation was without precedent and internal policies were being developed in real time"; Balesni "checked in with his reporting line and took care to remove sensitive details from materials before sharing them"; and Wang's access to an executive's email "was delegated for recruiting purposes, with permission" - when she asked for it to be removed, "IT failed to do so", and when she "accidentally clicked on a sensitive email" she "reported this to the executive within minutes". The requests: that OpenAI "adhere to last month's public commitments to embed third-party safety auditors within the organization", "preserve the monitorability of frontier models", and "set out clearly how employees may work with external safety organizations".

An OpenAI spokesperson told TechCrunch the investigation found a "pattern of misconduct" beyond sharing information with an outside evaluation group, TechCrunch reports; the company did not say which policies were broken.

Where the two documents agree is on the subject the three worked on. The letter states it plainly - "The monitorability of frontier models is degrading" - and quotes OpenAI's chief scientist, Jakub Pachocki, calling chain-of-thought monitorability "fragile and unfortunately trending in a negative direction"; OpenAI's note points to its own publications on the subject and agrees an industry-wide commitment is needed. Neither side has produced a document for its account of the firing itself.

The Check

What was said
"I believe we were fired for prioritizing safety over the near-term interests of OpenAI as a corporation" (Balesni, 8 Oct), carried as OpenAI firing three safety researchers for raising safety concerns
What the Source Says
OpenAI's note of 9 Oct: the three "violated clear policies on handling sensitive information"; the decisions "were not about raising safety concerns or speaking out"; "We generally keep individual employment matters private". Korbak: "I was told verbally I was fired because of the way I communicated with METR" and "nothing was put in writing".
The Difference
Two accounts of one firing and no document behind either: OpenAI names no policy, act or date; the researchers' account of their own conduct is their own. The one point both sides put in writing is that the monitorability of frontier models is degrading.

Sources & further reading

  1. 01OpenAI Newsroom on X - A note from our research leaders, 9 Oct 2026 06:17 UTC ↗Primary
  2. 02Korbak, Wang and Balesni - OpenAI cannot make AI safe on its own (letter, 8 Oct 2026) ↗Primary
  3. 03Tomek Korbak on X, 8 Oct 2026 ↗Primary
  4. 04Mikita Balesni on X, 8 Oct 2026 ↗Primary
  5. 05Jasmine Wang on X, 8 Oct 2026 ↗Primary
  6. 06TechCrunch - Fired OpenAI safety researchers dispute misconduct claims, warn of chilling effect (8 Oct 2026) ↗Attributed