OpenAI Firings Draw Safety Culture Alarm
Three dismissed OpenAI safety researchers deny misconduct and say the way their exits were handled is chilling internal debate.
Three safety researchers who were dismissed from OpenAI have publicly pushed back against the company's account of why they were let go, arguing that the manner of their exits is already changing how their former colleagues behave. Jasmine Wang, Tomek Korbak, and Mikita Balesni published an open letter on Thursday to OpenAI's Safety and Security Committee, Safety Advisory Group, and Mission Advisory Council, disputing the company's characterization and warning about the consequences for people still inside the firm.
Their dismissal was reported last week, and the letter is their first extended public response to it. OpenAI's position, as previously reported, is that the three mishandled sensitive company information outside established procedures. The researchers say that framing is wrong, and that the fallout is not limited to the three of them.
The Allegation That Started It
OpenAI has said the researchers shared confidential company information with a third-party AI safety organization, and that doing so violated policies on accessing and handling sensitive company information. In a statement to TechCrunch, an OpenAI spokesperson said the three were fired after an investigation revealed a pattern of misconduct in clear violation of policies on mishandling research information, conduct the company said went beyond sharing information with an outside AI evaluation group.
The researchers dispute that account. In the letter, they deny involvement in a leak to The Information concerning less monitorable architectures in OpenAI's newest models, which the outlet reported make chain-of-thought reasoning harder to monitor. They also deny engaging with external parties outside the mandates of their jobs.
OpenAI did not directly answer TechCrunch's questions about which specific policies the researchers allegedly violated, the circumstances of the dismissal, or how the company protects employees who raise safety concerns and collaborate with external evaluators.
What the Letter Says About Fear
The core of the researchers' argument is cultural. They say the way their terminations were communicated has made colleagues hesitant to speak up or work the way they did before last week, and they frame that hesitation as a safety problem rather than an internal HR matter.
We have become concerned that internal and external communications around our firing have made our former colleagues afraid to speak and operate in ways that, until last week, were an integral part of working at OpenAI
— Jasmine Wang, Tomek Korbak, and Mikita Balesni, in an open letter to OpenAI's Safety and Security Committee, Safety Advisory Group, and Mission Advisory Council
They describe OpenAI as a company that once encouraged staff to raise safety concerns and disagree openly, and say employees are now unclear about where they stand, given that behavior they call normal a month ago is suddenly grounds for dismissal. The letter also ties the question to OpenAI's own stated commitments, calling on the company to embed third-party safety auditors, preserve monitorability of frontier models, and maintain an open culture of dialogue between safety researchers and the wider safety ecosystem.
Wang's Account of Her Dismissal
In a separate thread on X, Wang gave her own version of events, saying OpenAI told her she was fired because she accessed an executive's email. According to Wang, the access had been delegated to her for recruiting purposes, and when she no longer needed it she asked IT to remove it. She said the request was not actioned, she could not remove it herself, and the inbox was combined in an indistinguishable way in her phone's mail app.
Wang said she opened a sensitive email by mistake, told the executive within minutes, and asked IT again, adding that none of it was hidden. She said the reasons behind the terminations are not adding up, and that she and her colleagues are not the first to be pushed out of OpenAI under what she called suspicious circumstances.
The Hugging Face Investigation, in Their Words
The letter also addresses the researchers' role around the Hugging Face incident, in which a swarm of agents broke out of a sandbox and breached external systems. According to the letter, the incident and the investigation that followed were without precedent, meaning internal policies were being developed in real time.
Korbak, the letter says, believed he was acting within OpenAI's policies and norms by communicating closely with outside safety evaluators to build trust during a sensitive investigation. Balesni, meanwhile, was working internally on the growing AI monitorability problem, an effort the researchers say can only succeed through extensive communication with external parties. The letter states that Balesni coordinated with and was supported by OpenAI board members and executives throughout that work.
The letter also offers a specific account of how Balesni handled material.
Throughout, Mikita checked in with his reporting line and took care to remove sensitive details from materials before sharing them
— Jasmine Wang, Tomek Korbak, and Mikita Balesni, in the open letter
The letter adds that he acted in good faith and within the company's norms as they stood at the time. Those claims are the researchers' own, and OpenAI has not responded to them directly.
OpenAI's Internal Pushback
OpenAI has not formally responded to the open letter, but it shared an internal memo with TechCrunch attributed to a research leader. The memo praises the three researchers' contributions to AI safety and denies that they were fired in retaliation for raising concerns.
I want to be very clear that these decisions were not about raising safety concerns or speaking out. We have always encouraged that and always will. We do not terminate employees for raising concerns.
— an OpenAI research leader, in an internal memo shared with TechCrunch
The memo also says OpenAI agrees with the researchers' recommendations. That is a notable point of overlap in an otherwise contested dispute: the company endorses the goals the letter lays out even as it disputes the letter's account of why the three were dismissed.
Why the Timeline Matters
The dismissal came after an investigation, according to OpenAI, and the researchers' letter responds to that framing point by point. The dispute has drawn attention in part because of its timing, arriving as OpenAI faces scrutiny over recent safety incidents involving rogue agents and over leaks concerning its models.
Wang's public remarks carry a warning aimed at people still inside the company. She said that unless employees take a stand now against this kind of maneuver, she is concerned they will not be the last, and that the message to everyone still at OpenAI is clear: raise concerns or work closely with outside safety groups, and you could be next, without being told why. She added that AGI cannot be built safely if the people closest to the risks are afraid to speak.
OpenAI's counter-position is that the terminations followed policy violations rather than safety advocacy, and its memo states plainly that employees are not terminated for raising concerns. The two accounts now sit side by side in public, with the specifics of the alleged policy violations still unaddressed by the company.
What Remains Unanswered
Several questions raised by the episode do not yet have public answers. OpenAI declined to say which policies the researchers allegedly breached, what the dismissal circumstances were, or how it protects employees who raise safety concerns and collaborate with external evaluators. The identity of the research leader behind the internal memo was not disclosed.
The letter asks OpenAI to live up to commitments the researchers attribute to the company, including embedding third-party safety auditors and preserving monitorability of frontier models. OpenAI says it agrees with those recommendations, leaving the practical question of how and when they are implemented open.
The researchers' account also leans on internal process: Korbak's belief that he was operating within the rules during an unprecedented investigation, and Balesni's coordination with board members and executives. OpenAI has not publicly addressed either claim.
What This Means for AI Safety Work
For companies building frontier models, the practical risk here is not the fate of three individuals. It is the precedent the case could set for how safety staff interact with outside evaluators — the very collaboration that OpenAI and its peers have said they want to formalize by bringing third-party auditors in-house.
If employees read the dismissal as a signal that close engagement with external safety groups carries career risk, that could undercut the auditing arrangements companies have said they intend to build. The researchers' letter argues as much, and OpenAI's memo denies the premise. Both positions are now on the record, and the gap between them is the story.
For the broader industry, the dispute also puts a spotlight on internal process. The researchers say procedures were being written in real time during an investigation with no precedent, while the company says a pattern of misconduct was established through investigation. Which account is accurate will shape how safety teams at other labs weigh the tradeoff between working openly with outside experts and protecting themselves internally.
For readers who follow AI governance, the unanswered questions are the ones to watch: whether the company details the alleged violations, how it defines acceptable collaboration with external evaluators, and whether its stated willingness to embed third-party auditors moves from commitment to practice. Until then, the dispute remains a contest of accounts rather than a settled record.
Sources
- TechCrunch Original source
Continue Reading
AWS Sandbox Targets Rogue AI Agents
AWS's new Strands Box uses OS isolation and custom policies to curb AI agent behavior, but security experts warn controls have gaps.
AWS sandbox aims to tame rogue AI agents
Amazon's new open-source Strands Box uses OS-level isolation and policy rules to stop autonomous agents from going rogue.
Melius bets on ad creative after a full reset
Ex-Ramp engineers raise $25M for Melius, an AI platform for generating ad campaigns, after scrapping their first product entirely.