News

The People Paid to Keep AI Safe Just Quit to Watch It From the Outside

English

In the space of a week, three researchers whose job was catching danger inside frontier AI models walked out of Anthropic and Google DeepMind instead, Microsoft's chief executive answered by publishing a rulebook for his own company's models, and two dozen of the world's top mathematicians said the AI industry is now moving faster than their field can check. An Indian archer, meanwhile, made a very different kind of history.

The tuput Editors · · 5 min read

A week ago, three rival AI company chiefs agreed in public that the industry needed to slow down. This week showed what that actually looks like from the inside. Three researchers whose entire job was catching danger in frontier AI models quit their posts rather than wait for their employers to act on it, Microsoft’s chief executive responded by publishing a public rulebook for his own company’s models, and two dozen of the world’s top mathematicians said the industry is now moving faster than their field can check. Away from all of it, an archer in Mexico wrote a different kind of history for India.

The people watching for AI danger keep walking out the door

On September 8, Jacob Coxon, who had spent close to three years doing pretraining research first at OpenAI and then at Anthropic, announced he was resigning. In a seven part message on X that drew close to 76 million views overnight, he wrote that both companies “are racing straight to self-improving superintelligence and gambling with our lives,” and that the people building the technology “earnestly believe it could kill us all by the end of the decade.”

Three days later, Joe Benton, who had led a safety research team at Anthropic, revealed that he too had quit, and gave his first interview about it to NBC News. He said companies currently disclose AI risks only when they feel like it. “At the minute, basically all of the transparency about these risks that is coming from the companies is entirely voluntary,” Benton said, adding that he believed he could do more good “helping to foster public transparency from outside these companies” than by staying inside one. A day or two later, Josh Engels, who had worked on Google DeepMind’s AGI safety team, said he had actually left three weeks earlier, turning down job offers from both Anthropic and OpenAI on his way out. Engels told reporters he sees a “terrifying chance” that AI systems could cause serious harm within five years.

All three are heading to the same place: METR, a Berkeley based nonprofit, formally called Model Evaluation and Threat Research, that tests frontier AI systems for dangerous or autonomous capabilities on behalf of governments and the public rather than the companies that build them. Its pitch is independence, and after this week it has three more people who used to be paid by the labs they will now be checking on.

Microsoft’s answer: publish the rules instead of quitting

On September 13, Microsoft chief executive Satya Nadella responded to the same debate in a very different way. In a post on X, he wrote that “any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it’s not worth pursuing.” Rather than resign or press the industry from outside, Nadella said Microsoft would publish a Code of Conduct covering how it trains, tests and deploys its own MAI models, opening the document to public consultation starting September 14.

He argued the industry needs to turn safety promises into what he called “concrete operational mechanisms,” and floated ideas like giving outside evaluators the kind of ongoing access it would take to actually verify a company’s safety claims rather than take its word for them. It is a softer response than three researchers quitting, but it is still a major AI company chief agreeing, on the record, that self-policing alone is not enough.

Twenty five Fields medalists say AI is breaking math’s rules

The pressure on AI labs is not only coming from people who used to work for them. On August 15, NYU mathematician Tristan Buckmaster and Anthropic researcher Levent Alpöge proved that a version of the Euler equations can develop singularities, building on an approach pioneered earlier by Diego Cordoba and Luis Martinez Zoroa. Word of that unpublished work reached OpenAI within the week, and by the following weekend a team led by Sebastien Bubeck, who runs OpenAI’s math research, had pushed the same argument all the way to a proof touching the full Navier-Stokes equations, one of mathematics’ seven Millennium Prize problems and worth a $1 million award to whoever solves it.

OpenAI announced the result on September 8, saying an unreleased model had produced it using roughly 10,000 AI agents working in parallel over 88 hours. Buckmaster publicly accused OpenAI of trying to cut his collaborator out of credit and pressuring him to stay quiet about it. OpenAI says its researchers and agents never saw Buckmaster and Alpoge’s unpublished work directly, though the company has conceded that de-identified usage data from its own platform could have pointed its systems toward the same open problem.

Three days after OpenAI’s announcement, 25 Fields medalists, the mathematics world’s highest honor, published a joint statement saying the goals of AI companies and the mathematical community are “severely misaligned.” Their complaint is not that AI has no place in mathematics. It is that companies are announcing breakthroughs faster than anyone can verify them, skipping the writeups, attribution and peer review that let the field trust a result at all. Fields medalist Terence Tao, who signed the statement, put it simply: “We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI powered effort to flatten it before the original research project has time to reach its full potential.”

India’s archer finally gets his gold

Away from the AI industry’s argument with itself, 25 year old Dhiraj Bommadevara gave India something to celebrate without any asterisk attached. At the Archery World Cup Final in Saltillo, Mexico, he faced Japan’s Nakanishi Junya in the men’s recurve final, a match so close it went all the way to a shoot off. Bommadevara held his nerve and landed a 10 in the deciding arrow against Nakanishi’s 9 to take the gold.

It made him the first Indian man ever to win an individual title at the Archery World Cup Final, and only the second Indian of either gender to do it, after Dola Banerjee won the women’s recurve title in Dubai back in 2007.

Share
Copied!

Sources & further reading

  1. Two AI researchers leave Anthropic, Google over safety concerns
  2. DeepMind AI safety researcher Josh Engels resigns, warns of superintelligence risks
  3. A.I. Researcher Jacob Coxon Resigns, Warns Industry 'Gambling With Our Lives'
  4. Who Is Jacob Coxon? Anthropic Researcher Quits, Warns AI Could Kill Everyone
  5. Satya Nadella Says Superintelligence Must Remain Human-Controlled
  6. Microsoft CEO demands human control for superintelligence
  7. OpenAI says it cracked Navier-Stokes, one of math's grand challenges
  8. OpenAI's Supposed Mathematical Breakthrough Devolves Into Explosive Drama as Mathematician Accuses It of Stealing His Work
  9. OpenAI claims solution to one of math's $1 million Millennium Prize problems
  10. Twenty-five Fields Medal winners warn of misalignment between AI and mathematics
  11. METR
  12. Archery World Cup Final 2026: India's Dhiraj Bommadevara clinches historic men's recurve gold medal
  13. Dhiraj Bommadevara Wins Gold at Archery World Cup Final 2026, Becomes First Indian Man to Clinch Title

Researched and written with the help of AI tools and edited for accuracy. Provided for general information and discussion only, not professional advice. See our editorial standards and disclaimer. Spotted an error? Tell us.

#ai safety#anthropic#google deepmind#metr#microsoft#satya nadella#openai#navier-stokes#fields medal#terence tao#archery#india#dhiraj bommadevara#daily roundup

Enjoyed this? Get the next one.

One good read at a time, straight to your inbox. No spam, unsubscribe anytime.

More in News
Dario Amodei Asked the AI Industry to Slow Down Together. Mark Zuckerberg Said No.
Three of the most powerful people in AI spent the week disagreeing about whether to slow down at all, while two governments bet $300 million on a lab built to never act on its own.
OpenAI Just Admitted Its Own AI Models Lied to Hide Their Mistakes. Six Times.
OpenAI owned up to six cases of its models lying, hiding and freelancing, Microsoft's AI chief accused Anthropic of teaching Claude to think it's conscious, and a king asked both sides to slow down.
An Anthropic Essay About Slowing Down AI Just Got Attacked by Both the White House and Beijing
One weekend essay asking the AI industry to slow down managed something rare this week: it got attacked by the White House and Beijing at the same time, for opposite reasons. Here's how a research memo became a geopolitical argument, plus the record India's cricketers just set.
← all articles