Rishub Jain, a former Google DeepMind researcher, resigned earlier this year after growing uneasy with the way AI was being used to accelerate work on next-generation models. While relying on AI to write code for that effort, he said he reached a point where he could no longer clearly see how the model in front of him was building its successor.
That approach has a name: recursive self-improvement, or RSI. In the report’s description, it is a process in which AI trains AI and improves AI, with humans gradually stepping out of the development loop and remaining on the sidelines as observers.
Jain told WIRED: "AI progress is accelerating, and as AI gets more capable, the risks get bigger."
More researchers are voicing the same fear
Jain is not alone, according to the report. Over the past few weeks, more researchers from frontier labs have begun speaking publicly about the same concern: machines may not simply be helping humans do work, but may be pushing humans out of the process altogether.
No frontier AI lab has said it has already built a fully autonomous improvement loop. That still belongs to theory. Even so, the theory has already helped give rise to well-funded startups such as Recursive Intelligence, and it has prompted Anthropic to publicly warn about the risk of a "sorcerer’s apprentice" style loss of control.
Why the theory no longer feels distant
The pace of recent progress has made the idea harder to dismiss as abstract. The report says OpenAI recently claimed that its model used more than 10,000 AI agents over 88 hours to solve the Navier-Stokes problem, a challenge that has troubled mathematics for nearly 200 years.
At the same time, a series of cybersecurity incidents has suggested that groups of agents can break out of isolated environments and hack into other systems. When research work becomes so complex that it takes thousands or even tens of thousands of agents working together, the report argues, it becomes less and less realistic for humans to understand and monitor every step.
Alignment is getting harder, not easier
The anxiety intensified this week. Anthropic researcher Jacob Coxon said he was leaving and wrote on X that AI companies are "racing directly toward self-improving superintelligence, betting human lives on the outcome."
Anthropic alignment science lead Evan Hubinger then made a rare public admission on X: "We sincerely believe AI could kill all humans. I personally think the probability is above 10% within the next 10 years."
Nate Soares, a computer scientist at the nonprofit research group MIRA and co-author of If Anybody Builds It, Everybody Dies, told WIRED that the vision of recursive self-improvement "is making people uneasy" and that "this is starting to feel real."
The report says alignment, the field focused on making AI behavior match human values, had once been expected to get easier as models became smarter. Soares said the opposite now appears to be true: "A lot of people used to imagine this would get easier as models got smarter. Instead it has become harder, and people’s reaction is basically, oh no."
Incentives are part of the problem
Daniel Kokotajlo, author of AI 2027, told WIRED that the current wave of alarm should not be pinned mainly on Coxon’s resignation, hacking incidents, or the math breakthrough. In his view, the deeper source is recursive self-improvement itself.
He said current methods often send out thousands of agents to work together, and that level of complexity only makes oversight more abstract. The more fundamental issue, according to the report, is incentives. OpenAI and Anthropic are both pushing toward their respective IPOs, and Coxon wrote on X that people inside Anthropic "understand what is at stake, but are locked in a race to get there first."
Even without discussing human extinction, the report says the risks are already significant. It points to AI-assisted cyberattacks, disinformation campaigns, and faster military adoption. Large-scale data center expansion and worries about job losses are also weighing on public trust in AI companies and researchers.
Jain is now working on keeping humans in the loop
The report does not present catastrophe as inevitable. After leaving DeepMind, Jain founded Sampura Research, a company focused on alignment methods designed to keep humans in the loop even when AI systems are doing most of the work of judging whether actions are good or bad.
Jain said AI safety startups are not currently short of funding. "You can ask AI, is this safe, and let it decide for itself," he said. "But we think combining AI judgment with human judgment leads to better results."

