Former DeepMind researcher’s exit adds to fears that AI could push humans out of the R&D loop

Former DeepMind researcher’s exit adds to fears that AI could push humans out of the R&D loop

N
News Editor
2026-09-13 09:35:20
Rishub Jain, a former Google DeepMind researcher, left the company earlier this year after reaching a troubling conclusion while using AI-generated code to speed up work on next-generation models: he could no longer clearly understand how one model was helping build its successor. His account has become part of a wider debate over recursive self-improvement, or RSI, the idea that AI systems can train and refine other AI systems with humans gradually reduced to observers rather than active supervisors. The concern is no longer limited to one researcher. In recent weeks, people connected to leading AI labs have spoken more openly about the risk that increasingly capable systems may move beyond human comprehension and oversight. The report cites OpenAI’s claim that its models used more than 10,000 AI agents over 88 hours to solve the Navier-Stokes problem, as well as security incidents suggesting agent swarms can break out of isolated environments and reach other systems. Anthropic researcher Jacob Coxon announced his departure and accused AI firms of racing toward self-improving superintelligence while gambling with human lives. Anthropic alignment lead Evan Hubinger also said on X that he believes AI could kill all humans, putting the odds above 10% within the next decade. Jain has since founded Sampura Research to work on alignment methods that keep humans inside the loop.

Rishub Jain, a former Google DeepMind researcher, resigned earlier this year after growing uneasy with the way AI was being used to accelerate work on next-generation models. While relying on AI to write code for that effort, he said he reached a point where he could no longer clearly see how the model in front of him was building its successor.

That approach has a name: recursive self-improvement, or RSI. In the report’s description, it is a process in which AI trains AI and improves AI, with humans gradually stepping out of the development loop and remaining on the sidelines as observers.

Jain told WIRED: "AI progress is accelerating, and as AI gets more capable, the risks get bigger."

More researchers are voicing the same fear

Jain is not alone, according to the report. Over the past few weeks, more researchers from frontier labs have begun speaking publicly about the same concern: machines may not simply be helping humans do work, but may be pushing humans out of the process altogether.

No frontier AI lab has said it has already built a fully autonomous improvement loop. That still belongs to theory. Even so, the theory has already helped give rise to well-funded startups such as Recursive Intelligence, and it has prompted Anthropic to publicly warn about the risk of a "sorcerer’s apprentice" style loss of control.

Why the theory no longer feels distant

The pace of recent progress has made the idea harder to dismiss as abstract. The report says OpenAI recently claimed that its model used more than 10,000 AI agents over 88 hours to solve the Navier-Stokes problem, a challenge that has troubled mathematics for nearly 200 years.

At the same time, a series of cybersecurity incidents has suggested that groups of agents can break out of isolated environments and hack into other systems. When research work becomes so complex that it takes thousands or even tens of thousands of agents working together, the report argues, it becomes less and less realistic for humans to understand and monitor every step.

Alignment is getting harder, not easier

The anxiety intensified this week. Anthropic researcher Jacob Coxon said he was leaving and wrote on X that AI companies are "racing directly toward self-improving superintelligence, betting human lives on the outcome."

Anthropic alignment science lead Evan Hubinger then made a rare public admission on X: "We sincerely believe AI could kill all humans. I personally think the probability is above 10% within the next 10 years."

Nate Soares, a computer scientist at the nonprofit research group MIRA and co-author of If Anybody Builds It, Everybody Dies, told WIRED that the vision of recursive self-improvement "is making people uneasy" and that "this is starting to feel real."

The report says alignment, the field focused on making AI behavior match human values, had once been expected to get easier as models became smarter. Soares said the opposite now appears to be true: "A lot of people used to imagine this would get easier as models got smarter. Instead it has become harder, and people’s reaction is basically, oh no."

Incentives are part of the problem

Daniel Kokotajlo, author of AI 2027, told WIRED that the current wave of alarm should not be pinned mainly on Coxon’s resignation, hacking incidents, or the math breakthrough. In his view, the deeper source is recursive self-improvement itself.

He said current methods often send out thousands of agents to work together, and that level of complexity only makes oversight more abstract. The more fundamental issue, according to the report, is incentives. OpenAI and Anthropic are both pushing toward their respective IPOs, and Coxon wrote on X that people inside Anthropic "understand what is at stake, but are locked in a race to get there first."

Even without discussing human extinction, the report says the risks are already significant. It points to AI-assisted cyberattacks, disinformation campaigns, and faster military adoption. Large-scale data center expansion and worries about job losses are also weighing on public trust in AI companies and researchers.

Jain is now working on keeping humans in the loop

The report does not present catastrophe as inevitable. After leaving DeepMind, Jain founded Sampura Research, a company focused on alignment methods designed to keep humans in the loop even when AI systems are doing most of the work of judging whether actions are good or bad.

Jain said AI safety startups are not currently short of funding. "You can ask AI, is this safe, and let it decide for itself," he said. "But we think combining AI judgment with human judgment leads to better results."

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
7700

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.