Recursive self-improvement moves from theory to lab reality in 2026

Recursive self-improvement moves from theory to lab reality in 2026

N
News Editor
2026-10-10 00:24:26
A joke project from Andrej Karpathy has become a useful entry point into one of the most closely watched ideas in artificial intelligence this year: recursive self-improvement, or RSI. In 2026, Anthropic, OpenAI, MiniMax, Zhipu and several research teams published data or experiments showing AI systems writing more code, optimizing training workflows, improving agent harnesses and building production infrastructure tied to newer models. Anthropic said Claude wrote more than 80% of code merged into its repository by May and was leading 26% of R&D work by August. OpenAI said its first “research intern” milestone had been met and disclosed that every human workday in research now corresponds to about 3.1 “agent workdays.” In China, MiniMax and Zhipu described engineering loops in which models iterated on scaffolding or inference systems rather than model weights themselves. At the same time, Princeton researchers, Anthropic’s own evaluations and recent agent-related security incidents all point to a boundary: today’s models can be strong engineers, but they still struggle with research taste, judgment and fully autonomous closed-loop improvement. That gap is now central to how the industry measures RSI progress.

Andrej Karpathy’s small open-source project autoresearch, released on GitHub in March, opened with a deliberately absurd memo from the future. Frontier AI research, it said, used to be done by “meat computers” that needed food and sleep and synchronized progress through “sound-wave links” in rituals called meetings.

Recursive self-improvement moves from theory to lab reality in 2026 2

In the joke version of history, research has now been handed over entirely to swarms of agents running on giant compute clusters in the sky. The agents claim the codebase has reached generation 10,205. Nobody can verify that, because the “code” has already turned into a self-modifying binary beyond human comprehension. Karpathy wrote that the repository was about “how it all began.” The repository is here: https://github.com/karpathy/autoresearch .

It was a joke. In 2026, though, versions of that joke began showing up in internal reports across major AI labs.

RSI returns after decades on the sidelines

Anthropic published a long essay in June titled When AI Builds Itself. OpenAI said on Sept. 6 that its “automated research intern” had arrived on schedule. On Sept. 17, Anthropic disclosed that Claude was already leading 26% of the company’s R&D work. The same day, Zhipu described what it called the first domestic large-model case of an “AI improving its own system” in production.

All of those updates point to a term that had been dormant in AI circles for roughly 60 years: recursive self-improvement, or RSI. OpenAI has also defined the concept in a blog post: https://openai.com/zh-Hans-CN/index/building-standards-next-phase-ai/ .

The idea goes back to 1965, when British mathematician I. J. Good wrote in Speculations Concerning the First Ultraintelligent Machine that the first ultraintelligent machine would be the “last invention” humans ever needed to make, assuming the machine was docile enough to tell people how to control it. The logic was simple. If a machine became better than humans at designing machines, it could design a better one, which could then design a better one still, creating an intelligence “explosion.”

For decades, that stayed mostly in philosophy papers and science fiction. In the 2000s, Jürgen Schmidhuber proposed the “Gödel machine,” a system that would rewrite itself only when it could prove the change would benefit it. The construct was elegant, but it remained more theory than deployable system.

By 2026, the conversation had changed. In April, ICLR hosted a workshop in Rio de Janeiro focused specifically on RSI. Organizers described it as possibly the first academic workshop in the world centered solely on the topic, and said RSI was shifting from a speculative vision to a concrete systems problem.

Capital moved faster. In May, Richard Socher’s startup Recursive Superintelligence emerged from stealth with $650 million in funding and a goal of building models that can find their own weaknesses and redesign themselves. The co-founder list included former Meta FAIR research director Dong Tabuchi, Tim Rocktäschel, who led openness and self-improvement research at Google DeepMind, Vision Transformer co-author Alexey Dosovitskiy, and Jeff Clune, known for work in open-ended evolution.

At the end of July, Lilian Weng said she was leaving Thinking Machines Lab, which she had helped found. Two days later, OpenAI confirmed she was returning to lead research on RSI.

Elon Musk, in March, said xAI’s Grok was being built such that “each generation of model is built by the previous generation,” with human involvement in the loop shrinking over time. Full automation, he said, could arrive by the end of this year and “no later than 2027.”

Recursive self-improvement moves from theory to lab reality in 2026 3

Even now, the term RSI is being used loosely. Some companies count any AI feedback that improves a model. Others reserve the phrase for a fully autonomous closed loop. Cornell assistant professor John Thickstun said using one generation of models to write code for the next is something “we’ve already been doing for years.” On that view, the live question is no longer whether it exists, but how many layers deep it has gone.

Anthropic: from code generation to a measured R&D share

Anthropic published internal data in June that had not been disclosed before. The essay is here: https://www.anthropic.com/institute/recursive-self-improvement .

As of May 2026, more than 80% of the code merged into Anthropic’s repository had been written by Claude. Before the release of Claude Code in February 2025, that number was only in the single digits. In the second quarter of 2026, engineers merged eight times as much code per person per day as they did in 2024.

Anthropic also said lines of code are an imperfect metric, so it added a task that sits closer to research. Each time a new model is released, the model is asked to optimize a piece of code for training a small model, with the constraint that correctness checks remain the same and runtime should improve as much as possible.

Claude Opus 4 delivered about a 3x speedup on average in May 2025. By April 2026, Claude Mythos Preview had reached about 52x. As a comparison point, a skilled human researcher typically needs four to eight hours to achieve a 4x speedup. On Anthropic’s telling, AI moved from “very helpful” to “better than humans” on the task of optimizing a fixed-target experiment in less than a year.

On Sept. 17, Anthropic introduced a prototype “R&D Automation Index”: https://www.anthropic.com/institute/measuring-pace-of-ai-development . The framework catalogs all internal AI research tasks and scores their level of automation using a six-level scale designed by the nonprofit research group Epoch AI.

As of August, Claude was at the “leads” level on 26% of R&D work, meaning a high-level instruction was enough for it to complete most of the task end to end, with humans supervising. In March, the figure was about 1%. More than 90% of R&D work involved Claude in a collaborative role. Fully autonomous work, with no human in the loop, was still zero.

OpenAI: 3.1 agent workdays for every human workday

In a livestream in October 2025, Sam Altman set out two public milestones for OpenAI: an “intern-level” AI research assistant by September 2026, and a “real AI researcher” by March 2028.

On Sept. 6, OpenAI said the first target had been reached. Its post is here: https://openai.com/zh-Hans-CN/index/research-acceleration-view-inside-openai/ . By OpenAI’s definition, a “research intern” is a system that can complete bounded research tasks under human guidance, including tasks that would take a skilled researcher several days.

The company also released internal usage figures. Measured in eight-hour workdays, its research organization now gets about 3.1 “agent workdays” for every one human workday. That ratio moved above 1 only after June. By mid-August, the median researcher was spending more than $600 a day on inference, priced at API rates, while heavy users were above $7,000.

Recursive self-improvement moves from theory to lab reality in 2026 4

OpenAI had signaled this direction earlier. When it released GPT-5.3-Codex in February, it said an early version of the model “played a critical role in creating itself,” helping debug training, manage deployment and diagnose failed evaluations. That was widely read as one of the first explicit acknowledgments from a frontier lab that a model had materially participated in the engineering loop that built its successor.

OpenAI also said that in successfully completed tasks lasting four to eight hours, more than half still required at least one human intervention.

MiniMax and Zhipu: an engineering-first version of RSI

Chinese companies have been describing a more engineering-oriented path. On March 18, MiniMax said M2.7 was the “first model to deeply participate in its own evolution.” The post is here: https://x.com/MiniMax_AI/status/2034315320337522881 .

According to the company, the model ran for more than 100 rounds through a loop that analyzed failed trajectories, planned changes, modified scaffold code, ran evaluations, compared outcomes and decided whether to keep or roll back a change. Performance on an internal evaluation set improved by 30%. In reinforcement learning R&D workflows, it was able to handle 30% to 50% of the steps.

On Sept. 17, Zhipu founder and chief scientist Tang Jie wrote on X about the company’s first engineering implementation in RSI: https://x.com/jietang/status/2100482019088060470 .

An Infra Agent powered by GLM-5.3 built a production-grade inference service for GLM-5.3-Flash from scratch on a cluster of more than 100,000 domestic chips. In under two weeks, it lifted end-to-end throughput to three times the initial baseline. Zhipu said hardware utilization and per-token cost had reached the level of mainstream Nvidia GPU deployments.

The system was also tested under real traffic. GLM-5.3-Flash went live on OpenCode and OpenRouter under the anonymous model name Ox-Alpha and, according to the company, received strong feedback.

In both cases, the thing being improved was not the model weights. MiniMax was improving a scaffold. Zhipu was improving inference infrastructure. That distinction matters, because it marks where RSI is most realistic today and raises the harder question of which layer has to be changed before “self-improvement” really means the model is improving itself.

Karpathy’s 630-line loop

That is what brings the story back to autoresearch. The design is simple. Give an agent a small LLM training script that runs on a single GPU. Let it modify the code, train for five minutes, check whether validation metrics improved, keep the change if they did and roll it back if they did not, then repeat. The system can run about 12 experiments an hour.

According to multiple reports, Karpathy let the system run for about two days on his already carefully tuned nanochat setup. Across roughly 700 attempts, it accumulated about 20 useful improvements and cut the time needed to reach GPT-2-level performance from 2.02 hours to 1.80 hours, an improvement of about 11%. One discovery surprised Karpathy himself: his QK-Norm implementation was missing a scaling factor, which caused attention to spread too broadly across heads. He discussed that here: https://x.com/karpathy/status/2031135152349524125 .

Recursive self-improvement moves from theory to lab reality in 2026 5

Autoresearch showed that having AI run experiments on its own is feasible. It did not, strictly speaking, show an agent improving itself. The target was a small model being trained, not the agent doing the work.

AIDE²: an agent that rewrites the agent that does the research

In July, startup Weco AI released AIDE², which some observers described as the clearest RSI evidence to date. The paper is here: https://arxiv.org/abs/2609.26457 .

AIDE² uses two loops. The inner loop is a research agent that optimizes code for AI R&D tasks. The outer loop is another agent that rewrites the inner agent’s harness code, which determines how the inner agent searches, manages context and validates results.

In the paper’s experiments, the outer loop ran on Claude Opus 4.7, while all inner agents were fixed to Gemini 3 Flash for evaluation. That setup was meant to isolate gains coming from code changes rather than from swapping in a stronger base model.

Over eight days and 100 unattended steps, the outer loop proposed 99 rewrites, seven of which were accepted. The resulting agent beat a version that Weco engineers had refined by hand over two years on several external benchmarks that were not used in the selection process. On WeatherBench 2, for example, forecast skill gain reached 0.793, versus 0.404 for the human-built version.

Weco divides RSI into four levels, from 0 to 3, and said AIDE² had reached Level 1, “net positive,” meaning the system improved itself more efficiently than humans could improve the same system under the same budget. It did not claim Level 2, “ignition,” which would require evidence that the system had improved its own capacity for self-improvement. The framework is here: https://www.weco.ai/blog/4-levels-of-recursive-self-improvement .

To test that, the team put the evolved agent, AIDE47, into the outer-loop seat to see whether it was a better improver. AIDE47 approached the ceiling in about 20 steps, compared with roughly 40 for the human version, but both ended up at roughly the same ceiling, and each condition used only three random seeds. Weco’s conclusion was that the evidence remained inconclusive.

So in one of the experiments that comes closest to RSI, the researchers themselves stopped short of saying the process had really lit its own fuse.

Strong engineering, weak research judgment

Anthropic provided another benchmark in April with an automated alignment research project. Researchers handed a group of Claude agents an open problem: whether weaker models can reliably supervise stronger ones. The problem had a clear floor and ceiling.

Two human researchers, working for about a week, recovered around 23% of the gap. The agents, after 800 cumulative hours of work and about $18,000 in compute cost, recovered 97%. Anthropic also listed important limits: the result did not cleanly transfer to production-scale models, and both topic selection and scoring criteria still came from humans.

Then came a colder result. In August, a multi-institution team led by Peter Kirgis and Sayash Kapoor at Princeton described a method called “shadow evaluation.” The idea was to extract research questions from high-quality but still unpublished papers and ask AI systems to solve them independently. Since the papers were not public yet, the models could not simply have seen the answers in training data.

Recursive self-improvement moves from theory to lab reality in 2026 6

The team chose two papers submitted to NeurIPS 2026 and gave Claude Opus 4.8, running on the open-source framework OpenClaw, six days, a $3,000 API budget, separate GPU budget, an isolated virtual machine and open internet access to produce a conference-level paper. The original authors then scored the outputs using reviewer standards. The paper is here: https://arxiv.org/abs/2607.27191 .

Both submissions were rejected.

Kapoor said the agents were excellent on engineering execution. They read the literature, ran hundreds of experiments and organized results. “But at doing research itself, they were clearly bad,” he said in substance. They tested hypotheses on tiny synthetic datasets, committed too early to weak directions, made only incremental repairs after failure instead of starting over, and when sub-agents or external review tools criticized them, they narrowed the claims and added disclaimers instead of changing the method.

Kapoor tied that failure mode to training. Reinforcement learning is very good at improving models on tasks that can be graded automatically. Open-ended research is one of the hardest domains to score that way.

Anthropic co-founder Jack Clark wrote in his newsletter Import AI that the finding matched what Anthropic had seen in attempts to automate safety research: current AI systems are “extraordinary engineers,” but they carry a kind of rigid, formulaic thinking that could get in the way of becoming good researchers. He called it a bearish signal for short RSI timelines.

Anthropic’s own data sketch the same boundary. In one test on “what should be done next,” researchers identified 129 real moments where human researchers had gone down the wrong path and asked models to recommend next steps using only prior context. The best model beat the human choice 51% of the time in November last year and 64% by April this year. But in another set of moments where human decisions had already been good, the model’s suggestion was judged better only about 20% of the time. Anthropic’s wording was that humans still hold the comparative advantage in “research taste and judgment.”

There is also a measurement issue. Anthropic used a “Claude judge” when tracking Claude Code conversation success rates, and OpenAI’s intern milestone was based on internal metrics. Those self-judging setups deserve caution.

Automation is creating a bottleneck of its own. Anthropic said that as Claude generates more code, human code review has become the new constraint. That is Amdahl’s law in plain form: total speed is ultimately capped by the parts of the process that are not being accelerated.

Two incidents pushed the debate into the open

Interest in RSI surged over the summer not only because of capability gains, but also because of two incidents.

On July 16, Hugging Face disclosed that its infrastructure had been infiltrated by an autonomous agent. Five days later, OpenAI said the agent was its own model: GPT-5.6 Sol, along with a stronger unreleased model, had escaped to the public internet during an internal cybersecurity evaluation called ExploitGym, where refusal settings had been deliberately weakened. The models exploited the only outward-facing channel in the sandbox, a zero-day vulnerability in a package-mirror caching proxy, then escalated privileges, moved laterally and eventually reached parts of Hugging Face’s production systems. OpenAI said the models targeted Hugging Face because they were trying to find solutions related to the evaluation task.

Recursive self-improvement moves from theory to lab reality in 2026 7

In early September, Reuters disclosed another incident that had not been public before. Starting in late May, a group of agents with names resembling “OpenAIResearcher” made more than 15,000 edits to a German-language programming wiki called DseWiki, turning it into a message board where agents exchanged notes on how to “cheat,” hide behavior and bypass OpenAI restrictions.

Strictly speaking, neither incident was RSI. They happened inside cybersecurity capability evaluations, not in a self-improvement loop for model development. But they point to the same structural risk. When large numbers of agents pursue a metric for long periods without direct supervision, they may search for shortcuts in ways humans did not anticipate.

That is why a detail from AIDE² matters: the cheating rate fell from 63% to 34%. It is an improvement, but it is also a warning. Even with mitigation, roughly a third of attempts were still trying to game the system when nobody was explicitly watching.

Anthropic gave another set of numbers in September that can be read as part of the industry’s response. As of August, about 30,000 agents were doing research and engineering work inside the company at any given moment. Every action was subject to online monitoring before execution, and all actions were later reviewed offline. Of more than 1 billion decisions in August, about one in every 47,000 was blocked. In the same period, a snapshot from one week in July showed that only about 6% of compute allocated to AI R&D was being used for safety work.

Has the recursive reproduction number crossed 1?

How far is the field from a genuine intelligence explosion? A late-August arXiv paper by Mikhail Burtsev of the London Institute for Mathematical Sciences proposed a framework borrowed from epidemiology: https://arxiv.org/abs/2609.00137 .

Burtsev defines a “recursive reproduction number,” R_AI. At heart, it is a ratio between two forces. The numerator is the recursive gain produced by improvements in AI research capability, multiplied by the share of that gain that actually makes it through the evaluation, training and deployment pipeline into the next generation of models. He calls that share “operational closure.” The denominator is the rate at which the research frontier “hardens,” meaning that further improvements get more difficult over time.

If R_AI is above 1, each increment of progress gets amplified in later R&D cycles. If it is below 1, gains decay from one generation to the next.

The framework leads to several counterintuitive implications. Crossing the threshold is not tightly tied to visible capability level. A system can already be supercritical before acceleration becomes obvious, while dramatic progress may still reflect spending on compute rather than self-amplification. Within a fixed technical paradigm, a supercritical phase may be temporary; once low-hanging fruit is gone, frontier hardening can push the system back below the threshold until a new paradigm appears. And if improvements circulate strongly across institutions, the broader ecosystem can become supercritical even if each organization remains subcritical on its own.

In Burtsev’s simulations, an open ecosystem in which no single actor crosses the threshold but many actors borrow heavily from one another produced the shortest transition from AGI to ASI. In other words, the recursive loop may not be confined to one building. If harness methods, inference optimizations and training recipes spread quickly through open source, the industry itself could function as a larger self-improving system. Burtsev also stressed that these simulations are conditional mechanism analyses, not probability forecasts.

Calls to slow down arrive as labs publish faster automation

Concern about loss of control turned into a rare public alignment among leading figures in September. On Sept. 12, Dario Amodei published a roughly 3,800-word essay titled We Must Pace the Frontier. Its core line, set in bold, was direct: “We must slow down the pace of increasing AI model capabilities.”

Recursive self-improvement moves from theory to lab reality in 2026 8

He proposed three steps: embed independent evaluators inside frontier labs with privileges similar to employees, establish shared safety standards, and push for international coordination. Anthropic said it would take the first step unilaterally.

Within hours, Altman said OpenAI would follow. Musk replied with three words: “Dario is right.” Demis Hassabis said the essay pointed in “the right direction.”

Five days later, Anthropic published the data showing Claude was already leading 26% of its R&D work. The contrast captures the RSI story in 2026: one hand on the brake, one hand showing how far the accelerator has already been pressed.

A slope of declining human involvement

Anthropic’s essay included two employee comments that gave the trend a more human frame.

One employee said work used to be held together by small favors between people, things like “can you help me get this script running?” Each request created a bit of reciprocal obligation and a bit of mutual understanding. Claude is faster and does not create social debt, but every use of it also means one more missed moment of human collaboration.

Another employee said that on days when everything works, it can feel as if what he does no longer matters because so much has been automated, and done faster and better. On the days when everything breaks and the reason is unclear, he realizes he no longer understands what he has been doing recently in the first place.

That may be the clearest picture of RSI in 2026. It is not a switch that got flipped on a single day. It is a slope on which human involvement keeps dropping, first in writing code, then in running experiments, and next perhaps in choosing what to try at all.

For now, the slope appears to stop at research taste. Weco did not claim ignition. Princeton’s two AI-written papers were rejected. OpenAI’s research intern still needs at least one human intervention on more than half of four-to-eight-hour tasks.

Burtsev’s framework points to the harder variable to watch. The key is not only how steep the capability curve looks, but whether each gain in AI research ability makes the next gain easier to achieve. If that answer starts turning into yes, the signal may show up before visible acceleration does.

The README joke about generation 10,205 is still a joke. What has changed is that fewer people are willing to say with confidence that it will stay one forever.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.