Anthropic’s June 4 research essay, When AI Builds Itself, opened with a striking claim: by May 2026, Claude had written more than 80% of the merged code in Anthropic’s product codebase. In the same document, the company argued that the world should have an option to verifiably slow or temporarily pause frontier AI development when needed. Those two messages landed together and immediately fueled debate.
According to Anthropic, Claude’s share of merged code was still in the single digits before Claude Code launched in early 2025. By 2026, the change was visible in internal engineering output. The company said the average number of lines of code merged per engineer per quarter rose by 8x from Q2 2024 to Q2 2026. In a separate internal survey of 130 research staff, the median estimate for productivity gains from Mythos Preview was 4x.
Software task duration expanded from 4 minutes to 12 hours
Anthropic laid out a timeline for Claude’s growing autonomy on software work. In March 2024, Claude Opus 3 could independently handle a software task that would take a human about 4 minutes. By March 2025, Claude Sonnet 3.7 pushed that to 90 minutes. By March 2026, Claude Opus 4.6 had reached 12 hours.
The company described that as more than steady growth. It said the doubling time for task duration had compressed from 7 months to 4 months. In one example, Anthropic said Claude independently completed more than 800 API bug fixes in April 2026 and cut one class of errors by a factor of 1,000. One engineer estimated the same workload would have taken humans four years.
Anthropic also highlighted research-side results. It said two human researchers spent one week recovering 23% of a performance gap on an AI safety problem. A group of Claude agents, after 800 cumulative hours and roughly $18,000 in compute, recovered 97%. As of May 2026, Anthropic said Claude’s code quality had reached parity with human engineers. The company’s wording was direct: at the end of 2025, Claude wrote worse code than humans; now it is at parity, and Anthropic expects it to be strictly better within a year.
The pause proposal arrived days after an IPO filing
Anthropic paired those capability claims with a call for a global, verifiable slowdown mechanism. The source material says the company filed for an IPO on June 1 at a $965 billion valuation, then on June 4 publicly argued for slowing or pausing frontier AI work if necessary. The timing drew immediate scrutiny.
Co-founder Jack Clark added a concrete estimate of his own: he puts the probability of AI reaching recursive self-improvement by the end of 2028 at 60%. Critics were not persuaded. Bentley University mathematics professor Noah Giansiracusa told Scientific American that he does not believe Anthropic truly wants to slow down, arguing that Dario Amodei’s practical stance is to move at full speed because a pause is impossible to enforce in the real world. Georgia Tech professor Mark Riedl was blunter, saying major AI firms have all climbed aboard the recursive self-improvement hype train.
The core dispute is what the data actually proves
New York University professor Gary Marcus argued that Anthropic was blending two separate ideas. One is AGI, where AI can autonomously do everything humans can do. The other is the present reality: AI as a very fast and capable coding tool that multiplies human output. Marcus’ point was that Anthropic’s evidence fits the second category. Claude may be writing most of the code, but humans still set goals, choose directions, and review results.
Anthropic’s own numbers partly support that view. The paper said Claude’s accuracy in choosing the next research direction improved from 51% in November 2025 to 64% in April 2026. That is progress, but it still means more than one wrong choice out of every three. An anonymous Anthropic employee quoted in the source said humans still hold an advantage in seeing the bigger picture and thinking beyond the immediate task.
Why a global AI brake is hard to verify
Anthropic compared its proposed mechanism to Cold War arms control, specifically the INF Treaty. Yet the same discussion makes clear why the analogy is hard to carry into AI. Missile silos can be monitored by satellite. Model training does not look like that. Training can happen in ordinary office settings while compute infrastructure may sit elsewhere, making detection far harder.
Anthropic also attached a condition to any slowdown: it expects other frontier developers to do the same in a verifiable way. That brings the issue back to strategy. If no one believes rivals will stop, each player has a reason to keep pushing ahead. In the material provided, that is the central tension: Anthropic is presenting rapid gains in Claude’s code generation while calling for a system-wide brake at the same moment.

