Anthropic published a report on June 4, 2026, showing that its latest model, Mythos Preview, outperformed human experts in 64% of AI research decision tests. Two years earlier, in 2024, the same metric stood at just 22%—a nearly threefold jump in judgment capability.
Research Decision Test: 64% Win Rate
The team conducted a test: they showed Claude transcripts where human researchers were about to head down a wrong path, then asked, "What should we do next?" Mythos Preview made better decisions 64% of the time. In contrast, the 2024 model managed only 22%, suggesting AI is gaining the ability to guide high-level research.
Code Optimization: 52x Speed Leap
Beyond decision-making, coding efficiency surged. Internal data shows engineers now deliver 8x more code per quarter than the 2021–2025 average. On open-ended coding tasks, Claude's success rate climbed 50 percentage points in six months to 76%. Many engineers consider its code quality nearly human-level.
In a standard benchmark—optimizing training code for a small AI model—a skilled human takes 4–8 hours to achieve roughly 4x speedup; the 2024 Claude Opus 4 averaged 3x. But Mythos Preview delivered an astonishing 52x speedup, shattering previous efficiency limits.
RSI Acceleration: Revolution and Risk
Despite impressive numbers, Anthropic cautioned that recursive self-improvement (RSI) is not guaranteed—it remains unclear if Claude truly chooses the "right research questions" autonomously. Yet if the trend holds, AI systems designing their own successors becomes plausible.
The company acknowledged potential revolutionary benefits in medicine, technology, and the economy, but also warned of intensified alignment problems, culminating in loss-of-control risks. To manage this, Anthropic announced the formation of the Anthropic Institute, in collaboration with external stakeholders, to deeply study the far-reaching impacts of powerful AI and ensure cautious human decision-making.

