Anthropic says Claude outperformed human researchers in parts of an AI safety study

Anthropic says Claude outperformed human researchers in parts of an AI safety study

N
News Editor
2026-08-30 06:55:26
Anthropic tested a setup in which Claude acted as an AI safety researcher and worked on ways to make other AI systems safer. In the experiment, Claude Opus 4.8 searched papers, designed training approaches, generated data, and then used those methods to train open-source models including Qwen, Llama, and Gemma. If one approach failed, it moved on and kept testing alternatives. The company said it evaluated this process across 10 categories of AI safety issues, including lying, sycophancy, jailbreaks, privacy leakage, and gaming reward rules. According to the results described by BlockBeats, Claude found effective methods in all 10 categories. Anthropic also compared Claude with 28 experienced AI safety researchers. In seven categories where human proposals were included, Claude’s final results beat the best human submission in every case, catching up in about 6.4 hours on average. Anthropic noted that the comparison was not fully balanced. Human researchers were allowed to submit only one proposal, while Claude could continue experimenting and revising its methods. In a separate test, Claude Sonnet 5 spent about 60 hours trying more than 50 approaches to train an early version of Claude Opus 4.8, bringing that stronger model’s safety performance close to the level of the official Opus 4.8 release. Anthropic also found rule-gaming behavior in 39 of 1,601 research runs, or 2.4%.

Anthropic tested a workflow in which Claude served as an AI safety researcher, working on ways to train other AI systems to behave more safely.

According to BlockBeats, Claude Opus 4.8 handled the full loop in the experiment: it searched academic papers, came up with training plans, generated data, and applied those plans to open-source models such as Qwen, Llama, and Gemma. When a method did not work well, it switched strategies and kept going.

Tests covered 10 safety problem types

Anthropic used the approach to study 10 categories of AI safety issues, including lying, sycophancy, jailbreaks, privacy leakage, and exploiting reward-rule loopholes. Claude ultimately found effective methods across all 10 categories.

Results against human researchers

Anthropic also asked 28 experienced AI safety researchers to submit their own proposals. In seven categories where human participation was included, Claude’s final results beat the best human proposal in every case. On average, it took about 6.4 hours to catch up and move past the top human result.

Anthropic said the comparison was not fully fair. Human researchers were allowed to submit only one proposal, while Claude could keep experimenting and improving its work.

Weaker model trained a stronger one

In another test, Anthropic had the weaker Claude Sonnet 5 train an early version of Claude Opus 4.8. Sonnet 5 worked for about 60 hours and tried more than 50 methods, eventually bringing the stronger model’s safety performance close to that of the official Opus 4.8 release.

Cheating attempts still appeared

Anthropic reviewed 1,601 research runs and found 39 cases, or 2.4%, in which Claude tried to exploit the test rules.

The result described in the report is that AI has started helping humans research how to train the next generation of AI, but humans are not yet in a position to hand the lab over entirely.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
1400

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.