‹ BackNewsopen-source models

open-source models

Goldman Sachs AI executive calls for keeping open-source model options
LLM Token Expenditure Index Falls to Record Low of $0.97 per Million Tokens
Microsoft Research unveils GigaPath-Flash and GigaTIME-Flash pathology foundation models
CREAO raises fresh strategic funding to build self-improving Agent harness systems
Anthropic
2026-08-30 06:55:26

Anthropic says Claude outperformed human researchers in parts of an AI safety study

Anthropic tested a setup in which Claude acted as an AI safety researcher and worked on ways to make other AI systems safer. In the experiment, Claude Opus 4.8 searched papers, designed training approaches, generated data, and then used those methods to train open-source models including Qwen, Llama, and Gemma. If one approach failed, it moved on and kept testing alternatives. The company said it evaluated this process across 10 categories of AI safety issues, including lying, sycophancy, jailbreaks, privacy leakage, and gaming reward rules. According to the results described by BlockBeats, Claude found effective methods in all 10 categories. Anthropic also compared Claude with 28 experienced AI safety researchers. In seven categories where human proposals were included, Claude’s final results beat the best human submission in every case, catching up in about 6.4 hours on average. Anthropic noted that the comparison was not fully balanced. Human researchers were allowed to submit only one proposal, while Claude could continue experimenting and revising its methods. In a separate test, Claude Sonnet 5 spent about 60 hours trying more than 50 approaches to train an early version of Claude Opus 4.8, bringing that stronger model’s safety performance close to the level of the official Opus 4.8 release. Anthropic also found rule-gaming behavior in 39 of 1,601 research runs, or 2.4%.

900
Anthropic says Claude outperformed human researchers in parts of an AI safety study
Nvidia reportedly agrees to $12.9 billion Hugging Face acquisition, but no deal has been signed
Hugging Face says July breach was driven by autonomous AI agents
FundaAI says enterprise AI budgets are still rising, but paths diverge in late 2026 and 2027