Nonprofit AI evaluation group ARC Prize has previewed ARC-AGI-4, saying the next benchmark will focus on "autonomous open-ended innovation" rather than only measuring performance on fixed tasks. The new test is intended to examine whether AI systems can independently explore, generate original ideas, and potentially produce new inventions or discoveries.
ARC Prize has not yet disclosed the benchmark’s specific tasks, scoring method, or release date. In its preview, the organization said humans still hold a clear lead over AI in open-ended innovation, which it described as one of the most important capabilities behind scientific and technological progress.
The announcement also cited a newly released AI slowdown essay by Dario Amodei. ARC Prize said open source remains a foundation for AI progress and pushed back against efforts to reduce openness in the name of coordinated slowdown, warning against concentrating frontier AI in the hands of a small number of institutions.
ARC-AGI has long been described as one of the world’s hardest AGI evaluations. When ARC-AGI-3 was released, humans scored 100% while frontier AI scored 0.51%. ARC Prize added that GPT-6 Astra has recently closed much of that gap, reaching 62.7% under a unified Standard harness and 99.9% only after using OpenAI’s own Provider Adapter.
ARC Prize, a nonprofit AI evaluation organization, has previewed ARC-AGI-4, saying the next benchmark will focus on "autonomous open-ended innovation" and test whether AI can independently explore, come up with new ideas, and even make new inventions or discoveries.
The group has not disclosed the benchmark’s specific tasks, scoring framework, or launch date. ARC Prize said humans still maintain a clear lead over AI in open-ended innovation, which it called one of the most important capabilities behind scientific and technological progress.
The preview also cited a newly published AI slowdown essay by Dario Amodei. ARC Prize said open source remains a core foundation of AI progress and opposed reducing openness in the name of coordinated slowdown, arguing against concentrating frontier AI within a small number of institutions.
Prior ARC-AGI results
ARC-AGI has long carried a reputation as one of the world’s hardest AGI evaluations. According to ARC Prize, when ARC-AGI-3 was released, humans scored 100% while frontier AI scored 0.51%. The organization added that the latest GPT-6 Astra has narrowed the gap sharply, but it still reached only 62.7% under a unified Standard harness. Its score rose to 99.9% only after connecting through OpenAI’s own Provider Adapter.
With ARC-AGI-4 now teased, ARC Prize is preparing to move the next stage of testing toward a harder question: whether AI can generate innovation on its own.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.