OpenAI and Paradigm Launch EVMbench to Assess AI as DeFi Security Guardian

OpenAI and Paradigm Launch EVMbench to Assess AI as DeFi Security Guardian

N
News Editor 01
2026-07-23 18:30:15
OpenAI and Paradigm unveiled EVMbench, a benchmark to test AI agents on detecting, patching, and exploiting vulnerabilities in Ethereum smart contracts. The dataset includes 120 high-severity bugs from 40 audit projects.
OpenAIParadigmEVMbenchsmart contract securityDeFi

OpenAI, led by Sam Altman, has teamed up with crypto investment giant Paradigm to release EVMbench, a new benchmark designed to rigorously evaluate how AI agents perform on smart contract security tasks across Ethereum Virtual Machine (EVM)-compatible chains. The initiative aims to establish clearer AI evaluation standards for blockchain security, responding to the rising asset protection demands in decentralized finance (DeFi).

Smart contract vulnerabilities remain acute amid $100B+ value under protection

Smart contracts power decentralized exchanges, lending platforms, and stablecoin payments. These immutable, on-chain programs currently guard over $100 billion in open-source crypto assets. Past exploits have caused massive losses, making security auditing one of the most pressing issues in the blockchain industry.

EVMbench draws from real-world cases: 120 high-severity vulnerabilities collected from 40 audit projects, most from public audit contests like Code4rena, plus additional bug scenarios from Paradigm-backed Tempo blockchain payments.

Three tasks test detection, patching, and exploitation skills

The benchmark covers three core capabilities. Detect – AI agents review smart contract code to locate known vulnerabilities, scored by severity and audit reward. Patch – AI must fix vulnerable contracts while preserving original functionality, verified by automated tests and attack simulations. Exploit – In a sandbox blockchain environment, AI attempts a full fund-draining attack, with success validated programmatically. The combined score provides a percentage-based metric for comparing AI models on security tasks.

OpenAI noted in its blog that as AI agents grow more capable of reading, writing, and executing code, their defensive role in high-stakes environments becomes increasingly critical. EVMbench not only challenges AI limits but also encourages the industry to apply AI for proactive auditing and hardening of live contracts. The benchmark aligns with OpenAI's Preparedness Framework, which identifies high-risk cybersecurity scenarios.

With DeFi and stablecoin payments expanding, reliable AI-assisted detection and patching could significantly improve ecosystem security. At the same time, EVMbench highlights the need to tightly regulate AI's exploitation ability to prevent misuse. As models advance, EVMbench may become a key yardstick for whether AI can truly safeguard digital assets.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.