OpenAI and Paradigm Unveil EVMbench to Test Ethereum Contract Security

OpenAI and Paradigm Unveil EVMbench to Test Ethereum Contract Security

N
News Editor 01
2026-07-23 09:05:14
OpenAI and Paradigm introduced EVMbench, a benchmark built from 120 real high-severity smart contract flaws to test AI in detection, patching, and exploitation. GPT-5.3-Codex scored 72.2% in exploit mode.
OpenAIParadigmEthereum securitysmart contractsAI auditing

OpenAI, in collaboration with Paradigm, has introduced EVMbench, a benchmark designed to measure how AI agents handle Ethereum smart contract security. The system tests whether models can detect vulnerabilities, patch them, and exploit them under controlled conditions. The release targets a large attack surface: smart contracts across EVM networks secure more than $100 billion in crypto assets.

Benchmark built from 120 real high-severity flaws

According to OpenAI, EVMbench is based on 120 high-severity vulnerabilities collected from 40 professional smart contract audits. The benchmark uses real issues instead of synthetic examples. A significant share of the cases came from open audit competitions, including Code4rena.

The dataset also includes scenarios tied to security work on Tempo, a payment-focused Layer 1 network designed for stablecoin transfers. Those examples bring payment logic risks into the benchmark, expanding the range of contract behaviors under test.

To keep the environment realistic, engineers reused exploit proof-of-concept scripts where those were available. Where documentation was incomplete, they manually rebuilt missing components. OpenAI said the goal was to preserve exploitability while making sure patch submissions could still compile correctly.

Detect, patch, and exploit modes measure different abilities

EVMbench evaluates agents in three modes: detect, patch, and exploit. In detect mode, agents scan repositories and are scored on confirmed vulnerability recall. In patch mode, the task is to fix the flaw without changing the original behavior of the contract. That makes the benchmark more demanding than a simple pass-fail bug fix test.

Exploit mode simulates a full fund-draining attack inside a sandbox blockchain. OpenAI said graders verify outcomes through transaction replay and on-chain state checks. For consistency across runs, the company built a Rust-based harness to support deterministic deployments.

Exploit testing runs locally, not on live networks

The exploit tests run in a local Anvil environment rather than on public chains. OpenAI stated that every vulnerability in the set is historical and already publicly disclosed. The harness also limits unsafe RPC calls, a constraint intended to reduce misuse.

In the published results, GPT-5.3-Codex scored 72.2% in exploit mode. That was well above the 31.9% reported for GPT-5, which had launched months earlier. OpenAI added that coverage in detection and patching is still incomplete, indicating that model performance remains uneven across different parts of the smart contract security workflow.

OpenAI also confirmed a new hire

Alongside the benchmark launch, OpenAI confirmed that OpenClaw founder Peter Steinberger has joined the company to work on agent development. Sam Altman later confirmed the move on X and said Steinberger would lead next-generation personal agent projects.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.