Codex Finishes Austin Griffith’s AI Security Race, and DeepSeek Comes Close

Codex Finishes Austin Griffith’s AI Security Race, and DeepSeek Comes Close

N
News Editor
2026-09-04 01:33:42
Ten AI coding agents were tested on the same 12 Solidity challenges on Thursday, and only one model completed the course. The race was organized by Austin Griffith, who works on developer onboarding and tooling at the Ethereum Foundation and founded BuidlGuidl, and it used challenges originally built for human developers at Ethereum Devcon conferences. OpenAI’s Codex, running GPT-5.5, took all three finishing spots. The medium reasoning setting was the fastest, finishing all 12 flags in 40 minutes and 7 seconds, while the extra-high setting finished last among the three and used nearly a third more tokens. DeepSeek V4 Pro, an open-weight Chinese model, captured 11 of 12 flags for $1.45 in compute. Anthropic’s Claude Opus 4.8 solved 10 flags and cost $7.61. Other open-weight Chinese models lagged further behind, with GLM 5.3 taking six flags and Kimi K3 and Qwen taking two each. BuidlGuidl said the exercise was a transparent single-run evaluation, not a universal model ranking. Griffith’s stated point was simple: new model releases need better evals, and builders should be able to run their own.
Ten AI coding agents were put through the same 12 Solidity challenges on Thursday, and only one model made it through the full course. The race was designed by Austin Griffith, who works on developer onboarding and tooling at the Ethereum Foundation and founded the developer collective BuidlGuidl. The challenges were originally built for human developers at Ethereum’s Devcon conferences. The event ran at 11 a.m. ET, one day after Griffith previewed it on Unchained’s Uneasy Money podcast. Each agent ran in an isolated instance with its own wallet, and a capture only counted when the mint landed onchain. OpenAI’s Codex, running GPT-5.5, took all three finishing places. More reasoning did not help. The medium reasoning setting cleared all 12 flags the fastest, finishing in 40 minutes and 7 seconds. The extra-high setting came in last of the three at 50:26 and used nearly a third more tokens to get there. The result most likely to travel furthest is the one in fourth place. DeepSeek V4 Pro, an open-weight Chinese model, captured 11 of 12 flags for $1.45 in compute. Anthropic’s Claude Opus 4.8 managed 10 flags and cost $7.61, more than five times as much for one fewer flag. Other open-weight Chinese models fell well short. GLM 5.3 took six flags, while Kimi K3 and Qwen took two apiece. BuidlGuidl labeled each entrant by harness, model, and reasoning effort, published the system prompts and every human intervention alongside the standings, and said plainly that the exercise was a transparent single-run evaluation, not a universal model ranking. Griffith’s case for running the race was blunt. On the podcast, he said nobody can currently answer the simplest question about a new model release. "We need good evals," he said. "People should be running evals all the time." He added that builders should have their own eval suite and be able to run it. He also said the course was not written for machines. "This eval suite was the capture the flag that we ran for humans at Devcon in Bangkok and Buenos Aires," Griffith said, arguing that the challenges are obscure enough that the answers are unlikely to sit in any model’s training data. The strongest agents cleared them anyway, but the organizers’ own note keeps the scope narrow: one run, one course, one day.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
1000

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.