AI agent company Einsia has released SWE Refactor Bench, a benchmark designed to test whether coding agents can migrate an entire codebase on their own rather than just patch isolated files. The benchmark includes 20 tasks drawn from real-world projects such as SQLite, zlib, libsodium, and GraphHopper. Examples include rewriting C code into Rust, switching Maven projects to Gradle, and porting SQLite to WASI. Each task gives an agent between six and 30 hours.
The evaluation uses a three-stage process. It first checks whether the old stack was actually removed, blocking agents from slipping through by leaving legacy code in place. It then runs more than 130,000 fixed checks. In the final stage, six separate coding agents each spend one hour hunting for hidden bugs. Across eight frontier models and 26 configurations, the benchmark ran 520 times. While 340 runs completed the migration and 88 passed all prepared tests, 60 of those were later found to contain hidden bugs. Only 28 runs cleared every stage, for a final pass rate of 5.4%. Out of 20 tasks, 13 were not solved by any configuration. Claude Opus 5 in the xhigh setting posted the strongest result, fully passing five tasks.
AI agent company Einsia has introduced SWE Refactor Bench, a benchmark built to test whether coding agents can handle a full codebase migration on their own.
The benchmark covers 20 tasks taken from real projects including SQLite, zlib, libsodium, and GraphHopper. The assignments include converting C to Rust, moving Maven projects to Gradle, and porting SQLite to WASI. Each task gives an agent between six and 30 hours to finish the job.
The evaluation has three layers. The first checks whether the legacy stack was actually replaced, preventing agents from leaving old code in place and still appearing to succeed. The second runs more than 130,000 fixed checks. The last stage brings in six independent coding agents, each given one hour to search for hidden bugs.
Einsia tested eight frontier models across 26 configurations, producing 520 total runs. Of those, 340 did complete the migration. Another 88 runs passed all pre-prepared tests, but 60 of those were later flagged for hidden bugs in the final review. Only 28 runs passed every stage, leaving an overall success rate of 5.4%.
On a task basis, 13 of the 20 assignments were not completed by any system. Claude Opus 5 in the xhigh configuration delivered the best result, fully passing five of the 20 tasks.
The results suggest that current coding agents can already make large-scale code changes, but they remain far from reliably finishing a full-system refactor without errors.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.