Specific Labs, an AI data company backed by Y Combinator, has launched Real-SWE, a programming benchmark designed to test coding agents against real private codebases from operating companies. The tasks come directly from production environments and include tax calculation, billing migration, API billing, and customer data migration. Because neither the code nor the answers are publicly available online, agents have to work out each company’s business rules and code structure on their own.
In the benchmark results shared in the report, Fable 5.1 ranked first with a 38.8% pass rate. GPT-6 Astra followed at 33.8%, while Gemini 3.8 Flash posted 31.2%. GLM 5.3 came in fourth at 28.8%, outperforming Grok 4.6, Kimi K3, and GPT-5.6 Sol. Among the 10 tasks currently disclosed, six had overall pass rates below 15%, and one task saw every model fail.
According to Xiaopu Peng, a researcher at Zhipu, the team has been training long-horizon tasks since GLM-5.1. He said Real-SWE is closer than public SWE leaderboards to measuring that type of capability, which may help explain GLM 5.3’s stronger showing in this benchmark.
Specific Labs, an AI data company backed by Y Combinator, has introduced Real-SWE, a programming benchmark that uses private code from real companies to test coding agents.
The tasks are drawn directly from enterprise production environments. They include tax calculation, billing migration, API billing, and customer data migration. Since neither the code nor the answers are publicly available online, agents have to figure out a company’s business rules and codebase structure by themselves.
In the published results, Fable 5.1 ranked first with a 38.8% pass rate. GPT-6 Astra posted 33.8%, and Gemini 3.8 Flash reached 31.2%. GLM 5.3 placed fourth at 28.8%, ahead of Grok 4.6, Kimi K3, and GPT-5.6 Sol.
Of the 10 tasks currently made public, six recorded overall pass rates below 15%. One task was not completed by any model.
BlockBeats said Real-SWE puts more weight on whether an agent can keep reading code inside an unfamiliar project, identify rules, and complete multi-step modifications. Xiaopu Peng, a researcher at Zhipu, said the team has trained long-horizon tasks since GLM-5.1. He added that Real-SWE is closer than public SWE rankings to measuring that ability, which may partly explain why GLM 5.3 stood out on this benchmark.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.