NeoCognition, an AI agent startup, has released ApprenticeBench, a benchmark designed to test whether an AI system can learn a job the way a new employee does after joining a company. Its first setup placed agents in a simulated California construction company and assigned them accounts payable work. The agents first reviewed six months of historical invoices, company manuals, and Odoo tutorials, then processed 100 invoices over a seven-month period while dealing with rule changes and manager feedback. In the reported results, Fable 5.1 posted a 72% success rate, GPT-6 Astra reached 68%, and the best human tester recorded 51%. The work included invoice entry, error detection, cost code assignment, supplier contact, approvals, and learning internal company rules from records and feedback. NeoCognition also said direct GUI use slightly outperformed API-based operation for both models. The company added two key caveats: the "better than humans" result applied only to this single simulated role, and the human sample included just two people. Cost was also higher for the model, with Fable 5.1 averaging $18.23 per task versus about $7.21 for humans.
NeoCognition has released ApprenticeBench, a benchmark built to measure whether an AI agent can learn a job after joining a company, rather than simply complete a fixed task.
In the first version, the agent was placed inside a simulated California construction company and assigned to accounts payable work. It first read six months of historical invoices, company manuals, and Odoo tutorials. It then processed 100 invoices over seven months, with rule changes and manager feedback introduced during the run.
Benchmark results
Fable 5.1 reached a 72% success rate in the test. GPT-6 Astra scored 68%, while the best human tester posted 51%.
The work went beyond invoice entry. It also included spotting mistakes, assigning cost codes, contacting suppliers, moving items through approval, and learning the company’s internal rules from past records and supervisor feedback.
GUI use outperformed API calls
NeoCognition said both models performed slightly better when operating the GUI directly than when calling APIs. According to the company, the performance penalty that had often appeared in earlier computer-use settings largely disappeared in this test.
Limits of the comparison
The company also attached clear limits to the result. The claim that the model outperformed humans applied only to this single simulated role, and the human sample size was just two people.
Cost was higher as well. Fable 5.1 averaged $18.23 per task, compared with about $7.21 for humans. The test also found that humans became faster over time, while the agent slowed down as the amount of remembered information increased.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.