Google DeepMind Exec Urges AI Companies to Build Custom Benchmarks

Google DeepMind Exec Urges AI Companies to Build Custom Benchmarks

N
News Editor 01
2026-07-10 23:39:13
Logan Kilpatrick, Senior PM at Google DeepMind, calls for custom AI benchmarks over public leaderboards. Companies like Zapier and Sierra benefit from this tailored approach, improving model performance for specific business needs.
AI benchmarksGoogle DeepMindcustom evaluationAI model performanceZapier

Logan Kilpatrick, Senior Product Manager at Google DeepMind, recently took to X to advocate for AI companies to develop their own benchmarks. He argued that relying on public leaderboards — such as MMLU or standard open-source evaluations — often fails to capture the metrics that truly matter for an organization's unique use cases.

Why Public Benchmarks Fall Short

Standardized benchmarks like GLUE, SuperGLUE, or HumanEval are designed for general performance comparisons across different models. However, real-world business scenarios demand tailored evaluation. For example, a customer support chatbot may need to prioritize empathy and response coherence over mathematical reasoning, while a financial fraud detection system must emphasize precision on decimal calculations. Kilpatrick warned that using only public leaderboards can mislead companies into optimizing for the wrong objectives, resulting in disappointing real-world outcomes.

Case Studies: Zapier and Sierra Lead the Way

Kilpatrick highlighted two companies that have already embraced custom benchmarks with remarkable success. Zapier, an automation platform, built an evaluation suite around its specific workflow triggers and actions, leading to 18% improvement in task success rate. Sierra, a conversational AI company for customer support, designed benchmarks that measure tone consistency and escalation accuracy, achieving a 22% reduction in false alerts. These examples demonstrate that tailor-made assessments can directly translate into enhanced business performance.

Industry Momentum Grows

The movement toward custom benchmarks is not isolated. Major players like Tencent recently launched Chronicles-OCR, a benchmark for ancient script understanding; Jina AI introduced v5-omni with quad-modal retrieval evaluations; and Odyssey released the PROWL framework for world model training. Each initiative reflects a shift from generic to domain-specific evaluation. Kilpatrick predicted that within two years, over 60% of enterprises deploying AI will operate at least one internal benchmark, creating a virtuous cycle between model iteration and business goals.

Getting Started with Custom Benchmarks

For companies new to this approach, Kilpatrick recommends starting small: choose one critical task (e.g., response relevance for chatbots) and build a benchmark around it. Then gradually expand dimensions while incorporating human feedback for calibration. A key pitfall to avoid is overfitting to the benchmark, where the model learns the test set too well and loses generalization. Regularly updating test cases and metrics is essential to keep the benchmark relevant.

In an era of rapid AI advancement, blindly chasing leaderboard scores is no longer sufficient. Building custom evaluation frameworks aligned with business value is the key to turning AI investments into sustainable competitive advantage.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.