Logan Kilpatrick, Senior Product Manager at Google DeepMind, recently took to X to advocate for AI companies to develop their own benchmarks. He argued that relying on public leaderboards — such as MMLU or standard open-source evaluations — often fails to capture the metrics that truly matter for an organization's unique use cases.
Why Public Benchmarks Fall Short
Standardized benchmarks like GLUE, SuperGLUE, or HumanEval are designed for general performance comparisons across different models. However, real-world business scenarios demand tailored evaluation. For example, a customer support chatbot may need to prioritize empathy and response coherence over mathematical reasoning, while a financial fraud detection system must emphasize precision on decimal calculations. Kilpatrick warned that using only public leaderboards can mislead companies into optimizing for the wrong objectives, resulting in disappointing real-world outcomes.
Case Studies: Zapier and Sierra Lead the Way
Kilpatrick highlighted two companies that have already embraced custom benchmarks with remarkable success. Zapier, an automation platform, built an evaluation suite around its specific workflow triggers and actions, leading to 18% improvement in task success rate. Sierra, a conversational AI company for customer support, designed benchmarks that measure tone consistency and escalation accuracy, achieving a 22% reduction in false alerts. These examples demonstrate that tailor-made assessments can directly translate into enhanced business performance.
Industry Momentum Grows
The movement toward custom benchmarks is not isolated. Major players like Tencent recently launched Chronicles-OCR, a benchmark for ancient script understanding; Jina AI introduced v5-omni with quad-modal retrieval evaluations; and Odyssey released the PROWL framework for world model training. Each initiative reflects a shift from generic to domain-specific evaluation. Kilpatrick predicted that within two years, over 60% of enterprises deploying AI will operate at least one internal benchmark, creating a virtuous cycle between model iteration and business goals.
Getting Started with Custom Benchmarks
For companies new to this approach, Kilpatrick recommends starting small: choose one critical task (e.g., response relevance for chatbots) and build a benchmark around it. Then gradually expand dimensions while incorporating human feedback for calibration. A key pitfall to avoid is overfitting to the benchmark, where the model learns the test set too well and loses generalization. Regularly updating test cases and metrics is essential to keep the benchmark relevant.
In an era of rapid AI advancement, blindly chasing leaderboard scores is no longer sufficient. Building custom evaluation frameworks aligned with business value is the key to turning AI investments into sustainable competitive advantage.

