Keenable AI open-sources NEEDLE, a live search benchmark that rebuilds queries every hour

Keenable AI open-sources NEEDLE, a live search benchmark that rebuilds queries every hour

N
News Editor
2026-08-31 23:45:51
Keenable AI has released NEEDLE, an open-source benchmark for real-time search systems. Instead of relying on a fixed evaluation dataset, NEEDLE rebuilds its query set every hour using public sources including RSS feeds and Google Trends. The stated aim is to reduce distortions that can appear when models read answers directly from static benchmarks or rely on parameter memory. The benchmark covers several query categories: news, finance, academic, legal, and rare-entity lookups. It evaluates 15 search APIs under a unified protocol and adds a "ceiling" metric designed to capture the best retrievable result available across the field. NEEDLE is available as an open-source evaluation tool that can be installed and run through a Python CLI. Judging requires an OpenRouter key, and the tool supports use on laptops or in CI environments. Results cited in the report show finance queries are close to being solved. Exa scored 0.910, followed by Keenable at 0.872, Perplexity at 0.871, and Google at 0.847. The reported "ceiling" score was 0.965. The item was cited by MarkTechPost.

AI search company Keenable AI has open-sourced NEEDLE, a benchmark for real-time search. According to Techub News, the benchmark rebuilds its query set every hour from public sources such as RSS feeds and Google Trends, rather than using a fixed dataset.

The setup is intended to address problems that can arise when a model reads answers directly or relies on parameter memory. NEEDLE covers news, finance, academic, legal, and rare-entity queries, and evaluates 15 search APIs through a unified protocol. It also introduces a "ceiling" metric to measure the best retrievable result available across the field.

As an open-source evaluation tool, NEEDLE can be installed and run through a Python CLI. It requires an OpenRouter key for judging and supports use on a laptop or in CI environments.

The reported results show finance queries are close to solved. Exa scored 0.910, Keenable scored 0.872, Perplexity scored 0.871, and Google scored 0.847, while the reported "ceiling" was 0.965.

The source item cited MarkTechPost.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.