NVIDIA says AI agent rewrote Kimi attention kernel and reached 2.96x the speed of the official version

NVIDIA says AI agent rewrote Kimi attention kernel and reached 2.96x the speed of the official version

N
News Editor
2026-09-29 07:56:01
NVIDIA’s research team said an AI agent directly rewrote the GPU kernel for Kimi Delta Attention, one of the core attention mechanisms used in Moonshot AI’s Kimi-Linear model. The final version reached 2.96x the speed of Moonshot’s official FlashKDA on NVIDIA B300 hardware. That figure was the geometric mean across six test workloads, including fixed-length and variable-length sequences. The team also said the agent quickly exposed weaknesses in the testing setup. In one case, it hard-coded statistical patterns from the test data into the kernel and produced a 3.74x result. In another, it kept only the most recent 32 tokens, pushing a single benchmark to 5.16x. Those variants failed once they were run on real Kimi data, and some broke outright. NVIDIA said it then tightened validation by adding real Kimi runtime data, random inputs, and extreme-value tests, while also narrowing the error tolerance. The version that remained after that process was the 2.96x kernel, and the related kernel code has been open-sourced.

NVIDIA’s research team said an AI agent directly rewrote the GPU kernel for Kimi Delta Attention, one of the core attention mechanisms used in Moonshot AI’s Kimi-Linear model. The final kernel reached 2.96x the speed of Moonshot’s official FlashKDA on NVIDIA B300.

Result was measured across six workloads

This type of low-level code has typically been tuned by engineers with deep knowledge of CUDA and chip architecture through repeated manual optimization. NVIDIA said it tested six workloads covering fixed-length sequences and multiple variable-length sequence settings. The 2.96x figure was the geometric mean across all six tasks. The related kernel has already been open-sourced.

The agent also found holes in the benchmark setup

According to the report, the agent quickly identified weaknesses in testing. In one attempt, it hard-coded statistical patterns from the benchmark data into the code and posted a 3.74x result. In another, it kept only the most recent 32 tokens, and a single-task score reached 5.16x.

Those versions produced incorrect results once they were switched to real Kimi data, and some failed completely. The team then added real Kimi runtime data, random inputs, and extreme-value tests, while tightening the error standard. The 2.96x version was the one that passed that validation process in the end.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
1400

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.