Moonshot’s Kimi K3 tops several AI benchmarks, with Claude Fable 5 still ahead on composite score

Moonshot’s Kimi K3 tops several AI benchmarks, with Claude Fable 5 still ahead on composite score

N
News Editor
2026-07-17 17:36:42
Moonshot AI has released Kimi K3, which Decrypt described as the largest Chinese open-source model yet made public. The model carries 2.8 trillion parameters in a mixture-of-experts design and was reported to have moved to the top of several benchmark rankings, including Towards AI’s Writing Elo and Arena AI’s Frontend Code Leaderboard. In those tests, K3 posted a 2,840 Writing Elo score versus Claude Fable 5’s 2,760, and 1,679 on the frontend coding board versus Fable 5’s 1,631. The broader picture is more mixed. On the Artificial Analysis Intelligence Index, which aggregates nine independent evaluations, K3 scored 57. Claude Fable 5 scored 60, GPT-5.6 Sol 59, and Claude Opus 4.8 56, leaving K3 in third place on that composite. Decrypt also cited BridgeBench and a web engineering benchmark where K3 was reported ahead of Fable in several head-to-head measures. Moonshot said K3 supports a 1 million-token context window, native image and video understanding, and “always-on reasoning.” The company plans to release weights on July 27 for large enterprises and businesses. The report also noted a rise in hallucination rate versus K2.6, from 39% to 51%, and said heavy traffic has made the free web version difficult to use consistently.
Moonshot AIKimi K3open-source AIlarge language modelsClaudeGPT-5.6 SolAI benchmarkstechnology

Moonshot AI has released Kimi K3, a new open-source model that Decrypt described as the largest Chinese open-source model yet publicly released. The report said K3 has posted leading results across several high-profile benchmarks, including writing and frontend coding tests where it moved ahead of Claude Fable 5.

Moonshot’s Kimi K3 tops several AI benchmarks, with Claude Fable 5 still ahead on composite score 2

On Towards AI’s Writing Elo benchmark, Kimi K3 scored 2,840, above Claude Fable 5’s 2,760. The benchmark measures models by having them write real scripts that are judged blind against published versions, using the same Elo system used in chess rankings. Decrypt cited a post saying Anthropic’s team had historically dominated that ranking, and that Kimi K3 is now in the top spot.

The same post said Kimi K3 jumped from No. 21 to No. 1 compared with its predecessor, Kimi K2.6, and runs at about $0.25 per script.

K3 also led a frontend coding leaderboard

Kimi K3 also ranked first on Arena AI’s Frontend Code Leaderboard. That ranking is built from thousands of pairwise human votes on code-generation tasks and uses Elo scoring as well. K3 scored 1,679, compared with Claude Fable 5 at 1,631, and placed first in six of seven frontend domains.

According to an Arena.ai post dated July 16, 2026, Kimi K3 climbed from No. 18 to No. 1 versus Kimi K2.6. The post listed domains including Brand & Marketing, Reference-Based Design, and Data & Analytics.

The Artificial Analysis Intelligence Index, which combines nine independent evaluations covering coding, reasoning, agentic work, and knowledge on a 0-to-100 scale, gave K3 a score of 57. Claude Fable 5 scored 60, GPT-5.6 Sol scored 59, and Claude Opus 4.8 scored 56. On that composite, K3 ranked third and trailed Fable 5 by 3%.

Decrypt also pointed to BridgeBench, where K3 was reported to be beating Fable 5 head to head. A related post said that under the same task setup and a blind judging panel, K3 won seven of eight arenas, including Refactoring by 9-0 and Debugging by 6-1. Fable 5’s only win in that comparison was speed.

Another post, from Guillermo Rauch on July 16, 2026, said Kimi K3 was the top-performing model on a comprehensive web engineering benchmark, ahead of Fable and reaching a comparable success rate in less time. That post described it as the first time an open model had moved ahead of all proprietary models on that benchmark.

The report also cited a zero-shot example in which K3 was prompted to build an iOS clone. Decrypt said the result was compared with what social media users had shared as the best approximation produced with GPT-5.6 Sol using a much more elaborate prompt.

2.8 trillion parameters in a mixture-of-experts system

Kimi K3 has 2.8 trillion parameters in a mixture-of-experts architecture. Decrypt explained that the setup splits those parameters into 896 “expert” subnetworks and activates only a fraction of them for any given task, aiming to deliver frontier-level capability without requiring all of the model at once.

Moonshot’s Kimi K3 tops several AI benchmarks, with Claude Fable 5 still ahead on composite score 3

Moonshot AI said, “It is the world’s first open-source model in the 3-trillion-parameter class, designed for frontier intelligence scenarios including long-horizon coding, knowledge work, and reasoning.” The report added that DeepSeek’s V4-Pro tops out at 1.6 trillion parameters, while Moonshot’s own K2 sits at 1 trillion, which puts K3 at roughly double the size of the nearest open-weight rival mentioned in the article.

K3 comes with a 1 million-token context window, native image and video understanding, and always-on reasoning.

Two architectural changes were cited for efficiency gains

Decrypt said two technical approaches sit behind K3’s efficiency gains.

  • Kimi Delta Attention is used to speed up decoding on long sequences, reaching up to 6.3x faster decoding at 1 million-token contexts.
  • Attention Residuals selectively route information across model layers instead of accumulating it uniformly. Moonshot said that adds about 25% training efficiency at less than 2% extra compute cost.

Taken together, those changes were said to deliver roughly 2.5x better scaling efficiency than K2.

Pricing sits at mid-tier levels

Kimi K3 is priced at $3 per million input tokens and $15 per million output tokens. Decrypt said that matches the pricing of Claude Sonnet 5, Anthropic’s mid-tier model.

The article contrasted that pricing with K3’s benchmark profile. On the Artificial Analysis composite, K3 sits three points below Fable 5. Across that nine-benchmark suite, Decrypt said K3 costs $0.94 per task, versus $1.04 for GPT-5.6 Sol and $1.80 for Opus 4.8.

Decrypt also referred back to its May coverage, which said the pricing gap between Chinese and American frontier AI earlier this year ran from 15x to 30x. The article said K3 does not price at DeepSeek-style discount levels, but instead at a Western mid-range rate while offering near-frontier performance. For teams building through an API, the report framed that as a major cost improvement.

The article added that if Anthropic follows through on plans to make Fable 5 available only through an API, K3 would become the nearest open-weight alternative to that tier of model, at about half the per-task cost of Opus 4.8.

Launch lands in the middle of the chip-control debate

Decrypt said K3’s launch raises questions that supporters of U.S. chip export controls would rather avoid. The U.S. restricted exports of Nvidia’s H800 GPUs to China in late 2023, and Moonshot had previously confirmed that it trained earlier models on those chips.

Moonshot’s Kimi K3 tops several AI benchmarks, with Claude Fable 5 still ahead on composite score 4

K3’s benchmark documentation refers to H200s and what Moonshot called “a GPGPU from an alternative vendor.” Decrypt said that has been widely interpreted as Huawei Ascend hardware, though the documentation does not specify where that hardware is located.

According to Silicon Republic, Moonshot AI president Yutong Zhang said at Davos this year: “We knew we didn't have the luxury to simply scale up compute… That forced us to focus on fundamental research and efficiency.”

Bank of America analysts, in a note after the launch, wrote that K3 shows “pre-training scaling, paired with architectural innovation, can still deliver step-change gains for flagship Chinese models” under those constraints.

Decrypt described Moonshot as one of the so-called AI Tiger startups that have shifted the global model landscape without access to the chips Washington said they would need. Whether that is an argument for tighter controls or evidence that the restrictions do not work remains, in the article’s framing, an unresolved policy question in Washington.

Hallucination rate increased from K2.6

The article also pointed to a clear drawback. On AA-Omniscience, a benchmark that measures how often a model confidently fabricates an answer it does not know, K3’s hallucination rate rose to 51% from 39% on K2.6. Decrypt summed that up as more correct answers overall, but more invented ones as well.

Moonshot’s own documentation also says K3 can be “excessively proactive,” making unexpected decisions on a user’s behalf during long autonomous tasks.

For teams already using tooling based on Kimi K2.6, the report said K3 looks like a meaningful upgrade on most fronts. It also said the change in hallucination behavior is something worth stress-testing before using the model in accuracy-sensitive workflows.

Free access is available, but the article said usage has been unstable

Decrypt said users can try K3 for free on Kimi’s official website. It also said traffic has been so heavy that tasks are often interrupted by capacity limits, leaving the free version hard to use reliably. The article suggested paid subscriptions or API access as alternatives.

Weights are scheduled for release on July 27 and will be available for large enterprises and businesses. Decrypt added that no domestic GPU currently available is able to handle a model of this size.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
400

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.