Moonshot AI has released Kimi K3, a new open-source model that Decrypt described as the largest Chinese open-source model yet publicly released. The report said K3 has posted leading results across several high-profile benchmarks, including writing and frontend coding tests where it moved ahead of Claude Fable 5.

On Towards AI’s Writing Elo benchmark, Kimi K3 scored 2,840, above Claude Fable 5’s 2,760. The benchmark measures models by having them write real scripts that are judged blind against published versions, using the same Elo system used in chess rankings. Decrypt cited a post saying Anthropic’s team had historically dominated that ranking, and that Kimi K3 is now in the top spot.
The same post said Kimi K3 jumped from No. 21 to No. 1 compared with its predecessor, Kimi K2.6, and runs at about $0.25 per script.
K3 also led a frontend coding leaderboard
Kimi K3 also ranked first on Arena AI’s Frontend Code Leaderboard. That ranking is built from thousands of pairwise human votes on code-generation tasks and uses Elo scoring as well. K3 scored 1,679, compared with Claude Fable 5 at 1,631, and placed first in six of seven frontend domains.
According to an Arena.ai post dated July 16, 2026, Kimi K3 climbed from No. 18 to No. 1 versus Kimi K2.6. The post listed domains including Brand & Marketing, Reference-Based Design, and Data & Analytics.
The Artificial Analysis Intelligence Index, which combines nine independent evaluations covering coding, reasoning, agentic work, and knowledge on a 0-to-100 scale, gave K3 a score of 57. Claude Fable 5 scored 60, GPT-5.6 Sol scored 59, and Claude Opus 4.8 scored 56. On that composite, K3 ranked third and trailed Fable 5 by 3%.
Decrypt also pointed to BridgeBench, where K3 was reported to be beating Fable 5 head to head. A related post said that under the same task setup and a blind judging panel, K3 won seven of eight arenas, including Refactoring by 9-0 and Debugging by 6-1. Fable 5’s only win in that comparison was speed.
Another post, from Guillermo Rauch on July 16, 2026, said Kimi K3 was the top-performing model on a comprehensive web engineering benchmark, ahead of Fable and reaching a comparable success rate in less time. That post described it as the first time an open model had moved ahead of all proprietary models on that benchmark.
The report also cited a zero-shot example in which K3 was prompted to build an iOS clone. Decrypt said the result was compared with what social media users had shared as the best approximation produced with GPT-5.6 Sol using a much more elaborate prompt.
2.8 trillion parameters in a mixture-of-experts system
Kimi K3 has 2.8 trillion parameters in a mixture-of-experts architecture. Decrypt explained that the setup splits those parameters into 896 “expert” subnetworks and activates only a fraction of them for any given task, aiming to deliver frontier-level capability without requiring all of the model at once.

Moonshot AI said, “It is the world’s first open-source model in the 3-trillion-parameter class, designed for frontier intelligence scenarios including long-horizon coding, knowledge work, and reasoning.” The report added that DeepSeek’s V4-Pro tops out at 1.6 trillion parameters, while Moonshot’s own K2 sits at 1 trillion, which puts K3 at roughly double the size of the nearest open-weight rival mentioned in the article.
K3 comes with a 1 million-token context window, native image and video understanding, and always-on reasoning.
Two architectural changes were cited for efficiency gains
Decrypt said two technical approaches sit behind K3’s efficiency gains.
- Kimi Delta Attention is used to speed up decoding on long sequences, reaching up to 6.3x faster decoding at 1 million-token contexts.
- Attention Residuals selectively route information across model layers instead of accumulating it uniformly. Moonshot said that adds about 25% training efficiency at less than 2% extra compute cost.
Taken together, those changes were said to deliver roughly 2.5x better scaling efficiency than K2.
Pricing sits at mid-tier levels
Kimi K3 is priced at $3 per million input tokens and $15 per million output tokens. Decrypt said that matches the pricing of Claude Sonnet 5, Anthropic’s mid-tier model.
The article contrasted that pricing with K3’s benchmark profile. On the Artificial Analysis composite, K3 sits three points below Fable 5. Across that nine-benchmark suite, Decrypt said K3 costs $0.94 per task, versus $1.04 for GPT-5.6 Sol and $1.80 for Opus 4.8.
Decrypt also referred back to its May coverage, which said the pricing gap between Chinese and American frontier AI earlier this year ran from 15x to 30x. The article said K3 does not price at DeepSeek-style discount levels, but instead at a Western mid-range rate while offering near-frontier performance. For teams building through an API, the report framed that as a major cost improvement.
The article added that if Anthropic follows through on plans to make Fable 5 available only through an API, K3 would become the nearest open-weight alternative to that tier of model, at about half the per-task cost of Opus 4.8.
Launch lands in the middle of the chip-control debate
Decrypt said K3’s launch raises questions that supporters of U.S. chip export controls would rather avoid. The U.S. restricted exports of Nvidia’s H800 GPUs to China in late 2023, and Moonshot had previously confirmed that it trained earlier models on those chips.

K3’s benchmark documentation refers to H200s and what Moonshot called “a GPGPU from an alternative vendor.” Decrypt said that has been widely interpreted as Huawei Ascend hardware, though the documentation does not specify where that hardware is located.
According to Silicon Republic, Moonshot AI president Yutong Zhang said at Davos this year: “We knew we didn't have the luxury to simply scale up compute… That forced us to focus on fundamental research and efficiency.”
Bank of America analysts, in a note after the launch, wrote that K3 shows “pre-training scaling, paired with architectural innovation, can still deliver step-change gains for flagship Chinese models” under those constraints.
Decrypt described Moonshot as one of the so-called AI Tiger startups that have shifted the global model landscape without access to the chips Washington said they would need. Whether that is an argument for tighter controls or evidence that the restrictions do not work remains, in the article’s framing, an unresolved policy question in Washington.
Hallucination rate increased from K2.6
The article also pointed to a clear drawback. On AA-Omniscience, a benchmark that measures how often a model confidently fabricates an answer it does not know, K3’s hallucination rate rose to 51% from 39% on K2.6. Decrypt summed that up as more correct answers overall, but more invented ones as well.
Moonshot’s own documentation also says K3 can be “excessively proactive,” making unexpected decisions on a user’s behalf during long autonomous tasks.
For teams already using tooling based on Kimi K2.6, the report said K3 looks like a meaningful upgrade on most fronts. It also said the change in hallucination behavior is something worth stress-testing before using the model in accuracy-sensitive workflows.
Free access is available, but the article said usage has been unstable
Decrypt said users can try K3 for free on Kimi’s official website. It also said traffic has been so heavy that tasks are often interrupted by capacity limits, leaving the free version hard to use reliably. The article suggested paid subscriptions or API access as alternatives.
Weights are scheduled for release on July 27 and will be available for large enterprises and businesses. Decrypt added that no domestic GPU currently available is able to handle a model of this size.

