Moonshot AI launches Kimi K3, a 2.8-trillion-parameter model set for open-weight release on July 27

Moonshot AI launches Kimi K3, a 2.8-trillion-parameter model set for open-weight release on July 27

N
News Editor
2026-07-18 12:44:50
Moonshot AI has introduced Kimi K3, its latest flagship large language model, with 2.8 trillion parameters, a Mixture of Experts design, a native 1 million-token context window and built-in multimodal support across text, images and video. The company’s official X account said full model weights will be released on July 27, alongside four product endpoints: Kimi.com, Kimi Work, Kimi Code and the Kimi API platform. The article says Kimi K3 is among the world’s largest frontier models with open weights. It also outlines two named architectural changes, Kimi Delta Attention and Attention Residuals, which Moonshot says improve long-context decoding speed and training efficiency. The model uses 896 experts, activates 16 per inference step, and stores weights in MXFP4 format, putting total storage needs at roughly 1.4 TB. Beyond the model release, the report revisits Moonshot AI’s corporate backdrop. Bloomberg previously reported that the company’s annual recurring revenue had exceeded $200 million as of April 2026. Moonshot is now seeking to raise $2 billion at a $30 billion valuation and is restructuring for a Hong Kong IPO, according to the article. Benchmark data cited from ChainCatcher and Artificial Analysis places Kimi K3 ahead in coding and agent-style tasks, while still trailing Claude Fable 5 and GPT-5.6 Sol in general-purpose user experience.
Moonshot AIKimi K3large language modelopen weightsmultimodal AIMoEAI codingHong Kong IPO

Moonshot AI has rolled out Kimi K3, the latest version of its flagship large language model and productivity platform. According to the source article, Kimi K3 was released on July 16, 2026, with 2.8 trillion parameters, native support for a 1 million-token context window, and built-in multimodal understanding across text, images and video.

The company’s official X account said full model weights will be released on July 27. At the same time, Moonshot is launching four product entry points: Kimi.com, Kimi Work, Kimi Code and the Kimi API platform. The article describes Kimi K3 as one of the world’s largest frontier models with open weights.

Moonshot AI’s growth, valuation and Hong Kong IPO plan

Moonshot AI was founded in early 2023 by Yang Zhilin, a former Tsinghua University professor. The article says he previously worked at Meta and Google, studied at Carnegie Mellon University, and published multiple highly cited papers in natural language processing.

Backed by venture capital, Moonshot’s valuation climbed rapidly to $20 billion, with Kimi becoming one of the best-known ChatGPT alternatives in China, according to the report.

Bloomberg reported that by April 2026, Moonshot AI’s annual recurring revenue had surpassed $200 million. The company is currently seeking to raise $2 billion at a $30 billion valuation and is restructuring in preparation for a Hong Kong IPO.

The piece also refers to a February report by ChainCatcher, which said Anthropic had accused Moonshot AI, DeepSeek and MiniMax of using Claude outputs through knowledge distillation to train their own models. Moonshot did not comment on the accusation, according to the article.

Kimi K3 architecture: 2.8T parameters, 896 experts and a native 1M context window

Kimi K3 uses a large-scale Mixture of Experts, or MoE, architecture. Total parameter count is 2.8 trillion. The model includes 896 experts, with only 16 activated for each inference through Stable LatentMoE, a design the article says cuts real inference costs sharply.

For weights and quantization, K3 uses MXFP4, a 4-bit floating-point format with an independent scale for each block, and MXFP8 during startup. Based on the figures in the article, storing the full 2.8T model requires roughly 1.4 TB. Overall efficiency is about 2.5 times higher than the previous Kimi K2.

Two named innovations: Kimi Delta Attention and Attention Residuals

The report highlights two core technical changes, Kimi Delta Attention, or KDA, and Attention Residuals, or AttnRes. It says this is the first time the Kimi team has publicly named these proprietary mechanisms.

KDA is described as a hybrid linear attention mechanism. It replaces standard quadratic attention in some layers while preserving full expressive power in key layers, lowering total compute cost for long-context workloads. Based on official data cited in the article, KDA improves decoding speed by as much as 6.3x under a 1 million-token context. The source says that matters in agent-style tasks and long-horizon coding work, where a single session can require very large context handling.

AttnRes replaces the conventional residual connection design. Instead of uniformly adding outputs across layers, AttnRes allows each layer to selectively retrieve representations from much earlier layers. The article says this is especially relevant for MoE systems, where different experts are often activated at different depths, and uniform accumulation can blur those differences. Official figures cited in the piece show AttnRes improving training efficiency by about 25%, with added compute cost below 2%.

K3 also introduces Quantile Balancing and Per-Head Muon. The first derives expert allocation directly from router-score quantiles without heuristic updates. The second optimizes each attention head independently.

Four product lines launch alongside the model

Moonshot is rolling out Kimi K3 across four access points:

  • Kimi.com: a browser-based chat interface for K3, aimed at general chat, writing and information organization.
  • Kimi Work: a desktop application for Windows and Apple Silicon, positioned as an AI assistant for work.
  • Kimi Code: a terminal and IDE-integrated coding agent focused on long-horizon autonomous programming.
  • Kimi API: a developer platform available at platform.kimi.ai.

The subscription tiers use classical music tempo terms, including Adagio, Andante and Moderato, spanning personal and enterprise plans. API pricing is listed as $0.30 per million tokens for cache hits, $3.00 for cache misses and $15.00 for output tokens. The article says cache-hit rates exceed 90% in coding workloads, which lowers real-world cost.

How Kimi K3 compares with Claude Fable 5, GPT-5.6 Sol and DeepSeek V4

Citing a July 17 ChainCatcher report and data from independent evaluation platform Artificial Analysis, the article places Kimi K3 between leading U.S. flagship models on a range of benchmarks.

  • Parameter scale: Kimi K3 has 2.8T parameters in an MoE architecture with 16 experts activated per step. Claude Fable 5 and GPT-5.6 Sol do not disclose parameter counts. DeepSeek V4 is also MoE.
  • Context window: Kimi K3, Claude Fable 5 and GPT-5.6 Sol each support 1 million tokens natively, while DeepSeek V4 supports 128K to 256K.
  • Multimodal support: Kimi K3, Claude Fable 5 and GPT-5.6 Sol are native multimodal models. DeepSeek V4 offers partial support.
  • Frontend Code Arena: Kimi K3 ranks first with a 76% win rate. Claude Fable 5 ranks second and GPT-5.6 Sol ranks third.
  • SWE Marathon and Program Bench: the article says Kimi K3 leads all compared models, with Claude Fable 5 and GPT-5.6 Sol behind it.
  • General task quality: Kimi K3 is slightly behind Claude Fable 5 and GPT-5.6 Sol, but ahead of DeepSeek V4.
  • Estimated cost for standard intelligence tasks: Kimi K3 is listed at $0.95, Claude Fable 5 at $2.75, GPT-5.6 Sol is not listed, and DeepSeek V4 Flash ranges from $0.02 to $0.33.

The source article concludes that Kimi K3 is strongest in coding and agent-based workloads. In general conversation quality, it trails Claude Fable 5 and GPT-5.6 Sol, but offers a much stronger cost-performance profile than U.S. flagships. Moonshot itself also acknowledged, according to the report, that K3 still has an objective gap in user experience versus Claude Fable 5 and GPT-5.6 Sol.

What the July 27 open-weight release means

The full Kimi K3 weights are scheduled for release on July 27, 2026. The article says it will become another frontier-scale open model with trillion-level parameters after GLM-5.2.

Deployment, however, remains demanding. A SemiAnalysis assessment cited in the piece says enterprises would need at least a 64-accelerator supernode to self-host K3, along with high-end GPUs, NVLink interconnects and HBM memory.

The article says that helps explain why K3’s open release is being seen as a long-term positive for Nvidia’s ecosystem: as Chinese frontier models get stronger, the global race for computing infrastructure intensifies.

For developers, open weights mean the community can fine-tune K3, quantize it and self-host services, placing it alongside Llama, Qwen and DeepSeek as an open-source option. For users who only need API access, Kimi.com and Kimi Work are described as the most direct ways in, while subscription pricing and API fees sit clearly below Claude or GPT, according to the article.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.