Moonshot AI has rolled out Kimi K3, the latest version of its flagship large language model and productivity platform. According to the source article, Kimi K3 was released on July 16, 2026, with 2.8 trillion parameters, native support for a 1 million-token context window, and built-in multimodal understanding across text, images and video.
The company’s official X account said full model weights will be released on July 27. At the same time, Moonshot is launching four product entry points: Kimi.com, Kimi Work, Kimi Code and the Kimi API platform. The article describes Kimi K3 as one of the world’s largest frontier models with open weights.
Moonshot AI’s growth, valuation and Hong Kong IPO plan
Moonshot AI was founded in early 2023 by Yang Zhilin, a former Tsinghua University professor. The article says he previously worked at Meta and Google, studied at Carnegie Mellon University, and published multiple highly cited papers in natural language processing.
Backed by venture capital, Moonshot’s valuation climbed rapidly to $20 billion, with Kimi becoming one of the best-known ChatGPT alternatives in China, according to the report.
Bloomberg reported that by April 2026, Moonshot AI’s annual recurring revenue had surpassed $200 million. The company is currently seeking to raise $2 billion at a $30 billion valuation and is restructuring in preparation for a Hong Kong IPO.
The piece also refers to a February report by ChainCatcher, which said Anthropic had accused Moonshot AI, DeepSeek and MiniMax of using Claude outputs through knowledge distillation to train their own models. Moonshot did not comment on the accusation, according to the article.
Kimi K3 architecture: 2.8T parameters, 896 experts and a native 1M context window
Kimi K3 uses a large-scale Mixture of Experts, or MoE, architecture. Total parameter count is 2.8 trillion. The model includes 896 experts, with only 16 activated for each inference through Stable LatentMoE, a design the article says cuts real inference costs sharply.
For weights and quantization, K3 uses MXFP4, a 4-bit floating-point format with an independent scale for each block, and MXFP8 during startup. Based on the figures in the article, storing the full 2.8T model requires roughly 1.4 TB. Overall efficiency is about 2.5 times higher than the previous Kimi K2.
Two named innovations: Kimi Delta Attention and Attention Residuals
The report highlights two core technical changes, Kimi Delta Attention, or KDA, and Attention Residuals, or AttnRes. It says this is the first time the Kimi team has publicly named these proprietary mechanisms.
KDA is described as a hybrid linear attention mechanism. It replaces standard quadratic attention in some layers while preserving full expressive power in key layers, lowering total compute cost for long-context workloads. Based on official data cited in the article, KDA improves decoding speed by as much as 6.3x under a 1 million-token context. The source says that matters in agent-style tasks and long-horizon coding work, where a single session can require very large context handling.
AttnRes replaces the conventional residual connection design. Instead of uniformly adding outputs across layers, AttnRes allows each layer to selectively retrieve representations from much earlier layers. The article says this is especially relevant for MoE systems, where different experts are often activated at different depths, and uniform accumulation can blur those differences. Official figures cited in the piece show AttnRes improving training efficiency by about 25%, with added compute cost below 2%.
K3 also introduces Quantile Balancing and Per-Head Muon. The first derives expert allocation directly from router-score quantiles without heuristic updates. The second optimizes each attention head independently.
Four product lines launch alongside the model
Moonshot is rolling out Kimi K3 across four access points:
- Kimi.com: a browser-based chat interface for K3, aimed at general chat, writing and information organization.
- Kimi Work: a desktop application for Windows and Apple Silicon, positioned as an AI assistant for work.
- Kimi Code: a terminal and IDE-integrated coding agent focused on long-horizon autonomous programming.
- Kimi API: a developer platform available at platform.kimi.ai.
The subscription tiers use classical music tempo terms, including Adagio, Andante and Moderato, spanning personal and enterprise plans. API pricing is listed as $0.30 per million tokens for cache hits, $3.00 for cache misses and $15.00 for output tokens. The article says cache-hit rates exceed 90% in coding workloads, which lowers real-world cost.
How Kimi K3 compares with Claude Fable 5, GPT-5.6 Sol and DeepSeek V4
Citing a July 17 ChainCatcher report and data from independent evaluation platform Artificial Analysis, the article places Kimi K3 between leading U.S. flagship models on a range of benchmarks.
- Parameter scale: Kimi K3 has 2.8T parameters in an MoE architecture with 16 experts activated per step. Claude Fable 5 and GPT-5.6 Sol do not disclose parameter counts. DeepSeek V4 is also MoE.
- Context window: Kimi K3, Claude Fable 5 and GPT-5.6 Sol each support 1 million tokens natively, while DeepSeek V4 supports 128K to 256K.
- Multimodal support: Kimi K3, Claude Fable 5 and GPT-5.6 Sol are native multimodal models. DeepSeek V4 offers partial support.
- Frontend Code Arena: Kimi K3 ranks first with a 76% win rate. Claude Fable 5 ranks second and GPT-5.6 Sol ranks third.
- SWE Marathon and Program Bench: the article says Kimi K3 leads all compared models, with Claude Fable 5 and GPT-5.6 Sol behind it.
- General task quality: Kimi K3 is slightly behind Claude Fable 5 and GPT-5.6 Sol, but ahead of DeepSeek V4.
- Estimated cost for standard intelligence tasks: Kimi K3 is listed at $0.95, Claude Fable 5 at $2.75, GPT-5.6 Sol is not listed, and DeepSeek V4 Flash ranges from $0.02 to $0.33.
The source article concludes that Kimi K3 is strongest in coding and agent-based workloads. In general conversation quality, it trails Claude Fable 5 and GPT-5.6 Sol, but offers a much stronger cost-performance profile than U.S. flagships. Moonshot itself also acknowledged, according to the report, that K3 still has an objective gap in user experience versus Claude Fable 5 and GPT-5.6 Sol.
What the July 27 open-weight release means
The full Kimi K3 weights are scheduled for release on July 27, 2026. The article says it will become another frontier-scale open model with trillion-level parameters after GLM-5.2.
Deployment, however, remains demanding. A SemiAnalysis assessment cited in the piece says enterprises would need at least a 64-accelerator supernode to self-host K3, along with high-end GPUs, NVLink interconnects and HBM memory.
The article says that helps explain why K3’s open release is being seen as a long-term positive for Nvidia’s ecosystem: as Chinese frontier models get stronger, the global race for computing infrastructure intensifies.
For developers, open weights mean the community can fine-tune K3, quantize it and self-host services, placing it alongside Llama, Qwen and DeepSeek as an open-source option. For users who only need API access, Kimi.com and Kimi Work are described as the most direct ways in, while subscription pricing and API fees sit clearly below Claude or GPT, according to the article.

