KDA

Moonshot AI
2026-07-18 12:44:50

Moonshot AI launches Kimi K3, a 2.8-trillion-parameter model set for open-weight release on July 27

Moonshot AI has introduced Kimi K3, its latest flagship large language model, with 2.8 trillion parameters, a Mixture of Experts design, a native 1 million-token context window and built-in multimodal support across text, images and video. The company’s official X account said full model weights will be released on July 27, alongside four product endpoints: Kimi.com, Kimi Work, Kimi Code and the Kimi API platform. The article says Kimi K3 is among the world’s largest frontier models with open weights. It also outlines two named architectural changes, Kimi Delta Attention and Attention Residuals, which Moonshot says improve long-context decoding speed and training efficiency. The model uses 896 experts, activates 16 per inference step, and stores weights in MXFP4 format, putting total storage needs at roughly 1.4 TB. Beyond the model release, the report revisits Moonshot AI’s corporate backdrop. Bloomberg previously reported that the company’s annual recurring revenue had exceeded $200 million as of April 2026. Moonshot is now seeking to raise $2 billion at a $30 billion valuation and is restructuring for a Hong Kong IPO, according to the article. Benchmark data cited from ChainCatcher and Artificial Analysis places Kimi K3 ahead in coding and agent-style tasks, while still trailing Claude Fable 5 and GPT-5.6 Sol in general-purpose user experience.

1690
Moonshot AI launches Kimi K3, a 2.8-trillion-parameter model set for open-weight release on July 27
Fable 5
2026-07-07 09:02:12

Fable 5 Tops KernelBench-Mega With 18.71x Speedup and the First Single-Launch Megakernel

Fable 5 has taken the top spot in the latest KernelBench-Mega benchmark by delivering an 18.71x speedup on RTX PRO 6000 using a fully hand-written CUDA kernel. The reported result places it well ahead of Claude Opus 4.8 at 14.4x, GPT-5.5 at 4.34x, and Sonnet 5 at 4.0x. The tested workload was 02_kimi_linear_decode, a Kimi-Linear W4A16 mixed decoding task with 4-bit weights and bf16 activations, under a strict setup that allowed only one autonomous session and a three-hour wall-clock limit. What makes the result stand out is not only the score but the implementation: according to the report, Fable 5 is the first model in KernelBench-Mega to produce a true end-to-end megakernel, compressing the inference path into a single GPU kernel launch per decoded token. Anthropic co-founder Jack Clark said the development could mark the beginning of a recursive self-improvement loop, arguing that once models can optimize the low-level systems used to train and run future models, the feedback cycle may accelerate substantially.

200
Fable 5 Tops KernelBench-Mega With 18.71x Speedup and the First Single-Launch Megakernel