DeepSeek2026-10-01 08:42:13DeepSeek open-sources Ascend infrastructure components with APIs aligned to Nvidia versionsDeepSeek said it open-sourced a batch of infrastructure components for Huawei’s Ascend platform on Sept. 30, publishing the code on GitHub and describing the new releases as one-to-one counterparts to components it had previously released for Nvidia GPU systems. The package includes DeepGEMM Ascend, a matrix multiplication library whose documentation says the first version supports the Ascend 950 series and remains fully compatible with the original DeepGEMM API, allowing developers to keep the same workflow across platforms. It also includes DeepEP-Ascend, a communication library for training and inference that provides expert-parallel communication operations for mixture-of-experts models and exposes a buffer API aligned with the Nvidia version of DeepEP. According to QbitAI, the open-source release also covers TileLang, FlashMLA, TileKernel and DeepSelect. Documentation for DeepEP-Ascend lists benchmark results from tests on Ascend 950DT with CANN 9.2.0, showing dispatch bandwidth of 373-375 GB/s and combine bandwidth of 345-347 GB/s at eight expert-parallel units, with lower figures at 128 units. The documents also note that full bandwidth depends on a Huawei commercial firmware release for Atlas 850E expected around mid-October 2026.50
NaiveAI2026-09-28 03:40:43NaiveAI open-sources first model as AI helps drive 151 optimization runs in six daysNaiveAI, founded by Tsinghua University associate professor Dai Jifeng, has released Naive-N0.5-Flash and opened its model weights under the MIT license. The model is based on Xiaomi’s MiMo-V2.5 Base and keeps the original mixture-of-experts, or MoE, architecture while changing the attention design. According to the team, global attention was replaced with sliding-window attention and DeepSeek sparse attention, followed by continued training on 3.25 trillion tokens. The updated model natively supports a 1 million-token context window, which the company said cuts compute costs for very long text processing. NaiveAI also said AI systems were directly involved in development, handling coding, experiments, result analysis and later optimization, while human researchers set direction and made key decisions. Using its NaiveRT inference system as an example, the team said it completed 151 optimization rounds over six days, with 63 adopted. The company also disclosed a peak inference speed of 2,122 tok/s under a specific setup, while saying standard mode runs at about 50 tok/s per user.210
Vitalik Buter2026-09-26 09:45:23Vitalik says offline mobile AI apps have improved, but still trail laptop modelsEthereum co-founder Vitalik Buterin said mobile offline AI knowledge apps have improved sharply from what he was able to build two months ago, but they still struggle on harder questions and remain slower and less effective than models that can run on a laptop. In a Sept. 26 post on X, Buterin said the weakest results came from travel-style queries. He tested the prompt asking for the best vegan restaurant in his current city, and said every app he tried performed poorly. The apps came out of a community-led bounty on decentralized bounty platform poidh, where contributors locked up a total of 1.098 ETH. Buterin himself supplied 1 ETH, more than 90% of the pool, even though the bounty description said it was independently created in response to his earlier post and that he would not judge entries unless he explicitly chose to do so. The challenge asks builders to create the best offline AI research app for Android under tight constraints: support for Android and GrapheneOS, a 12GB memory cap, a 50GB total limit for app, model and database, no network requests after installation, and open-source code on GitHub. As of Sept. 26, three submissions had been filed, and none had been publicly confirmed by Buterin as meeting the target.270
SemiAnalysis2026-09-22 04:50:22SemiAnalysis says memory bandwidth matters more than capacity in AI inference, with scheduling emerging as a core layerSemiAnalysis has published a report breaking down the underlying architecture of large-model inference services, arguing that the rise of mixture-of-experts, or MoE, models has turned inference into a multi-stage pipeline rather than a single compute task. The report describes that pipeline as consisting of Prefill, Midfill, Decode Attention, and Decode Experts, with each stage placing different demands on compute, memory bandwidth, and networking. Its central conclusion is that, in most inference scenarios, memory bandwidth carries more economic value than raw capacity. According to the report, high-bandwidth memory can improve token generation efficiency, while idle HBM mainly adds cost. SemiAnalysis also projects that by 2027, a single pipeline stage may require roughly 400 GB to 500 GB of local fast memory, though processed KV cache should be moved to CPU DRAM and lower-cost network storage to avoid tying up scarce HBM resources. The report also identifies the scheduling layer as a critical part of inference infrastructure. It says Prefill and Midfill are relatively predictable in runtime, while decode latency is more variable and can create backlog and delay-feedback oscillation. SemiAnalysis further compares integrated and disaggregated architectures, and includes simulator-based projections for Kimi K3 on Nvidia B200, B300, and GB200 systems.400
DeepSeek2026-09-22 01:51:56DeepSeek says it is training a 2 trillion-parameter model and planning an 8 trillion-parameter versionDeepSeek CEO Liang Wenfeng told investors that the company is training a 2 trillion-parameter model and plans to develop an 8 trillion-parameter model afterward. The company’s current flagship, V4-Pro, has 1.6 trillion total parameters and has been open-sourced with 49 billion activated parameters disclosed. Among ultra-large models with publicly available weights, Moonshot AI’s Kimi K3 stands at 2.8 trillion parameters. Kimi K3 uses a mixture-of-experts, or MoE, architecture, with 104 billion parameters activated per token out of its 2.8 trillion total. On that basis, DeepSeek’s planned 8 trillion-parameter model would be close to nearly three times that scale. DeepSeek also completed a fundraising round of more than 50 billion yuan in July, reaching a valuation above $50 billion, according to ChainCatcher.370
StepFun2026-09-20 02:10:49StepFun unveils Step 5 Preview, with full model weights set for Oct. 15 releaseStepFun has released its flagship model, Step 5 Preview, and said the full model weights will be made available on Oct. 15. The model uses a sparse mixture-of-experts, or MoE, architecture with 600 billion total parameters while activating 27 billion parameters per inference. It supports a 1 million-token context window and multimodal image-text input, while its API and Studio access are already fully open. Artificial Analysis had completed its benchmark run ahead of the launch. In its Intelligence Index, Step 5 Preview scored 44 points, matching Kimi K3 Max. Artificial Analysis also measured the model at roughly 100 tokens per second, with a per-task cost of $0.71, compared with $2 for Kimi K3 Max. StepFun highlighted the model’s performance on long-running tasks. In one experiment, Step 5 Preview autonomously optimized an H100 GPU kernel for as long as 24 hours by rewriting code, running tests, comparing results, and iterating. After about 22 hours, it reached 508 TFLOPS, versus 493 TFLOPS for Claude Opus 5 in the same test. In another 24-hour experiment, the model designed post-training data on its own and improved Qwen3-30B-A3B’s AIME24 accuracy from 53.3% to 60%, matching Claude Opus 5 while using fewer labeled tokens.400
Tencent2026-09-11 16:32:27Tencent open-sources Hunyuan Hy4 preview with 770B parameters and 1 million-token contextTencent Hunyuan released and open-sourced its next-generation large language model, Hy4 preview, on Aug. 28. The model has 770 billion total parameters, activates 49 billion parameters per token, and supports a context window of more than 1 million tokens. Its weights have been published on Hugging Face under the Apache 2.0 license. According to Xinhua, Tencent positions Hy4 preview as a model built for productivity, targeting software engineering, office analysis, game development, and scientific research. The release adds to a recent run of open model launches by major Chinese technology companies, following Zhipu’s GLM-5.3 and the DeepSeek V4 series. Tencent’s model card says Hy4 preview uses a 78-layer mixture-of-experts architecture, includes FP8 quantization and speculative decoding support, and can be deployed in an example setup using eight GPUs for tensor parallelism. Tencent also disclosed benchmark and internal blind-test results, while noting that third-party independent verification has not yet been completed. The model is available through Tencent Cloud Tokenhub and OpenRouter, and has been integrated into WorkBuddy and CodeBuddy.1170
Tencent2026-09-11 10:20:30Tencent unveils T1 for long-horizon terminal Agent tasks, with 300-plus tool calls per runTencent has introduced T1, a terminal-focused Agent model trained from Qwen3.5-122B-A10B for complex Linux command-line tasks. The model can call tools for more than 300 rounds within a single task, targeting longer and more demanding terminal workflows. On Terminal-Bench 2.1, T1 scored 64.0%, up from the base model’s 43.8%, a gain of 20.2 points. Tencent said most of that improvement came from reinforcement learning. Supervised fine-tuning, or SFT, lifted the score to 49.4%, and reinforcement learning added another 14.6 points after that. The team also prepared about 15,000 terminal tasks, each paired with automated tests. During training, the Agent could still receive rewards for completing only part of a task’s requirements rather than the full assignment. Tencent also described a training-inference mismatch in long tasks: the same content may be tokenized differently between execution and later training, while a mixture-of-experts model may route the sequence to a different set of experts. To address that, T1 records the generated tokens and the experts selected at the time of execution, then replays them during training. Tencent said this reduced the mismatch by about one-third.840