Alibaba2026-08-26 08:05:56Alibaba open-sources Qwen3.8-Flash-Next as an early look at the Qwen 4 architectureAlibaba’s Qwen team has released Qwen3.8-Flash-Next on August 26, presenting it as a preview of the Qwen 4 architecture before the full Qwen 4 family arrives. According to the model page on Hugging Face, the release gives developers an early chance to test and prepare for the next-generation architecture rather than wait for the formal rollout. The model is described as an open-source multimodal system built on a mixture-of-experts, or MoE, design. Alibaba said it has about 125 billion total parameters, while activating roughly 6 billion parameters per token. The stated goal is to keep the knowledge scale of a large model while lowering inference costs in practice. Model weights are already available on Hugging Face, and Alibaba said the model is also planned for China’s ModelScope platform. That means developers can access it without relying on a proprietary API. At the same time, Alibaba has not released benchmark results for Qwen3.8-Flash-Next, and the published parameter specifications have not yet been independently verified. Full performance still awaits broader evaluation.890
Nvidia2026-08-25 12:37:13Nvidia says Groq 3 LPX is in full production, hits 3,400 tokens per second on Gemma 4 31BNvidia said its Groq 3 LPX inference chip has entered full production and reached an inference speed of 3,400 tokens per second on the Gemma 4 31B model. The company said that figure is four times faster than Cerebras. The comparison, though, is not straightforward. According to the report, Nvidia needs at least 64 accelerators to reach that throughput, while Cerebras requires only one to two. That makes raw tokens-per-second figures harder to compare without accounting for system scale. The report also flagged an open question around architecture scalability. While Nvidia presented the production milestone and throughput result, its ability to scale on large mixture-of-experts, or MoE, models remains to be seen. The report was cited by Techub News, with The Decoder named as the source.960
Stanford2026-08-24 13:07:10Stanford professor Percy Liang opens Marin 535B model training to the publicStanford professor and Simile AI founder Percy Liang has kicked off training for Marin 535B-A23B and is making the process public in real time. The model is described as having 535 billion total parameters and 23 billion active parameters, trained on 18.75 trillion tokens across 11 GB200 NVL72 systems, or about 792 GB200 GPUs. Liang said the run is expected to last roughly three months and consume about 2.7e24 FLOPs, followed by a post-training phase. Before starting the main run, the team completed a four-stage scaling ladder ranging from 1.6B-A61M with 48 billion tokens to 27.7B-A1.2B with 926 billion tokens. Liang said those smaller runs were used to surface and debug issues early and to forecast the behavior of the larger “hero run.” The project has drawn attention because it exposes details that are usually hidden in frontier model development, including data mix, processing methods, training configuration, code, experiment design, live loss curves, and model-state tracking. The team has also published links to a data overview page, a Weights & Biases dashboard, and a GitHub issue page where the work can be followed in detail.1420
Z.ai2026-08-19 07:28:20Z.ai founder Jie Tang says bigger parameter counts no longer tell the full story of model strengthJie Tang, founder of Z.ai and a professor at Tsinghua University, argues that asking only how many parameters a model has no longer says much about how strong it is. In his review of the evolution of scaling laws—from GPT-3 to Chinchilla and then Mixture of Experts (MoE)—he says model capability depends on more than parameter count. Training data volume, where compute is spent, and how a model is actually used all matter. Tang’s point is that the old training-first view of scaling is less useful once commercial AI systems are deployed and called billions of times a day. Under that setup, inference cost changes the optimization target. A smaller model trained for longer may make more sense than a larger one trained less efficiently. He cited Llama-2-7B and Gemma-2-9B as examples of models trained far beyond the classic Chinchilla ratio. He also said MoE makes headline parameter numbers even less informative, because total parameters and activated parameters describe different things. For reasoning-heavy workloads, Tang argued that effective depth in a single inference pass and post-training may now be more important scaling dimensions. He described GLM-5.3 as a controlled test of that idea, keeping the base model and parameter counts unchanged from GLM-5.2 while expanding long-horizon environments and reinforcement learning over a month.1510
HBF2026-08-04 11:43:12HBF standard goes public, but near-term focus stays on SanDisk earnings and NAND pricingWhiteLine Daily said the launch of a public HBF specification matters because it turns the concept from a single-company proposal into a shared industry standard. SK hynix and SanDisk published the first HBF technical specification through the Open Compute Project, defining 8-layer and 16-layer NAND stacks, capacities of up to 512GB per package, bandwidth tiers of roughly 0.4 TB/s to 3.0 TB/s, and UCIe connectivity to CPUs, GPUs and other accelerators. Google and Tenstorrent have also joined the HBF consortium. The report argues that HBF is not meant to replace HBM. Instead, it adds a new memory tier between HBM and SSD, placing larger-capacity storage closer to xPU compute. High-frequency, latency-sensitive data would still remain in HBM or DRAM, while model weights with lower access frequency could sit in HBF. That setup may be especially relevant for MoE models, where many expert weights stay idle most of the time yet still consume capacity. WhiteLine Daily said the short-term investment angle is still not HBF revenue, since first HBF samples are scheduled for the second half of 2026 and the first inference device samples using HBF are expected in early 2027. The nearer catalysts are SanDisk’s earnings, NAND pricing, data center demand, margins, customer agreements, and enterprise SSD progress.1860
Alibaba2026-08-04 01:51:01Alibaba unveils Qwen3.8-Max with self-reported benchmark lead and aggressive token pricingAlibaba’s Qwen team has introduced Qwen3.8-Max, a new flagship model that the company says scored 86.1 on OSWorld-Verified, ahead of GPT-5.6 Sol Max at 83.2, Fable 5 at 85.0, and Gemini 3.1 Pro at 76.2. The release also included a broader slate of benchmark claims, such as 93.0 on PaperBench and 86.6 on TerminalBench 2.1, alongside positioning the model for long-running autonomous work rather than standard chatbot use. The model uses a mixture-of-experts architecture with 2.4 trillion total parameters and about 95 billion active during inference, built on the Qwen3.5 architecture with a 1 million-token context window. Alibaba also said Qwen3.8-Max is suited for extended coding tasks, desktop software operation, experiment reproduction, and industrial workflows that feed visual input back into a planning loop. Pricing appears to be a central part of the launch. According to QwenCloud pricing cited in the report, Qwen3.8-Max costs $2 per million input tokens and $6 per million output tokens overseas, bringing the combined total to $8 per million tokens. That is below one-third of Claude Opus 5’s combined $30 and below one-quarter of GPT-5.6 Sol standard mode at $35. Still, the benchmarks and capability demonstrations were all disclosed by Alibaba and have not been independently verified, while the company has yet to publish the licensing terms for the promised open-weight release next week on Hugging Face and ModelScope.1940
Kimi.ai2026-07-27 15:28:55Kimi.ai open-sources MoonEP, AgentENV and FlashKDA alongside Kimi K3 releaseKimi.ai has continued its open-source push as it released several pieces of underlying infrastructure alongside the opening of Kimi K3 model weights and its technical report. The newly disclosed components include MoonEP, a high-performance communication library built for distributed mixture-of-experts, or MoE, training; AgentENV, a distributed environment system designed for large-scale agent workflows and developed in collaboration with kvcache-ai; and FlashKDA, a high-performance Kimi Delta Attention kernel based on CUTLASS. According to the official description, these components are intended to reduce communication and inference overhead in large-scale MoE training and agent reinforcement learning workloads. The team also said the tools can be used as a plug-and-play backend for flash-linear-attention. The announcement was cited by ChainCatcher.1910
AMD2026-07-25 03:18:25AMD unveils fully open-source Instella-MoE, a 16B-parameter model aimed at top-tier open modelsAMD on July 25 introduced Instella-MoE, a fully open-source mixture-of-experts language model with 16 billion total parameters and 2.8 billion active parameters per token. The company said the model was trained from scratch entirely on its own AMD Instinct MI300X and MI325X GPUs, using the ROCm software stack, and incorporates architectural designs including Gated Multi-head Latent Attention, or Gated MLA, and FarSkip-Collective to improve training and inference efficiency. According to AMD, the Instella-MoE-16B-A3B base model posted an average score of 76.7 in performance tests, placing it among the leading fully open-source models and ahead of models such as SmolLM3-3B and OLMo-3-7B. AMD also said the model can compete with larger systems while activating only 2.8 billion parameters. Instella-MoE supports 64K-token long-context processing and has gone through a full training pipeline that includes pretraining, mid-training, long-context extension, supervised fine-tuning, direct preference optimization, and reinforcement learning. AMD said it is releasing the full model weights, training configurations, data mix, intermediate checkpoints, and inference code.3830