Space Bunny2026-09-29 09:13:31Anonymous AI model Space Bunny tops OpenCode usage six days after launchOpen-source coding agent platform OpenCode introduced an anonymous model called Space Bunny on Sept. 23 and made it free for one week. Six days later, the model had reached about 201,000 unique users on OpenCode, with 12,307,734 sessions completed and a 5.3% share of the platform’s weekly token usage, putting it in first place there. On OpenRouter, the parallel listing under the name Space Bunny Alpha climbed to No. 2 in the weekly ranking with 18.2 trillion tokens processed, behind only DeepSeek V4.1 Flash. OpenCode described the model with three headline specs: a 1 million-token context window, multimodal input support, and zero data retention. OpenRouter listed free input and output pricing as well, though its data terms differ from OpenCode’s. The developer has not identified itself. A tokenizer test by model-tracking site YFarmX found that all 50 tested strings matched MiniMax-family token counts, pointing to MiniMax as a possible family match, but not confirming a specific version. MiniMax has not publicly acknowledged any link to Space Bunny.230
Kuaishou2026-09-29 07:28:19Kuaishou teases Kling 4.0 with 30-second native video generation, October launch setKuaishou has previewed Kling 4.0, with the full release and API both scheduled for October. Ahead of that launch, Kling 4.0 Flash has already been opened to Black Gold annual membership users for early access. The update raises the maximum length of a single native video generation to 30 seconds, up from the 15-second cap in Kling 3.0. The new version also expands control options. It supports up to 10 keyframes, allowing users to preset character states, scene transitions, and plot beats. Image, video, and subject inputs can be combined, with support for as many as 15 multimodal references in one generation. Editing tools have also been extended, letting users change people, backgrounds, weather, colors, materials, and camera shots in an existing video, or use another clip as a reference for camera movement, motion, and narrative pacing. Kling 4.0 adds dual-channel stereo and improves handling of complex camera movement, lip sync, and multilingual dialogue. Kuaishou said 4K and 1080p 10-bit HDR output are planned but not yet available. Multi-step continuation, which could extend a video to as long as two minutes, has also not gone live. The company has not disclosed pricing or speed figures for Kling 4.0 Flash.490
StepFun2026-09-21 06:40:16StepFun unveils Step 5 Preview with 600 billion parameters and a 1 million-token context windowStepFun has introduced Step 5 Preview, its flagship artificial intelligence model built as a sparse mixture-of-experts system with 600 billion total parameters and roughly 27 billion parameters activated per token. The model supports a context window of up to 1 million tokens and can take text, image, and video inputs, according to Techub News. StepFun said the model is aimed at agent workloads in software engineering, professional knowledge, and finance, while lowering task costs without giving up a comparable level of intelligence. The company is making the model available through a hosted API and the StepFun platform. An open-weight self-hosted version is scheduled for release on Oct. 15, 2026. Third-party evaluator Artificial Analysis gave Step 5 Preview a score of 44 on its Intelligence Index, compared with a median score of 24 for models in the same price range. API pricing is listed at $1 per million input tokens and $2.7 per million output tokens. StepFun also shared benchmark results for FrontierFinance and DRACO, along with examples showing H100 kernel optimization in a 24-hour experiment and model capability gains from automated post-training.320
StepFun2026-09-19 16:15:43Step 5 preview appears on Artificial Analysis before launch, matching Kimi K3 Max with far lower output pricingStepFun has not officially announced Step 5, but a preview version has already surfaced on Artificial Analysis. The listing describes Step 5 Preview as a multimodal reasoning model with roughly 600B total parameters, 27B active parameters per inference, image input support, and a 1 million-token context window. In Artificial Analysis’ latest Intelligence Index, the model scored 44, the same as Kimi K3 Max. On Terminal-Bench 4.0, Step 5 Preview posted 33% versus 13% for Kimi K3, while both models scored 59% on SciCode. Pricing on the page shows Step 5 Preview at $1 per million input tokens and $2.7 per million output tokens, compared with $3 and $15 for Kimi K3. That puts Step 5’s output cost at less than one-fifth of K3’s, while the cost of running the full Intelligence Index is about one-quarter of K3’s, based on Artificial Analysis figures. The report noted that Artificial Analysis had already completed its evaluation through the StepFun API, even though Step 5 has not yet been formally unveiled. The listed parameters and pricing were also described as not being the company’s final official figures.320
PrismML2026-09-18 18:15:27PrismML releases Ternary Bonsai 2 27B, shrinking Qwen3.8 27B to 5.93 GB with 98.2% average retentionPrismML has released Ternary Bonsai 2 27B, a ternary-weight version of Qwen3.8 27B that cuts model size to 5.93 GB from 53.80 GB for the FP16 build while keeping an average 98.2% of performance across 20 benchmarks. The model supports both text and image input, offers a 262K-token context window, and is designed to run on consumer hardware including the RTX 5090. PrismML said the weights are open-sourced under the Apache 2.0 license. The architecture matches Qwen3.8 27B, with 27.36B total parameters split across a 24.35B language backbone, 2.54B embedding and language modeling head, and a 0.47B vision module. According to the reported evaluation, retention topped 99% on math and coding tasks, while performance dropped more noticeably on long-horizon agent workloads such as Terminal-Bench 2.1. Reported decoding speeds reached 142.5 tokens per second on an RTX 5090, 96.7 token/s on an RTX 4090, and 46.8 token/s on an Apple M5 Max laptop.470
DeepSeek2026-09-10 02:40:36DeepSeek rolls out V4.1 Flash in its app, merging three modes into one entry pointDeepSeek has started rolling out V4.1 Flash in its app, replacing the previous "Fast," "Expert," and "Image" modes with a single unified entry point. The change removes the need for users to manually switch between modes. Instead, one model now handles everyday conversations, image understanding, and more complex queries inside the app. The company had previously said V4.1 Flash comes with native multimodal support. DeepSeek also said the new model has already surpassed V4 Pro across performance, cost, speed, and total time-used metrics. Before V4.1 Pro goes live, V4.1 Flash is set to take over all V4 Pro requests, according to the company’s earlier statement. The update signals a product-level shift in how DeepSeek packages its model capabilities for app users, centering multiple use cases in a single interface rather than separate mode selection.780
National Anti2026-09-06 10:39:12China Launches 'National Anti-Fraud AI' App to Detect ScamsDeveloped by the Shanghai Public Security Bureau under the guidance of the Ministry of Public Security's Criminal Investigation Bureau, China's 'National Anti-Fraud AI' app has officially launched. Users can input suspicious situations to receive scam risk assessments, pattern recognition, and preventive advice along with relevant cases. The app leverages large language models, multimodal models, and agent technology, though the underlying base model has not been disclosed. It is an upgraded version of the '803 Anti-Fraud Agent,' which was first introduced by the Shanghai Anti-Fraud Center in late 2025 and made available nationwide, offering smart Q&A, anti-fraud information, and a glossary.770
B.AI2026-09-04 04:27:00B.AI Platform Integrates Google DeepMind's Gemini 3.8 Flash Model via APIOn September 4, the B.AI platform announced the official integration of Google DeepMind's latest Gemini 3.8 Flash model via API. The model is the latest flagship in the Gemini 3 family, featuring a 1M token context window and native multimodal input. It maintains the Flash series' low latency while improving performance in software engineering, agent tasks, and multi-step reasoning, particularly in finance and law. The API is now open for use.820