Meituan has quietly rolled out LongCat-2.0-Preview, a new large model now available through the LongCat API platform on an invite-only basis. The update log is dated April 20, yet the company has not issued an official press release, technical report, or open-source release alongside the launch.
A model built around AI agents and automation
Based on the limited official description, LongCat-2.0-Preview is aimed directly at AI agent development. It natively supports tool calling, multi-step reasoning, and long-context tasks, with a focus on code generation, automated workflows, and execution of complex instructions. The platform page also lists integrations with Claude Code, OpenClaw, OpenCode, and Kilo Code.
The release approach stands out. According to the source material, earlier Meituan models in the LongCat line, including Flash-Chat, Flash-Thinking, and Omni, were typically accompanied by detailed blog posts, technical papers, and open-source code on Hugging Face and GitHub. This time, the company has limited access to API-based invite testing and kept public communication to a minimum.
Reports point to trillion-scale parameters and a 1M context window
More details surfaced on April 24 through media reports and comments from people said to be familiar with the model. The source says LongCat-2.0-Preview uses a mixture-of-experts, or MoE, architecture, with total parameters exceeding the trillion level. Its overall scale and activated parameter count were described as being in the same range as DeepSeek V4, which was released the same day.
Context length is another headline feature. The model is said to support a 1M-token context window, allowing it to process extremely large inputs suited to long-document understanding, chained tasks, and extended agent workflows. For current invite testers, the official site offers 10 million free tokens per day.
Domestic compute infrastructure draws the most attention
Beyond the model itself, outside attention has centered on the compute stack behind it. The source cites people familiar with the matter as saying that both training and inference were completed entirely on domestic Chinese compute clusters. Meituan reportedly used 50,000 to 60,000 domestic accelerator cards, a figure described as the largest large-model training run completed on domestic compute to date.
Meituan has not published fuller technical documentation so far, including details on training methods, benchmark results, chip models, or any open-source timeline. What is publicly visible at this stage remains limited to the invite-only API test, the model’s AI agent focus, its million-token context capacity, and the free token quota.

