Alibaba has released Qwen3.7-Plus, a multimodal large language model priced far below its previous generation. Official pricing puts input at $0.40 per million tokens and output at $1.60, for a combined $2.00. That compares with Qwen3.7-Max at $2.50 for input and $7.50 for output, or $10.00 in total, making the new model roughly 80% cheaper.
The price cut is only part of the shift. Qwen3.7-Plus does not come with open model weights. Like the earlier Max version, it is available only through Alibaba Cloud Model Studio’s commercial API and Qwen Chat. Alibaba is pushing multimodal capability and lower usage costs at the same time, but local deployment is no longer on the table.
Cached input pricing targets repetitive enterprise workloads
One figure likely to stand out for enterprise users is cached input pricing. If an agent repeatedly reads the same codebase or a fixed enterprise UI, the follow-up cost can drop to just $0.04 per million tokens. That pricing is aimed squarely at high-frequency, repetitive workflows such as developer operations, RPA, and data engineering pipelines.
Alibaba’s positioning is straightforward: Qwen3.7-Plus is meant to handle recurring development tasks and process-heavy jobs without invoking the most expensive frontier model every time. It is built for volume and repetition. That makes cost control a core part of the pitch rather than a side benefit.
Benchmark claims show strength in specific tasks
Based on Alibaba’s self-reported results, Qwen3.7-Plus scored 70.3 on Terminal Bench 2.0-Terminus, ahead of DeepSeek-V4-Pro Max at 67.9 and Gemini-3.1 Pro at 63.5. On ScreenSpot Pro, which measures computer vision and interface understanding, it posted 79.0, above GPT-5.4 xhigh at 67.4 and Claude-Opus-4.6 at 49.5.
Those numbers still come with limits. The source material notes that strong point performance does not mean the model has broadly caught up with leading closed-source US systems. Benchmarks can indicate where a model starts, but they do not settle questions around runtime speed, reliability, or production behavior. Buyers would still need to validate those factors on their own workloads.
Closed delivery changes the tradeoff for compliance and ecosystem growth
Qwen had long been one of the clearest names in China’s open-source AI camp. Apache 2.0 licensing and downloadable weights allowed companies and developers to fine-tune and deploy the models directly, helping build a wider ecosystem through platforms such as Hugging Face. Qwen3.7-Plus takes a different route.
The consequences are practical. Companies cannot deploy the model inside isolated internal networks, and API traffic must pass through Alibaba Cloud international nodes, meaning data moves outside a customer’s own servers. Industries facing data sovereignty or compliance constraints, including healthcare, defense, and government, may find that model access path hard to approve. At the same time, the closed API approach lowers infrastructure burdens because customers do not need to build and maintain large multi-GPU clusters, and the OpenAI-compatible format keeps integration changes relatively small.
On pricing, Alibaba is placing the model close to domestic competition. The source notes that rival MiniMax-M3 is running a limited-time offer at a combined $1.50, while Qwen3.7-Plus comes in at $2.00. The model is cheaper than before, but the tradeoff is explicit: the open-weight strategy that once helped Qwen build developer trust is giving way to a more conventional closed API business model.

