Ant Ling releases Ling-3.0-flash-VL, its first native multimodal open-source model

Ant Ling releases Ling-3.0-flash-VL, its first native multimodal open-source model

N
News Editor
2026-09-09 04:17:23
Ant Ling on Sept. 9 officially open-sourced Ling-3.0-flash-VL, described as its first native multimodal model. The model had already been available through an API, and the latest release includes BF16 and FP8 weights under the MIT license, with FP4 and INT4 versions planned later. According to the announcement cited by BlockBeats, the model can read images, videos, documents and software interfaces, then use tools to complete tasks and keep checking and revising its work based on execution results. In one example, it can build a webpage by following a reference image, generating code, reviewing the rendered page and adjusting it on its own. Ling-3.0-flash-VL uses a Mixture-of-Experts, or MoE, architecture with about 124 billion total parameters and roughly 5.5 billion activated per inference. Its context window can scale up to 1 million tokens. Ant Ling said the model carries over the core strength of Ling-3.0-flash as an efficient execution node in Agent workflows, balancing output quality and execution efficiency inside a visual feedback loop while pushing full tasks forward at lower cost and in less time.

Ant Ling has officially open-sourced Ling-3.0-flash-VL, its first native multimodal model, according to BlockBeats on Sept. 9.

The model had already been launched through an API. In this release, Ant Ling published BF16 and FP8 weights under the MIT license, while FP4 and INT4 versions are set to follow later.

Ling-3.0-flash-VL can read images, videos, documents and software interfaces, then call tools to carry out tasks and continue checking and revising its work based on the results. For example, when building a webpage, the model can write code from a reference image, review the actual page output, and adjust it on its own.

The model uses a Mixture-of-Experts, or MoE, architecture, with about 124 billion total parameters and around 5.5 billion activated in a single inference. Its context window can scale to as much as 1 million tokens.

Ant Ling said Ling-3.0-flash-VL inherits the core strengths of Ling-3.0-flash as an efficient execution node in Agent workflows, balancing output quality and execution efficiency in a visual feedback loop while moving complete tasks forward with lower cost and less time.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
1200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.