Qwen has introduced Qwen3.8-Omni-Flash, its first multimodal model designed specifically for AI agents. According to Techub, the model can process audio and video at the same time and use tools on its own to handle tasks such as editing video blogs, translating video clips, and summarizing movie content. The report, citing The Decoder, said Qwen3.8-Omni-Flash performs nearly on par with Google’s Gemini 3.8 Flash in audio and video benchmark tests. Its API pricing, however, is only a fraction of the cost of Google’s model. The launch positions Qwen with a lower-cost offering in multimodal AI, centered on agent use cases that require both media understanding and tool use.
Qwen has launched Qwen3.8-Omni-Flash, its first multimodal model designed specifically for AI agents, according to Techub.
The model can handle audio and video at the same time. It can also use tools independently to edit video blogs, translate video clips, and summarize movie content.
Techub, citing The Decoder, said the model delivers audio and video benchmark performance that is nearly comparable to Google’s Gemini 3.8 Flash, while API call costs are only a fraction of the latter.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.