Microsoft launches MAI-Image-2.5-Pro and MAI-Voice-2-Flash, says GPU costs fell by as much as 89%

Microsoft launches MAI-Image-2.5-Pro and MAI-Voice-2-Flash, says GPU costs fell by as much as 89%

N
News Editor
2026-07-24 02:30:06
Microsoft has put two in-house AI models, MAI-Image-2.5-Pro and MAI-Voice-2-Flash, into public preview and paired the launch with internal figures meant to show how far its self-developed stack has moved into production. The company said a customer service deployment that switched to MAI-Voice-2-Flash cut compute costs by 89%, while PowerPoint image generation reduced GPU costs by 84% compared with OpenAI’s GPT-Image-2. Microsoft also said Bing Image Creator has fully moved to MAI-Image-2.5, OneDrive improved storage rate by 26% and lowered latency by about 25% after changing models, and Dragon Copilot reduced transcription and language-recognition error rates by 50% after adopting MAI-Transcribe-1.5 across workflows spanning 58 languages. Satya Nadella said on X that Microsoft will route traffic to MAI whenever its own models match or beat frontier alternatives, though he added that models from OpenAI and Anthropic remain part of the company’s broader orchestration system. Microsoft is also turning its internal training approach, which it calls a “hill-climbing” strategy, into Azure products through Foundry and Frontier Tuning for enterprise customers.

Microsoft has launched two in-house AI models, MAI-Image-2.5-Pro and MAI-Voice-2-Flash, in public preview and released internal figures showing sharp reductions in GPU costs across several products. The company said a customer service center that switched to MAI-Voice-2-Flash cut compute costs by 89%.

Microsoft lays out a more modular AI stack

The announcement adds to a shift that had been visible for months. Reuters reported in April that Microsoft’s once-exclusive licensing arrangement with OpenAI had been changed to a non-exclusive setup. The Information also reported in September last year that Microsoft had started bringing Anthropic models into some of its products.

On July 23, Microsoft CEO Satya Nadella posted on X under the title “Frontier diffusion and control.” He said frontier capabilities have become saturated enough that they can now be delivered at lower cost through models optimized for high-frequency use cases at scale. His wording was direct: once Microsoft’s own model is on par with or better than a frontier alternative, traffic starts moving to MAI.

Nadella also said frontier models from OpenAI and Anthropic remain part of Microsoft’s overall orchestration system. At the same time, he described a principle of model independence: the evaluation system should keep improving even if any one model is removed. In practice, that means the supporting system, memory, and skills sit outside the model itself, leaving Microsoft with control at the system layer.

That produces a clear structure. Microsoft acts as the orchestrator. Frontier models from OpenAI and Anthropic become replaceable components, while Microsoft’s own models absorb a larger share of routine traffic.

Internal training methods are being turned into Azure products

Microsoft tied that strategy to a separate methodological note built around what it calls a “hill-climbing” approach. The idea, as described by the company, is to create a feedback loop linking data, models, and product-side supporting systems, instead of treating one training run as the final answer.

Its main example is MAI-Code-1-Flash, a lightweight model used in GitHub Copilot. Microsoft said the model delivered roughly 10% higher code adoption in VS Code than GPT-5.4 Mini and Claude Haiku 4.5, while using 10% fewer tokens.

Microsoft then placed that model into an Excel reinforcement learning environment for additional training. The company said the result was a model small enough to run on older H100s and even A100 GPUs, yet able to match GPT-5.6 on most Excel tasks. That, in Microsoft’s telling, removes the need to spend the latest-generation chips on those workloads and frees newer GB200 clusters for training instead of serving requests.

Microsoft is not keeping that process in-house. Through Foundry and what it calls Frontier Tuning, enterprise customers can train custom models on their own data. In effect, Microsoft is packaging its internal cost-cutting playbook as new Azure products.

The company also stressed that the models were trained on “clean, traceable enterprise-grade data” and were not distilled from third-party models. That line speaks to enterprise buyers and regulators at a time when scrutiny over AI training data sources has grown.

Microsoft highlights results across Bing, PowerPoint, OneDrive, and healthcare

Microsoft backed the strategy with product metrics from internal evaluations. Bing Image Creator has fully moved to MAI-Image-2.5. In PowerPoint, GPU costs fell 84% compared with OpenAI’s GPT-Image-2. In OneDrive, switching models improved storage rate by 26% and cut latency by about 25%.

The most sensitive deployment mentioned was in healthcare. Microsoft said Dragon Copilot, which serves 170,000 clinicians and processed 28 million patient records last quarter, shifted its transcription workflow across 58 languages to MAI-Transcribe-1.5. After that change, transcription and language-recognition error rates fell by 50% on a relative basis.

As for the two newly announced models, MAI-Image-2.5-Pro is priced at $5 per million text input tokens and $106 per million image output tokens. Rob Reilly, global chief creative officer at WPP, said in Microsoft’s announcement that the model marks “a major leap forward in generative media tools.”

Microsoft said MAI-Voice-2-Flash is twice as fast as its predecessor and costs 32% less. It is aimed at high-volume settings such as customer service centers, where throughput matters more than top-end model performance.

Reaction is split, and the metrics are Microsoft’s own

Response on social platforms has been mixed. One designer wrote on X that Microsoft has long been poor at listening to users and argued that this would cost the company in the AI race. Another user mocked Microsoft’s model-independence framing, saying they would like to see whether the system keeps improving once Microsoft itself is removed.

Others backed the strategy. One developer asked why changing an Excel column should require an all-knowing large model at all. Another summed up the pitch more bluntly: do not send every task to the biggest model, and both cost and performance improve.

Still, the skepticism has a basis. The adoption rates, storage figures, and GPU savings cited in the release all come from Microsoft’s internal testing rather than third-party benchmarks, and Microsoft chose which numbers to publish.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
6300

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.