Microsoft has launched two in-house AI models, MAI-Image-2.5-Pro and MAI-Voice-2-Flash, in public preview and released internal figures showing sharp reductions in GPU costs across several products. The company said a customer service center that switched to MAI-Voice-2-Flash cut compute costs by 89%.
Microsoft lays out a more modular AI stack
The announcement adds to a shift that had been visible for months. Reuters reported in April that Microsoft’s once-exclusive licensing arrangement with OpenAI had been changed to a non-exclusive setup. The Information also reported in September last year that Microsoft had started bringing Anthropic models into some of its products.
On July 23, Microsoft CEO Satya Nadella posted on X under the title “Frontier diffusion and control.” He said frontier capabilities have become saturated enough that they can now be delivered at lower cost through models optimized for high-frequency use cases at scale. His wording was direct: once Microsoft’s own model is on par with or better than a frontier alternative, traffic starts moving to MAI.
Nadella also said frontier models from OpenAI and Anthropic remain part of Microsoft’s overall orchestration system. At the same time, he described a principle of model independence: the evaluation system should keep improving even if any one model is removed. In practice, that means the supporting system, memory, and skills sit outside the model itself, leaving Microsoft with control at the system layer.
That produces a clear structure. Microsoft acts as the orchestrator. Frontier models from OpenAI and Anthropic become replaceable components, while Microsoft’s own models absorb a larger share of routine traffic.
Internal training methods are being turned into Azure products
Microsoft tied that strategy to a separate methodological note built around what it calls a “hill-climbing” approach. The idea, as described by the company, is to create a feedback loop linking data, models, and product-side supporting systems, instead of treating one training run as the final answer.
Its main example is MAI-Code-1-Flash, a lightweight model used in GitHub Copilot. Microsoft said the model delivered roughly 10% higher code adoption in VS Code than GPT-5.4 Mini and Claude Haiku 4.5, while using 10% fewer tokens.
Microsoft then placed that model into an Excel reinforcement learning environment for additional training. The company said the result was a model small enough to run on older H100s and even A100 GPUs, yet able to match GPT-5.6 on most Excel tasks. That, in Microsoft’s telling, removes the need to spend the latest-generation chips on those workloads and frees newer GB200 clusters for training instead of serving requests.
Microsoft is not keeping that process in-house. Through Foundry and what it calls Frontier Tuning, enterprise customers can train custom models on their own data. In effect, Microsoft is packaging its internal cost-cutting playbook as new Azure products.
The company also stressed that the models were trained on “clean, traceable enterprise-grade data” and were not distilled from third-party models. That line speaks to enterprise buyers and regulators at a time when scrutiny over AI training data sources has grown.
Microsoft highlights results across Bing, PowerPoint, OneDrive, and healthcare
Microsoft backed the strategy with product metrics from internal evaluations. Bing Image Creator has fully moved to MAI-Image-2.5. In PowerPoint, GPU costs fell 84% compared with OpenAI’s GPT-Image-2. In OneDrive, switching models improved storage rate by 26% and cut latency by about 25%.
The most sensitive deployment mentioned was in healthcare. Microsoft said Dragon Copilot, which serves 170,000 clinicians and processed 28 million patient records last quarter, shifted its transcription workflow across 58 languages to MAI-Transcribe-1.5. After that change, transcription and language-recognition error rates fell by 50% on a relative basis.
As for the two newly announced models, MAI-Image-2.5-Pro is priced at $5 per million text input tokens and $106 per million image output tokens. Rob Reilly, global chief creative officer at WPP, said in Microsoft’s announcement that the model marks “a major leap forward in generative media tools.”
Microsoft said MAI-Voice-2-Flash is twice as fast as its predecessor and costs 32% less. It is aimed at high-volume settings such as customer service centers, where throughput matters more than top-end model performance.
Reaction is split, and the metrics are Microsoft’s own
Response on social platforms has been mixed. One designer wrote on X that Microsoft has long been poor at listening to users and argued that this would cost the company in the AI race. Another user mocked Microsoft’s model-independence framing, saying they would like to see whether the system keeps improving once Microsoft itself is removed.
Others backed the strategy. One developer asked why changing an Excel column should require an all-knowing large model at all. Another summed up the pitch more bluntly: do not send every task to the biggest model, and both cost and performance improve.
Still, the skepticism has a basis. The adoption rates, storage figures, and GPU savings cited in the release all come from Microsoft’s internal testing rather than third-party benchmarks, and Microsoft chose which numbers to publish.

