Quick Facts

  • Microsoft’s MAI-Voice-2-Flash cuts GPU costs up to 89% versus OpenAI models in Dynamics 365 Contact Center deployments.
  • MAI-Image-2.5-Pro reduces GPU costs by up to 84% in PowerPoint compared to OpenAI’s GPT-Image-2, and is now the default image tool in OneDrive and Bing Image Creator.
  • MAI models are now deployed across more than half of Microsoft’s products, with replacements underway in Excel and Outlook.

Microsoft released two new in-house AI models into public preview on July 23: MAI-Image-2.5-Pro, its highest-fidelity image generator to date, and MAI-Voice-2-Flash, a speech model built for high-volume enterprise workloads. The releases mark the latest step in Microsoft’s accelerating push to build proprietary AI and reduce spending on outside vendors including OpenAI and Anthropic.

MAI-Voice-2-Flash now powers Dynamics 365 Contact Center, used by T-Mobile and EasyJet, where Microsoft reports GPU cost reductions of up to 89% compared to OpenAI equivalents. The model also integrates with Azure Voice Live for developers building speech-to-speech agents.

On the image side, Bing Image Creator now runs entirely on MAI-Image-2.5, end to end. In OneDrive, where the model is now the default for key image-editing scenarios, Microsoft reports a 26% increase in save rates, 25% lower P95 latency, and 2.5 times greater efficiency under medium-utilization production workloads.

Mustafa Suleyman, CEO of Microsoft AI, was direct about the financial motivation. “We pay a lot of money to Anthropic, so our goal is to reduce and ultimately eliminate that cost,” he wrote on X. He also noted that after optimizing models for McKinsey’s consulting workflows, Microsoft outperformed OpenAI’s GPT-5.5 with 10 times better cost efficiency.

The MAI Superintelligence team, which Suleyman leads, was formed in November 2025. The first MAI models shipped in August 2025, followed by MAI-Image-1 in October 2025. At Build 2026 in June, Microsoft launched seven in-house models across reasoning, coding, image generation, voice, and transcription. The pace of releases has picked up sharply since.

The contractual foundation for this independence was set in October 2025, when Microsoft and OpenAI restructured their partnership. Under the revised agreement, Microsoft gained the right to pursue artificial general intelligence independently or with other partners. The original deal had effectively barred Microsoft from building competing AI systems on its own.

Microsoft’s transcription model, MAI-Transcribe, now covers 58 languages in Dragon Copilot and cuts the error rate on multilingual clinical transcription in half. The company says it runs five times faster than competing models, with domain-specific terminology support across 43 languages.

Microsoft and Mayo Clinic also announced a strategic collaboration to develop a frontier AI model for healthcare. The model will be owned by Mayo Clinic and built using the clinic’s de-identified clinical data and longitudinal records. Microsoft plans to make it available through Azure Foundry APIs. “Frontier medical intelligence is around the corner,” Suleyman said of the partnership.

Microsoft CEO Satya Nadella framed the broader shift at Build 2026. “We believe the time has come for every company to just move from consuming a frontier model to fully participating at the frontier,” he said onstage.

For software executives watching their own AI infrastructure costs, the numbers Microsoft is reporting are significant. An 89% reduction in GPU costs on voice workloads, if replicated at scale, changes the math on AI-powered products across contact centers, productivity tools, and developer platforms. Microsoft is demonstrating that vertical integration on AI models is not just technically viable but economically compelling.

Read more: Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI

This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.