AI IndustryMicrosoftJul 25, 2026 21:18 UTC

Microsoft Releases Two In-House Developed AI Models

Microsoft unveiled two new self-developed AI models on Wednesday as public previews: the image generation model 'MAI-Image-2.5-Pro' and the voice model 'MAI-Voice-2-Flash'. Concurrently, the company released operational data showing a maximum 89% reduction in GPU costs through adoption of its own models. The announcement revealed that these models are already running in major Microsoft services such as Bing, PowerPoint, and GitHub Copilot, demonstrating that the company is actively building its own independent foundation without reliance on OpenAI models.

Microsoft Releases Two In-House Developed AI Models

Microsoft unveiled two newly developed AI models on Wednesday as public previews. The two models are the image generation model 'MAI-Image-2.5-Pro' and the voice model 'MAI-Voice-2-Flash', and the company simultaneously released detailed data demonstrating actual deployment across its own products. The announcement was made by Microsoft AI's 'Superintelligence Team', and the simultaneous introduction of two models covering different use cases from image generation to voice processing demonstrates that the company's proprietary model strategy has entered a full-scale deployment phase.

This development occurs approximately one year after Microsoft began committing seriously to in-house model development. Previously, the company had pursued an approach of leveraging state-of-the-art models through investment in OpenAI; however, in this announcement, major products and services including Bing, PowerPoint, OneDrive, Dynamics 365, Excel, GitHub Copilot, and Azure were specifically named as deployment targets for proprietary models. In its official blog, the company stated that the goal is to 'run Microsoft products with Microsoft models', clearly demonstrating that proprietary models have transitioned from the research stage to serve as foundational infrastructure for actual operations.

MAI-Image-2.5-Pro is a model specialized in high-quality image generation, positioned for use cases including commercial primary image creation, fine-grained editing, and accurate rendering of text within images. Text rendering accuracy within images is widely known as an area where traditional image generation models have struggled, and addressing this aspect is being emphasized. Pricing is set at 5 dollars per million text input tokens, 8 dollars per million image input tokens, and 106 dollars per million image output tokens. Meanwhile, MAI-Image-2.5, the base model, achieved second place in the image editing category of Arena, a generative media evaluation platform.

MAI-Voice-2-Flash, on the other hand, is designed with emphasis on cost efficiency and processing speed. This model, first unveiled at Microsoft's 'Build' conference, offers processing speed twice that of its predecessor MAI-Voice-2, with costs 32% lower, and is priced at 15 dollars per million characters. Primary targets are business use cases requiring massive voice processing such as call centers, voice agents, and real-time voice processing, aiming at markets where response speed and per-transaction cost determine competitive advantage.

Rob Reilly, Global Chief Creative Officer at advertising giant WPP, described MAI-Image-2.5-Pro as 'a significant advance as a GenMedia tool' in his comments on Microsoft's announcement, stating that 'Microsoft has established a solid position among generative AI leaders'. According to Microsoft's explanation, the simultaneous release of models with different characteristics—one optimized for high quality and one for high-volume processing—stems from recognition that users seeking maximum quality in creative domains and enterprises needing to handle massive processing volumes at low cost have fundamentally different requirements.

What deserves particular attention in this announcement is the disclosure of operational data rather than the models themselves. Microsoft released data showing that adoption of its proprietary models achieved maximum GPU cost reductions of 89%, serving as evidence of its claim that proprietary products can adequately support the company's business without relying on OpenAI's frontier models. Substantial cost reduction is an extremely practical value proposition for enterprise customers, and beyond merely demonstrating technological capability, this messaging functions as an impetus for transition to proprietary models from the perspective of economic rationality.

For Microsoft, this development cannot be separated from the context of its relationship with OpenAI. While the company is OpenAI's largest investor and joint development partner, this proactive positioning of proprietary model competitiveness and economic efficiency can be viewed as signaling a direction toward reducing that dependency over time. Going forward, which product domains proprietary models will replace OpenAI models in, and how far that scope will expand, will be key points of interest.

#GenerativeAI#Microsoft#ImageGeneration#VoiceAI#AICost#EnterpriseAI#OpenAI
AI issue Staff

This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.

Comments

Log in to comment