Microsoft unveils MAI-Image-2.5-Pro and MAI-Voice-2-Flash AI models for image generation and speech recognition
Microsoft
OpenAI
Anthropic
Microsoft has released two new proprietary AI models: MAI-Image-2.5-Pro for high-quality image generation and MAI-Voice-2-Flash for low-latency speech recognition in enterprise workloads. The company highlights significant cost savings and performance improvements when replacing OpenAI models with its own across products like Bing, PowerPoint, and Dynamics 365.
Microsoft has launched public previews of two new internally developed AI models. MAI-Image-2.5-Pro is the flagship image generator, offering high-quality generation, detailed editing, and accurate text rendering. MAI-Voice-2-Flash is optimized for high-throughput voice applications like call centers, running twice as fast and costing 32% less than the base model. The company reports that Bing Image Creator now fully runs on MAI-Image-2.5, and PowerPoint has reduced GPU resource costs by 84% compared to using OpenAI's GPT-Image-2. MAI-Voice-2-Flash deployed in Dynamics 365 Contact Center cut GPU costs by up to 89% and is also integrated into Azure Voice Live. Microsoft's Dragon Copilot, used by 170,000 healthcare organizations, was migrated to MAI-Transcribe-1.5 supporting 58 languages. The lightweight MAI-Code-1-Flash model on GitHub Copilot shows 10% better code acceptance than GPT-5.4 Mini and Claude Haiku 4.5, with lower token consumption, and was further trained for Excel tasks where it performs at GPT-5.6 level while running on older Nvidia H100 and A100 accelerators.
- Сокращения
- GPU = Graphics Processing Unit — графический процессор
- API = Application Programming Interface — интерфейс программирования приложений
Source: 3DNews —
original
