Robots Atlas>ROBOTS ATLAS
Microsoft MAI

Microsoft MAI

MAI-1-preview, MAI-Voice-1 (2025) · Family: MAI
Microsoft's first in-house AI models, announced 28 August 2025: MAI-1-preview (MoE text model) and MAI-Voice-1 (speech generation model).
⏳ Preview⏳ Limited accessLLMAudio📁 MAI
Release date
28 August 2025
Access:APIHostedDeployment:☁ Cloud

Overview

Microsoft MAI is the family of the first artificial-intelligence models built in-house by the Microsoft AI division, unveiled on 28 August 2025. The launch comprised two models: MAI-1-preview and MAI-Voice-1.

MAI-1-preview

MAI-1-preview is a text model based on a mixture-of-experts (MoE) architecture, pre-trained and post-trained on roughly 15,000 NVIDIA H100 GPUs. It is designed to handle everyday queries and follow user instructions. The model was made available for public evaluation on LMArena, and Microsoft said it would roll out to selected text use cases in Copilot; API access requires an application.

MAI-Voice-1

MAI-Voice-1 is a speech-generation model. According to Microsoft it can generate a full minute of audio in under a second on a single GPU and supports both single- and multi-speaker scenarios. It powers the Copilot Daily and Podcasts features and is available in Copilot Labs.

Context

The MAI models (named after "Microsoft AI") are part of Microsoft's strategy to build its own models alongside its existing partnership with OpenAI. The Microsoft AI division is led by Mustafa Suleyman (EVP and CEO of Microsoft AI). Microsoft has not disclosed the parameter count or context-window size of these models.

Classification
LLMAudio
Family: MAI
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📥 Input: text

Technical specification

License
Proprietary (closed weights)
Modalities
⬇ Input
text
⬆ Output
textaudio

Capabilities and applications

Native model capabilities
Instruction following
Precisely following instructions contained in the prompt: response format, length, style, constraints (e.g. 'reply in six words'). GPT-5.1 significantly improved this capability compared to GPT-5.
Category: language
Language modeling
Ability to predict subsequent tokens and generate coherent natural-language text based on the preceding context.
Category: language
Text to speech
Category: speech
Natural conversation
Conducting a conversation with a tone close to human: a warmer voice, empathy in emotional responses, humour, and avoiding the stiff 'AI assistant jargon'. Introduced as a deliberate improvement in GPT-5.1 Instant.
Category: language
Real-time inference
The model's ability to generate responses with very low latency (>1000 tokens/sec) on specialized inference hardware (e.g. Cerebras WSE), enabling interactive, turn-by-turn collaboration with a human.
Category: coding

Technical architecture