Skip to main content
Models & Technology

Microsoft Tests Its First Self-Developed Full-Duplex AI Voice MAI Realtime Model: Supports 16 Languages Including Chinese and 2 Styles

Technology media outlet TestingCatalog published a blog post on August 2, reporting that Microsoft is testing its first native real-time voice model, MAI Realtime. It supports 16 languages and 2 voices, can listen and speak simultaneously during conversations, and supports automatic language detection and switching languages mid-conversation.

Microsoft Tests Its First Self-Developed Full-Duplex AI Voice MAI Realtime Model: Supports 16 Languages Including Chinese and 2 Styles

Technology media outlet TestingCatalog published a blog post on August 2, reporting that Microsoft is testing its first native real-time voice model, MAI Realtime. It supports 16 languages and 2 voices, can listen and speak simultaneously during conversations, and supports automatic language detection and switching languages mid-conversation.

For testing, Microsoft has reportedly invited some partners to test the model on the MAI Playground platform. It is not yet clear when it will launch on Microsoft Foundry or expand to Copilot's voice features.

In terms of operation, MAI Realtime uses bidirectional, full-duplex interaction. It does not require strict turn-taking in which the user must finish speaking before the model responds, and can listen and respond simultaneously.

In terms of supported languages, the MAI Realtime model supports 16 languages, including Chinese. Citing the blog post, ITHome lists the supported languages as follows:

English

German

Spanish

French

Italian

Portuguese

Japanese

Korean

Chinese

Dutch

Hindi

Indonesian

Arabic

Russian

Turkish

Vietnamese

Thai

In terms of voice styles, the model reportedly offers two voices in testing: Victoria and Grant, which are more natural than the voice modes currently available in Copilot. Users can specify a language manually or enable automatic detection.

Microsoft Tests Its First Self-Developed Full-Duplex AI Voice MAI Realtime Model: Supports 16 Languages Including Chinese and 2 Styles
Microsoft Tests Its First Self-Developed Full-Duplex AI Voice MAI Realtime Model: Supports 16 Languages Including Chinese and 2 Styles

The model does not support singing or generating non-speech sound effects; it remains positioned as a conversational system. The debugging panel can display real-time latency, the model's thoughts, and processing steps. The sample-sharing feature may become available to platform users after testing access is expanded.

Microsoft Tests Its First Self-Developed Full-Duplex AI Voice MAI Realtime Model: Supports 16 Languages Including Chinese and 2 Styles

Note: Full-duplex voice means that a system can receive the user's voice and output a response at the same time, without waiting for either party to finish speaking completely. It relies on endpoint detection to identify when speech starts, pauses, and ends, and must also adjust its output when the user interrupts.

Microsoft's previously released MAI voice models were all unidirectional, including MAI-Voice-2 and its Flash version for speech synthesis, as well as MAI-Transcribe-1.5 for speech recognition.