
Technology media outlet TestingCatalog published a blog post on August 2, reporting that Microsoft is testing its first native real-time voice model, MAI Realtime. It supports 16 languages and 2 voices, can listen and speak simultaneously during conversations, and supports automatic language detection and switching languages mid-conversation.
For testing, Microsoft has reportedly invited some partners to test the model on the MAI Playground platform. It is not yet clear when it will launch on Microsoft Foundry or expand to Copilot's voice features.
In terms of operation, MAI Realtime uses bidirectional, full-duplex interaction. It does not require strict turn-taking in which the user must finish speaking before the model responds, and can listen and respond simultaneously.
In terms of supported languages, the MAI Realtime model supports 16 languages, including Chinese. Citing the blog post, ITHome lists the supported languages as follows:
English
German
Spanish
French
Italian
Portuguese
Japanese
Korean
Chinese
Dutch
Hindi
Indonesian
Arabic
Russian
Turkish
Vietnamese
Thai
In terms of voice styles, the model reportedly offers two voices in testing: Victoria and Grant, which are more natural than the voice modes currently available in Copilot. Users can specify a language manually or enable automatic detection.


The model does not support singing or generating non-speech sound effects; it remains positioned as a conversational system. The debugging panel can display real-time latency, the model's thoughts, and processing steps. The sample-sharing feature may become available to platform users after testing access is expanded.

Note: Full-duplex voice means that a system can receive the user's voice and output a response at the same time, without waiting for either party to finish speaking completely. It relies on endpoint detection to identify when speech starts, pauses, and ends, and must also adjust its output when the user interrupts.
Microsoft's previously released MAI voice models were all unidirectional, including MAI-Voice-2 and its Flash version for speech synthesis, as well as MAI-Transcribe-1.5 for speech recognition.
