News

Microsoft ships MAI-Transcribe-2-Streaming and two new MAI-Voice models

Microsoft added its first streaming speech-to-text model plus two text-to-speech models to the MAI family, aimed at voice agents.

Dispatch news card about Microsoft MAI streaming transcription and voice models. Source: siliconangle.com

On October 1, Microsoft expanded its MAI model family with three models for voice agents: MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. SiliconANGLE describes them as built for developers who want agents that listen and reply almost instantly, like a human conversation.

The three models

  • MAI-Transcribe-2-Streaming is Microsoft's first streaming transcription model. It accepts speech over a WebSocket and keeps updating the transcript while the person talks, then confirms when it is final. That allows live captions or starting a response before someone finishes speaking. It supports more than 60 languages with automatic language detection, and delivers first transcript hypotheses within about 320 ms on average. Microsoft says it cannot guarantee that in every case, because network and the reply system also affect speed.
  • MAI-Voice-2.1 is the text-to-speech model for more expressive, higher-fidelity output.
  • MAI-Voice-2.1-Flash trades some of that expressiveness for faster responses and lower cost.

Both voice models support 23 languages.

Prices

According to Vercel AI Gateway listings cited by SiliconANGLE, the streaming model costs 54 cents per audio hour. That is more than five times the 10 cents of the non-streaming MAI-Transcribe-2, which waits until the speaker is done before it starts. MAI-Voice-2.1 costs $22 per million characters, and the Flash version $15.

The bigger picture

A voice agent needs three steps: understand speech, decide what to do, and speak. Microsoft now offers a model for the first and third, and its reasoning model MAI-Thinking-1 can handle the middle. SiliconANGLE also links the release to Microsoft's wish to rely less on outside model providers: in July, AI chief Mustafa Suleyman reportedly told his researchers to double down on MAI, and he said "we pay a lot of money to Anthropic, so our goal is to reduce and ultimately eliminate that cost".

Why it matters

Latency is what makes voice agents feel natural or awkward. A 320 ms first guess lets the system prepare an answer while you are still talking. And with its own full stack, Microsoft gains more control over cost and quality.

Dany's take

Voice is where agents become something normal people use. I care less about the exact millisecond figure and more about whether it holds up on a bad connection and with accents. Test it with your own audio before you build on it, and compare the price of streaming against batch for your use case.

Source: SiliconANGLE, Microsoft targets ultra-realistic voice agents

Source: siliconangle.com

Newsletter

The AI news that matters, in your inbox.