Microsoft launches streaming transcription and voice models
Microsoft expanded its MAI model family with three new releases: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash, all aimed at developers building conversational voice agents. The streaming transcription model accepts speech via WebSocket, delivers first transcript hypotheses within 320 milliseconds on average, supports over 60 languages, and is priced at 54 cents per audio hour. The two text-to-speech models support 23 languages, priced at $22 and $15 per million characters respectively. Together with Microsoft's existing MAI-Thinking-1 reasoning model, the trio gives developers a complete stack for building humanlike voice agents.
