Microsoft MAI-Transcribe-2-Streaming tops Artificial Analysis: 2.5% WER with first partials ~100ms after speech — Japanese streamer demos it live: "the meeting-notes app era is over"

Microsoft AI released its first streaming speech-to-text model, MAI-Transcribe-2-Streaming, on Oct 1, taking the #1 spot on Artificial Analysis AA-WER Streaming for both final and first-partial transcript accuracy: 2.5% WER (2.5 errors per 100 words) with a stable transcript committed 0.13s after end of speech, dethroning Grok Voice Transcribe 2.0 (2.7%, 0.49s). It supports 60 languages with continuous automatic language detection, produces first partials ~100ms after receiving audio, and revises as context arrives — voice agents can start reasoning or calling tools mid-sentence; internal dictation/subtitling tests show words appearing 2x faster than the closest competitor. Priced at $0.54/audio-hour (~$9 per 1,000 minutes, introductory through year-end) — the premium end (ElevenLabs Scribe v2 Realtime and Deepgram Flux charge $6.50). Also launched: MAI-Voice-2.1 ($22/M chars) and MAI-Voice-2.1-Flash ($15/M chars). Japanese streamer studio_yebisu demoed it live (34s video, ElevenLabs v4 TTS fed via mic): very high accuracy and speed on common vocabulary, imperfect on whispers and proper nouns; he argues meeting-notes transcription apps are under threat though custom-noun dictionaries retain value. Another in-house win for Microsoft''s audio stack — transcription commoditization is squeezing standalone speech-AI vendors'' pricing power.




