AI Tools2026-10-038 min read

Microsoft MAI-Transcribe-2-Streaming Speaks the OpenAI Realtime Protocol. Check the Subprocessor Before You Swap

Microsoft released MAI-Transcribe-2-Streaming and two MAI-Voice-2.1 speech models on 1 October 2026, in public preview, with promotional pricing that ends on 31 December. Because the streaming API is compatible with the OpenAI Realtime protocol, moving your call audio to a new vendor is now a config change. Here is what to check first.

Microsoft AI released MAI-Transcribe-2-Streaming on 1 October 2026, its first real time speech to text model, alongside two text to speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. The transcription model takes audio as a continuous stream and returns text while the caller is still talking, in 60 languages with automatic language detection. Artificial Analysis placed it first on its streaming word error benchmark at a 2.5 percent final transcript error rate, returned about 0.13 seconds after the end of speech, ahead of the previous leader at 2.7 percent and roughly half a second. All three models are available in Microsoft Foundry, Azure AI Speech, the MAI Playground and the Vercel AI Gateway, with OpenRouter listed as well.

The detail that matters most for engineering teams is not the benchmark. Microsoft exposes the streaming model over a WebSocket API that is compatible with the OpenAI Realtime protocol. If you built a voice agent, a call summariser or a live captioning feature against OpenAI, as many teams did after the GPT Live 1 launch we covered in September, you can point it at Microsoft by changing an endpoint and a key. That is good for cost and resilience. It also means a developer can move every customer call you process to a new subprocessor in a single pull request, without anyone in security, privacy or procurement noticing.

Read the pricing as a launch offer rather than a budget line. Streaming transcription costs 0.54 US dollars per hour of audio, or 9 dollars per 1,000 minutes, and that rate is promotional through 31 December 2026. The batch model, MAI-Transcribe-2, stays at 0.10 dollars per hour, so streaming is about five times the price for the same words, and you are paying for latency and a held open connection. MAI-Voice-2.1 is 22 dollars per million characters and the Flash variant, with about 150 milliseconds to first audio across 23 languages, is 15 dollars per million characters. Permanent prices have not been published. If your business case only works at the promotional rate, write that assumption down now, because it expires inside the current quarter for most finance teams.

All three models are in public preview. In Azure terms that means no service level agreement and a recommendation not to run production workloads on them until general availability. For a contact centre, a voice agent or anything customer facing, that is an availability decision, not a footnote. Under ISO 27001 Annex A 5.23 on information security for cloud services and 5.30 on ICT readiness for business continuity, and under SOC 2 availability criteria A1.1 and A1.2, you should be able to show a fallback path: a second transcription vendor, a batch reprocess, or a human queue when the stream drops.

The launch coverage says nothing specific about data retention, data residency or whether audio is used to improve the models, and speaker diarization for the streaming model has not been confirmed. Do not fill those gaps from blog posts, including this one. Get the answers from the Microsoft Product Terms and Data Protection Addendum that apply to your Azure agreement, confirm which region your Foundry deployment actually runs in, and check whether preview services carry different terms from generally available ones. Call audio is some of the most sensitive data a company holds. It contains names, health details, addresses and, in sales and billing calls, payment card numbers read aloud.

That last point catches people. Under PCI DSS 4.0 requirement 3.3.1, sensitive authentication data such as the card verification code must not be stored after authorisation, and a transcript that captures a caller reading out a card number and its security code is storage, whether or not anyone meant it to be. If a streaming transcript lands in your CRM, your data warehouse or an LLM prompt log, your cardholder data environment has just grown. Use pause and resume or DTMF masking during payment capture, redact numbers before transcripts are persisted, and keep the transcription service out of scope by design rather than by hope. Our PCI DSS requirement pages walk through how scoping decisions like this are evidenced.

The voice models raise a separate set of questions. Both MAI-Voice-2.1 variants support voice cloning from a few seconds of reference audio, and Microsoft describes access as gated, with built in consent verification aligned with the FTC 2025 Voice Cloning Rule. Treat that as the vendor side of the control, not yours. Keep your own record of who consented, to what use, and for how long, and do not clone staff voices for outbound agents without a written agreement that survives their departure. If you serve EU customers, the EU AI Act Article 50 transparency duties require that people are told when they are interacting with an AI system, and synthetic audio that imitates a real person is exactly what those duties target. A natural sounding cloned voice makes the disclosure more important, not less.

Map it to the frameworks you already report against. ISO 27001 Annex A 5.19 to 5.22 on supplier relationships and 5.34 on privacy and protection of PII apply to any new speech vendor, and 8.11 on data masking applies to card numbers and identifiers in transcripts. ISO 42001 Annex A controls on third party and customer relationships and on the AI system lifecycle expect you to record the model, version and intended use, which a preview model with no published permanent price will change under you. For SOC 2, CC9.2 on vendor risk management and the privacy criteria are where an auditor will look. In Vanta, Drata, Secureframe or Sprinto, that means adding Microsoft AI speech as a distinct vendor entry with its own data categories, even if Azure is already on your list.

A short list for this week. Search your codebase and environment variables for Realtime endpoints so you know which services could be repointed, and require security review for any change to them. Decide whether preview speech models are allowed in production, and write the answer into your AI acceptable use policy. Confirm retention, training use and region in writing before any customer audio is sent. Add payment card redaction to every transcript path. Our view is that the MAI models look like a strong, cheap option for voice agents, and protocol compatibility is a real win against lock in. The same compatibility removes the friction that used to force a vendor review, so put that review back on purpose.

MicrosoftMAI-Transcribe-2-StreamingMAI-Voice-2.1Azure AI Foundryvoice agentsspeech to textvoice cloningsubprocessorsPCI DSSEU AI ActISO 27001ISO 42001SOC 2VantaDrata

Editorial note: AES Tech reviews are independent. Some outbound links are affiliate links and are marked sponsored; they never change our rankings. See our disclosure.

// Signal, not noise

Get the next post by email

One short email when something worth knowing ships. No spam, unsubscribe anytime.

Loading comments...

Add a comment

Corrections and first-hand experience are the most useful things you can leave. Comments are screened automatically and reviewed by a human; see the moderation policy.

0/4000 · plain text · links are held for review

More from the blog