Skip to main content

Vercel AI Gateway Adds Microsoft MAI Voice and Streaming Transcription

3 min read

Vercel AI Gateway now offers Microsoft MAI voice generation and streaming transcription. Builders should test latency, audio format, language quality and retention terms.

Vercel AI Gateway Adds Microsoft MAI Voice and Streaming Transcription

Vercel has added three Microsoft audio models to AI Gateway. They share an access layer, but they solve different parts of the voice stack.

Vercel announced the Microsoft AI models on October 1. The additions are MAI-Voice-2.1, MAI-Voice-2.1-Flash, and MAI-Transcribe-2 Streaming. You can call them through Vercel AI Gateway and AI SDK 7.

Choose the model by interaction pattern.

MAI-Voice-2.1 is the general speech-generation option and is described as supporting 23 languages. MAI-Voice-2.1-Flash is positioned for lower latency. MAI-Transcribe-2 Streaming returns partial transcription results as audio arrives. A produced narration, conversational agent, and live caption system therefore need different tests.

Do not select Flash from the name alone. Measure time to first audio, complete response time, pronunciation, and stability from the same region and client used in production. For transcription, measure partial-result churn as well as the final word error rate.

The audio format is part of the API contract.

Vercel’s transcription example uses 16 kHz, 16-bit PCM audio. A browser or telephony source may provide a different rate, channel layout, or codec. Conversion can add latency and can reduce recognition quality when implemented carelessly.

Preserve a small set of consented test recordings with known transcripts. Run them through the complete capture, conversion, and streaming path. This catches errors that a model-only benchmark cannot show.

No platform markup is not the same as free.

Vercel says AI Gateway does not add a platform fee or markup to inference prices. The listed model rates still apply, and an application can incur additional costs for storage, streaming infrastructure, retries, and observability. Budget accepted audio minutes, not raw requests alone.

For generated speech, count discarded takes and post-processing. For transcription, count audio minutes sent more than once after reconnects. A low per-unit rate can still produce waste when the client retries an entire stream.

Vercel presents the models as supporting Zero Data Retention through the Gateway. Treat that as a service statement to verify against the account configuration and current terms. Voice data can identify people and may contain sensitive content. Obtain consent, minimize storage, and prevent raw audio from entering general-purpose logs.

  • Match the model to narration, conversation, or live transcription.
  • Test all production languages, accents, names, and noisy conditions.
  • Validate the sample rate, bit depth, channel layout and reconnect behavior.
  • Measure first output, final output, retries, and accepted-quality cost.
  • Document consent, retention, deletion and access controls for audio.

Our Eleven v4 API guide provides a useful comparison framework: endpoint, latency, request shape, and accepted-output cost matter more than a model name. Apply the same discipline to the new MAI audio routes.

Leave a comment

Your email address will not be published. Required fields are marked *