microsoft logo
microsoft

Microsoft AI: MAI-Voice-2.1-Flash

Source description (English)

MAI-Voice-2.1-Flash is a low-latency text-to-speech model from Microsoft AI, optimized for real-time responsiveness. It produces natural, expressive speech across 23 languages, with human-like intonation, rhythm, and emotional nuance. It is...

Overview

Model specifications

Context
Not provided
Maximum output
Not provided
Architecture
text->speech
Tokenizer
Other
Knowledge cutoff
Not provided
Moderated
No
OPENROUTER

Complete pricing

Synchronized OpenRouter rates. Token prices are shown per one million tokens.

Characters
$15
per million characters
API

Model configuration

Supported voices

cs-CZ-Grant:MAI-Voice-2.1-Flashcs-CZ-Harper:MAI-Voice-2.1-Flashda-DK-Grant:MAI-Voice-2.1-Flashda-DK-Harper:MAI-Voice-2.1-Flashde-DE-Grant:MAI-Voice-2.1-Flashde-DE-Harper:MAI-Voice-2.1-Flashde-DE-Klaus:MAI-Voice-2.1-Flashde-DE-Mia:MAI-Voice-2.1-Flashen-AU-Isla:MAI-Voice-2.1-Flashen-GB-Emily:MAI-Voice-2.1-Flashen-GB-Harry:MAI-Voice-2.1-Flashen-IN-Dhruv:MAI-Voice-2.1-Flashen-IN-Priya:MAI-Voice-2.1-Flashen-US-Ethan:MAI-Voice-2.1-Flashen-US-Grant:MAI-Voice-2.1-Flashen-US-Harper:MAI-Voice-2.1-Flashen-US-Iris:MAI-Voice-2.1-Flashen-US-Jasper:MAI-Voice-2.1-Flashen-US-Olivia:MAI-Voice-2.1-Flashen-US-Sage:MAI-Voice-2.1-Flashes-ES-Marta:MAI-Voice-2.1-Flashes-MX-Alejo:MAI-Voice-2.1-Flashes-MX-Grant:MAI-Voice-2.1-Flashes-MX-Harper:MAI-Voice-2.1-Flashes-MX-Valeria:MAI-Voice-2.1-Flashfi-FI-Grant:MAI-Voice-2.1-Flashfi-FI-Harper:MAI-Voice-2.1-Flashfr-FR-Grant:MAI-Voice-2.1-Flashfr-FR-Harper:MAI-Voice-2.1-Flashfr-FR-Marc:MAI-Voice-2.1-Flashfr-FR-Soleil:MAI-Voice-2.1-Flashhi-IN-Arjun:MAI-Voice-2.1-Flashhi-IN-Dhruv:MAI-Voice-2.1-Flashhi-IN-Grant:MAI-Voice-2.1-Flashhi-IN-Harper:MAI-Voice-2.1-Flashhi-IN-Kavya:MAI-Voice-2.1-Flashhi-IN-Priya:MAI-Voice-2.1-Flashhu-HU-Bence:MAI-Voice-2.1-Flashhu-HU-Grant:MAI-Voice-2.1-Flashhu-HU-Harper:MAI-Voice-2.1-Flashhu-HU-Levente:MAI-Voice-2.1-Flashhu-HU-Lilla:MAI-Voice-2.1-Flashhu-HU-Reka:MAI-Voice-2.1-Flashid-ID-Grant:MAI-Voice-2.1-Flashid-ID-Harper:MAI-Voice-2.1-Flashit-IT-Grant:MAI-Voice-2.1-Flashit-IT-Harper:MAI-Voice-2.1-Flashit-IT-Luca:MAI-Voice-2.1-Flashit-IT-Rosa:MAI-Voice-2.1-Flashko-KR-Grant:MAI-Voice-2.1-Flashko-KR-Haena:MAI-Voice-2.1-Flashko-KR-Harper:MAI-Voice-2.1-Flashko-KR-Junho:MAI-Voice-2.1-Flashnb-NO-Grant:MAI-Voice-2.1-Flashnb-NO-Harper:MAI-Voice-2.1-Flashnl-NL-Grant:MAI-Voice-2.1-Flashnl-NL-Harper:MAI-Voice-2.1-Flashnl-NL-Sander:MAI-Voice-2.1-Flashpl-PL-Grant:MAI-Voice-2.1-Flashpl-PL-Harper:MAI-Voice-2.1-Flashpt-BR-Caio:MAI-Voice-2.1-Flashpt-BR-Grant:MAI-Voice-2.1-Flashpt-BR-Harper:MAI-Voice-2.1-Flashpt-BR-Luana:MAI-Voice-2.1-Flashpt-BR-Pedro:MAI-Voice-2.1-Flashpt-BR-Rafael:MAI-Voice-2.1-Flashpt-PT-Grant:MAI-Voice-2.1-Flashpt-PT-Harper:MAI-Voice-2.1-Flashpt-PT-Rui:MAI-Voice-2.1-Flashro-RO-Andrei:MAI-Voice-2.1-Flashro-RO-Elena:MAI-Voice-2.1-Flashro-RO-Grant:MAI-Voice-2.1-Flashro-RO-Harper:MAI-Voice-2.1-Flashro-RO-Ioana:MAI-Voice-2.1-Flashro-RO-Radu:MAI-Voice-2.1-Flashru-RU-Grant:MAI-Voice-2.1-Flashru-RU-Harper:MAI-Voice-2.1-Flashru-RU-Lev:MAI-Voice-2.1-Flashru-RU-Masha:MAI-Voice-2.1-Flashsv-SE-Grant:MAI-Voice-2.1-Flashsv-SE-Harper:MAI-Voice-2.1-Flashth-TH-Grant:MAI-Voice-2.1-Flashth-TH-Harper:MAI-Voice-2.1-Flashth-TH-Krit:MAI-Voice-2.1-Flashth-TH-Nattapong:MAI-Voice-2.1-Flashtr-TR-Aydin:MAI-Voice-2.1-Flashtr-TR-Elif:MAI-Voice-2.1-Flashtr-TR-Grant:MAI-Voice-2.1-Flashtr-TR-Harper:MAI-Voice-2.1-Flashvi-VN-Grant:MAI-Voice-2.1-Flashvi-VN-Harper:MAI-Voice-2.1-Flashzh-CN-Bo:MAI-Voice-2.1-Flashzh-CN-Grant:MAI-Voice-2.1-Flashzh-CN-Harper:MAI-Voice-2.1-Flashzh-CN-Lan:MAI-Voice-2.1-Flashzh-CN-Mei:MAI-Voice-2.1-Flashzh-CN-Wei:MAI-Voice-2.1-Flash
API

Quick start

Set OPENROUTER_API_KEY locally. Python requires requests; JavaScript runs in Node.js. Keep the key on the server.

Choose a supported voice. Replace YOUR_VOICE_ID if no voice is listed.

API documentation

curl --fail-with-body https://openrouter.ai/api/v1/audio/speech \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"microsoft/mai-voice-2.1-flash","input":"Hello!","voice":"cs-CZ-Grant:MAI-Voice-2.1-Flash","response_format":"mp3"}' \
  --output speech.mp3
Capabilities

Input → Output

Input
text
Output
speech
API

Supported API parameters

Available providers

1 Provider

Checked: October 2, 2026

Live providers on OpenRouter

Provider availability, latency, throughput and routing can change continuously. Open the source page for current operational data.

OpenRouter

Azure

Not provided
Context
Not provided
Maximum output
Not provided
microsoft

4 models

All models
microsoft logomicrosoft

Microsoft: Microsoft-Decision-1

Microsoft-Decision-1 is a small model built for fast decision-making. Instead of generating text, it reads the provided content and returns a calibrated probability for each fixed answer option, so the...

Context
33K
Input
text
Output
decisions
Input: $0.042 per 1M tokensOutput: Free per 1M tokens
View model details →
microsoft logomicrosoft

Microsoft AI: MAI-Voice-2.1

MAI-Voice-2.1 is Microsoft AI's highest-fidelity, most expressive text-to-speech model. It produces natural, studio-grade speech across 23 languages, with detailed prosody, nuanced expressiveness, and speaker consistency over long-form content. It is...

Context
Not provided
Input
text
Output
speech
Characters: $22 per million characters
View model details →
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.6

MAI-Image-2.6 is an image generation and editing model from Microsoft AI, the precision tier of the MAI-Image-2.6 family alongside the faster MAI-Image-2.6 Flash. It is suited for design-ready visuals and...

Context
4K
Input
text · image
Output
image
Input: $5 per 1M tokensImage output: $38 per 1M tokens
View model details →
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.6 Flash

MAI-Image-2.6 Flash is the lower-latency, lower-cost member of the MAI-Image-2.6 family from Microsoft AI, built for latency-sensitive, high-throughput production image generation and editing at comparable quality to the precision tier....

Context
4K
Input
text · image
Output
image
Input: $1.75 per 1M tokensImage output: $19 per 1M tokens
View model details →