microsoft logo
microsoft

Microsoft AI: MAI-Voice-2.1-Flash

Descripción de la fuente (inglés)

MAI-Voice-2.1-Flash is a low-latency text-to-speech model from Microsoft AI, optimized for real-time responsiveness. It produces natural, expressive speech across 23 languages, with human-like intonation, rhythm, and emotional nuance. It is...

Resumen

Especificaciones del modelo

Contexto
No indicado
Salida máxima
No indicado
Arquitectura
text->speech
Tokenizador
Other
Corte de conocimiento
No indicado
Moderado
No
OPENROUTER

Precios completos

Tarifas sincronizadas desde OpenRouter, por millón de tokens.

Caracteres
$15
por millón de caracteres
API

Configuración del modelo

Voces compatibles

cs-CZ-Grant:MAI-Voice-2.1-Flashcs-CZ-Harper:MAI-Voice-2.1-Flashda-DK-Grant:MAI-Voice-2.1-Flashda-DK-Harper:MAI-Voice-2.1-Flashde-DE-Grant:MAI-Voice-2.1-Flashde-DE-Harper:MAI-Voice-2.1-Flashde-DE-Klaus:MAI-Voice-2.1-Flashde-DE-Mia:MAI-Voice-2.1-Flashen-AU-Isla:MAI-Voice-2.1-Flashen-GB-Emily:MAI-Voice-2.1-Flashen-GB-Harry:MAI-Voice-2.1-Flashen-IN-Dhruv:MAI-Voice-2.1-Flashen-IN-Priya:MAI-Voice-2.1-Flashen-US-Ethan:MAI-Voice-2.1-Flashen-US-Grant:MAI-Voice-2.1-Flashen-US-Harper:MAI-Voice-2.1-Flashen-US-Iris:MAI-Voice-2.1-Flashen-US-Jasper:MAI-Voice-2.1-Flashen-US-Olivia:MAI-Voice-2.1-Flashen-US-Sage:MAI-Voice-2.1-Flashes-ES-Marta:MAI-Voice-2.1-Flashes-MX-Alejo:MAI-Voice-2.1-Flashes-MX-Grant:MAI-Voice-2.1-Flashes-MX-Harper:MAI-Voice-2.1-Flashes-MX-Valeria:MAI-Voice-2.1-Flashfi-FI-Grant:MAI-Voice-2.1-Flashfi-FI-Harper:MAI-Voice-2.1-Flashfr-FR-Grant:MAI-Voice-2.1-Flashfr-FR-Harper:MAI-Voice-2.1-Flashfr-FR-Marc:MAI-Voice-2.1-Flashfr-FR-Soleil:MAI-Voice-2.1-Flashhi-IN-Arjun:MAI-Voice-2.1-Flashhi-IN-Dhruv:MAI-Voice-2.1-Flashhi-IN-Grant:MAI-Voice-2.1-Flashhi-IN-Harper:MAI-Voice-2.1-Flashhi-IN-Kavya:MAI-Voice-2.1-Flashhi-IN-Priya:MAI-Voice-2.1-Flashhu-HU-Bence:MAI-Voice-2.1-Flashhu-HU-Grant:MAI-Voice-2.1-Flashhu-HU-Harper:MAI-Voice-2.1-Flashhu-HU-Levente:MAI-Voice-2.1-Flashhu-HU-Lilla:MAI-Voice-2.1-Flashhu-HU-Reka:MAI-Voice-2.1-Flashid-ID-Grant:MAI-Voice-2.1-Flashid-ID-Harper:MAI-Voice-2.1-Flashit-IT-Grant:MAI-Voice-2.1-Flashit-IT-Harper:MAI-Voice-2.1-Flashit-IT-Luca:MAI-Voice-2.1-Flashit-IT-Rosa:MAI-Voice-2.1-Flashko-KR-Grant:MAI-Voice-2.1-Flashko-KR-Haena:MAI-Voice-2.1-Flashko-KR-Harper:MAI-Voice-2.1-Flashko-KR-Junho:MAI-Voice-2.1-Flashnb-NO-Grant:MAI-Voice-2.1-Flashnb-NO-Harper:MAI-Voice-2.1-Flashnl-NL-Grant:MAI-Voice-2.1-Flashnl-NL-Harper:MAI-Voice-2.1-Flashnl-NL-Sander:MAI-Voice-2.1-Flashpl-PL-Grant:MAI-Voice-2.1-Flashpl-PL-Harper:MAI-Voice-2.1-Flashpt-BR-Caio:MAI-Voice-2.1-Flashpt-BR-Grant:MAI-Voice-2.1-Flashpt-BR-Harper:MAI-Voice-2.1-Flashpt-BR-Luana:MAI-Voice-2.1-Flashpt-BR-Pedro:MAI-Voice-2.1-Flashpt-BR-Rafael:MAI-Voice-2.1-Flashpt-PT-Grant:MAI-Voice-2.1-Flashpt-PT-Harper:MAI-Voice-2.1-Flashpt-PT-Rui:MAI-Voice-2.1-Flashro-RO-Andrei:MAI-Voice-2.1-Flashro-RO-Elena:MAI-Voice-2.1-Flashro-RO-Grant:MAI-Voice-2.1-Flashro-RO-Harper:MAI-Voice-2.1-Flashro-RO-Ioana:MAI-Voice-2.1-Flashro-RO-Radu:MAI-Voice-2.1-Flashru-RU-Grant:MAI-Voice-2.1-Flashru-RU-Harper:MAI-Voice-2.1-Flashru-RU-Lev:MAI-Voice-2.1-Flashru-RU-Masha:MAI-Voice-2.1-Flashsv-SE-Grant:MAI-Voice-2.1-Flashsv-SE-Harper:MAI-Voice-2.1-Flashth-TH-Grant:MAI-Voice-2.1-Flashth-TH-Harper:MAI-Voice-2.1-Flashth-TH-Krit:MAI-Voice-2.1-Flashth-TH-Nattapong:MAI-Voice-2.1-Flashtr-TR-Aydin:MAI-Voice-2.1-Flashtr-TR-Elif:MAI-Voice-2.1-Flashtr-TR-Grant:MAI-Voice-2.1-Flashtr-TR-Harper:MAI-Voice-2.1-Flashvi-VN-Grant:MAI-Voice-2.1-Flashvi-VN-Harper:MAI-Voice-2.1-Flashzh-CN-Bo:MAI-Voice-2.1-Flashzh-CN-Grant:MAI-Voice-2.1-Flashzh-CN-Harper:MAI-Voice-2.1-Flashzh-CN-Lan:MAI-Voice-2.1-Flashzh-CN-Mei:MAI-Voice-2.1-Flashzh-CN-Wei:MAI-Voice-2.1-Flash
API

Inicio rápido

Configura OPENROUTER_API_KEY localmente. Python requiere requests; JavaScript se ejecuta en Node.js. Mantén la clave en el servidor.

Elige una voz compatible. Sustituye YOUR_VOICE_ID si no se indica ninguna voz.

Documentación de la API

curl --fail-with-body https://openrouter.ai/api/v1/audio/speech \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"microsoft/mai-voice-2.1-flash","input":"Hello!","voice":"cs-CZ-Grant:MAI-Voice-2.1-Flash","response_format":"mp3"}' \
  --output speech.mp3
Capacidades y modalidades

Entrada → Salida

Entrada
text
Salida
speech
API

Parámetros API compatibles

Proveedores disponibles

1 Proveedor

Verificado: 2 de octubre de 2026

Proveedores en vivo en OpenRouter

Disponibilidad, latencia, rendimiento y enrutamiento cambian continuamente. Consulta la fuente para datos actuales.

OpenRouter

Azure

No indicado
Contexto
No indicado
Salida máxima
No indicado
microsoft

4 modelos

Todos los modelos
microsoft logomicrosoft

Microsoft: Microsoft-Decision-1

Microsoft-Decision-1 is a small model built for fast decision-making. Instead of generating text, it reads the provided content and returns a calibrated probability for each fixed answer option, so the...

Contexto
33K
Entrada
text
Salida
decisions
Entrada: $0.042 por millón de tokensSalida: Gratis por millón de tokens
Ver detalles del modelo →
microsoft logomicrosoft

Microsoft AI: MAI-Voice-2.1

MAI-Voice-2.1 is Microsoft AI's highest-fidelity, most expressive text-to-speech model. It produces natural, studio-grade speech across 23 languages, with detailed prosody, nuanced expressiveness, and speaker consistency over long-form content. It is...

Contexto
No indicado
Entrada
text
Salida
speech
Caracteres: $22 por millón de caracteres
Ver detalles del modelo →
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.6

MAI-Image-2.6 is an image generation and editing model from Microsoft AI, the precision tier of the MAI-Image-2.6 family alongside the faster MAI-Image-2.6 Flash. It is suited for design-ready visuals and...

Contexto
4K
Entrada
text · image
Salida
image
Entrada: $5 por millón de tokensSalida de imagen: $38 por millón de tokens
Ver detalles del modelo →
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.6 Flash

MAI-Image-2.6 Flash is the lower-latency, lower-cost member of the MAI-Image-2.6 family from Microsoft AI, built for latency-sensitive, high-throughput production image generation and editing at comparable quality to the precision tier....

Contexto
4K
Entrada
text · image
Salida
image
Entrada: $1.75 por millón de tokensSalida de imagen: $19 por millón de tokens
Ver detalles del modelo →