microsoft logo
microsoft

Microsoft AI: MAI-Voice-2.1-Flash

Descrizione della fonte (inglese)

MAI-Voice-2.1-Flash is a low-latency text-to-speech model from Microsoft AI, optimized for real-time responsiveness. It produces natural, expressive speech across 23 languages, with human-like intonation, rhythm, and emotional nuance. It is...

Panoramica

Specifiche del modello

Contesto
Non indicato
Output massimo
Non indicato
Architettura
text->speech
Tokenizer
Other
Limite di conoscenza
Non indicato
Moderato
No
OPENROUTER

Prezzi completi

Tariffe sincronizzate da OpenRouter, per milione di token.

Caratteri
$15
per milione di caratteri
API

Configurazione del modello

Voci supportate

cs-CZ-Grant:MAI-Voice-2.1-Flashcs-CZ-Harper:MAI-Voice-2.1-Flashda-DK-Grant:MAI-Voice-2.1-Flashda-DK-Harper:MAI-Voice-2.1-Flashde-DE-Grant:MAI-Voice-2.1-Flashde-DE-Harper:MAI-Voice-2.1-Flashde-DE-Klaus:MAI-Voice-2.1-Flashde-DE-Mia:MAI-Voice-2.1-Flashen-AU-Isla:MAI-Voice-2.1-Flashen-GB-Emily:MAI-Voice-2.1-Flashen-GB-Harry:MAI-Voice-2.1-Flashen-IN-Dhruv:MAI-Voice-2.1-Flashen-IN-Priya:MAI-Voice-2.1-Flashen-US-Ethan:MAI-Voice-2.1-Flashen-US-Grant:MAI-Voice-2.1-Flashen-US-Harper:MAI-Voice-2.1-Flashen-US-Iris:MAI-Voice-2.1-Flashen-US-Jasper:MAI-Voice-2.1-Flashen-US-Olivia:MAI-Voice-2.1-Flashen-US-Sage:MAI-Voice-2.1-Flashes-ES-Marta:MAI-Voice-2.1-Flashes-MX-Alejo:MAI-Voice-2.1-Flashes-MX-Grant:MAI-Voice-2.1-Flashes-MX-Harper:MAI-Voice-2.1-Flashes-MX-Valeria:MAI-Voice-2.1-Flashfi-FI-Grant:MAI-Voice-2.1-Flashfi-FI-Harper:MAI-Voice-2.1-Flashfr-FR-Grant:MAI-Voice-2.1-Flashfr-FR-Harper:MAI-Voice-2.1-Flashfr-FR-Marc:MAI-Voice-2.1-Flashfr-FR-Soleil:MAI-Voice-2.1-Flashhi-IN-Arjun:MAI-Voice-2.1-Flashhi-IN-Dhruv:MAI-Voice-2.1-Flashhi-IN-Grant:MAI-Voice-2.1-Flashhi-IN-Harper:MAI-Voice-2.1-Flashhi-IN-Kavya:MAI-Voice-2.1-Flashhi-IN-Priya:MAI-Voice-2.1-Flashhu-HU-Bence:MAI-Voice-2.1-Flashhu-HU-Grant:MAI-Voice-2.1-Flashhu-HU-Harper:MAI-Voice-2.1-Flashhu-HU-Levente:MAI-Voice-2.1-Flashhu-HU-Lilla:MAI-Voice-2.1-Flashhu-HU-Reka:MAI-Voice-2.1-Flashid-ID-Grant:MAI-Voice-2.1-Flashid-ID-Harper:MAI-Voice-2.1-Flashit-IT-Grant:MAI-Voice-2.1-Flashit-IT-Harper:MAI-Voice-2.1-Flashit-IT-Luca:MAI-Voice-2.1-Flashit-IT-Rosa:MAI-Voice-2.1-Flashko-KR-Grant:MAI-Voice-2.1-Flashko-KR-Haena:MAI-Voice-2.1-Flashko-KR-Harper:MAI-Voice-2.1-Flashko-KR-Junho:MAI-Voice-2.1-Flashnb-NO-Grant:MAI-Voice-2.1-Flashnb-NO-Harper:MAI-Voice-2.1-Flashnl-NL-Grant:MAI-Voice-2.1-Flashnl-NL-Harper:MAI-Voice-2.1-Flashnl-NL-Sander:MAI-Voice-2.1-Flashpl-PL-Grant:MAI-Voice-2.1-Flashpl-PL-Harper:MAI-Voice-2.1-Flashpt-BR-Caio:MAI-Voice-2.1-Flashpt-BR-Grant:MAI-Voice-2.1-Flashpt-BR-Harper:MAI-Voice-2.1-Flashpt-BR-Luana:MAI-Voice-2.1-Flashpt-BR-Pedro:MAI-Voice-2.1-Flashpt-BR-Rafael:MAI-Voice-2.1-Flashpt-PT-Grant:MAI-Voice-2.1-Flashpt-PT-Harper:MAI-Voice-2.1-Flashpt-PT-Rui:MAI-Voice-2.1-Flashro-RO-Andrei:MAI-Voice-2.1-Flashro-RO-Elena:MAI-Voice-2.1-Flashro-RO-Grant:MAI-Voice-2.1-Flashro-RO-Harper:MAI-Voice-2.1-Flashro-RO-Ioana:MAI-Voice-2.1-Flashro-RO-Radu:MAI-Voice-2.1-Flashru-RU-Grant:MAI-Voice-2.1-Flashru-RU-Harper:MAI-Voice-2.1-Flashru-RU-Lev:MAI-Voice-2.1-Flashru-RU-Masha:MAI-Voice-2.1-Flashsv-SE-Grant:MAI-Voice-2.1-Flashsv-SE-Harper:MAI-Voice-2.1-Flashth-TH-Grant:MAI-Voice-2.1-Flashth-TH-Harper:MAI-Voice-2.1-Flashth-TH-Krit:MAI-Voice-2.1-Flashth-TH-Nattapong:MAI-Voice-2.1-Flashtr-TR-Aydin:MAI-Voice-2.1-Flashtr-TR-Elif:MAI-Voice-2.1-Flashtr-TR-Grant:MAI-Voice-2.1-Flashtr-TR-Harper:MAI-Voice-2.1-Flashvi-VN-Grant:MAI-Voice-2.1-Flashvi-VN-Harper:MAI-Voice-2.1-Flashzh-CN-Bo:MAI-Voice-2.1-Flashzh-CN-Grant:MAI-Voice-2.1-Flashzh-CN-Harper:MAI-Voice-2.1-Flashzh-CN-Lan:MAI-Voice-2.1-Flashzh-CN-Mei:MAI-Voice-2.1-Flashzh-CN-Wei:MAI-Voice-2.1-Flash
API

Avvio rapido

Imposta OPENROUTER_API_KEY localmente. Python richiede requests; JavaScript viene eseguito in Node.js. Conserva la chiave sul server.

Scegli una voce supportata. Sostituisci YOUR_VOICE_ID se non sono elencate voci.

Documentazione API

curl --fail-with-body https://openrouter.ai/api/v1/audio/speech \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"microsoft/mai-voice-2.1-flash","input":"Hello!","voice":"cs-CZ-Grant:MAI-Voice-2.1-Flash","response_format":"mp3"}' \
  --output speech.mp3
Capacità e modalità

Input → Output

Input
text
Output
speech
API

Parametri API supportati

Provider disponibili

1 Provider

Verificato: 2 ottobre 2026

Provider live su OpenRouter

Disponibilità, latenza, throughput e routing cambiano continuamente. Consulta la fonte per i dati correnti.

OpenRouter

Azure

Non indicato
Contesto
Non indicato
Output massimo
Non indicato
microsoft

4 modelli

Tutti i modelli
microsoft logomicrosoft

Microsoft: Microsoft-Decision-1

Microsoft-Decision-1 is a small model built for fast decision-making. Instead of generating text, it reads the provided content and returns a calibrated probability for each fixed answer option, so the...

Contesto
33K
Input
text
Output
decisions
Input: $0.042 per milione di tokenOutput: Gratuito per milione di token
Vedi i dettagli del modello →
microsoft logomicrosoft

Microsoft AI: MAI-Voice-2.1

MAI-Voice-2.1 is Microsoft AI's highest-fidelity, most expressive text-to-speech model. It produces natural, studio-grade speech across 23 languages, with detailed prosody, nuanced expressiveness, and speaker consistency over long-form content. It is...

Contesto
Non indicato
Input
text
Output
speech
Caratteri: $22 per milione di caratteri
Vedi i dettagli del modello →
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.6

MAI-Image-2.6 is an image generation and editing model from Microsoft AI, the precision tier of the MAI-Image-2.6 family alongside the faster MAI-Image-2.6 Flash. It is suited for design-ready visuals and...

Contesto
4K
Input
text · image
Output
image
Input: $5 per milione di tokenOutput immagine: $38 per milione di token
Vedi i dettagli del modello →
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.6 Flash

MAI-Image-2.6 Flash is the lower-latency, lower-cost member of the MAI-Image-2.6 family from Microsoft AI, built for latency-sensitive, high-throughput production image generation and editing at comparable quality to the precision tier....

Contesto
4K
Input
text · image
Output
image
Input: $1.75 per milione di tokenOutput immagine: $19 per milione di token
Vedi i dettagli del modello →