microsoft logo
microsoft

Microsoft AI: MAI-Voice-2

Description de la source (anglais)

MAI-Voice-2 is an expressive text-to-speech model from Microsoft AI. It is suited for conversational assistants, media narration, accessibility, education, and other long-form voice applications. It supports 15 languages across 18...

Présentation

Caractéristiques du modèle

Contexte
Non indiqué
Sortie maximale
Non indiqué
Architecture
text->speech
Tokenizer
Other
Limite des connaissances
Non indiqué
Modéré
Non
OPENROUTER

Tarification complète

Tarifs synchronisés depuis OpenRouter, par million de tokens.

Caractères
$22
par million de caractères
API

Configuration du modèle

Voix prises en charge

en-US-Harper:MAI-Voice-2es-MX-Valeria:MAI-Voice-2fr-FR-Soleil:MAI-Voice-2de-DE-Klaus:MAI-Voice-2
API

Démarrage rapide

Définissez OPENROUTER_API_KEY localement. Python nécessite requests ; JavaScript s’exécute dans Node.js. Gardez la clé côté serveur.

Choisissez une voix compatible. Remplacez YOUR_VOICE_ID si aucune voix n’est indiquée.

Documentation de l’API

curl --fail-with-body https://openrouter.ai/api/v1/audio/speech \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"microsoft/mai-voice-2","input":"Hello!","voice":"en-US-Harper:MAI-Voice-2","response_format":"mp3"}' \
  --output speech.mp3
Capacités et modalités

EntréeSortie

Entrée
text
Sortie
speech
API

Paramètres API pris en charge

Fournisseurs disponibles

1 Fournisseur

Vérifié: 16 septembre 2026

Fournisseurs en direct sur OpenRouter

Disponibilité, latence, débit et routage évoluent en continu. Consultez la source pour les données en temps réel.

OpenRouter

Azure

Non indiqué
Contexte
Non indiqué
Sortie maximale
Non indiqué
microsoft

4 modèles

Tous les modèles
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.6

MAI-Image-2.6 is an image generation and editing model from Microsoft AI, the precision tier of the MAI-Image-2.6 family alongside the faster MAI-Image-2.6 Flash. It is suited for design-ready visuals and...

Contexte
4K
Entrée
text · image
Sortie
image
Entrée: $5 par million de tokensImage en sortie: $38 par million de tokens
Voir la fiche du modèle
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.6 Flash

MAI-Image-2.6 Flash is the lower-latency, lower-cost member of the MAI-Image-2.6 family from Microsoft AI, built for latency-sensitive, high-throughput production image generation and editing at comparable quality to the precision tier....

Contexte
4K
Entrée
text · image
Sortie
image
Entrée: $1.75 par million de tokensImage en sortie: $19 par million de tokens
Voir la fiche du modèle
microsoft logomicrosoft

Microsoft AI: MAI-Transcribe 2

MAI-Transcribe 2 is a multilingual speech-to-text model from Microsoft AI, ranked #1 on the FLEURS multilingual benchmark. It supports 60 languages with automatic language identification, code switching for mixed-language speech,...

Contexte
Non indiqué
Entrée
audio
Sortie
transcription
Durée audio: $0.1 par heure
Voir la fiche du modèle
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.5 Pro

Microsoft AI's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.

Contexte
4K
Entrée
text · image
Sortie
image
Entrée: $5 par million de tokensImage en sortie: $108 par million de tokens
Voir la fiche du modèle