microsoft logo
microsoft

MicrosoftAI: MAI-Transcribe 2

Source description (English)

MAI-Transcribe 2 is a multilingual speech-to-text model from Microsoft AI, ranked #1 on the FLEURS multilingual benchmark. It supports 60 languages with automatic language identification, code switching for mixed-language speech,...

Overview

Model specifications

Context
Not provided
Maximum output
Not provided
Architecture
audio->transcription
Tokenizer
Other
Knowledge cutoff
Not provided
Moderated
No
OPENROUTER

Complete pricing

Synchronized OpenRouter rates. Token prices are shown per one million tokens.

Audio seconds
$0.1
per hour
API

Quick start

Set OPENROUTER_API_KEY locally. Python requires requests; JavaScript runs in Node.js. Keep the key on the server.

API documentation

curl --fail-with-body https://openrouter.ai/api/v1/audio/transcriptions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -F 'model=microsoft/mai-transcribe-2' \
  -F 'file=@audio.wav'
Capabilities

InputOutput

Input
audio
Output
transcription
API

Supported API parameters

max_completion_tokensmax_tokenstemperaturetop_p
Available providers

1 Provider

Checked: September 7, 2026

Live providers on OpenRouter

Provider availability, latency, throughput and routing can change continuously. Open the source page for current operational data.

OpenRouter

Azure

Not provided
Context
Not provided
Maximum output
Not provided
microsoft

4 models

All models
microsoft logomicrosoft

MicrosoftAI: MAI-Image-2.6

Microsoft's MAI-Image-2.6 is an image generation and editing model available via Azure AI Foundry. It creates images from text prompts and supports image-guided editing across multiple aspect ratios.

Context
4K
Input
text · image
Output
image
Input: $5 per 1M tokensImage output: $38 per 1M tokens
View model details
microsoft logomicrosoft

MicrosoftAI: MAI-Image-2.6 Flash

Microsoft's MAI-Image-2.6 Flash is the lower-latency variant of MAI-Image-2.6, available via Azure AI Foundry. It supports image generation and image-guided editing across multiple aspect ratios.

Context
4K
Input
text · image
Output
image
Input: $1.75 per 1M tokensImage output: $19 per 1M tokens
View model details
microsoft logomicrosoft

MicrosoftAI: MAI-Image-2.5 Pro

Microsoft's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.

Context
4K
Input
text · image
Output
image
Input: $5 per 1M tokensImage output: $108 per 1M tokens
View model details
microsoft logomicrosoft

MicrosoftAI: MAI-Voice-2-Flash

MAI-Voice-2-Flash is a low-latency text-to-speech model from Microsoft for voice agents, assistants, call centers, accessibility, narration, and other interactive applications. It generates expressive 24 kHz mono speech across 15 languages...

Context
Not provided
Input
text
Output
speech
Characters: $15 per million characters
View model details