microsoft logo
microsoft

Microsoft AI: MAI-Voice-2

出典の説明(英語)

MAI-Voice-2 is an expressive text-to-speech model from Microsoft AI. It is suited for conversational assistants, media narration, accessibility, education, and other long-form voice applications. It supports 15 languages across 18...

概要

モデル仕様

コンテキスト
情報なし
最大出力
情報なし
アーキテクチャ
text->speech
トークナイザー
Other
知識カットオフ
情報なし
モデレーション
いいえ
OPENROUTER

料金の詳細

OpenRouterと同期した料金です。トークン料金は100万トークン単位です。

文字
$22
100万文字あたり
API

モデル設定

対応音声

en-US-Harper:MAI-Voice-2es-MX-Valeria:MAI-Voice-2fr-FR-Soleil:MAI-Voice-2de-DE-Klaus:MAI-Voice-2
API

クイックスタート

ローカルにOPENROUTER_API_KEYを設定してください。Pythonにはrequestsが必要です。JavaScriptはNode.jsで実行し、キーはサーバー側で管理してください。

対応する音声を選択してください。一覧がない場合はYOUR_VOICE_IDを置き換えてください。

APIドキュメント

curl --fail-with-body https://openrouter.ai/api/v1/audio/speech \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"microsoft/mai-voice-2","input":"Hello!","voice":"en-US-Harper:MAI-Voice-2","response_format":"mp3"}' \
  --output speech.mp3
機能とモダリティ

入力出力

入力
text
出力
speech
API

対応APIパラメータ

利用可能なプロバイダー

1 プロバイダー

確認日: 2026年9月16日

OpenRouterでプロバイダーを見る

可用性、遅延、スループット、ルーティングは常時変化します。最新情報は出典ページで確認してください。

OpenRouter

Azure

情報なし
コンテキスト
情報なし
最大出力
情報なし
microsoft

4 モデル

すべてのモデル
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.6

MAI-Image-2.6 is an image generation and editing model from Microsoft AI, the precision tier of the MAI-Image-2.6 family alongside the faster MAI-Image-2.6 Flash. It is suited for design-ready visuals and...

コンテキスト
4K
入力
text · image
出力
image
入力: $5 100万トークンあたり画像出力: $38 100万トークンあたり
モデル詳細を見る
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.6 Flash

MAI-Image-2.6 Flash is the lower-latency, lower-cost member of the MAI-Image-2.6 family from Microsoft AI, built for latency-sensitive, high-throughput production image generation and editing at comparable quality to the precision tier....

コンテキスト
4K
入力
text · image
出力
image
入力: $1.75 100万トークンあたり画像出力: $19 100万トークンあたり
モデル詳細を見る
microsoft logomicrosoft

Microsoft AI: MAI-Transcribe 2

MAI-Transcribe 2 is a multilingual speech-to-text model from Microsoft AI, ranked #1 on the FLEURS multilingual benchmark. It supports 60 languages with automatic language identification, code switching for mixed-language speech,...

コンテキスト
情報なし
入力
audio
出力
transcription
音声の長さ: $0.1 1時間あたり
モデル詳細を見る
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.5 Pro

Microsoft AI's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.

コンテキスト
4K
入力
text · image
出力
image
入力: $5 100万トークンあたり画像出力: $108 100万トークンあたり
モデル詳細を見る