microsoft logo
microsoft

Microsoft AI: MAI-Voice-2-Flash

来源介绍(英文)

MAI-Voice-2-Flash is a low-latency text-to-speech model from Microsoft AI for voice agents, assistants, call centers, accessibility, narration, and other interactive applications. It generates expressive 24 kHz mono speech across 15...

模型概览

模型规格

上下文
暂未提供
最大输出
暂未提供
架构
text->speech
分词器
Other
知识截止时间
暂未提供
内容审核
OPENROUTER

完整价格

同步自 OpenRouter。Token 价格均按每百万 Token 展示。

字符
$15
每百万字符
API

模型配置

支持的声音

en-US-Harper:MAI-Voice-2es-MX-Valeria:MAI-Voice-2fr-FR-Soleil:MAI-Voice-2de-DE-Klaus:MAI-Voice-2
API

快速调用

在本地设置 OPENROUTER_API_KEY。Python 需安装 requests;JavaScript 在 Node.js 中运行。密钥应保存在服务端。

请选择此模型支持的声音;若未列出声音,请替换 YOUR_VOICE_ID。

API 文档

curl --fail-with-body https://openrouter.ai/api/v1/audio/speech \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"microsoft/mai-voice-2-flash","input":"Hello!","voice":"en-US-Harper:MAI-Voice-2","response_format":"mp3"}' \
  --output speech.mp3
能力与模态

输入输出

输入
text
输出
speech
API

支持的 API 参数

可用提供方

1 提供方

核验日期: 2026年9月16日

在 OpenRouter 查看实时 Provider

Provider 可用性、延迟、吞吐量和路由会持续变化,请前往来源页面查看实时运行数据。

OpenRouter

Azure

暂未提供
上下文
暂未提供
最大输出
暂未提供
microsoft

4 个模型

全部模型
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.6

MAI-Image-2.6 is an image generation and editing model from Microsoft AI, the precision tier of the MAI-Image-2.6 family alongside the faster MAI-Image-2.6 Flash. It is suited for design-ready visuals and...

上下文
4K
输入
text · image
输出
image
输入: $5 每百万 Token图片输出: $38 每百万 Token
查看模型详情
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.6 Flash

MAI-Image-2.6 Flash is the lower-latency, lower-cost member of the MAI-Image-2.6 family from Microsoft AI, built for latency-sensitive, high-throughput production image generation and editing at comparable quality to the precision tier....

上下文
4K
输入
text · image
输出
image
输入: $1.75 每百万 Token图片输出: $19 每百万 Token
查看模型详情
microsoft logomicrosoft

Microsoft AI: MAI-Transcribe 2

MAI-Transcribe 2 is a multilingual speech-to-text model from Microsoft AI, ranked #1 on the FLEURS multilingual benchmark. It supports 60 languages with automatic language identification, code switching for mixed-language speech,...

上下文
暂未提供
输入
audio
输出
transcription
音频时长: $0.1 每小时
查看模型详情
microsoft logomicrosoft

Microsoft AI: MAI-Image-2.5 Pro

Microsoft AI's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.

上下文
4K
输入
text · image
输出
image
输入: $5 每百万 Token图片输出: $108 每百万 Token
查看模型详情