microsoft logo
microsoft

Microsoft: MAI-Transcribe 1.5

MAI-Transcribe 1.5 is a multilingual speech-to-text model from Microsoft AI. Priced at $0.36 per hour. 0 token context window.

模型概览

模型规格

上下文
暂未提供
最大输出
暂未提供
架构
audio->transcription
分词器
Other
知识截止时间
暂未提供
内容审核
OPENROUTER

完整价格

同步自 OpenRouter。Token 价格均按每百万 Token 展示。

Audio Hours
$0.36
/hour
API

快速调用

通过 OpenRouter 的 OpenAI 兼容 API 调用这个确切的模型 ID。

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"microsoft/mai-transcribe-1.5","messages":[{"role":"user","content":"Hello!"}]}'
能力与模态

输入输出

输入
audio
输出
transcription
API

支持的 API 参数

max_completion_tokensmax_tokenstemperaturetop_p
可用提供方

1 提供方

在 OpenRouter 查看实时 Provider

Provider 可用性、延迟、吞吐量和路由会持续变化,请前往来源页面查看实时运行数据。

OpenRouter

Azure

unknown
上下文
暂未提供
最大输出
暂未提供
输入
$360000.00
输出
免费
缓存读取
暂未提供
缓存写入
暂未提供
microsoft

个模型

全部模型
microsoft logomicrosoft

Microsoft: MAI-Image-2.5 Pro

Microsoft's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.

上下文
4K
输入
text · image
输出
image
输入: $5.00输出: 免费每百万 Token
查看模型详情
microsoft logomicrosoft

Microsoft: MAI-Voice-2-Flash

MAI-Voice-2-Flash is a low-latency text-to-speech model from Microsoft for voice agents, assistants, call centers, accessibility, narration, and other interactive applications. It generates expressive 24 kHz mono speech across 15 languages...

上下文
暂未提供
输入
text
输出
speech
输入: $15.00输出: 免费每百万 Token
查看模型详情
microsoft logomicrosoft

Microsoft: MAI-Voice-2

MAI-Voice-2 is an expressive text-to-speech model from Microsoft. It is suited for conversational assistants, media narration, accessibility, education, and other long-form voice applications. It supports 15 languages across 18 locales,...

上下文
暂未提供
输入
text
输出
speech
输入: $22.00输出: 免费每百万 Token
查看模型详情
microsoft logomicrosoft

Microsoft: MAI-Image-2.5

Microsoft's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.

上下文
4K
输入
text · image
输出
image
输入: $5.00输出: 免费每百万 Token
查看模型详情