mistralai logo
mistralai

Mistral: Voxtral Mini 3B 2507

来源介绍(英文)

Voxtral Mini 3B 2507 is a speech and audio understanding model from Mistral AI. It is suited for transcription, translation, and compact audio processing workloads.

模型概览

模型规格

上下文
暂未提供
最大输出
暂未提供
架构
audio->transcription
分词器
Mistral
知识截止时间
暂未提供
内容审核
OPENROUTER

完整价格

同步自 OpenRouter。Token 价格均按每百万 Token 展示。

音频时长
$0.000017
每秒
API

快速调用

在本地设置 OPENROUTER_API_KEY。Python 需安装 requests;JavaScript 在 Node.js 中运行。密钥应保存在服务端。

API 文档

curl --fail-with-body https://openrouter.ai/api/v1/audio/transcriptions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -F 'model=mistralai/voxtral-mini-3b-2507' \
  -F 'file=@audio.wav'
能力与模态

输入输出

输入
audio
输出
transcription
API

支持的 API 参数

可用提供方

1 提供方

核验日期: 2026年9月16日

在 OpenRouter 查看实时 Provider

Provider 可用性、延迟、吞吐量和路由会持续变化,请前往来源页面查看实时运行数据。

OpenRouter

DeepInfra

bf16
上下文
暂未提供
最大输出
暂未提供
mistralai

4 个模型

全部模型
mistralai logomistralai

Mistral: Voxtral Small 24B 2507 STT

Voxtral Small 24B 2507 STT is a speech transcription model from Mistral AI. It is suited for transcription, translation, and audio understanding workloads that benefit from its larger model capacity.

上下文
暂未提供
输入
audio
输出
transcription
音频时长: $0.00005 每秒
查看模型详情
mistralai logomistralai

Mistral: Voxtral Mini Transcribe

Voxtral Mini Transcribe is Mistral's speech-to-text model, derived from the Voxtral Mini family. It accepts audio input and returns transcribed text via the standard transcription API. Suited for transcribing meetings,...

上下文
暂未提供
输入
audio
输出
transcription
音频时长: $0.00005 每秒
查看模型详情
mistralai logomistralai

Mistral: Mistral Medium 3.5

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

上下文
262K
输入
text · image · file
输出
text
输入: $1.5输出: $7.5每百万 Token
查看模型详情
mistralai logomistralai

Mistral: Mistral Medium 3.5 (batch)

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

上下文
262K
输入
text · image · file
输出
text
输入: $0.75输出: $3.75每百万 Token
查看模型详情