qwen logo
qwen

Qwen: Qwen3 ASR 1.7B

来源介绍(英文)

Qwen3 ASR 1.7B is an automatic speech recognition model from Qwen. It supports multilingual language identification and transcription across 30 languages and 22 Chinese dialects, with streaming and offline inference...

模型概览

模型规格

上下文
暂未提供
最大输出
暂未提供
架构
audio->transcription
分词器
Qwen3
知识截止时间
暂未提供
内容审核
OPENROUTER

完整价格

同步自 OpenRouter。Token 价格均按每百万 Token 展示。

音频时长
$0.000008
每秒
API

快速调用

在本地设置 OPENROUTER_API_KEY。Python 需安装 requests;JavaScript 在 Node.js 中运行。密钥应保存在服务端。

API 文档

curl --fail-with-body https://openrouter.ai/api/v1/audio/transcriptions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -F 'model=qwen/qwen3-asr-1.7b' \
  -F 'file=@audio.wav'
能力与模态

输入输出

输入
audio
输出
transcription
API

支持的 API 参数

可用提供方

1 提供方

核验日期: 2026年9月16日

在 OpenRouter 查看实时 Provider

Provider 可用性、延迟、吞吐量和路由会持续变化,请前往来源页面查看实时运行数据。

OpenRouter

DeepInfra

暂未提供
上下文
暂未提供
最大输出
暂未提供
qwen

4 个模型

全部模型
qwen logoqwen

Qwen: Qwen3.8 Omni Flash

Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding. It is suited for audio-video analysis and summarization,...

上下文
1M
输入
text · image · audio · video
输出
text
输入: $0.15输出: $0.47每百万 Token
查看模型详情
qwen logoqwen

Qwen: Qwen3.8 Max (0902)

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...

上下文
1M
输入
text · image · video
输出
text
输入: $2输出: $6每百万 Token
查看模型详情
qwen logoqwen

Qwen: Qwen3.8 Flash

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

上下文
1M
输入
text · image · video
输出
text
输入: $0.15输出: $0.47每百万 Token
查看模型详情
qwen logoqwen

Qwen: Qwen3.8 27B

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

上下文
1M
输入
text · image · video
输出
text
输入: $0.42输出: $3每百万 Token
查看模型详情