qwen logo
qwen

Qwen: Qwen3 ASR 1.7B

出典の説明(英語)

Qwen3 ASR 1.7B is an automatic speech recognition model from Qwen. It supports multilingual language identification and transcription across 30 languages and 22 Chinese dialects, with streaming and offline inference...

概要

モデル仕様

コンテキスト
情報なし
最大出力
情報なし
アーキテクチャ
audio->transcription
トークナイザー
Qwen3
知識カットオフ
情報なし
モデレーション
いいえ
OPENROUTER

料金の詳細

OpenRouterと同期した料金です。トークン料金は100万トークン単位です。

音声の長さ
$0.000008
1秒あたり
API

クイックスタート

ローカルにOPENROUTER_API_KEYを設定してください。Pythonにはrequestsが必要です。JavaScriptはNode.jsで実行し、キーはサーバー側で管理してください。

APIドキュメント

curl --fail-with-body https://openrouter.ai/api/v1/audio/transcriptions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -F 'model=qwen/qwen3-asr-1.7b' \
  -F 'file=@audio.wav'
機能とモダリティ

入力出力

入力
audio
出力
transcription
API

対応APIパラメータ

利用可能なプロバイダー

1 プロバイダー

確認日: 2026年9月16日

OpenRouterでプロバイダーを見る

可用性、遅延、スループット、ルーティングは常時変化します。最新情報は出典ページで確認してください。

OpenRouter

DeepInfra

情報なし
コンテキスト
情報なし
最大出力
情報なし
qwen

4 モデル

すべてのモデル
qwen logoqwen

Qwen: Qwen3.8 Omni Flash

Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding. It is suited for audio-video analysis and summarization,...

コンテキスト
1M
入力
text · image · audio · video
出力
text
入力: $0.15出力: $0.47100万トークンあたり
モデル詳細を見る
qwen logoqwen

Qwen: Qwen3.8 Max (0902)

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...

コンテキスト
1M
入力
text · image · video
出力
text
入力: $2出力: $6100万トークンあたり
モデル詳細を見る
qwen logoqwen

Qwen: Qwen3.8 Flash

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

コンテキスト
1M
入力
text · image · video
出力
text
入力: $0.15出力: $0.47100万トークンあたり
モデル詳細を見る
qwen logoqwen

Qwen: Qwen3.8 27B

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

コンテキスト
1M
入力
text · image · video
出力
text
入力: $0.42出力: $3100万トークンあたり
モデル詳細を見る