qwen logo
qwen

Qwen: Qwen3 ASR 1.7B

출처 설명 (영어)

Qwen3 ASR 1.7B is an automatic speech recognition model from Qwen. It supports multilingual language identification and transcription across 30 languages and 22 Chinese dialects, with streaming and offline inference...

개요

모델 사양

컨텍스트
정보 없음
최대 출력
정보 없음
아키텍처
audio->transcription
토크나이저
Qwen3
지식 기준일
정보 없음
검토됨
아니요
OPENROUTER

전체 가격

OpenRouter에서 동기화한 요금이며 토큰 가격은 100만 토큰 기준입니다.

오디오 길이
$0.000008
초당
API

빠른 시작

로컬에 OPENROUTER_API_KEY를 설정하세요. Python에는 requests가 필요하며 JavaScript는 Node.js에서 실행됩니다. 키는 서버에 보관하세요.

API 문서

curl --fail-with-body https://openrouter.ai/api/v1/audio/transcriptions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -F 'model=qwen/qwen3-asr-1.7b' \
  -F 'file=@audio.wav'
기능 및 모달리티

입력출력

입력
audio
출력
transcription
API

지원 API 매개변수

사용 가능한 제공업체

1 제공업체

확인일: 2026년 9월 16일

OpenRouter 실시간 제공업체

가용성, 지연 시간, 처리량과 라우팅은 계속 바뀝니다. 최신 운영 데이터는 출처 페이지에서 확인하세요.

OpenRouter

DeepInfra

정보 없음
컨텍스트
정보 없음
최대 출력
정보 없음
qwen

4 개 모델

모든 모델
qwen logoqwen

Qwen: Qwen3.8 Omni Flash

Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding. It is suited for audio-video analysis and summarization,...

컨텍스트
1M
입력
text · image · audio · video
출력
text
입력: $0.15출력: $0.47백만 토큰당
모델 상세 보기
qwen logoqwen

Qwen: Qwen3.8 Max (0902)

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...

컨텍스트
1M
입력
text · image · video
출력
text
입력: $2출력: $6백만 토큰당
모델 상세 보기
qwen logoqwen

Qwen: Qwen3.8 Flash

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

컨텍스트
1M
입력
text · image · video
출력
text
입력: $0.15출력: $0.47백만 토큰당
모델 상세 보기
qwen logoqwen

Qwen: Qwen3.8 27B

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

컨텍스트
1M
입력
text · image · video
출력
text
입력: $0.42출력: $3백만 토큰당
모델 상세 보기