fish-audio logo
fish-audio

Fish Audio: Transcribe 1

출처 설명 (영어)

Transcribe 1 is a speech-to-text model from Fish Audio. It is suited for audio transcription with automatic language detection and can return timestamped word-level segments when alignment details are requested.

개요

모델 사양

컨텍스트
정보 없음
최대 출력
정보 없음
아키텍처
audio->transcription
토크나이저
Other
지식 기준일
정보 없음
검토됨
아니요
OPENROUTER

전체 가격

OpenRouter에서 동기화한 요금이며 토큰 가격은 100만 토큰 기준입니다.

오디오 길이
$0.0001
초당
API

빠른 시작

로컬에 OPENROUTER_API_KEY를 설정하세요. Python에는 requests가 필요하며 JavaScript는 Node.js에서 실행됩니다. 키는 서버에 보관하세요.

API 문서

curl --fail-with-body https://openrouter.ai/api/v1/audio/transcriptions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -F 'model=fish-audio/transcribe-1' \
  -F 'file=@audio.wav'
기능 및 모달리티

입력출력

입력
audio
출력
transcription
API

지원 API 매개변수

사용 가능한 제공업체

1 제공업체

확인일: 2026년 9월 7일

OpenRouter 실시간 제공업체

가용성, 지연 시간, 처리량과 라우팅은 계속 바뀝니다. 최신 운영 데이터는 출처 페이지에서 확인하세요.

OpenRouter

Fish Audio

정보 없음
컨텍스트
정보 없음
최대 출력
정보 없음
fish-audio

4 개 모델

모든 모델
fish-audio logofish-audio

Fish Audio: S1

S1 is a multilingual text-to-speech model from Fish Audio. It is suited for voice applications that need broad emotional expression, using parenthetical controls to guide speaking style across its supported...

컨텍스트
정보 없음
입력
text
출력
speech
UTF-8 바이트: $15 백만 UTF-8 바이트당
모델 상세 보기
fish-audio logofish-audio

Fish Audio: S2 Pro

S2 Pro is a multilingual text-to-speech model from Fish Audio. It is suited for expressive narration and multi-speaker dialogue, with natural-language controls for speaking style and emotion.

컨텍스트
정보 없음
입력
text
출력
speech
UTF-8 바이트: $15 백만 UTF-8 바이트당
모델 상세 보기
fish-audio logofish-audio

Fish Audio: S2.1 Pro Free (free)

S2.1 Pro Free is the no-cost variant of Fish Audio S2.1 Pro, intended for testing, prototyping, and low-volume applications. It provides the same synthesis capabilities without production latency or availability...

컨텍스트
정보 없음
입력
text
출력
speech
입력: 무료 백만 토큰당출력: 무료 백만 토큰당
모델 상세 보기
fish-audio logofish-audio

Fish Audio: S2.1 Pro

S2.1 Pro is a production-oriented text-to-speech model from Fish Audio. It is suited for multilingual voice applications, expressive narration, and dialogue synthesis, with open-ended natural-language controls for speaking style and...

컨텍스트
정보 없음
입력
text
출력
speech
UTF-8 바이트: $15 백만 UTF-8 바이트당
모델 상세 보기