FA
fish-audio

Fish Audio: Transcribe 1

Transcribe 1 is a speech-to-text model from Fish Audio. It is suited for audio transcription with automatic language detection and can return timestamped word-level segments when alignment details are requested.

Przegląd

Specyfikacja modelu

Context
Not provided
Maksymalne wyjście
Not provided
Architektura
audio->transcription
Tokenizer
Other
Granica wiedzy
Not provided
Moderowany
Nie
Możliwości i modalności

InputOutput

Input
audio
Output
transcription
API

Obsługiwane parametry API

Dostępni dostawcy

1 Dostawca

Fish Audio

unknown
Context
Not provided
Maksymalne wyjście
Not provided
Input
$100.00
Output
Free
Odczyt pamięci podręcznej
Not provided
Zapis pamięci podręcznej
Not provided
fish-audio

models

Wszystkie modele
FAfish-audio

Fish Audio: S1

S1 is a multilingual text-to-speech model from Fish Audio. It is suited for voice applications that need broad emotional expression, using parenthetical controls to guide speaking style across its supported...

Context
Not provided
Input
text
Output
speech
Input: $15.00Output: Freeper 1M tokens
Zobacz szczegóły modelu
FAfish-audio

Fish Audio: S2 Pro

S2 Pro is a multilingual text-to-speech model from Fish Audio. It is suited for expressive narration and multi-speaker dialogue, with natural-language controls for speaking style and emotion.

Context
Not provided
Input
text
Output
speech
Input: $15.00Output: Freeper 1M tokens
Zobacz szczegóły modelu
FAfish-audio

Fish Audio: S2.1 Pro Free (free)

S2.1 Pro Free is the no-cost variant of Fish Audio S2.1 Pro, intended for testing, prototyping, and low-volume applications. It provides the same synthesis capabilities without production latency or availability...

Context
Not provided
Input
text
Output
speech
Input: FreeOutput: Freeper 1M tokens
Zobacz szczegóły modelu
FAfish-audio

Fish Audio: S2.1 Pro

S2.1 Pro is a production-oriented text-to-speech model from Fish Audio. It is suited for multilingual voice applications, expressive narration, and dialogue synthesis, with open-ended natural-language controls for speaking style and...

Context
Not provided
Input
text
Output
speech
Input: $15.00Output: Freeper 1M tokens
Zobacz szczegóły modelu