fish-audio logo
fish-audio

Fish Audio: S2 Pro

Source description (English)

S2 Pro is a multilingual text-to-speech model from Fish Audio. It is suited for expressive narration and multi-speaker dialogue, with natural-language controls for speaking style and emotion.

Overview

Model specifications

Context
Not provided
Maximum output
Not provided
Architecture
text->speech
Tokenizer
Other
Knowledge cutoff
Not provided
Moderated
No
OPENROUTER

Complete pricing

Synchronized OpenRouter rates. Token prices are shown per one million tokens.

UTF-8 bytes
$15
per million UTF-8 bytes
API

Quick start

Set OPENROUTER_API_KEY locally. Python requires requests; JavaScript runs in Node.js. Keep the key on the server.

Choose a supported voice. Replace YOUR_VOICE_ID if no voice is listed.

API documentation

curl --fail-with-body https://openrouter.ai/api/v1/audio/speech \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"fish-audio/s2-pro","input":"Hello!","voice":"YOUR_VOICE_ID","response_format":"mp3"}' \
  --output speech.mp3
Capabilities

InputOutput

Input
text
Output
speech
API

Supported API parameters

Available providers

1 Provider

Checked: September 7, 2026

Live providers on OpenRouter

Provider availability, latency, throughput and routing can change continuously. Open the source page for current operational data.

OpenRouter

Fish Audio

Not provided
Context
Not provided
Maximum output
Not provided
fish-audio

4 models

All models
fish-audio logofish-audio

Fish Audio: Transcribe 1

Transcribe 1 is a speech-to-text model from Fish Audio. It is suited for audio transcription with automatic language detection and can return timestamped word-level segments when alignment details are requested.

Context
Not provided
Input
audio
Output
transcription
Audio seconds: $0.0001 per second
View model details
fish-audio logofish-audio

Fish Audio: S1

S1 is a multilingual text-to-speech model from Fish Audio. It is suited for voice applications that need broad emotional expression, using parenthetical controls to guide speaking style across its supported...

Context
Not provided
Input
text
Output
speech
UTF-8 bytes: $15 per million UTF-8 bytes
View model details
fish-audio logofish-audio

Fish Audio: S2.1 Pro Free (free)

S2.1 Pro Free is the no-cost variant of Fish Audio S2.1 Pro, intended for testing, prototyping, and low-volume applications. It provides the same synthesis capabilities without production latency or availability...

Context
Not provided
Input
text
Output
speech
Input: Free per 1M tokensOutput: Free per 1M tokens
View model details
fish-audio logofish-audio

Fish Audio: S2.1 Pro

S2.1 Pro is a production-oriented text-to-speech model from Fish Audio. It is suited for multilingual voice applications, expressive narration, and dialogue synthesis, with open-ended natural-language controls for speaking style and...

Context
Not provided
Input
text
Output
speech
UTF-8 bytes: $15 per million UTF-8 bytes
View model details