Public Versions and Endpoints
| Model ID | Endpoints | Price | Catalog |
|---|---|---|---|
s1
|
/api/v1/fish_audio/text_to_speech
|
$0.02 / 1K UTF-8 bytes | Model detail |
s2-pro
|
/api/v1/fish_audio/text_to_speech
|
$0.02 / 1K UTF-8 bytes | Model detail |
s2.1-pro
|
/api/v1/fish_audio/text_to_speech
|
$0.02 / 1K UTF-8 bytes | Model detail |
Verify
Poll until the task reaches a terminal status
Select <model-id> to generate verification commands.
Configuration
Guide endpoint: <endpoint>
Select <model-id> to generate a request with the endpoint's public input contract.
Additional operations
These operations do not use a model ID.
/api/v1/fish_audio/list_voices/api/v1/fish_audio/create_voice/api/v1/fish_audio/get_voice
Get Started in 3 Steps
-
Choose a model ID
Select a public catalog model ID and review its endpoint and current starting price.
-
Configure RunAPI
Set RUNAPI_API_KEY before making the endpoint request.
-
Verify the result
For asynchronous endpoints, poll the same endpoint until the Task reaches a terminal status.
What to Build with Hermes Agent + Fish Audio
-
Voiceover and narration
Turn scripts into spoken audio for videos, courses, and product demos with models that generate speech.
-
Music and soundtracks
Create background tracks and jingles from a description of genre, mood, and tempo with models that generate music.
-
Transcription and audio processing
Transcribe recordings or clean up and transform existing audio with models that accept audio input.
Why Use Fish Audio Through RunAPI + Hermes Agent
-
3 variants, one API key
Use one RunAPI connection to choose among the live model variants without changing your integration.
-
Clear usage pricing
See current catalog pricing before you send a request, with no subscription or minimum spend required.
-
Direct responses
Synchronous calls return the result in the same response, so your agent can use it immediately without task polling.
Hermes Agent + Fish Audio Questions
What is the difference between s1, s2-pro, and s2.1-pro?
s1 targets expressive conversational and narrative speech. s2.1-pro is recommended for production TTS with 83 languages and natural-language expression control; s2-pro remains available for previous-generation workflows.
What audio format does Fish Audio return?
Requests return MP3 by default and can select WAV. The RunAPI response includes accurate format, MIME type, and byte-size metadata.
Can I guide the voice with a reference sample?
You can try a returned RunAPI voice_id in later text-to-speech requests, but availability is not guaranteed. For one request only, pass reference samples with base64 audio and exact transcripts through references.
Is the generated audio hosted by RunAPI?
Yes. RunAPI validates and stores the result before returning a managed audio URL.
Do I need to poll for completion?
No. The endpoint is synchronous and returns the completed audio result in the successful response.
Which model ID should I use?
Choose a public model ID from the version table. Each ID exposes the endpoints shown for that version.
Does this guide configure a chat model?
No. This Model Line uses the endpoint workflow shown here and is not presented as an agent chat model.