ElevenLabs · Audio & Music

Hermes Agent x ElevenLabs

ElevenLabs is a voice AI company whose models cover TTS, dialogue, sound effects, transcription, and audio isolation. Through RunAPI, all ElevenLabs endpoints share one key and per-call billing.

6 variants from $0.04 / minute Commercial OK

Prerequisite: npx runapi mcp install

Prompt

Prompt models

Use audio-isolation when it matches the task: Vocal extraction from mixed audio sources.

You have access to RunAPI task tools for ElevenLabs.

Available ElevenLabs models:
- audio-isolation: Vocal extraction from mixed audio sources
  endpoints: /api/v1/elevenlabs/isolate_audio
  request fields: source_audio_url
- sound-effect-v2: Text-to-sound effects for games, video, and podcasts
  endpoints: /api/v1/elevenlabs/text_to_sound
  request fields: text, loop, duration_seconds, prompt_influence, output_format
- speech-to-text: Transcription across 29+ languages with speaker diarization
  endpoints: /api/v1/elevenlabs/speech_to_text
  request fields: source_audio_url, language_code, diarize
- text-to-dialogue-v3: Multi-speaker dialogue generation with natural turn-taking
  endpoints: /api/v1/elevenlabs/text_to_dialogue
  request fields: dialogue, stability, language_code
- text-to-speech-multilingual-v2: 29 languages; most lifelike emotional expression; ideal for audiobooks
  endpoints: /api/v1/elevenlabs/text_to_speech
  request fields: model, text, voice, language_code
- text-to-speech-turbo-v2.5: 32 languages; 3× faster non-English; 40K char limit
  endpoints: /api/v1/elevenlabs/text_to_speech
  request fields: model, text, voice, language_code

Use the model ID and endpoint that match the user's request.
Example prompt: Convert this paragraph into natural speech with a warm British male voice, moderate pace.
/api/v1/elevenlabs/isolate_audio: Submit the task, poll for its status, then verify the completed output. POST /api/v1/elevenlabs/isolate_audio and save the returned Task id.
GET /api/v1/elevenlabs/isolate_audio/<task-id> until status is completed or failed.
Verify the completed response contains the endpoint-specific output documented in the API reference.
/api/v1/elevenlabs/text_to_sound: Submit the task, poll for its status, then verify the completed output. POST /api/v1/elevenlabs/text_to_sound and save the returned Task id.
GET /api/v1/elevenlabs/text_to_sound/<task-id> until status is completed or failed.
Verify the completed response contains the endpoint-specific output documented in the API reference.
/api/v1/elevenlabs/speech_to_text: Submit the task, poll for its status, then verify the completed output. POST /api/v1/elevenlabs/speech_to_text and save the returned Task id.
GET /api/v1/elevenlabs/speech_to_text/<task-id> until status is completed or failed.
Verify the completed response contains the endpoint-specific output documented in the API reference.
/api/v1/elevenlabs/text_to_dialogue: Submit the task, poll for its status, then verify the completed output. POST /api/v1/elevenlabs/text_to_dialogue and save the returned Task id.
GET /api/v1/elevenlabs/text_to_dialogue/<task-id> until status is completed or failed.
Verify the completed response contains the endpoint-specific output documented in the API reference.
/api/v1/elevenlabs/text_to_speech: Submit the task, poll for its status, then verify the completed output. POST /api/v1/elevenlabs/text_to_speech and save the returned Task id.
GET /api/v1/elevenlabs/text_to_speech/<task-id> until status is completed or failed.
Verify the completed response contains the endpoint-specific output documented in the API reference.

Public Versions and Endpoints

Model ID Endpoints Price Catalog
audio-isolation
/api/v1/elevenlabs/isolate_audio
$0.12 / minute Model detail
sound-effect-v2
/api/v1/elevenlabs/text_to_sound
$0.15 / minute Model detail
speech-to-text
/api/v1/elevenlabs/speech_to_text
$0.04 / minute Model detail
text-to-dialogue-v3
/api/v1/elevenlabs/text_to_dialogue
$0.14 / 1K chars Model detail
text-to-speech-multilingual-v2
/api/v1/elevenlabs/text_to_speech
$0.12 / 1K chars Model detail
text-to-speech-turbo-v2.5
/api/v1/elevenlabs/text_to_speech
$0.06 / 1K chars Model detail

Verify

Poll until the task reaches a terminal status

Select <model-id> to generate verification commands.

Configuration

Guide endpoint: <endpoint>

Select <model-id> to generate a request with the endpoint's public input contract.
How it works

Get Started in 3 Steps

  1. Choose a model ID

    Select a public catalog model ID and review its endpoint and current starting price.

  2. Configure RunAPI

    Set RUNAPI_API_KEY before making the endpoint request.

  3. Verify the result

    For asynchronous endpoints, poll the same endpoint until the Task reaches a terminal status.

What to Build with Hermes Agent + ElevenLabs

  • Conversational voice agents

    Build voice agents that speak naturally, generating low-latency speech for customer service bots, assistants, or phone interfaces.

  • YouTube content narration

    Produce voiceover for YouTube videos in consistent character voices across an entire series.

  • Text-to-spoken-video pipelines

    Chain ElevenLabs speech with a talking-avatar model in a Hermes Agent workflow to go from text to a narrated video.

Why Use ElevenLabs Through RunAPI + Hermes Agent

  • 6 variants, one API key

    Use one RunAPI connection to choose among the live model variants without changing your integration.

  • Clear usage pricing

    See current catalog pricing before you send a request, with no subscription or minimum spend required.

  • Automatic task workflows

    Submit, poll, and collect asynchronous results through a consistent task workflow without writing manual polling code.

Hermes Agent + ElevenLabs Questions

Can I use ElevenLabs in Hermes Agent?

Yes. Configure RunAPI as a provider in Hermes Agent, then call any ElevenLabs endpoint shown on this page, including speech, transcription, dialogue, sound effects, and audio isolation.

Can I transcribe audio with ElevenLabs in Hermes Agent?

Yes. Call the speech-to-text endpoint with an audio URL. It can separate speakers and tag audio events, and results are returned asynchronously.

Can Hermes Agent chain ElevenLabs with video generation?

Yes. Hermes Agent can generate speech with ElevenLabs, then pass the audio URL to a talking-avatar or speech-to-video model in the same run.

Which model ID should I use?

Choose a public model ID from the version table. Each ID exposes the endpoints shown for that version.

Does this guide configure a chat model?

No. This Model Line uses the endpoint workflow shown here and is not presented as an agent chat model.

Start using ElevenLabs with Hermes Agent

Building with a team?

We're here to help with enterprise setup, integrations, and technical questions.

Contact Us