ElevenLabs · 音频与音乐

Hermes Agent x ElevenLabs

ElevenLabs 是一家语音 AI 公司,模型覆盖 TTS、对话、音效、转录和音频隔离。通过 RunAPI,所有 ElevenLabs 端点共享一把 key 和按次计费。

6 variants 起价 $0.04 / 分钟 可商用

Prerequisite: npx runapi mcp install

Prompt

Prompt 模型

任务匹配时使用 audio-isolation:从混合音频中提取人声。

You have access to RunAPI task tools for ElevenLabs.

Available ElevenLabs models:
- audio-isolation: 从混合音频中提取人声
  endpoints: /api/v1/elevenlabs/isolate_audio
  request fields: source_audio_url
- sound-effect-v2: 文字生成音效,适用于游戏、视频和播客
  endpoints: /api/v1/elevenlabs/text_to_sound
  request fields: text, loop, duration_seconds, prompt_influence, output_format
- speech-to-text: 29+ 种语言转录,支持说话人分离
  endpoints: /api/v1/elevenlabs/speech_to_text
  request fields: source_audio_url, language_code, diarize
- text-to-dialogue-v3: 多说话人对话生成,自然轮流交替
  endpoints: /api/v1/elevenlabs/text_to_dialogue
  request fields: dialogue, stability, language_code
- text-to-speech-multilingual-v2: 29 种语言;最逼真的情感表达;适合有声书和配音
  endpoints: /api/v1/elevenlabs/text_to_speech
  request fields: model, text, voice, language_code
- text-to-speech-turbo-v2.5: 32 种语言;非英语快 3 倍;4 万字符上限
  endpoints: /api/v1/elevenlabs/text_to_speech
  request fields: model, text, voice, language_code

Use the model ID and endpoint that match the user's request.
Example prompt: 将这段文字转换为自然语音,使用温暖的英式男声,语速适中。
/api/v1/elevenlabs/isolate_audio: Submit the task, poll for its status, then verify the completed output. POST /api/v1/elevenlabs/isolate_audio,并保存返回的 Task ID。
GET /api/v1/elevenlabs/isolate_audio/<task-id>,直到状态为 completed 或 failed。
确认 completed 响应包含 API 参考中记录的端点专属输出。
/api/v1/elevenlabs/text_to_sound: Submit the task, poll for its status, then verify the completed output. POST /api/v1/elevenlabs/text_to_sound,并保存返回的 Task ID。
GET /api/v1/elevenlabs/text_to_sound/<task-id>,直到状态为 completed 或 failed。
确认 completed 响应包含 API 参考中记录的端点专属输出。
/api/v1/elevenlabs/speech_to_text: Submit the task, poll for its status, then verify the completed output. POST /api/v1/elevenlabs/speech_to_text,并保存返回的 Task ID。
GET /api/v1/elevenlabs/speech_to_text/<task-id>,直到状态为 completed 或 failed。
确认 completed 响应包含 API 参考中记录的端点专属输出。
/api/v1/elevenlabs/text_to_dialogue: Submit the task, poll for its status, then verify the completed output. POST /api/v1/elevenlabs/text_to_dialogue,并保存返回的 Task ID。
GET /api/v1/elevenlabs/text_to_dialogue/<task-id>,直到状态为 completed 或 failed。
确认 completed 响应包含 API 参考中记录的端点专属输出。
/api/v1/elevenlabs/text_to_speech: Submit the task, poll for its status, then verify the completed output. POST /api/v1/elevenlabs/text_to_speech,并保存返回的 Task ID。
GET /api/v1/elevenlabs/text_to_speech/<task-id>,直到状态为 completed 或 failed。
确认 completed 响应包含 API 参考中记录的端点专属输出。

公开版本与端点

模型 ID 端点 起步价格 模型目录
audio-isolation
/api/v1/elevenlabs/isolate_audio
$0.12 / 分钟 模型详情
sound-effect-v2
/api/v1/elevenlabs/text_to_sound
$0.15 / 分钟 模型详情
speech-to-text
/api/v1/elevenlabs/speech_to_text
$0.04 / 分钟 模型详情
text-to-dialogue-v3
/api/v1/elevenlabs/text_to_dialogue
$0.14 / 1K 个字符 模型详情
text-to-speech-multilingual-v2
/api/v1/elevenlabs/text_to_speech
$0.12 / 1K 个字符 模型详情
text-to-speech-turbo-v2.5
/api/v1/elevenlabs/text_to_speech
$0.06 / 1K 个字符 模型详情

验证

轮询,直到 Task 进入终态

选择 <model-id> 后生成验证命令。

配置

指南端点: <endpoint>

选择 <model-id> 后,按端点的公开输入契约生成请求。
使用流程

三步开始使用

  1. 选择模型 ID

    选择公开模型 ID,并查看其端点与当前起步价格。

  2. 配置 RunAPI

    调用端点前设置 RUNAPI_API_KEY。

  3. 验证结果

    对于异步端点,轮询同一端点,直到 Task 进入终态。

使用 Hermes Agent + ElevenLabs 可以构建什么

  • 对话式语音 Agent

    搭建说话自然的语音 Agent,用低延迟语音支撑客服机器人、助手或电话交互。

  • YouTube 内容旁白

    为 YouTube 视频制作旁白,整个系列保持一致的角色音色。

  • 文字转口播视频流水线

    在 Hermes Agent 工作流中把 ElevenLabs 语音与数字人模型串联,从文字直接得到带旁白的视频。

为什么通过 RunAPI + Hermes Agent 使用 ElevenLabs

  • 6 个变体,一个 API 密钥

    通过一个 RunAPI 连接选择可用的模型变体,无需更改现有集成。

  • 清晰的用量价格

    发送请求前查看目录中的当前价格,无需订阅或最低消费。

  • 自动化任务流程

    通过一致的任务流程提交任务、查询状态并收集异步结果,无需编写手动轮询代码。

Hermes Agent + ElevenLabs 常见问题

可以在 Hermes Agent 中使用 ElevenLabs 吗?

可以。在 Hermes Agent 中把 RunAPI 配置为提供方,然后调用本页列出的任意 ElevenLabs 端点,包括语音合成、转写、对话、音效和人声分离。

可以在 Hermes Agent 中用 ElevenLabs 转写音频吗?

可以。用音频 URL 调用语音转文字端点。它支持区分说话人并标注音频事件,结果以异步方式返回。

Hermes Agent 能把 ElevenLabs 与视频生成串联吗?

可以。Hermes Agent 可以先用 ElevenLabs 生成语音,再在同一次运行中把音频 URL 交给数字人或语音生视频模型。

应该使用哪个模型 ID?

从版本表选择公开模型 ID。每个 ID 可使用的端点均显示在对应行。

本指南会配置聊天模型吗?

不会。该 Model Line 使用页面所示的端点流程,不会被描述为 Agent 聊天模型。

开始在 Hermes Agent 中使用 ElevenLabs

需要团队接入支持?

我们可以协助团队接入、完成系统集成并解答技术问题。

联系我们