Alibaba · Video

Hermes Agent x Wan

Wan is Alibaba's full-spectrum video and image model suite, one of the most complete open-weight video families available. Through RunAPI, all Wan versions share a single key and unified billing.

18 variants from $0.05 / call Commercial OK

Prerequisite: npx runapi mcp install

Prompt

Prompt models

Use wan-2.2-a14b-image-to-video-turbo when it matches the task: Image-anchored video on the 2.2 architecture.

You have access to RunAPI task tools for Wan.

Available Wan models:
- wan-2.2-a14b-image-to-video-turbo: Image-anchored video on the 2.2 architecture
  endpoints: /api/v1/wan/image_to_video
  request fields: model, prompt, first_frame_image_url
- wan-2.2-a14b-speech-to-video-turbo: Audio input drives video motion and lip-sync
  endpoints: /api/v1/wan/speech_to_video
  request fields: model, source_image_url, source_audio_url, prompt
- wan-2.2-a14b-text-to-video-turbo: Open-source A14B MoE; baseline quality
  endpoints: /api/v1/wan/text_to_video
  request fields: model, prompt
- wan-2.2-animate-move: Specialized motion animation on 2.2 base
  endpoints: /api/v1/wan/animate
  request fields: model, source_image_url, reference_video_url
- wan-2.2-animate-replace: Subject replacement animation on 2.2 base
  endpoints: /api/v1/wan/animate
  request fields: model, source_image_url, reference_video_url
- wan-2.5-image-to-video: Image-anchored with native audio on 2.5
  endpoints: /api/v1/wan/image_to_video
  request fields: model, prompt, first_frame_image_url, duration_seconds, output_resolution
- wan-2.5-text-to-video: Native audio generation + better motion realism vs 2.2
  endpoints: /api/v1/wan/text_to_video
  request fields: model, prompt, output_resolution
- wan-2.6-edit-video: Video editing with subject + camera control
  endpoints: /api/v1/wan/edit_video
  request fields: model, source_video_urls, prompt
- wan-2.6-flash-edit-video: Speed-optimized 2.6 video editing
  endpoints: /api/v1/wan/edit_video
  request fields: model, source_video_urls, prompt, audio
- wan-2.6-flash-image-to-video: Speed-optimized 2.6; lower cost, faster inference
  endpoints: /api/v1/wan/image_to_video
  request fields: model, prompt, first_frame_image_url, audio
- wan-2.6-image-to-video: Image-anchored with 2.6 scene continuity
  endpoints: /api/v1/wan/image_to_video
  request fields: model, prompt, first_frame_image_url
- wan-2.6-text-to-video: Improved scene continuity + complex scene support
  endpoints: /api/v1/wan/text_to_video
  request fields: model, prompt
- wan-2.7-edit-video: wan-2.7-edit-video
  endpoints: /api/v1/wan/edit_video
  request fields: model, source_video_urls, prompt, source_video_url
- wan-2.7-image: Standard image generation on 2.7 architecture
  endpoints: /api/v1/wan/text_to_image
  request fields: model, prompt
- wan-2.7-image-pro: 3×3 multi-angle image grid; highest quality image output
  endpoints: /api/v1/wan/text_to_image
  request fields: model, prompt
- wan-2.7-image-to-video: Image-anchored with 2.7 multi-ref support
  endpoints: /api/v1/wan/image_to_video
  request fields: model, prompt, first_frame_image_url
- wan-2.7-r2v: R2V text-to-video with character appearance + voice references
  endpoints: /api/v1/wan/text_to_video
  request fields: model, prompt
- wan-2.7-text-to-video: First-and-last-frame control; 5 simultaneous video refs
  endpoints: /api/v1/wan/text_to_video
  request fields: model, prompt

Use the model ID and endpoint that match the user's request.
Example prompt: Generate a 5-second video of a cat jumping onto a bookshelf, natural indoor lighting, handheld camera feel.
/api/v1/wan/image_to_video: Submit the task, poll for its status, then verify the completed output. POST /api/v1/wan/image_to_video and save the returned Task id.
GET /api/v1/wan/image_to_video/<task-id> until status is completed or failed.
Verify the completed response contains the endpoint-specific output documented in the API reference.
/api/v1/wan/speech_to_video: Submit the task, poll for its status, then verify the completed output. POST /api/v1/wan/speech_to_video and save the returned Task id.
GET /api/v1/wan/speech_to_video/<task-id> until status is completed or failed.
Verify the completed response contains the endpoint-specific output documented in the API reference.
/api/v1/wan/text_to_video: Submit the task, poll for its status, then verify the completed output. POST /api/v1/wan/text_to_video and save the returned Task id.
GET /api/v1/wan/text_to_video/<task-id> until status is completed or failed.
Verify the completed response contains the endpoint-specific output documented in the API reference.
/api/v1/wan/animate: Submit the task, poll for its status, then verify the completed output. POST /api/v1/wan/animate and save the returned Task id.
GET /api/v1/wan/animate/<task-id> until status is completed or failed.
Verify the completed response contains the endpoint-specific output documented in the API reference.
/api/v1/wan/edit_video: Submit the task, poll for its status, then verify the completed output. POST /api/v1/wan/edit_video and save the returned Task id.
GET /api/v1/wan/edit_video/<task-id> until status is completed or failed.
Verify the completed response contains the endpoint-specific output documented in the API reference.
/api/v1/wan/text_to_image: Submit the task, poll for its status, then verify the completed output. POST /api/v1/wan/text_to_image and save the returned Task id.
GET /api/v1/wan/text_to_image/<task-id> until status is completed or failed.
Verify the completed response contains the endpoint-specific output documented in the API reference.

Public Versions and Endpoints

Model ID Endpoints Price Catalog
wan-2.2-a14b-image-to-video-turbo
/api/v1/wan/image_to_video
$0.40 / call Model detail
wan-2.2-a14b-speech-to-video-turbo
/api/v1/wan/speech_to_video
$0.24 / second Model detail
wan-2.2-a14b-text-to-video-turbo
/api/v1/wan/text_to_video
$0.40 / call Model detail
wan-2.2-animate-move
/api/v1/wan/animate
$0.13 / second Model detail
wan-2.2-animate-replace
/api/v1/wan/animate
$0.13 / second Model detail
wan-2.5-image-to-video
/api/v1/wan/image_to_video
$0.12 / second Model detail
wan-2.5-text-to-video
/api/v1/wan/text_to_video
$0.12 / second Model detail
wan-2.6-edit-video
/api/v1/wan/edit_video
$0.14 / second Model detail
wan-2.6-flash-edit-video
/api/v1/wan/edit_video
$0.30 / call Model detail
wan-2.6-flash-image-to-video
/api/v1/wan/image_to_video
$9.44 / call Model detail
wan-2.6-image-to-video
/api/v1/wan/image_to_video
$0.58 / second Model detail
wan-2.6-text-to-video
/api/v1/wan/text_to_video
$0.58 / second Model detail
wan-2.7-edit-video
/api/v1/wan/edit_video
$0.16 / second Model detail
wan-2.7-image
/api/v1/wan/text_to_image
$0.05 / call Model detail
wan-2.7-image-pro
/api/v1/wan/text_to_image
$0.12 / call Model detail
wan-2.7-image-to-video
/api/v1/wan/image_to_video
$0.16 / second Model detail
wan-2.7-r2v
/api/v1/wan/text_to_video
$0.16 / second Model detail
wan-2.7-text-to-video
/api/v1/wan/text_to_video
$0.16 / second Model detail

Verify

Poll until the task reaches a terminal status

Select <model-id> to generate verification commands.

Configuration

Guide endpoint: <endpoint>

Select <model-id> to generate a request with the endpoint's public input contract.
How it works

Get Started in 3 Steps

  1. Choose a model ID

    Select a public catalog model ID and review its endpoint and current starting price.

  2. Configure RunAPI

    Set RUNAPI_API_KEY before making the endpoint request.

  3. Verify the result

    For asynchronous endpoints, poll the same endpoint until the Task reaches a terminal status.

What to Build with Hermes Agent + Wan

  • Branded content at volume

    Use Wan's character consistency to produce branded video at scale, with Hermes Agent dispatching tasks for different product lines in parallel.

  • Dialogue content with lip sync

    Chain a text-to-speech model with Wan's speech-to-video endpoint in one Hermes Agent workflow to go from script to talking video.

  • Pre-visualization for filmmakers and agencies

    Generate pre-vis clips with anchored keyframes, setting first and last frames to control scene transitions for client review.

Why Use Wan Through RunAPI + Hermes Agent

  • 18 variants, one API key

    Use one RunAPI connection to choose among the live model variants without changing your integration.

  • Clear usage pricing

    See current catalog pricing before you send a request, with no subscription or minimum spend required.

  • Automatic task workflows

    Submit, poll, and collect asynchronous results through a consistent task workflow without writing manual polling code.

Hermes Agent + Wan Questions

Which Wan endpoints can I call from Hermes Agent?

Every Wan endpoint listed on this page. Configure RunAPI as a provider once, then switch endpoint and model version per request.

Do I need extra configuration for Wan in Hermes Agent?

No. The RunAPI provider you already use for chat reaches every Wan endpoint and every other RunAPI model without additional plugins.

Can Hermes Agent use Wan speech-to-video with a TTS model?

Yes. Hermes Agent can generate speech with a text-to-speech model, then pass the audio URL to Wan's speech-to-video endpoint in the same workflow.

Which model ID should I use?

Choose a public model ID from the version table. Each ID exposes the endpoints shown for that version.

Does this guide configure a chat model?

No. This Model Line uses the endpoint workflow shown here and is not presented as an agent chat model.

Start using Wan with Hermes Agent

Building with a team?

We're here to help with enterprise setup, integrations, and technical questions.

Contact Us