Moonshot AI · Text

Hermes Agent x Kimi

kimi-k3 is Kimi's current flagship and uses always-on reasoning. kimi-k2.5 and kimi-k2.6 remain available; their older K2 architecture has 1T parameters and 256K context, and the 58.6% SWE-bench Pro result belongs to kimi-k2.6. Kimi is available through OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, and Gemini native APIs. RunAPI's initial kimi-k3 release supports basic text requests with synchronous and SSE streaming responses.

4 variants from $0.57 / 1M tokens Commercial OK

Prerequisite: npx runapi mcp install

Prompt

Prompt models

Use kimi-k2.5 when it matches the task: 1T / 32B active; 256K context; native multimodal input.

You have access to RunAPI chat models for Kimi.

Available Kimi models:
- kimi-k2.5: 1T / 32B active; 256K context; native multimodal input
  chat endpoints: /v1/chat/completions
- kimi-k2.6: 1T / 32B active; 256K context; 58.6% SWE-bench Pro; 300-agent swarm
  chat endpoints: /v1/chat/completions
- kimi-k2.7-code: Dedicated coding model; always-on thinking; 256K context
  chat endpoints: /v1/chat/completions
- kimi-k3: Current Kimi flagship with always-on reasoning
  chat endpoints: /v1/chat/completions

Use the model ID and chat endpoint that match the user's request.
Example request: Plan and implement a small CLI tool: scaffold the project, write the commands, add tests, and run them until they pass.
Verify the active model with: hermes model;hermes doctor

Public Versions and Endpoints

Model ID Endpoints Price Catalog
kimi-k2.5
/v1/chat/completions
$0.60 / 1M tokens Model detail
kimi-k2.6
/v1/chat/completions
$0.95 / 1M tokens Model detail
kimi-k2.7-code
/v1/chat/completions
$0.57 / 1M tokens Model detail
kimi-k3
/v1/chat/completions
$3.00 / 1M tokens Model detail

Verify

Poll until the task reaches a terminal status

Select <model-id> to generate verification commands.

Configuration

Guide endpoint: <endpoint>

Select <model-id> to generate a request with the endpoint's public input contract.
How it works

Get Started in 3 Steps

  1. Choose a model ID

    Select a public catalog model ID and review its endpoint and current starting price.

  2. Configure RunAPI

    Add the provider configuration for this agent.

  3. Verify the result

    Run the agent status and model-selection commands shown below.

What to Build with Hermes Agent + Kimi

  • Agent coding and reasoning

    Let your agent plan, write, and review code or work through multi-step problems with the model's reasoning.

  • Long document analysis and extraction

    Summarize reports, contracts, or codebases and pull structured fields out of long documents.

  • Tool-calling workflows

    Connect the model to functions and APIs so your agent can look up data and take actions between replies.

Why Use Kimi Through RunAPI + Hermes Agent

  • 4 variants, one API key

    Use one RunAPI connection to choose among the live model variants without changing your integration.

  • Clear usage pricing

    See current catalog pricing before you send a request, with no subscription or minimum spend required.

  • Direct responses

    Synchronous calls return the result in the same response, so your agent can use it immediately without task polling.

Hermes Agent + Kimi Questions

Which Kimi model should I start with?

Start with kimi-k3, Kimi's current flagship with always-on reasoning. kimi-k2.5 and kimi-k2.6 remain available for existing K2 use cases.

What is the difference between kimi-k2.5 and kimi-k2.6?

Same 1T MoE architecture. kimi-k2.5 added native multimodal input (vision + text). kimi-k2.6 is a post-training refinement with stronger long-horizon stability — SWE-bench Pro jumps from 50.7% to 58.6%, and Agent Swarm scales from 100 to 300 sub-agents.

Which APIs can call Kimi through RunAPI?

Use OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, or Gemini native APIs with a Kimi model id. RunAPI's initial kimi-k3 release supports basic text requests with synchronous and SSE streaming responses.

How is Kimi billed?

See each Kimi model's page in the RunAPI catalog for current pricing.

Which model ID should I use?

Choose a public model ID from the version table. Each ID exposes the endpoints shown for that version.

Does this guide configure a chat model?

Yes. This Model Line supports a client-facing LLM protocol.

Start using Kimi with Hermes Agent

Building with a team?

We're here to help with enterprise setup, integrations, and technical questions.

Contact Us