Text Moonshot AI

Kimi API

Use the Kimi API via RunAPI with a model skill, unified auth, and pay-as-you-go pricing.

runapi.ai
# Base URL
https://runapi.ai

# Endpoints
POST /v1/chat/completions
POST /v1/responses
POST /v1/messages
POST /v1beta/models/{model}:generateContent
POST /v1beta/models/{model}:streamGenerateContent
curl https://runapi.ai/v1/chat/completions \
  -H "Authorization: Bearer $RUNAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "kimi-k3",
  "messages": [
    {
      "role": "user",
      "content": "Plan and implement a small CLI tool: scaffold the project, write the commands, add tests, and run them until they pass."
    }
  ]
}'
from openai import OpenAI

client = OpenAI(
    base_url="https://runapi.ai/v1",
    api_key="your-runapi-key"
)

response = client.chat.completions.create(
    model="kimi-k3",
    messages=[{"role": "user", "content": "Plan and implement a small CLI tool: scaffold the project, write the commands, add tests, and run them until they pass."}]
)
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://runapi.ai/v1",
  apiKey: "your-runapi-key"
});

const response = await client.chat.completions.create({
  model: "kimi-k3",
  messages: [{ role: "user", content: "Plan and implement a small CLI tool: scaffold the project, write the commands, add tests, and run them until they pass." }]
});
https://runapi.ai 5 endpoints
OVERVIEW

About Kimi

Kimi includes kimi-k3, the current flagship with always-on reasoning, alongside kimi-k2.5 and kimi-k2.6. The older K2 models use a 1T-parameter Mixture-of-Experts architecture with 32B active parameters per token and 384 experts per layer. kimi-k2.5 added native multimodal input; kimi-k2.6 improved long-horizon agent stability and reached 58.6% on SWE-bench Pro. RunAPI's initial kimi-k3 release supports basic text requests with synchronous and SSE streaming responses.

Provider
Moonshot AI
Modality
Text

Kimi API endpoints

EndpointProtocol
POST /v1/chat/completionsOpenAI compatible
POST /v1/responsesOpenAI Responses
POST /v1/messagesAnthropic compatible
POST /v1beta/models/{model}:generateContentGemini generateContent
POST /v1beta/models/{model}:streamGenerateContentGemini streamGenerateContent

Available Versions

Variant Billing Pricing
kimi-k2.5 1K tokens Input $0.60 / 1M tokens | Output $3.00 / 1M tokens View →
kimi-k2.6 1K tokens Input $0.95 / 1M tokens | Output $4.00 / 1M tokens View →
kimi-k2.7-code 1K tokens Input $0.57 / 1M tokens | Output $2.40 / 1M tokens View →
kimi-k3 1K tokens Input $3.00 / 1M tokens | Output $15.00 / 1M tokens View →
PRICING

Kimi Pricing

Endpoint Resolution Duration Price
chat_completion — — Input $0.60 / 1M tokens | Output $3.00 / 1M tokens

Prices shown for kimi-k2.5. Other variants retain their own endpoint pricing.

Agent integration

Run Kimi From an Agent

HOW IT WORKS

Get Started with Kimi

  1. Create an API key

    Sign up and generate an API key from the dashboard.

  2. Pick a version

    Choose the version that best matches your quality and cost requirements.

  3. Send a request

    Submit your prompt and parameters to the model's documented endpoint.

  4. Get your result

    Poll the task or set a webhook, then download the completed output.

CONTEXT

About Kimi on RunAPI

kimi-k3 is Kimi's current flagship and uses always-on reasoning. kimi-k2.5 and kimi-k2.6 remain available; their older K2 architecture has 1T parameters and 256K context, and the 58.6% SWE-bench Pro result belongs to kimi-k2.6. Kimi is available through OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, and Gemini native APIs. RunAPI's initial kimi-k3 release supports basic text requests with synchronous and SSE streaming responses.

Provider
Moonshot AI
See all →
Modality
Text
Browse models →

Why Use Kimi Through RunAPI

One auth, every provider

A single RunAPI API key unlocks the whole model catalog across all providers. No separate accounts to create, no API keys to rotate per integration, and no credential management overhead. Add a new model to your app by changing one parameter.

Unified pricing & billing

Per-call pricing in USD, billed monthly into a single invoice. No subscription tiers, no minimum spend, and failed generations are never charged. The pricing page and check_pricing API show exact costs before you commit to a model.

Schema-first SDK

Typed schemas, parameter constraints, and setup notes are packaged in the model skill so your implementation starts from the right contract. The skill loads into Claude Code, Codex, Gemini CLI, Cursor, and VS Code — your agent knows the correct request shape before you write a line of code.

Kimi Questions

Which Kimi model should I start with?

Start with kimi-k3, Kimi's current flagship with always-on reasoning. kimi-k2.5 and kimi-k2.6 remain available for existing K2 use cases.

What is the difference between kimi-k2.5 and kimi-k2.6?

Same 1T MoE architecture. kimi-k2.5 added native multimodal input (vision + text). kimi-k2.6 is a post-training refinement with stronger long-horizon stability — SWE-bench Pro jumps from 50.7% to 58.6%, and Agent Swarm scales from 100 to 300 sub-agents.

Which APIs can call Kimi through RunAPI?

Use OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, or Gemini native APIs with a Kimi model id. RunAPI's initial kimi-k3 release supports basic text requests with synchronous and SSE streaming responses.

How is Kimi billed?

See each Kimi model's page in the RunAPI catalog for current pricing.

Which variant should I start with?

Pick the cheapest variant that meets your quality bar. Most teams start on the fast variant and graduate to pro for production.

Is there a free tier?

Creating an account and API key is free. API calls use prepaid, pay-as-you-go billing; add funds before making requests.

SIMILAR MODELS

Similar Text Models

Start generating with Kimi