PROVIDER

Z.ai AI Models

Z.ai's GLM — MIT-licensed MoE LLMs from 128K to 200K context, top open-weight SWE-bench scores, via one RunAPI key.

1 models · 8 variants · from $0.0001
OVERVIEW

Z.ai builds the GLM family of MIT-licensed Mixture-of-Experts language models for coding and agentic workflows. The line spans GLM-4.5 (355B / 32B active, 128K context) through GLM-5.1 (754B / 40B active, 200K context), which holds the top open-weight SWE-bench Pro score at 58.4%. All are available through RunAPI from the OpenAI and Anthropic SDKs with per-token billing.

  • Single API key shared across all providers
  • No separate Z.ai account required
  • Model skills carry docs, schemas, and setup steps into your workspace
  • Per-call billing in USD, no subscription or minimum spend
  • Failed generations are never charged
  • Switch models by changing one parameter
  • Billing consolidated into one monthly invoice
FEATURES

What stands out

MODELS

All Z.ai models available through RunAPI

QUICKSTART

Install a Z.ai model skill for your app.

Pick a model and add its skill so your coding tool has docs, schemas, pricing notes, and setup steps. Skills work with Claude Code, Codex, Gemini CLI, Cursor, and VS Code. Install once, then switch models by changing one parameter.

runapi.ai
# Base URL
https://runapi.ai

# Endpoints
POST /v1/chat/completions
POST /v1/responses
POST /v1/messages
POST /v1beta/models/{model}:generateContent
POST /v1beta/models/{model}:streamGenerateContent
curl https://runapi.ai/v1/chat/completions \
  -H "Authorization: Bearer $RUNAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5.2",
  "messages": [
    {
      "role": "user",
      "content": "Read this multi-file repository, find the failing integration test, and propose a patch with an explanation of the root cause."
    }
  ]
}'
from openai import OpenAI

client = OpenAI(
    base_url="https://runapi.ai/v1",
    api_key="your-runapi-key"
)

response = client.chat.completions.create(
    model="glm-5.2",
    messages=[{"role": "user", "content": "Read this multi-file repository, find the failing integration test, and propose a patch with an explanation of the root cause."}]
)
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://runapi.ai/v1",
  apiKey: "your-runapi-key"
});

const response = await client.chat.completions.create({
  model: "glm-5.2",
  messages: [{ role: "user", content: "Read this multi-file repository, find the failing integration test, and propose a patch with an explanation of the root cause." }]
});
https://runapi.ai 5 endpoints
REFERENCE

Every Z.ai variant with pricing and model IDs

Full pricing table →
Model Variant Billing Pricing
GLM
glm-4.5 1K tokens Input $0.41 / 1M tokens | Output $1.61 / 1M tokens View →
glm-4.5-air 1K tokens Input $0.10 / 1M tokens | Output $0.55 / 1M tokens View →
glm-4.6 1K tokens Input $0.60 / 1M tokens | Output $1.22 / 1M tokens View →
glm-4.7 1K tokens Input $0.60 / 1M tokens | Output $1.19 / 1M tokens View →
glm-5 1K tokens Input $1.00 / 1M tokens | Output $1.60 / 1M tokens View →
glm-5-turbo 1K tokens Input $1.20 / 1M tokens | Output $2.00 / 1M tokens View →
glm-5.1 1K tokens Input $1.40 / 1M tokens | Output $2.20 / 1M tokens View →
glm-5.2 1K tokens Input $0.70 / 1M tokens | Output $2.20 / 1M tokens View →
FAQ

Frequently asked questions about Z.ai

Is this an official Z.ai integration?

RunAPI exposes a managed API surface with transparent per-call pricing, fully documented capability and parameters, and clear error behavior. You get the same model output quality without managing a direct provider relationship or provider-side account.

Do I need a Z.ai account?

No. Your RunAPI API key is enough for managed access to all Z.ai models. You do not need to create a separate account, manage provider-specific credentials, or handle provider-side billing.

What's the latency overhead from proxying through RunAPI?

Typically under 20 ms. RunAPI keeps the proxy layer close to model execution regions to minimize added latency. Media generation time is dominated by the model itself, not the proxy.

Are images / videos cached?

Generated outputs are stored and retrievable by task ID. You can fetch completed results at any time using the task status endpoint or the RunAPI dashboard. Output URLs remain accessible for the retention period shown in the API docs. Inputs are not cached or stored.

Can I bring my own key?

Not currently. Calls use RunAPI-managed access, which simplifies authentication and lets RunAPI handle rate limiting, retries, and billing consolidation on your behalf.

How is billing consolidated?

All API calls across all providers appear on a single monthly USD invoice. There is no per-provider billing, no subscription, and no minimum spend. Failed generations are never charged.

What SDKs can I use with Z.ai models?

Official SDKs are available for Python, Node.js, PHP, Java, Ruby, and Go. Each SDK handles authentication, async task polling, and typed responses. For LLM models, the OpenAI and Anthropic SDKs also work by pointing the base URL to RunAPI.

What are model skills and how do they work?

Model skills are installable packages that load a model's docs, typed schemas, pricing notes, and setup steps directly into your coding workspace. Install a skill with one command and your agent has the right context before you write integration code. Skills work with Claude Code, Codex, Gemini CLI, Cursor, and VS Code.

How do I switch between Z.ai models?

Change the model parameter in your API request. All Z.ai models share the same API key, the same request shape, and the same billing. No code changes, no re-authentication, and no separate billing setup are required when switching between models or between variants of the same model. You can also switch to models from other providers by changing the same parameter — the API surface is unified across the entire catalog.

START NOW

Start building with Z.ai models.