---
title: "GLM API — Variants, pricing & model skill | RunAPI"
url: "https://runapi.ai/models/glm.md"
canonical: "https://runapi.ai/models/glm.md"
locale: "en"
model: "GLM"
provider: "Z.ai"
modality: "text"
variant_count: 7
price_from_cents: 1
---

# GLM API

Z.ai GLM API access via RunAPI — MIT-licensed MoE models with up to 200K context, leading open-weight coding benchmarks.

**Provider:** Z.ai
**Modality:** Text
**Catalog:** 7 variants

GLM is Z.ai&#39;s family of MIT-licensed Mixture-of-Experts language models. GLM-4.5 (355B total / 32B active, 128K context) introduced the open-weight MoE line with a flagship and a lighter Air tier. GLM-4.6 and 4.7 extend to 200K context with stronger code generation — 4.7 reaches 73.8% on SWE-bench. The GLM-5 series (744B / 40B active, 200K context) pushes further to 77.8% SWE-bench Verified, and GLM-5.1 holds the top open-weight score on SWE-bench Pro at 58.4%. All are available through RunAPI with one key and per-token billing.

## Variants

| Version | Variant | Pricing | Billing | URL |
|---|---|---|---|---|
| glm-4.5 | `4.5` | $0.020 | 1K tokens | https://runapi.ai/models/glm/4.5.md |
| glm-4.5-air | `4.5-air` | $0.010 | 1K tokens | https://runapi.ai/models/glm/4.5-air.md |
| glm-4.6 | `4.6` | $0.020 | 1K tokens | https://runapi.ai/models/glm/4.6.md |
| glm-4.7 | `4.7` | $0.020 | 1K tokens | https://runapi.ai/models/glm/4.7.md |
| glm-5 | `5` | $0.020 | 1K tokens | https://runapi.ai/models/glm/5.md |
| glm-5-turbo | `5-turbo` | $0.020 | 1K tokens | https://runapi.ai/models/glm/5-turbo.md |
| glm-5.1 | `5.1` | $0.030 | 1K tokens | https://runapi.ai/models/glm/5.1.md |


## API endpoints

Base URL: `https://runapi.ai`

- `POST /v1/chat/completions`

Use the OpenAI or Anthropic SDK with your RunAPI API key. No extra SDK required.

## Context

GLM models from Z.ai are MIT-licensed MoE LLMs spanning 128K–200K context. GLM-5.1 leads open-weight models on SWE-bench Pro. Through RunAPI they share a single API key with pay-as-you-go token billing, callable from the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages surfaces.

## FAQ

### Which variant should I start with?

Pick the cheapest variant that meets your quality bar. Most teams start on the fast variant and graduate to pro for production.

### Is there a free tier?

New accounts get free first calls on every model. After that, pay per call.

### Do you stream results?

Where streaming is available, RunAPI streams end-to-end.

### How are failures billed?

Failed generations are not charged.

### Are outputs cached?

Generated outputs are stored and retrievable by task ID. Inputs are not cached.

### Can I use commercially?

Yes — commercial use is included for every variant unless a model license explicitly restricts it, which is called out on the variant page.

### What about rate limits?

Per-key rate limits scale with usage tier. See pricing page for current limits.

### Where can I report issues?

Open an issue on the public GitHub repo or email support.

