MiniMax · Text

Hermes Agent x MiniMax

MiniMax M-series text models are sparse MoE LLMs with 200K–1M context, delivering frontier coding scores at a fraction of the cost of dense models. Through RunAPI they share a single API key with pay-as-you-go token billing, callable from the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages surfaces. These are MiniMax's text models, distinct from MiniMax Hailuo video generation.

7 variants from $0.18 / 1M tokens Commercial OK

Prerequisite: npx runapi mcp install

Prompt

Prompt models

Use MiniMax-M2 when it matches the task: 230B / 10B active; 200K context; baseline M-series coding model.

You have access to RunAPI chat models for MiniMax.

Available MiniMax models:
- MiniMax-M2: 230B / 10B active; 200K context; baseline M-series coding model
  chat endpoints: /v1/chat/completions
- MiniMax-M2.1: 200K context; polyglot programming; 74% SWE-bench Verified
  chat endpoints: /v1/chat/completions
- MiniMax-M2.5: 200K context; stronger agentic tool calling and search
  chat endpoints: /v1/chat/completions
- MiniMax-M2.5-highspeed: M2.5 at ~100 tokens/sec; same weights, higher throughput
  chat endpoints: /v1/chat/completions
- MiniMax-M2.7: 200K context; 56.2% SWE-bench Pro; self-evolving training
  chat endpoints: /v1/chat/completions
- MiniMax-M2.7-highspeed: M2.7 at ~100 tokens/sec; same weights, higher throughput
  chat endpoints: /v1/chat/completions
- MiniMax-M3: 1M context; 80.5% SWE-bench Verified; frontier open-weight coding + long context
  chat endpoints: /v1/chat/completions

Use the model ID and chat endpoint that match the user's request.
Example request: Given this API spec, generate a typed client, write integration tests against a mock server, and iterate until they pass.
Verify the active model with: hermes model;hermes doctor

Public Versions and Endpoints

Model ID Endpoints Price Catalog
MiniMax-M2
/v1/chat/completions
$0.19 / 1M tokens Model detail
MiniMax-M2.1
/v1/chat/completions
$0.19 / 1M tokens Model detail
MiniMax-M2.5
/v1/chat/completions
$0.19 / 1M tokens Model detail
MiniMax-M2.5-highspeed
/v1/chat/completions
$0.37 / 1M tokens Model detail
MiniMax-M2.7
/v1/chat/completions
$0.19 / 1M tokens Model detail
MiniMax-M2.7-highspeed
/v1/chat/completions
$0.37 / 1M tokens Model detail
MiniMax-M3
/v1/chat/completions
$0.18 / 1M tokens Model detail

Verify

Poll until the task reaches a terminal status

Select <model-id> to generate verification commands.

Configuration

Guide endpoint: <endpoint>

Select <model-id> to generate a request with the endpoint's public input contract.
How it works

Get Started in 3 Steps

  1. Choose a model ID

    Select a public catalog model ID and review its endpoint and current starting price.

  2. Configure RunAPI

    Add the provider configuration for this agent.

  3. Verify the result

    Run the agent status and model-selection commands shown below.

What to Build with Hermes Agent + MiniMax

  • Agent coding and reasoning

    Let your agent plan, write, and review code or work through multi-step problems with the model's reasoning.

  • Long document analysis and extraction

    Summarize reports, contracts, or codebases and pull structured fields out of long documents.

  • Tool-calling workflows

    Connect the model to functions and APIs so your agent can look up data and take actions between replies.

Why Use MiniMax Through RunAPI + Hermes Agent

  • 7 variants, one API key

    Use one RunAPI connection to choose among the live model variants without changing your integration.

  • Clear usage pricing

    See current catalog pricing before you send a request, with no subscription or minimum spend required.

  • Direct responses

    Synchronous calls return the result in the same response, so your agent can use it immediately without task polling.

Hermes Agent + MiniMax Questions

Are these the same as MiniMax Hailuo video models?

No. These are MiniMax's text language models for coding and chat; Hailuo is MiniMax's separate video generation line.

Which MiniMax text model should I pick?

MiniMax-M3 is the strongest — 1M context, 80.5% SWE-bench Verified, and the first open-weight model to combine frontier coding with million-token context. M2.7 is the best 200K-context option. Highspeed variants (M2.5 and M2.7) run the same weights at ~100 tokens/sec for lower latency at higher token cost.

Which SDKs can call MiniMax text through RunAPI?

Use the OpenAI SDK (Chat Completions or Responses) or the Anthropic Messages SDK against RunAPI with the MiniMax model id; the proxy adapts the protocol.

How is MiniMax text billed?

Per token at RunAPI's published input and output rates for each model, pay-as-you-go. Highspeed variants are billed at their own published rates for the same output quality at higher throughput.

Which model ID should I use?

Choose a public model ID from the version table. Each ID exposes the endpoints shown for that version.

Does this guide configure a chat model?

Yes. This Model Line supports a client-facing LLM protocol.

Start using MiniMax with Hermes Agent

Building with a team?

We're here to help with enterprise setup, integrations, and technical questions.

Contact Us