Google · Text

Hermes Agent x Gemini

Gemini is Google's flagship multimodal LLM, available in Flash (fast) and Pro (frontier reasoning) variants. Through RunAPI, all Gemini models share the same API shape and billing.

11 variants from $0.10 / 1M tokens Commercial OK

Prerequisite: npx runapi mcp install

Prompt

Prompt models

Use gemini-2.5-flash when it matches the task: Speed/cost optimized; 1M context; older generation baseline.

You have access to RunAPI chat models for Gemini.

Available Gemini models:
- gemini-2.5-flash: Speed/cost optimized; 1M context; older generation baseline
  chat endpoints: /v1/chat/completions
- gemini-2.5-flash-lite: Gemini 2.5 Flash-Lite; 1M context, thinking off by default; OpenAI-compatible Chat Completions
  chat endpoints: /v1/chat/completions
- gemini-2.5-pro: Best reasoning in 2.5 gen; 1M context
  chat endpoints: /v1/chat/completions
- gemini-3-flash-preview: gemini-3-flash-preview
  chat endpoints: /v1/chat/completions, /v1beta/models/gemini-3-flash-preview:streamGenerateContent
- gemini-3.1-flash-lite: Gemini 3.1 Flash-Lite; 1M context, adjustable thinking; text requests through OpenAI-compatible Chat Completions
  chat endpoints: /v1/chat/completions
- gemini-3.1-pro-preview: gemini-3.1-pro-preview
  chat endpoints: /v1/chat/completions
- gemini-3.5-flash: Fast multimodal streaming for high-volume production workloads
  chat endpoints: /v1/chat/completions, /v1beta/models/gemini-3.5-flash:streamGenerateContent
- gemini-3.5-flash-lite: Lightweight Gemini for text generation through Chat Completions or native content streaming
  chat endpoints: /v1/chat/completions, /v1beta/models/gemini-3.5-flash-lite:streamGenerateContent
- gemini-3.6-flash: Fast multimodal chat, tools, and streaming for production workloads
  chat endpoints: /v1/chat/completions, /v1beta/models/gemini-3.6-flash:streamGenerateContent
- gemini-3.7-flash: Fast streaming chat and native content output with usage reporting
  chat endpoints: /v1/chat/completions, /v1beta/models/gemini-3.7-flash:streamGenerateContent
- gemini-3.8-flash: Text chat with function calling and native content streaming for agent workloads
  chat endpoints: /v1/chat/completions, /v1beta/models/gemini-3.8-flash:streamGenerateContent

Use the model ID and chat endpoint that match the user's request.
Example request: Analyze this codebase and suggest three performance improvements with before/after examples.
Verify the active model with: hermes model;hermes doctor

Public Versions and Endpoints

Model ID Endpoints Price Catalog
gemini-2.5-flash
/v1/chat/completions
$0.30 / 1M tokens Model detail
gemini-2.5-flash-lite
/v1/chat/completions
$0.10 / 1M tokens Model detail
gemini-2.5-pro
/v1/chat/completions
$1.25 / 1M tokens Model detail
gemini-3-flash-preview
/v1/chat/completions /v1beta/models/gemini-3-flash-preview:streamGenerateContent
$0.30 / 1M tokens Model detail
gemini-3.1-flash-lite
/v1/chat/completions
$0.25 / 1M tokens Model detail
gemini-3.1-pro-preview
/v1/chat/completions
$1.00 / 1M tokens Model detail
gemini-3.5-flash
/v1/chat/completions /v1beta/models/gemini-3.5-flash:streamGenerateContent
$0.90 / 1M tokens Model detail
gemini-3.5-flash-lite
/v1/chat/completions /v1beta/models/gemini-3.5-flash-lite:streamGenerateContent
$0.15 / 1M tokens Model detail
gemini-3.6-flash
/v1/chat/completions /v1beta/models/gemini-3.6-flash:streamGenerateContent
$0.75 / 1M tokens Model detail
gemini-3.7-flash
/v1/chat/completions /v1beta/models/gemini-3.7-flash:streamGenerateContent
$0.45 / 1M tokens Model detail
gemini-3.8-flash
/v1/chat/completions /v1beta/models/gemini-3.8-flash:streamGenerateContent
$0.45 / 1M tokens Model detail

Verify

Poll until the task reaches a terminal status

Select <model-id> to generate verification commands.

Configuration

Guide endpoint: <endpoint>

Select <model-id> to generate a request with the endpoint's public input contract.
How it works

Get Started in 3 Steps

  1. Choose a model ID

    Select a public catalog model ID and review its endpoint and current starting price.

  2. Configure RunAPI

    Add the provider configuration for this agent.

  3. Verify the result

    Run the agent status and model-selection commands shown below.

What to Build with Hermes Agent + Gemini

  • Multimodal agents

    Use Gemini's multimodal input to build Hermes Agent workflows that reason over text, images, audio, and video together.

  • Long document analysis

    Give Gemini large codebases, contracts, or research collections in one request and let Hermes Agent ask follow-up questions over the same material.

  • Cost-efficient tool-calling chains

    Run a Flash version of Gemini for fast, low-cost tool-calling loops where the agent makes many sequential calls.

Why Use Gemini Through RunAPI + Hermes Agent

  • 11 variants, one API key

    Use one RunAPI connection to choose among the live model variants without changing your integration.

  • Clear usage pricing

    See current catalog pricing before you send a request, with no subscription or minimum spend required.

  • Direct responses

    Synchronous calls return the result in the same response, so your agent can use it immediately without task polling.

Hermes Agent + Gemini Questions

Can I use Gemini in Hermes Agent without Google Cloud credentials?

Yes. RunAPI serves Gemini through its OpenAI-compatible endpoint. Configure RunAPI as a provider in Hermes Agent with your RunAPI API key; no Google Cloud project or Vertex AI setup is required.

How does context caching help with long documents?

When the same large context is sent across many requests, caching lowers the input cost of later calls. This helps agent loops where instructions and reference material repeat.

Can Hermes Agent switch between Gemini and other models mid-session?

Yes. All RunAPI language models share the same provider and API key, so you can change the model during a session without changing the provider configuration.

Which model ID should I use?

Choose a public model ID from the version table. Each ID exposes the endpoints shown for that version.

Does this guide configure a chat model?

Yes. This Model Line supports a client-facing LLM protocol.

Start using Gemini with Hermes Agent

Building with a team?

We're here to help with enterprise setup, integrations, and technical questions.

Contact Us