Public Versions and Endpoints
| Model ID | Endpoints | Price | Catalog |
|---|---|---|---|
gemini-2.5-flash
|
/v1/chat/completions
|
$0.30 / 1M tokens | Model detail |
gemini-2.5-flash-lite
|
/v1/chat/completions
|
$0.10 / 1M tokens | Model detail |
gemini-2.5-pro
|
/v1/chat/completions
|
$1.25 / 1M tokens | Model detail |
gemini-3-flash-preview
|
/v1/chat/completions
/v1beta/models/gemini-3-flash-preview:streamGenerateContent
|
$0.30 / 1M tokens | Model detail |
gemini-3.1-flash-lite
|
/v1/chat/completions
|
$0.25 / 1M tokens | Model detail |
gemini-3.1-pro-preview
|
/v1/chat/completions
|
$1.00 / 1M tokens | Model detail |
gemini-3.5-flash
|
/v1/chat/completions
/v1beta/models/gemini-3.5-flash:streamGenerateContent
|
$0.90 / 1M tokens | Model detail |
gemini-3.5-flash-lite
|
/v1/chat/completions
/v1beta/models/gemini-3.5-flash-lite:streamGenerateContent
|
$0.15 / 1M tokens | Model detail |
gemini-3.6-flash
|
/v1/chat/completions
/v1beta/models/gemini-3.6-flash:streamGenerateContent
|
$0.75 / 1M tokens | Model detail |
gemini-3.7-flash
|
/v1/chat/completions
/v1beta/models/gemini-3.7-flash:streamGenerateContent
|
$0.45 / 1M tokens | Model detail |
gemini-3.8-flash
|
/v1/chat/completions
/v1beta/models/gemini-3.8-flash:streamGenerateContent
|
$0.45 / 1M tokens | Model detail |
Verify
Poll until the task reaches a terminal status
Select <model-id> to generate verification commands.
Configuration
Guide endpoint: <endpoint>
Select <model-id> to generate a request with the endpoint's public input contract.
Get Started in 3 Steps
-
Choose a model ID
Select a public catalog model ID and review its endpoint and current starting price.
-
Configure RunAPI
Add the provider configuration for this agent.
-
Verify the result
Run the agent status and model-selection commands shown below.
What to Build with Hermes Agent + Gemini
-
Multimodal agents
Use Gemini's multimodal input to build Hermes Agent workflows that reason over text, images, audio, and video together.
-
Long document analysis
Give Gemini large codebases, contracts, or research collections in one request and let Hermes Agent ask follow-up questions over the same material.
-
Cost-efficient tool-calling chains
Run a Flash version of Gemini for fast, low-cost tool-calling loops where the agent makes many sequential calls.
Why Use Gemini Through RunAPI + Hermes Agent
-
11 variants, one API key
Use one RunAPI connection to choose among the live model variants without changing your integration.
-
Clear usage pricing
See current catalog pricing before you send a request, with no subscription or minimum spend required.
-
Direct responses
Synchronous calls return the result in the same response, so your agent can use it immediately without task polling.
Hermes Agent + Gemini Questions
Can I use Gemini in Hermes Agent without Google Cloud credentials?
Yes. RunAPI serves Gemini through its OpenAI-compatible endpoint. Configure RunAPI as a provider in Hermes Agent with your RunAPI API key; no Google Cloud project or Vertex AI setup is required.
How does context caching help with long documents?
When the same large context is sent across many requests, caching lowers the input cost of later calls. This helps agent loops where instructions and reference material repeat.
Can Hermes Agent switch between Gemini and other models mid-session?
Yes. All RunAPI language models share the same provider and API key, so you can change the model during a session without changing the provider configuration.
Which model ID should I use?
Choose a public model ID from the version table. Each ID exposes the endpoints shown for that version.
Does this guide configure a chat model?
Yes. This Model Line supports a client-facing LLM protocol.