# Google: gemini-2.5-flash

- Model ID: `gemini-2.5-flash`
- Provider: Google
- Web version: https://www.moleapi.com/en/models/google/gemini-2.5-flash
- Content status: verified

## Model introduction

Gemini 2.5 Flash is Google's mature model for speed, cost, and scaled multimodal processing, retaining 1M context and controllable thinking.

## Model capabilities

- OpenAI compatible
- Gemini compatible
- Vision
- Prompt cache
- Reasoning

### Verified specifications

- Official positioning: Efficient, mature multimodal Flash model
- Context window: 1,048,576 tokens
- Maximum output: 65,536 tokens
- Input / output: Text, image, audio, video, PDF / Text

## Model pricing and access

Pricing is supplied dynamically by the MoleAPI console API.

### Standard (default, x1)

- Input: $0.3 / 1M tokens
- Output: $2.5 / 1M tokens
- Cache read: $0.075 / 1M tokens

### Discount (discount, x0.8)

- Input: $0.24 / 1M tokens
- Output: $2 / 1M tokens
- Cache read: $0.06 / 1M tokens

### Access protocols

| Protocol | Method | Endpoint |
| --- | --- | --- |
| openai | POST | /v1/chat/completions |
| gemini | POST | /v1beta/models/{model}:generateContent |

- Live pricing source: https://home.moleapi.com/api/pricing

## About gemini-2.5-flash

Gemini 2.5 Flash is Google's mature model for speed, cost, and scaled multimodal processing, retaining 1M context and controllable thinking.

### Best for

Responsive assistants, document and media batches, extraction, classification, moderation, and cost-sensitive multimodal applications.

### Core strengths

- Provides a stable balance of speed, cost, and general capability.
- Supports 1M context plus text, image, audio, video, and PDF input.
- A tunable thinking budget helps control latency by task.

### Limitations

- Complex reasoning and accuracy-first coding favor Pro.
- It is no longer the latest Flash generation, so new projects should compare Gemini 3 Flash.

### Selection and production evaluation

Start a gemini-2.5-flash evaluation by mapping its official positioning to real work: Responsive assistants, document and media batches, extraction, classification, moderation, and cost-sensitive multimodal applications. The first pass should exercise both its main strength, "Provides a stable balance of speed, cost, and general capability.", and its known limitation, "Complex reasoning and accuracy-first coding favor Pro.", instead of relying on a single subjective general-chat comparison.

For access, MoleAPI currently lists openai, gemini protocols and the Standard, Discount billing groups for this Google model; the default price summary is Input $0.3 / 1M tokens · Output $2.5 / 1M tokens. Pricing, protocols, and groups come from the live catalog, so production planning should still price representative requests using real context length, output size, and cache-hit assumptions.

No independent ranking is shown unless it matches this exact model ID and reasoning profile, so nearby variants are not used as a proxy. Before launch, pin the model ID, prompt, and sample set, then compare task accuracy, structured-output validity, tool-call success, and timeout rates under the same conditions.

The model material on this page was last checked on 2026-07-24. When upstream model cards, context limits, or tool support change, update the cited bilingual facts before changing the recommendation; live MoleAPI price changes remain separate and update from the catalog automatically.

### Sources

- [B.AI Gemini 2.5 Flash model guide](https://docs.b.ai/llmservice/models/gemini-2.5-flash/) — B.AI
- [Google Gemini model documentation](https://ai.google.dev/gemini-api/docs/models) — Google
- Last verified: 2026-07-24

## Code examples

### cURL

```bash
curl https://api.moleapi.com/v1/chat/completions \
  -H "Authorization: Bearer $MOLEAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gemini-2.5-flash","messages":[{"role":"user","content":"Hello"}]}'
```

## More related models

- [gemini-3.5-flash](https://www.moleapi.com/en/models/google/gemini-3.5-flash) — Google
- [gemini-3.1-pro](https://www.moleapi.com/en/models/google/gemini-3.1-pro) — Google
- [gemini-3-flash](https://www.moleapi.com/en/models/google/gemini-3-flash) — Google

## Frequently asked questions

### How is gemini-2.5-flash priced?

Prices are read from the MoleAPI console API and update with model prices, context tiers, and account groups.

### How can I access gemini-2.5-flash?

The catalog currently lists openai, gemini.

### How do I switch an existing project to gemini-2.5-flash?

Keep the MoleAPI API address and key, replace the model parameter, and check protocol-specific parameters.

### Where do the gemini-2.5-flash details come from?

Capabilities and limitations are checked against B.AI and the other cited pages; pricing and protocols come from the MoleAPI console.
