Skip to content
GoogleChat

Gemini 2.5 Flash Preview 09-2025

gemini-2.5-flash-preview-09-2025

Gemini 2.5 Flash Preview (Sept 2025) is Google's high-performance workhorse model built for advanced reasoning, coding, math, and scientific tasks. With built-in “thinking” capabilities, it delivers more accurate, context-aware answers across complex problems.

Context
1.0M tokens
Endpoint

Service Status

Status information temporarily unavailable

Apertis cannot confirm the current service state. This is not a report that the model is down.

Get API KeyCompare

Pricing

Input$0.15 / 1M
Output$1.25 / 1M
Cache Write (5m)$0.15 / 1M
Cache Write (1h)$0.15 / 1M
Cache Read$0.15 / 1M
Web Search$0 / 1M

Quick Start

Select an endpoint and copy a working example for this model.

Endpoint
python
from openai import OpenAI client = OpenAI(    api_key="YOUR_API_KEY",    base_url="https://api.apertis.ai/v1") response = client.chat.completions.create(    model="gemini-2.5-flash-preview-09-2025",    messages=[        {"role": "user", "content": "Hello!"}    ],    max_tokens=1024,    temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(#     model="gemini-2.5-flash-preview-09-2025",#     messages=[{"role": "user", "content": "Hello!"}],#     extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )

Supported Parameters

API docs
Common7 params
modelmessagesmax_tokenstemperaturetop_pstreamtools
Extended4 params
reasoning_effortstream_optionsthinkingextra_body

Cursor IDE Model IDs

Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.

gemini-2.5-flash-preview-09-2025

Compare with Other Models

See how this model compares to others from the same provider.