Gemini 2.5 Pro Preview
gemini-2.5-pro-preview-03-25Gemini 2.5 Pro is Google's top reasoning model for coding, math, and scientific work. It uses built-in “thinking” to deliver more accurate, context-aware answers and ranks at the top of major benchmarks like LMArena, showing strong alignment and problem-solving ability.
- Context
- 1.0M tokens
- Endpoint
Service Status
Status information temporarily unavailable
Apertis cannot confirm the current service state. This is not a report that the model is down.
Pricing
Quick Start
Select an endpoint and copy a working example for this model.
from openai import OpenAI client = OpenAI( api_key="YOUR_API_KEY", base_url="https://api.apertis.ai/v1") response = client.chat.completions.create( model="gemini-2.5-pro-preview-03-25", messages=[ {"role": "user", "content": "Hello!"} ], max_tokens=1024, temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(# model="gemini-2.5-pro-preview-03-25",# messages=[{"role": "user", "content": "Hello!"}],# extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )Supported Parameters
API docsmodelmessagesmax_tokenstemperaturetop_pstreamtoolsreasoning_effortstream_optionsthinkingextra_bodyCursor IDE Model IDs
Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.
Compare with Other Models
See how this model compares to others from the same provider.
Gemini 2.5 Flash Preview 05-20 (thinking)
Gemini 2.5 Flash Preview (May 2025) is Google's high-performance general model built for advanced reasoning, coding, math, and science. It includes built-in “thinking” features to deliver more accurate, context-aware answers.
- Context
- 1.0M
- Input
- $0.075/M
- Output
- $1.75/M
Gemini 2.5 Flash Lite Preview 09-2025
Gemini 2.5 Flash-Lite is a lightweight, low-latency model focused on speed and cost efficiency. It generates tokens quickly and outperforms earlier Flash models on common benchmarks. “Thinking” (multi-pass reasoning) is off by default for maximum speed, but can be turned on through the Reasoning API when deeper reasoning is needed.
- Context
- 1.0M
- Input
- $0.05/M
- Output
- $0.20/M
Gemini 2.5 Flash Lite Preview 06-17
Gemini 2.5 Flash-Lite is a smaller, low-latency model focused on speed and cost efficiency. It delivers faster generation and better benchmark performance than earlier Flash models. Thinking mode is off by default for maximum speed, but developers can enable it when they want deeper reasoning at a higher cost.
- Context
- 1.0M
- Input
- $0.025/M
- Output
- $0.10/M
Gemini 2.5 Flash Preview (thinking)
Gemini 2.5 Flash Preview (May 2025) is Google's high-performance general model built for advanced reasoning, coding, math, and science. It includes built-in “thinking” features to deliver more accurate, context-aware answers.
- Context
- 1.0M
- Input
- $0.075/M
- Output
- $1.75/M