Gemini 2.5 Flash Lite
Google · Proprietary · Released Jul 22, 2025 · google/gemini-2.5-flash-lite
Best $0.10 in / $0.40 out via Google
1.05M contextVisionToolsReasoningStructured outputPrompt cachingAudio input
Available from 5 providersUSD / MTOK · PRICES 25 MINUTES AGO · SPEED 18 DAYS AGO
EST. MONTHLY COSTclick a row for detail · click headers to sort
| PROVIDER | INPUT | OUTPUT | CACHE | CONTEXT | QUANT | TTFT | TPS | UPTIME |
|---|---|---|---|---|---|---|---|---|
| $0.05 | $0.20 | $0.005 | 1.05M | — | 0.31s | 233 tok/s | 100.0% | |
| $0.10 | $0.40 | $0.01 | 1.05M | — | 1.3s | 164 tok/s | 100.0% | |
| $0.10 | $0.40 | $0.01 | 1.05M | — | — | — | 99.6% | |
| $0.10 | $0.40 | $0.01 | 1.05M | — | — | — | 99.9% | |
| $0.18 | $0.72 | $0.018 | 1.05M | — | — | — | 100.0% |
TTFT/TPS: median (P50) over trailing 504h · chat10k profile (~10K in / 200 out) · measured via OpenRouter routing · methodology v1.1 · measured 18 days ago
SERVE THIS MODEL AND NOT LISTED? LIST YOUR INFERENCE →
Pricing variants · USD / MTok in / out
STANDARD
$0.10 / $0.40
per MTok in / out
BATCH
$0.05 / $0.20
async batch
CACHE READ $0.01CACHE WRITE $0.083WEB SEARCH $0.014 / search
About
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...
Specifications
Context window1,048,576
Max output65,535
Modalities intext, image, file, audio, video
Modalities outtext
LicenseProprietary
ReleasedJul 22, 2025
Knowledge cutoff2025-01-31
TokenizerGemini
Deprecation2026-10-20
Related
Embed badge
[](https://modelindex.ai/models/google/gemini-2.5-flash-lite)Sources
OpenRouter API25 minutes ago
AGGREGATORLiteLLM dataset25 minutes ago
AGGREGATORModelIndex benchmarks20 days ago
MEASUREDVerified 25 minutes ago · list prices, not negotiated rates.