Model · Google · Text
Gemini 3.5 Flash-Lite
Google's fastest and most cost-effective Gemini tier — the high-throughput option for classification, extraction, and routing. Google's named replacement for Gemini 3.1 Flash-Lite.
- Modality
- Text
- License
- Proprietary (Proprietary)
- Context window
- Context window is not published on Google's models or pricing pages as of 2026-09-08.
- Released
- July 21, 2026
- Last verified
- September 8, 2026
- Runs locally
- No
- Also handles
- vision
Strengths
- Fastest tier in the lineup — around 350 output tokens per second on the Artificial Analysis Index
- Cheapest Gemini 3-series tier at $0.30 input / $2.50 output per MTok
- Holds up on long-context tasks (72.2%) and agentic terminal work (54% on Terminal-Bench 2.1) despite the price
Weaknesses
- The cheap tier — Google positions Gemini 3.6/3.7/3.8 Flash above it for coding and reasoning
- Context window is not published by Google
- Closed weights
Try it
| Where | Type | Notes |
|---|---|---|
| Google AI Studio | hosted-api | Free tier |
| Vertex AI | hosted-api | GCP enterprise |
Used in solutions
Version history
- Gemini 3.8 Flash Sep 2026
- Gemini 3.5 Flash-Lite Jul 2026 Current
- Gemini 3.5 Flash May 2026 Deprecated
- Gemini 3.1 Flash-Lite May 2026 Deprecated
- Gemini 3.1 Pro Preview May 2026
Official sources
- Model docs docs
- Announcement announcement
Change log
- — Initial entry. Released 2026-07-21; named on Google's deprecations page as the replacement for Gemini 3.1 Flash-Lite (shutdown 2027-05-07). Context window left unset — Google publishes none.
Esc