Gemini 3.5 Flash-Lite
Google's cheapest current Flash-class API tier for high-volume, latency-sensitive agent loops.
Catalog record checked August 6, 2026; individual provider fields may change.
| Specification | Gemini 3.5 Flash-Lite |
|---|---|
| Provider | |
| Tier | Budget |
| Context window | 1.05M |
| Max output | 66K |
| Input / 1M tokens | $0.30 |
| Output / 1M tokens | $2.50 |
| Weights | Closed |
| Parameters | Not disclosed |
| Modalities | text, image, video, audio, pdf |
| Released | July 21, 2026 |
Verified evidence
Published benchmark results
Each result keeps its source and measurement date visible. A missing benchmark is not treated as a zero.
Artificial Analysis Intelligence Index
2026-08-08 · Artificial Analysis
Best for
- High-volume agents
- Document pipelines
- Cost-sensitive multimodal work
Watch out
Common questions
Gemini 3.5 Flash-Lite
Answered from the verified figures on this page rather than general guidance.
How much does Gemini 3.5 Flash-Lite cost per million tokens?
Gemini 3.5 Flash-Lite is listed at $0.30 per million input tokens and $2.50 per million output tokens at standard rates. Output tokens usually dominate real bills, so weigh the output rate more heavily than the input rate.
What is Gemini 3.5 Flash-Lite's context window?
Gemini 3.5 Flash-Lite accepts about 1.05M tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.
What is Gemini 3.5 Flash-Lite best for?
Gemini 3.5 Flash-Lite is a budget tier from Google. It suits high-volume agents, document pipelines, cost-sensitive multimodal work. Cheaper than 3.6 and 3.7 Flash, not stronger on hard reasoning — pick it for throughput, not frontier coding.
Compare it
Head-to-head model comparisons
These are the published pairings that put this model against a plausible alternative.