Skip to main content

Google

Gemini 3.5 Flash-Lite

Google's cheapest current Flash-class API tier for high-volume, latency-sensitive agent loops.

Catalog record checked August 6, 2026; individual provider fields may change.

AI model specification details
SpecificationGemini 3.5 Flash-Lite
ProviderGoogle
TierBudget
Context window1.05M
Max output66K
Input / 1M tokens$0.30
Output / 1M tokens$2.50
WeightsClosed
ParametersNot disclosed
Modalitiestext, image, video, audio, pdf
ReleasedJuly 21, 2026

Verified evidence

Published benchmark results

Each result keeps its source and measurement date visible. A missing benchmark is not treated as a zero.

Artificial Analysis Intelligence Index

2026-08-08 · Artificial Analysis

Best for

  • High-volume agents
  • Document pipelines
  • Cost-sensitive multimodal work

Watch out

Cheaper than 3.6 and 3.7 Flash, not stronger on hard reasoning — pick it for throughput, not frontier coding.

Common questions

Gemini 3.5 Flash-Lite

Answered from the verified figures on this page rather than general guidance.

How much does Gemini 3.5 Flash-Lite cost per million tokens?

Gemini 3.5 Flash-Lite is listed at $0.30 per million input tokens and $2.50 per million output tokens at standard rates. Output tokens usually dominate real bills, so weigh the output rate more heavily than the input rate.

What is Gemini 3.5 Flash-Lite's context window?

Gemini 3.5 Flash-Lite accepts about 1.05M tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.

What is Gemini 3.5 Flash-Lite best for?

Gemini 3.5 Flash-Lite is a budget tier from Google. It suits high-volume agents, document pipelines, cost-sensitive multimodal work. Cheaper than 3.6 and 3.7 Flash, not stronger on hard reasoning — pick it for throughput, not frontier coding.

Compare it

Head-to-head model comparisons

These are the published pairings that put this model against a plausible alternative.