Gemini 3.5 Flash-Lite
gemini-3-5-flash-litegemini-3.5-flash-lite- - ⚡ Fastest, lowest-cost model in Google's Gemini 3.5 family.
- - 📏 Handles a 1M-token context window.
- - 👁️ Multimodal input across text, images, and audio.
- - 🧠 Adjustable thinking levels, defaulting to minimal for speed.
- - 🔧 Native tool use including function calling and search.
- - 🎯 Aimed at everyday questions, summarization, and lightweight coding.
- - 🆕 Reported near 350 output tokens/s by Artificial Analysis.
- - 🏢 Available for production on Google Cloud.
Google is an American multinational technology corporation and one of the world's most valuable brands. A subsidiary of parent company Alphabet Inc., Google operates across search, cloud computing, consumer electronics, and artificial intelligence. Its DeepMind and Google…
Explore 17 more models by Google →Gemini 3.5 Flash-Lite, released in July 2026 alongside [[sibling:gemini-3-6-flash|Gemini 3.6 Flash]], is Google's most cost-efficient and lowest-latency model in the Gemini 3.5 generation. It targets high-volume, cost-sensitive workloads such as agentic retrieval, classification, extraction, and document processing, while retaining a 1M-token context window and native multimodal input across text, images, and audio. Google's documentation positions it as a fit for less complex [[sibling:gemini-3-5-flash|Gemini 3.5 Flash]] workloads that prioritize throughput.
Compared with its direct predecessor, the earlier Gemini 3.1 Flash-Lite, this release continues the Flash-Lite line as the fastest tier of the family, now built on the newer 3.5 generation architecture. As measured by the independent evaluator Artificial Analysis, output throughput is reported at roughly 350 tokens per second.
A defining feature is configurable thinking levels: it defaults to minimal thinking to optimize speed and cost for latency-sensitive tasks, but can scale up reasoning effort when quality matters. It also carries the 3.5-series tool suite, including function calling and built-in search.
Relative to the heavier [[sibling:gemini-3-5-flash|Gemini 3.5 Flash]] and the newer [[sibling:gemini-3-6-flash|Gemini 3.6 Flash]], Flash-Lite trades peak reasoning depth for throughput and lower cost, making it the family's default choice for everyday questions, summarization, and lightweight coding.
This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.
| Seller | Reputation↓ | Routing | Input $/M | Cached $/M | Output $/M | Categories | API |
|---|---|---|---|---|---|---|---|
| ▲ Apex Ant 0x73b4…e736 | 83 | #1 | $0.2599 | $0.026 | $2.1661 | chat,fast,premium,vision,multimodal,reasoning,long-context,agents | openai-chat-completions |
| Fire Ant 🔥🐜 0xbe05…bc5d | 37 | gated | $0.3229 | $0.0323 | $2.6905 | chat | — |
| D5V1N2 0xd5e7…7be0 | 26 | gated | $0.37 | $0.37 | $3.07 | chat,reasoning,coding,fast,cheap,vision,large-context,gemini,router,fallback | openai-chat-completions |
| antseed-opal-badger-2580 0xc85d…2580 | 2 | gated | $0.1875 | $0.1875 | $1.5625 | chat | openai-chat-completions |
| Apex TEE Test 0xe672…7955 | 0 | gated | $0.2747 | $0.2747 | $2.2896 | chat,fast,premium,vision,multimodal,reasoning,long-context,agents | openai-chat-completions |
"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens — or $ per generated image for unit-billed image models — not settled amounts). Reputation = on-chain trust (0-100). "Routing" = the SDK's default buyer routing (what the VPR desktop app ships with): a trust ≥ 60 gate on the effective reputation, then cheapest-first among routable sellers; live failover state (per-peer cooldowns) is buyer-side runtime and not included. Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase. "Usage on AntSeed" counts only settlements whose buyers share the per-model split on-chain (metadata v2/v3, opt-in), so every usage figure is a lower bound.