Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite is a high-efficiency multimodal large language model released by Google on July 21, 2026. Positioned as the most cost-effective and fastest entry in the Gemini 3.5 series, it is specifically engineered for high-throughput, latency-sensitive applications. It is primarily used for agentic search, bulk document processing, and lightweight automation workflows where rapid execution is prioritized over maximum reasoning depth.
Technical performance is a core focus of the model, with benchmarks from Artificial Analysis indicating a generation speed of approximately 350 output tokens per second. Despite its "Lite" designation, the model shows significant gains in coding and agentic reasoning compared to previous generations, such as Gemini 3.1 Flash-Lite, and maintains a massive 1-million-token context window to support long-form data analysis and complex multi-step tasks.
The model is natively multimodal, supporting inputs across text, images, video, audio, and PDF documents. It incorporates Google’s "thinking levels" framework, allowing developers to choose between minimal, low, medium, and high reasoning effort. By default, Gemini 3.5 Flash-Lite utilizes a minimal thinking level to ensure the lowest possible latency and cost while maintaining reliability for structured output and function calling.
Key features include enhanced precision in code generation with reduced execution loops and improved performance in machine learning engineering tasks. It also supports advanced platform capabilities such as context caching, computer use, and integrated web search tools, making it a versatile sub-agent for larger orchestrations.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Gemini 3.5 Flash-Lite ranks
Gemini 3.5 Flash-Lite is highlighted in the table below. Switch the metric to see how the ordering changes.