3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Developers and customers building production AI agents require higher token efficiency, lower latency, and more reliable performance. Our Flash series of models is built to meet the sweet spot of efficiency and quality to enable scaling agentive workflows. Building on Gemini 3.5 Flash, we are introducing new Gemini models:

  • 3.6 Flash: Our workhorse model that delivers superior coding, knowledge work, and multimodal performance. according to artificial analysis indexThis reduces output token usage by 17% compared to 3.5 flash, and in some benchmarks like DeepSWE datacurveWe observe up to 65% lower cost per output token.
  • 3.5 Flash-Light: Our fastest, most cost-effective 3.5-tier model delivers up to 350 output tokens per second according to the Artificial Analytics Index, while also outperforming previous flash-light generations in agentic workflows.
  • Flash Cyber ​​3.5 in CodeMender: Successful cybersecurity applications require careful planning of the agent infrastructure as well as the model. We are introducing the combination of a new, highly efficient, specialized cyber-focused model combined with our Codemender code security agent that delivers competitive performance at the limit.

In addition to today’s release, Gemini 3.5 Pro is currently in testing with partners and we plan to make it widely available as soon as it is ready. In parallel, our team is already focusing on creating the next generation models. We have begun our most ambitious pre-training run to date for Gemini 4, and are excited by the progress.

3.6 Flash: More efficient and better quality than 3.5 flash

Gemini 3.6 Flash is based directly on developer and customer feedback from 3.5 Flash. 3.6 Flash not only takes a step forward in coding and knowledge work, but it does so while significantly improving token efficiency. For example, on the synthetic analytics index, we see that 3.6 flash consumes 17% fewer output tokens than 3.5 flash. It requires fewer logical steps and tool calls to complete multi-step workflows.

This increased efficiency is also coupled with a price tag of less than 3.5 flashes. At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the total cost per agent task, making agents more cost-effective to build and run.