New Gemini 3.6 Flash and 3.5 Flash-Lite models unveiled

The latest Gemini models have been introduced, offering improved efficiency, latency, and reliability for the development of AI agents at scale. These new models, part of the Flash series, aim to optimize agentic workflows by achieving a balance of efficiency and quality. Building on the Gemini 3.5 Flash, the new Gemini models include the Gemini 3.6 Flash, which has been designed based on extensive developer and customer feedback.

Gemini 3.6 Flash offers enhanced performance in coding and knowledge tasks by improving token efficiency. For instance, it uses 17% fewer output tokens on the Artificial Analysis Index compared to its predecessor, 3.5 Flash. This model also reduces the number of reasoning steps and tool calls required for multi-step workflows, enhancing overall efficiency. Additionally, 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens, making it a more cost-effective solution for building and operating agents.

The new model demonstrates better token efficiency and reduced verbosity than the previous version in OSWorld-verified tasks. Across various use cases, 3.6 Flash outperforms 3.5 Flash, offering improved capabilities in parsing financial data and transcripts, executing code migrations with lower latency, and developing tools for 3D workflows. It also features enhanced safeguards in potentially hazardous domains like Chemical, Biological, Radiological, and Nuclear areas, making it more resilient to misuse.

Alongside 3.6 Flash, the Gemini 3.5 Flash-Lite model is released, designed for tasks that require low latency and high throughput, such as agentic search and document processing. It operates at a speed of 350 output tokens per second, offering a strong price-to-performance ratio for high-volume production traffic. At $0.3 per million input tokens and $2.5 per million output tokens, 3.5 Flash-Lite provides significant cost savings and efficiency.

3.5 Flash-Lite supports the efficient scaling of agentic systems and excels in coding and real-world tasks, outperforming previous models in various benchmarks. It is capable of generating unique web design concepts, translating and summarizing receipts, and even building games by iterating through multiple options.

Furthermore, the 3.5 Flash Cyber model, based on 3.5 Flash, is tailored for cybersecurity applications. It is priced lower per token than larger models and offers capabilities for detecting and fixing cybersecurity vulnerabilities. This model is initially available through a limited-access pilot program, targeting governments and trusted partners to address critical vulnerabilities effectively.

Both Gemini 3.6 Flash and 3.5 Flash-Lite are available now. Users are encouraged to provide feedback as they start utilizing these models, helping to refine future versions. Updates and special offers can be received by subscribing to newsletters.