Gemini 3.6 Flash: Cut Costs, Code Faster
Unlock Google's new Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber models for faster code, lower costs, and automated security defense.
Jul 21, 2026 (Updated Jul 21, 2026) - Written by Lorenzo Pellegrini
Source: Google.
Lose the Friction: Generate Branded QR Codes for Your Bio Page
Struggling to direct live event audiences to your complex digital product URLs? Generate a fully customizable QR code linked straight to your profile.
Google Launches Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: A New Era of AI Efficiency
Google has officially expanded its artificial intelligence lineup with the launch of three new specialized models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. These updates represent a strategic shift toward extreme efficiency, delivering faster code generation and reduced token consumption while maintaining frontier-level performance for complex, multi-step workflows.
The announcement marks a significant evolution in the Flash series, which is now the default model for the Gemini app and AI Mode in Search globally. By prioritizing cost reduction and speed, Google aims to make high-volume AI agent systems more scalable for developers and enterprises alike.
Gemini 3.6 Flash: The Efficiency Upgrade
Gemini 3.6 Flash serves as the primary evolution of the previous 3.5 Flash model, offering measurable improvements in programming, knowledge management, and computer use capabilities. The core advantage of this model is its ability to accomplish tasks with significantly fewer resources.
- Token Reduction: It consumes 17% fewer output tokens compared to Gemini 3.5 Flash according to the Artificial Analysis Index.
- Workflow Efficiency: The model requires fewer reasoning steps and tool calls to complete multi-step workflows, directly lowering operational costs.
- Performance Focus: It excels in coding tasks, advanced reasoning, and multimodal understanding while maintaining the high-speed characteristics of the Flash family.
Developers can access Gemini 3.6 Flash immediately through Google Antigravity, AI Studio, and Android Studio. Enterprises can also utilize the model via the Gemini Enterprise Agent Platform.
Gemini 3.5 Flash-Lite: Speed at the Lowest Cost
Gemini 3.5 Flash-Lite is designed specifically for tasks requiring high throughput and ultra-low latency, such as document search and processing. It positions itself as the fastest and least expensive model in the entire Gemini 3.5 family.
Key Specifications and Pricing
The Lite model delivers a blistering output speed of 350 tokens per second, making it ideal for smaller tasks within larger AI-agent systems. Google has also aggressively reduced pricing to support high-volume workloads:
- Input Cost: $0.30 per 1 million tokens.
- Output Cost: $2.50 per 1 million tokens.
- Comparison: Previous pricing for the predecessor was $9 per 1 million output tokens, representing a massive cost reduction.
This model is now available in the Gemini app and for developers via the same channels as the standard Flash version.
Gemini 3.5 Flash Cyber: Specialized Security Defense
Addressing the growing demand for automated cybersecurity, Google introduced Gemini 3.5 Flash Cyber, a specialized model built to find and fix security vulnerabilities in code. This release serves as a direct competitive response to Anthropic's lead in the cybersecurity AI sector.
The model is capable of detecting, validating, and correcting software security issues at a large scale. It operates at a lower price per token compared to larger, general-purpose models, making security auditing more accessible for developers.
Availability Note: Unlike the other two models, Gemini 3.5 Flash Cyber will initially be available only to governments and trusted partners through a limited-access pilot program.
Strategic Impact and Future Outlook
The launch of these three models underscores Google's commitment to balancing performance with economic efficiency. By reducing token usage and lowering costs per million tokens, Google enables developers to scale AI agents without the prohibitive expenses often associated with frontier models.
These updates also signal the continued momentum of the Gemini ecosystem, with the company simultaneously teasing the upcoming Gemini 4 and confirming that training for the next generation is underway. The Flash family remains the backbone of Google's consumer and developer offerings, now optimized for a new standard of speed and cost-effectiveness.
Conclusion
Google's new trio of Gemini models—3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber—delivers a powerful combination of reduced costs, increased speed, and specialized security capabilities. These tools are now available to the public and developers, marking a pivotal moment for scalable and efficient artificial intelligence adoption.
Google’s launch of Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber reveals a strategic pivot: the Flash tier is no just a cost-saving fallback but now the primary engine for agentic work, effectively decoupling frontier performance from high latency and transforming AI efficiency into a competitive moat rather than a compromise.
By specializing Flash models for coding, throughput, and security while positioning them as the default for consumer and enterprise agents, Google is implicitly admitting that the next breakthrough in AI won’t come from bigger models, but from smarter, cheaper, faster inference architectures that make large-scale automation economically viable.
How much does Gemini 3.5 Flash-Lite cost to use?
