Google DeepMind launches Gemini 3.6 Flash with cost-efficient AI agents
Google's new Gemini 3.6 Flash cuts output token costs by 16.7% for AI agents, targeting real-world deployment over benchmark scores.
Google launched three new Gemini models—Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber—to optimize AI agents for production use. The 3.6 Flash model reduces output token costs by roughly 16.7% ($7.50 per million) and reports up to 65% fewer tokens on certain coding tasks, while Flash-Lite prioritizes speed and cost. Cyber targets security, initially limited to governments and selected trusted partners.
- Gemini 3.6 Flash reduces output token costs by 16.7% ($7.50 per million) and claims up to 65% fewer tokens on coding tasks.
- Gemini 3.5 Flash-Lite prioritizes speed and cost for high-volume parsing, while 3.5 Flash Cyber targets cybersecurity for governments.
- All models are designed for production deployment, with integrations into Android Studio, Google AI Studio, and Gemini Enterprise.
Why It Matters
Google’s focus on token efficiency and agentic workflows signals a shift toward practical, cost-effective AI deployment for enterprises.