VentureBeat
Follow
Google's Gemini 3.6 Flash model cuts AI agent token costs by up to 65% on long horizon engineering tasks —and 3.5 Pro is on the way
Google DeepMind has launched three new proprietary AI models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. These models are designed to be more token-efficient, making AI agents faster, smarter, and cheaper to operate at scale. Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens, while Gemini 3.5 Flash-Lite is significantly cheaper at $0.30 and $2.50 respectively. For comparison, previous models like Gemini 3.1 Flash-Lite remain the most cost-efficient but are slower. The new Gemini 3.5 Flash-Lite offers improved speed for enterprises prioritizing performance over absolute lowest cost. Gemini 3.6 Flash and 3.5 Flash-Lite achieve notable efficiency gains, reducing token usage by up to 65% in certain benchmarks. These models feature a 1-million-token input context window and a 64,000-token output limit. Gemini 3.6 Flash is suited for complex coding and knowledge work, while 3.5 Flash-Lite excels in high-throughput, low-latency applications. Gemini 3.5 Flash Cyber is a specialized model for cybersecurity research, available to select partners. All these models are proprietary and closed-source, accessible only through Google's API. Notably, the highly anticipated Gemini 3.5 Pro flagship model is still undergoing partner testing. The release signals a focus on agentic AI capabilities, with the Flash series likened to efficient delivery vans compared to older, fuel-hungry models.