Google says Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are now generally available through the Gemini API. The company positions Gemini 3.6 Flash as a token-efficient multimodal reasoning model for coding, knowledge work and agentic workflows.
Google has made Gemini 3.6 Flash and Gemini 3.5 Flash-Lite generally available through the Gemini API, according to the company’s API release notes dated July 21, 2026.
Google describes Gemini 3.6 Flash as a more token-efficient model with stronger code-generation and agentic-planning capabilities than Gemini 3.5 Flash, while also listing it at a lower price. The announcement gives developers two additional production-oriented options in the Gemini API lineup, including a lower-cost Flash-Lite variant.
In its model card, Google DeepMind describes Gemini 3.6 Flash as a natively multimodal reasoning model intended for coding, knowledge work and agentic workflows. This means the model is designed to reason over multiple kinds of inputs, rather than operating solely on text.
Google lists Gemini 3.6 Flash API pricing at $1.50 per million input tokens and $7.50 per million output tokens. Input tokens represent material submitted in a request, while output tokens are the material generated by the model in response.
The Gemini API release notes say the model delivers stronger coding and planning performance with reduced token use compared with Gemini 3.5 Flash. The available source material does not provide task-level benchmark results in the release-note summary, so developers will need to assess those claims against their own workloads and evaluation criteria.
Gemini 3.5 Flash-Lite was also moved to general availability, according to Google’s release notes. Google’s supplied release-note summary provides fewer technical details about Flash-Lite than it does for Gemini 3.6 Flash, but its inclusion broadens the company’s lower-cost model offering.
TechCrunch reported that Google released a third model, Gemini 3.5 Flash Cyber, in the same announcement. The publication characterized Gemini 3.6 Flash as targeting coding, knowledge-work and multimodal tasks, and noted that Google did not introduce a Gemini 3.5 Pro model alongside the release.
For teams selecting a Gemini API model, Google’s pricing and positioning make token consumption, coding performance and planning capabilities central considerations. A model that uses fewer tokens may reduce spending for applications with large prompts or frequent requests, although total costs will also depend on output volume and the design of each application.
The release does not establish a single model as the right choice for every use case. Developers evaluating Gemini 3.6 Flash or Gemini 3.5 Flash-Lite should compare latency, multimodal requirements, tool use, output quality and pricing on representative tasks before making deployment decisions.
Google expands its Gemini API model range Google has made Gemini 3.6 Flash and Gemini 3.5 Flash Lite generally available through the Gemini API, according to the company’s API release notes dated July 21, 2026.
Google describes Gemini 3.6 Flash as a more token efficient model with stronger code generation and agentic planning capabilities than Gemini 3.5 Flash, while also listing it at a lower price.
The announcement gives developers two additional production oriented options in the Gemini API lineup, including a lower cost Flash Lite variant.
Continue reading