Google has released Gemini 3.6 Flash, an optimization update that cuts output costs by 16.7% and boosts performance across coding, reasoning, agent workflows, and long-context processing. It keeps the same input pricing at $1.50 while reducing output from $9.00 to $7.50 per unit.
Lower API Cost
Gemini 3.6 Flash keeps the same input pricing while reducing output costs.
Model | Input | Output |
|---|---|---|
Gemini 3.5 Flash | $1.50 | $9.00 |
Gemini 3.6 Flash | $1.50 | $7.50 |
This represents approximately a 16.7% reduction in output cost.
Coding Performance
Gemini 3.6 Flash outperforms its predecessor across multiple software engineering benchmarks.
Benchmark | Gemini 3.5 Flash | Gemini 3.6 Flash |
|---|---|---|
SWE-Bench Pro | 55.1% | 58.7% |
DeepSWE v1.1 | 37% | 49% |
Terminal-Bench 2.1 | 76.2% | 78.0% |
MLE-Bench | 49.7% | 63.9% |
The biggest gains are in DeepSWE (+12 points) and MLE-Bench (+14.2 points), highlighting improvements in long-horizon software engineering and machine learning engineering tasks.
Agent & Tool Calling
On the OSWorld-Verified benchmark, Gemini 3.6 Flash improves from 78.4% to 83.0%.
These gains benefit AI agents using Function Calling, MCP, API integrations, and autonomous workflows. For a broader comparison of AI models, see our AI models comparison.
Reasoning
The model also improves its reasoning capabilities.
Without tools: 84.2 → 85.2
With tools: 84.9 → 89.4
This makes Gemini 3.6 Flash more reliable in tool-assisted reasoning tasks.
Long Context Performance
One of the most significant improvements comes in long-context processing.
Benchmark | Gemini 3.5 Flash | Gemini 3.6 Flash |
|---|---|---|
GDM-MRCR v2 (128K) | 77.3 | 91.8 |
GDM-MRCR v2 (1M) | 26.6 | 54.0 |
The near doubling of performance on the 1M token benchmark makes the model more suitable for large-scale document analysis, repository understanding, and RAG systems. For more on AI model developments, check out Kimi K3: Moonshot AI's Open Frontier Model.
What is Gemini 3.6 Flash?
Gemini 3.6 Flash is Google's latest optimization release of its Flash model line, designed for production AI workloads. It focuses on reducing costs and improving performance in coding, reasoning, agent tasks, and long-context processing without introducing new architectural capabilities.
Key Improvements
16.7% lower output cost
Better software engineering benchmarks
Improved agent execution and tool calling
Stronger reasoning performance
Major gains in long-context understanding
For a deeper dive into AI model comparisons, see our AI models compared article.
