Quick Facts
- Gemini 3.7 Flash launches at $0.75 per million input tokens and $3.75 per million output tokens through Dec. 31, 2026, half the standard rate.
- The model scores 43.6% on FrontierCode 1.1 Main, up from 34.4% for Gemini 3.6 Flash, and ranks first among 186 models on output speed at 340.1 tokens per second.
- Introductory pricing expires Jan. 1, 2027, when rates double to $1.50 per million input tokens and $7.50 per million output tokens.
Google released Gemini 3.7 Flash on Aug. 13, 2026, just three weeks after Gemini 3.6 Flash shipped. The fast turnaround was not planned. Developer backlash over frontend code regressions and weak benchmark scores pushed Google to accelerate the release.
Google credits algorithmic changes to the reasoning core, not a new pretraining run, for the improvements. Tulsee Doshi, senior director of product management at Google, said the release builds on developer feedback and introduces innovations the company plans to carry into future models.
Pricing Window
Through Dec. 31, Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens. On Jan. 1, 2027, those rates double. Google is applying the same introductory pricing to Gemini 3.6 Flash as well.
For teams running agentic workflows that consume between one million and three million tokens per task, the deadline is a real decision point. A task consuming two million output tokens costs $7.50 at current rates. After Jan. 1, the same task costs $15.00.
Benchmark Results
On FrontierCode 1.1 Main, Gemini 3.7 Flash scored 43.6%, ahead of Claude Sonnet 5 at 42.7% and GPT-5.6 Terra at 41.3%. On DeepSWE v1.1, a long-horizon software engineering test, it reached 65.3%, up from 49.0% for 3.6 Flash, though GPT-5.6 Terra leads that benchmark at 69.6%.
On AutomationBench, a test published by Zapier measuring enterprise workflow automation, Gemini 3.7 Flash scored 30.4%. That compares with 10.7% for Claude Sonnet 5 and 23.6% for GPT-5.6 Terra. The previous version scored 17.0%.
On GDP.PDF, which tests complex document comprehension, the model scored 34.0%, higher than Claude Sonnet 5 at 28.0% and GPT-5.6 Terra at 24.7%. On Arena.ai’s WebDev Arena, it posted an Elo score of 1588, up from 1538 for Gemini 3.6 Flash.
Speed and Specifications
Artificial Analysis ranks Gemini 3.7 Flash first among 186 models on output speed at 340.1 tokens per second, scoring it 56 against 52 for 3.6 Flash. The model supports a one-million token context window, 64,000 max output tokens, and tunable thinking levels set to low, medium, or high.
Google’s official documentation states the model adapts better to roadblocks, follows instructions with greater accuracy, and requires less manual oversight across multi-step engineering workflows than its predecessor.
What This Means for Engineering Teams
The combination of lower inference costs and stronger coding scores makes Gemini 3.7 Flash a credible option for teams evaluating production deployments before year end. The Dec. 31 pricing deadline gives engineering leaders roughly four months to test and migrate before costs reset.
Teams running high-volume agentic workflows, where tool calls average 30 to 50 per session, face the sharpest cost exposure after Jan. 1. The pricing structure rewards early adoption but does not lock in the introductory rate beyond the calendar year.
Read more: Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
