Google releases Gemini 3.7 Flash with major coding gains weeks after 3.6
The new coding-centric model boosts software engineering benchmark scores and introduces adjustable reasoning depth.
Google has released Gemini 3.7 Flash, a specialized model focused on coding, arriving just three weeks after the launch of Gemini 3.6 Flash. The rapid release underscores Google's push to dominate the agentic AI space by prioritizing leaner, high-performance models for software engineering.
The new model demonstrates a substantial leap in technical capabilities. According to data from Ars Technica and MarkTechPost, Gemini 3.7 Flash scored 65.3% on the DeepSWE v1.1 benchmark, a significant increase from the 49% recorded by Gemini 3.6 Flash. Similarly, the model achieved 43.6% on the FrontierCode 1.1 Main benchmark, surpassing the 34.4% score of its predecessor.
Balancing Speed and Reasoning
To give developers more control over performance, Google introduced "thinking_level" settings. These allow users to toggle between low, medium, and high settings, enabling a direct balance between latency and the depth of the model's reasoning. This feature is designed for "long-horizon" software engineering, where AI agents must maintain state and recover from errors across extended, multi-step workflows.
Market Strategy and Pricing
Google is pairing the technical launch with aggressive introductory pricing to accelerate developer adoption. According to the Google Antigravity blog, the model is currently priced at $0.75 per 1 million input tokens and $3.75 per 1 million output tokens. This pricing is temporary; the rates are scheduled to double on January 1, 2027, rising to $1.50 for input and $7.50 for output tokens.
Industry Implications
The jump in DeepSWE scores suggests a meaningful improvement in the ability of AI agents to handle complex software tasks without stalling. By releasing specialized "Flash" models in such quick succession, Google is signaling a shift away from relying solely on massive flagship models in favor of agile, task-specific versions that can be deployed more efficiently in production environments.
What to Watch
While the performance gains are concrete, the volatility of the release schedule has drawn attention. The New Stack noted that the model developers actually need does not always arrive on the promised timeline, pointing to a high-pressure race to capture the developer market. Observers will be watching to see if this accelerated cadence continues and how the model performs in real-world enterprise repositories beyond synthetic benchmarks.