TechNewsReel
Live

Google Ships Three Gemini Flash Models as 3.5 Pro Remains Delayed

DeepMind releases efficiency-focused variants while its flagship model misses June deadline with no new launch date.

TechNewsReel Newsroom · July 25, 2026

Flash Models Arrive, Pro Waits

Google DeepMind released three AI models on July 21, 2026: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The launches prioritize efficiency and specialized applications while the anticipated flagship, Gemini 3.5 Pro, remains unavailable past its promised June release window.

Efficiency Gains and Pricing

Gemini 3.6 Flash delivers measurable improvements over its predecessor. The model reduces output token usage by 17% compared to 3.5 Flash on the Artificial Analysis Index, with reductions reaching up to 65% on certain DeepSWE tasks. It scores 49% on the DeepSWE benchmark, up from 37% for 3.5 Flash. Output pricing is set at $7.50 per million tokens.

Cybersecurity Model Restricted

Gemini 3.5 Flash Cyber targets vulnerability detection and security analysis. Access is limited to governments and trusted partners through a restricted pilot program. The specialized release signals Google's push into enterprise and institutional markets where security concerns dominate purchasing decisions.

Pro Delay Continues

Gemini 3.5 Pro was announced at Google I/O on May 19, 2026, with a target general availability date of June 2026. That deadline passed without a public release. Logan Kilpatrick, Google DeepMind product lead, stated the company is testing Gemini 3.5 Pro with partners and hopes to "land soon."

When it arrives, Gemini 3.5 Pro will feature a 2M token context window and a "Deep Think" reasoning mode. The Deep Think capability is gated to Ultra subscribers at $250 per month. Standard API pricing is approximately $15 per million input tokens and $60 per million output tokens.

Strategic Positioning

The staggered release strategy reveals Google's approach to the AI race. By shipping Flash variants first, the company addresses developer needs for cost-effective, high-throughput inference while continuing to refine its flagship reasoning model. Reports have suggested the Gemini launch faced delays as the technology fell short of internal goals, though Google has not publicly detailed specific performance gaps.

The Flash lineup competes in a crowded efficiency segment where Anthropic and OpenAI have established positions. Google's ability to convert token efficiency into market share will depend on whether developers prioritize cost savings over the advanced reasoning capabilities that the delayed Pro model promises to deliver.

Sources

Get a notification when a big story breaks. A few a day at most — no spam.