When Google DeepMind announced three new Gemini models on Tuesday, the obvious question wasnt about what they could do. It was about what wasnt in the room.
Gemini 3.5 Pro, the flagship model Google promised would arrive in June, still hasnt shipped. The company says it is testing with partners and will arrive as soon as it is ready. Reports from last week suggested 3.5 Pro missed internal coding targets. That is an awkward place to be when Anthropics Claude Opus 4.8 and Sonnet 5 are shipping on schedule and OpenAI is rolling out GPT-5.6.

Meanwhile, the models Google did release tell a more interesting story than they look at first glance.
The star of the show is Gemini 3.6 Flash, which replaces the now-deprecated 3.5 Flash that was only announced at I/O two months ago. The update was driven by developer feedback. Users found 3.5 Flash good but not great at coding, and the pricing was a concern for teams building agentic workflows at scale. Google responded by dropping output token costs by 17 percent (from $9 to $7.50 per million output tokens) while keeping input pricing flat at $1.50. The model itself takes 17 percent fewer tokens to complete the same tasks.
The numbers on DeepSWE, a coding benchmark that measures real repository-level problem solving, jump from 37 percent on 3.5 Flash to 49 percent on 3.6 Flash. That is a real improvement and puts it ahead of many models from six months ago. The computer use benchmark OSWorld shows a more modest bump from 78.4 to 83 percent. More importantly, computer use is now a standard API feature rather than an experimental add-on. This matters for teams building agents that need to interact with desktop applications.

Then there is Gemini 3.5 Flash-Lite, which at 350 tokens per second is Googles fastest model. Priced at $0.30 per million input tokens and $2.50 for output, it is cheap enough to deploy at scale in agentic systems where every millisecond of latency adds up. Googles AI Overview product in search is expected to adopt this model. When you are serving billions of queries, a few hundred milliseconds per query translates into real money.
The third model, 3.5 Flash Cyber, is the most unusual. It is a specialized variant fine-tuned for finding and fixing security vulnerabilities. Google considers it dangerous enough to restrict to governments and trusted partners on a limited pilot basis. The approach mirrors what Anthropic did with its Mythos model: acknowledge the dual-use problem and gate access rather than releasing openly. Independent benchmarks suggest Flash Cyber approaches Mythos-level performance on common vulnerability classes while running at a fraction of the inference cost.
The HN thread on the announcement is characteristically mixed. Some users appreciate the lower pricing and faster speeds. Others are frustrated with Googles product fragmentation. MCP support is locked behind Google Spark rather than available in the core chat product, and enterprise Workspace accounts often get fewer features than free consumer accounts. One developer who spent over $4,000 on the Gemini API said they gave up and moved to Anthropic. Another described Googles approach as shipping their org chart. Anyone who has tried to navigate Googles consumer-versus-enterprise product maze will recognize the complaint.
The deeper story here isnt about benchmark scores or token pricing. Its about a company whose workhorse models keep getting better at exactly the moment its flagship is stuck. Gemini 3.5 Pro was supposed to be Googles answer to Claude Opus 4.8 and GPT-5.6. But those models are already shipping, and Gemini 4 pre-training has just begun. Google is running two races at once: one with Flash models that are genuinely improving, and one with a Pro model that cant seem to cross the finish line.
Sources
- TechCrunch: Google releases three new Gemini models: detailed breakdown of the three-model launch and Pro delay
- Ars Technica: Google Gemini 3.6 Flash and cybersecurity AI: benchmark numbers, pricing, and Googles AI roadmap
- 9to5Google: Gemini 3.6 Flash and 3.5 Flash-Lite launch: images, benchmark charts, and Gemini 4 teaser
- Hacker News discussion (232 points, 165 comments): community reactions including product fragmentation complaints
- Reddit r/GeminiAI: 3.6 Flash launch discussion: community analysis and comparisons
- Bloomberg: Google Gemini launch delayed: original report on Pros internal struggles