Skip to content
news12 min read

Google Shipped Three New Gemini Models — but Not the Flagship 3.5 Pro

On July 21, 2026, Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and a government-only 3.5 Flash Cyber — three Flash-tier models — while the promised Gemini 3.5 Pro flagship stayed in partner testing. Google also confirmed it has begun pre-training Gemini 4. The Flash tier is becoming the product; the Pro tier is still a promise.

Author
Anthony M.
12 min readVerified July 23, 2026Tested hands-on
Google's July 21, 2026 Gemini launch: three Flash-tier models shipped and the 3.5 Pro slot still empty
Google shipped three Flash-tier Gemini models on July 21, 2026 — and left the flagship 3.5 Pro slot empty while Gemini 4 pre-training begins.

On July 21, 2026, Google released three new Gemini models at once — Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, both generally available, plus Gemini 3.5 Flash Cyber, a security model restricted to governments and trusted partners. All three are additions to the Flash tier. The flagship Gemini 3.5 Pro, promised earlier in the year, did not ship: Google says it is "currently testing with partners." Gemini 3.6 Flash lists at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, down from $9.00 output on 3.5 Flash, with a March 2026 knowledge cutoff. In the same announcement, Google confirmed it has begun pre-training Gemini 4.

Key Takeaways

  • Three models, all Flash. Gemini 3.6 Flash (generally available), 3.5 Flash-Lite (generally available), and 3.5 Flash Cyber (government-only pilot) all landed on July 21, 2026. Not one of them is the long-awaited 3.5 Pro flagship.
  • Cheaper output, fresher brain. Gemini 3.6 Flash drops output pricing from $9.00 to $7.50 per 1 million tokens while holding input at $1.50, and its knowledge cutoff advances from January 2025 to March 2026.
  • Every benchmark here is Google's own. The launch numbers — DeepSWE 49 percent, MLE-Bench 63.9 percent, OSWorld-Verified 83 percent — are self-reported by Google. No independent evaluation of 3.6 Flash had been published at launch.
  • The flagship is still a promise. Google's only public word on Gemini 3.5 Pro is that it is "currently testing with partners." The model was expected months ago; the Flash tier keeps shipping in its place.
  • Gemini 4 is already training. Google says it has "started our most ambitious pre-training run yet, for Gemini 4" — a sign the company may be looking past 3.5 Pro toward the next full generation.

What Google Actually Shipped on July 21

Google released three Gemini models in a single announcement. Two are generally available to developers today: Gemini 3.6 Flash (API id gemini-3.6-flash) and Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite). Both carry a 1-million-token context window and a 64,000-token maximum output, and both are marked stable in Google's model documentation. The third, Gemini 3.5 Flash Cyber, is a specialized security model that is not for sale to the public — more on that below.

The distribution was wide from day one. Google says 3.6 Flash and 3.5 Flash-Lite are live across Google AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, the Gemini Enterprise app, the Gemini app, and — for Flash-Lite — Google Search. Separately, GitHub confirmed on the same day that Gemini 3.6 Flash is available inside GitHub Copilot for Pro, Pro+, Max, Business, and Enterprise plans, with Business and Enterprise administrators needing to switch on a "Gemini 3.6 Flash Preview" policy first. For context on where this tier sits, we have covered the current lineup in our reviews of Gemini 3.5 Flash and Gemini 3 Flash.

The through-line is unmistakable. Three launches in one day, and all three live in the Flash family — the fast, cheap, high-volume tier — while the flagship reasoning model stays offstage.

The three Gemini models Google launched July 21 2026: 3.6 Flash pricing and cutoff, Flash-Lite pricing and throughput, Flash Cyber government-only
The July 21 lineup at a glance: two generally available Flash models and one government-restricted security model.

Gemini 3.6 Flash: Cheaper Output, Fewer Tokens, a 2026 Brain

Gemini 3.6 Flash is positioned as Google's new workhorse — the default model most developers will reach for. According to Google's own pricing page, it costs $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, with output billing including thinking tokens. The output rate is the headline: 3.5 Flash charged $9.00 per 1 million output tokens, so 3.6 Flash cuts that by roughly 17 percent while keeping the input price flat. Google's API changelog frames the model as offering "improved token efficiency and code/agentic planning capabilities at a lower price point than 3.5 Flash."

The other quiet upgrade is freshness. As 9to5Google reported, the knowledge cutoff "finally advances from January 2025 to March 2026" — closing a gap that had made the earlier Flash models feel dated for anything touching recent events or libraries.

A note on the benchmarks, because it matters. Every performance figure Google published for 3.6 Flash is self-reported and measured against its own 3.5 Flash, not against rival models and not by an outside lab. With that label firmly attached: Google reports 3.6 Flash scoring 49 percent on DeepSWE (a coding benchmark by Datacurve) versus 37 percent for 3.5 Flash, 63.9 percent on MLE-Bench versus 49.7 percent, and 83 percent on OSWorld-Verified versus 78.4 percent. Google also says the model uses 17 percent fewer output tokens than 3.5 Flash — a reduction it attributes to the Artificial Analysis Index — and up to 65 percent fewer on the DeepSWE test specifically. These are real, plausible gains, but they are vendor numbers. At launch, no independent Artificial Analysis Intelligence Index score for 3.6 Flash had been published, so there is no neutral read to set beside them yet.

Flash-Lite and the Price Floor

Gemini 3.5 Flash-Lite is the cheapest of the pair and the clearest signal of where the volume war is heading. Google prices it at $0.30 per 1 million input tokens and $2.50 per 1 million output tokens, and describes it in the changelog as "a low-latency, highly cost-effective subagent option." Google also cites the Artificial Analysis Index for a throughput figure of roughly 350 output tokens per second — fast enough to make Flash-Lite the model you fan out across dozens of parallel sub-agent calls without watching the meter.

This is a continuation, not a surprise. We wrote in March 2026 about how Google's Flash-Lite pricing lit the sub-$0.25 LLM price war on fire, and the 3.5 generation keeps the pressure on. The strategy is coherent: make the bottom of the lineup so cheap and so fast that high-throughput agent workloads default to Gemini, then upsell reasoning where it is actually needed. The catch is that "where it is actually needed" is exactly the job of the Pro tier — which is the tier that did not ship.

Flash Cyber: A Vulnerability Hunter Only Governments Get

The most unusual release of the three is Gemini 3.5 Flash Cyber. Per Google DeepMind, it is a lightweight model "fine-tuned to find, validate, and patch vulnerabilities quickly and efficiently." Rather than sell it, Google is gating it: as part of what DeepMind calls a "limited-access pilot program," 3.5 Flash Cyber "will be exclusively available to governments and trusted partners via CodeMender soon, expanding over time." CodeMender is Google's autonomous code-security agent, and readers who followed our coverage of Google's AI Threat Defense platform will recognize the plumbing.

Google's headline evidence is an internal test on the V8 JavaScript engine that powers Chrome. In that evaluation, DeepMind reports, "3.5 Flash Cyber found 55 unique confirmed issues, compared to 47 found by mainline 3.5 Flash and 36 found by Opus 4.6, including 10 issues that the other two models tested did not catch." Two honesty flags belong on that sentence. First, it is Google's own internal benchmark, not a third-party audit. Second, the model it beats by the widest margin — Anthropic's Claude Opus 4.6 — is a prior-generation model; Anthropic's current flagship is Opus 4.8, which was not in the comparison. A vendor benchmark that outscores a competitor's older release is a weaker claim than the raw numbers suggest.

The dual-use framing is the interesting part. "Given the dual-use nature of this technology, we have taken an intentional approach to how we deploy 3.5 Flash Cyber," DeepMind writes — an acknowledgment that a model tuned to discover exploits is as useful to attackers as defenders. Restricting it to governments and vetted partners is Google's answer to that. Meanwhile, Google says CodeMender's foundational capabilities are also coming to ordinary customers through generally available Gemini models on the Gemini Enterprise Agent Platform, so the defensive tooling reaches the market even if the specialized model does not.

Google-reported internal V8 vulnerability test: Flash Cyber 55, mainline 3.5 Flash 47, prior-generation model 36
Google's internal, self-reported V8 test. The 36-vulnerability bar is a prior-generation competitor model, not a current flagship.

The Flagship That Isn't There: Where Is Gemini 3.5 Pro?

The most conspicuous thing about the July 21 launch is what was absent. Google introduced Gemini 3.5 Flash at I/O in May — we covered that launch and the Anthropic-procurement anxiety it caused, and later how 3.5 Flash gained native computer use. Throughout, the 3.5 Pro model was the headliner-in-waiting. Two months later, it is still waiting.

Google's only official statement on it is a single sentence: "Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it's ready." That is the whole disclosure — no date, no benchmark, no preview. We dug into the reasons the flagship keeps slipping in a separate piece on why Gemini 3.5 Pro is late and what the delay says about Google, so we will not re-litigate the causes here. The point for this launch is narrower: the Pro delay is now long enough that Google has shipped a whole new Flash generation — 3.6 — before the previous generation's Pro model appeared at all.

That inversion is the story. In a normal release cadence, the flagship leads and the cheaper variants follow. Google has run it backward: Flash 3.5, Flash-Lite, Flash Cyber, and now Flash 3.6 have all shipped while Pro 3.5 sits in private testing. The Flash tier has stopped being the follow-up act and become the product.

Gemini 4 Is Already Pre-Training

Then came the line that reframes everything. In the same announcement, Google wrote: "We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress." Pre-training is the earliest and most expensive phase of building a model — the months-long run that produces the base weights before any fine-tuning. Announcing it publicly is a deliberate signal.

It also raises an uncomfortable question about 3.5 Pro. If Gemini 4 is already in its most ambitious pre-training run, how much strategic attention is left for a 3.5 Pro that has been stuck in partner testing for months? One reading is benign: 3.5 Pro is nearly done and Google is simply telegraphing the roadmap. A less comfortable reading is that Google may be tempted to let the next full generation absorb the flagship slot, quietly folding 3.5 Pro into a footnote. Google has not said which it is — and the absence of any 3.5 Pro timeline, set against a confident Gemini 4 tease, is exactly what makes the second reading plausible.

A Developer Footnote: temperature, top_p and top_k Are Deprecated

Buried in the same changelog is a change worth flagging for anyone building on the API. Google states plainly: "The sampling parameters temperature, top_p and top_k are now deprecated." If your prompts or SDK calls set these on the newest Gemini models, expect them to stop mattering — or to break — over time. This is not a Google quirk; it is an industry pattern. The top-tier 2026 models from other labs already reject custom temperature outright, and Google is now following the same path of taking sampling control away from the caller. If you maintain a fallback chain across providers, this is the kind of change that silently poisons every request in the chain, so it is worth auditing now rather than after an outage.

Our Take: The Flash Tier Is the Product Now

We have followed Google's Gemini cadence closely all year, and July 21 clarified a shift that had been building for months: for Google, Flash is no longer the budget option beneath the real model. Flash is the real model. The company is shipping speed, price, and breadth — a cheaper 3.6 Flash, an ultra-cheap Flash-Lite, a specialized Flash Cyber, and day-one distribution everywhere from Android Studio to GitHub Copilot — while the reasoning flagship it promised in the spring stays in a testing loop with no public finish line.

There is a coherent strategy in that, and there is a risk. The strategy: own the high-volume agentic workloads that will define AI spending, where cheap-and-fast beats slow-and-brilliant, and let Gemini 4 leapfrog the whole 3.5 Pro question. The risk: developers who need a top-tier reasoning model today are not waiting politely — they are reaching for the models that actually shipped, from Anthropic and others, and procurement decisions made in a vacuum tend to stick. Shipping three Flash models is impressive volume. It is not a substitute for the flagship that keeps not arriving.

What would change our read? A dated 3.5 Pro release, or an independent Artificial Analysis score that shows 3.6 Flash closing the reasoning gap with rival flagships rather than just beating its own predecessor. Until one of those lands, the honest description of Google's position is the one the launch itself wrote: the Flash tier is the product, and the Pro tier is still a promise. For the record, we have no affiliate relationship with Google, Anthropic, or any model provider named here; this is editorial analysis, not a recommendation to buy.

Frequently Asked Questions

What did Google launch on July 21, 2026?

Google released three Gemini models on July 21, 2026: Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, both generally available to developers, and Gemini 3.5 Flash Cyber, a security-focused model restricted to governments and trusted partners. All three belong to the Flash tier. The flagship Gemini 3.5 Pro was not part of the launch and remains in partner testing.

How much does Gemini 3.6 Flash cost?

Gemini 3.6 Flash costs $1.50 per 1 million input tokens and $7.50 per 1 million output tokens on Google's standard paid tier, with output billing including thinking tokens. That is a drop from the $9.00 per 1 million output tokens that 3.5 Flash charged, while the input price stays at $1.50 per 1 million tokens.

Is Gemini 3.6 Flash cheaper than Gemini 3.5 Flash?

Yes, on output. Gemini 3.6 Flash keeps input pricing at $1.50 per 1 million tokens but cuts output from $9.00 to $7.50 per 1 million tokens, roughly a 17 percent reduction. Google also says the model uses about 17 percent fewer output tokens than 3.5 Flash on average, which compounds the savings on real workloads.

What is the context window of Gemini 3.6 Flash?

Gemini 3.6 Flash has a 1-million-token context window and a maximum output of 64,000 tokens, according to Google's model documentation. Gemini 3.5 Flash-Lite carries the same 1-million-token window and 64,000-token output ceiling. The knowledge cutoff for 3.6 Flash advances to March 2026, up from January 2025 on the earlier Flash models.

What is Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite is Google's fastest and cheapest 3.5 model, priced at $0.30 per 1 million input tokens and $2.50 per 1 million output tokens. Google describes it as a low-latency, cost-effective subagent option and cites the Artificial Analysis Index for a throughput of roughly 350 output tokens per second, aimed at high-volume, parallel agent workloads.

What is Gemini 3.5 Flash Cyber and who can use it?

Gemini 3.5 Flash Cyber is a lightweight Gemini model fine-tuned to find, validate, and patch software vulnerabilities. Google is not selling it to the public: it is a limited-access pilot available exclusively to governments and trusted partners through CodeMender, Google's autonomous code-security agent. Google cites the dual-use nature of vulnerability-hunting AI as the reason for the restricted distribution.

Are Gemini 3.6 Flash's benchmark scores independently verified?

No. Every benchmark Google published for Gemini 3.6 Flash — including DeepSWE at 49 percent, MLE-Bench at 63.9 percent, and OSWorld-Verified at 83 percent — is self-reported by Google and measured against its own 3.5 Flash. At launch, no independent evaluation such as an Artificial Analysis Intelligence Index score had been published for 3.6 Flash, so there is no neutral third-party read to compare against yet.

How did Flash Cyber compare to Claude Opus 4.6 in Google's test?

In Google's internal test on the V8 JavaScript engine, DeepMind reports that 3.5 Flash Cyber found 55 unique confirmed vulnerabilities, versus 47 for mainline 3.5 Flash and 36 for Claude Opus 4.6, including 10 issues the other two models missed. Two caveats apply: the benchmark is Google's own, and Opus 4.6 is a prior-generation model — Anthropic's current flagship, Opus 4.8, was not part of the comparison.

Why is Gemini 3.5 Pro still not available?

Google has not given a detailed reason. Its only public statement is that "Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it's ready," with no date, benchmark, or preview attached. The flagship was expected earlier in 2026, and the delay has now stretched long enough that Google shipped a newer Flash generation, 3.6, before 3.5 Pro appeared.

Is Google skipping Gemini 3.5 Pro in favor of Gemini 4?

Google has not said so, but the possibility is real. In the same July 21 announcement, Google confirmed it has "started our most ambitious pre-training run yet, for Gemini 4." With 3.5 Pro stuck in partner testing and Gemini 4 already in pre-training, one plausible reading is that the next full generation could absorb the flagship slot. Google has offered no timeline that would rule this out.

Is Gemini 3.6 Flash available in GitHub Copilot?

Yes. GitHub confirmed on July 21, 2026 that Gemini 3.6 Flash is available in GitHub Copilot for Pro, Pro+, Max, Business, and Enterprise plans, on a gradual rollout. Business and Enterprise administrators must enable a "Gemini 3.6 Flash Preview" policy in Copilot settings before their organizations can select it, and usage is billed at provider list pricing.

Does Gemini 3.6 Flash still support the temperature parameter?

Not going forward. Google's API changelog states that "the sampling parameters temperature, top_p and top_k are now deprecated" on the newest Gemini models. Developers who set these parameters should expect them to stop having an effect over time. It mirrors a broader 2026 trend, with other top-tier models already rejecting custom temperature values outright.

Sources

Related Articles

Was this review helpful?
Anthony M. — Founder & Lead Reviewer
Anthony M.Verified Builder

We're developers and SaaS builders who use these tools daily in production. Every review comes from hands-on experience building real products — DealPropFirm, ThePlanetIndicator, PropFirmsCodes, and many more. We don't just review tools — we build and ship with them every day.

Written and tested by developers who build with these tools daily.