Skip to content

Claude Opus 5 vs Gemini 3.1 Pro Preview: Fifteen Points and a Preview Label (2026)

Claude Opus 5 scores 61 to Gemini 3.1 Pro Preview’s 46. Gemini is cheaper and more accurate, but 159 days into preview with no SLA. July 28, 2026.

Claude Opus 5 scoring 61 at max effort for $2.03 per task beside Gemini 3.1 Pro Preview scoring 46 at default configuration for $0.29 per task, labeled public preview
Fifteen index points separate them, but the label under Google's model is the fact that changes a procurement decision.

Feature Comparison

FeatureClaude Opus 5Gemini 3.1 Pro Preview
Artificial Analysis Intelligence Index v4.1, top measured configuration61 (max effort)46 (default configuration, level not disclosed)
Index at the cheapest measured configuration51 (low effort), $0.36 per task46, $0.29 per task — only row published
Cost per task, measured July 28, 2026$0.36 at low effort, $2.03 at max$0.29
Release stageProduction model, no preview label"Public preview" since February 19, 2026
Long-context pricing above 200,000 input tokensFlat, no surcharge at any lengthAll tokens repriced: input $2.00 to $4.00, output $12.00 to $18.00
Input price per million tokens, standard tier$5.00$2.00 up to 200k, $4.00 above
AA-Omniscience factual reliability31 (max effort)33
Output throughputAbout 55 tokens per second at max effortAbout 129 tokens per second
Time to first token3.32 seconds at low effort42.32 seconds
Reasoning ladderFive levels via output_config.effortThree levels via thinking_level, no minimal
Input modalitiesText, image, PDF and plain-text documentsText, image, video, audio, PDF
Maximum output tokens128,000 synchronously65,536
Knowledge cutoffMay 2026January 2025

Pricing Comparison

Claude Opus 5

$5 in / $25 out per M tokens
paid

Gemini 3.1 Pro Preview

$2 in / $12 out per M tokens
Free trial available
paid

Detailed Comparison

Claude Opus 5 scores 61 on the Artificial Analysis Intelligence Index v4.1 at max effort against 46 for Gemini 3.1 Pro Preview, a gap of fifteen points. The label matters more than the gap: Google's flagship reasoning model has carried the launch stage "Public preview" since February 19, 2026, and Google's contract terms exclude Pre-GA offerings from any SLA and reserve the right to discontinue them without prior notice, while Claude Opus 5 shipped as an ordinary production model on July 24, 2026. Gemini is the cheaper and faster model, at $0.29 per task against $0.36 for Opus 5's lowest rung and about 129 output tokens per second against 55, and it beats Opus 5 on factual reliability by 33 to 31 on AA-Omniscience. We researched both from vendor documentation and independent evaluations; index figures are v4.1 and sliding measurements were read on July 28, 2026.

Which model should you choose?

Pick Claude Opus 5 if you need capability above index 46, a knowledge cutoff from 2026 rather than January 2025, flat pricing across a one-million-token window, or a model covered by an ordinary support commitment. Pick Gemini 3.1 Pro Preview if you need audio or video input, roughly 2.3 times the throughput, a lower bill per task, or measurably better factual recall.

Claude Opus 5 wins this comparison, but not the way the headline numbers suggest and not without real concessions. Fifteen index points is the widest measured gap on the page, and Opus 5's cheapest configuration still scores five points above the only score Artificial Analysis publishes for Gemini 3.1 Pro Preview. That part is a clean win.

The part that is not clean: Gemini is cheaper per task, roughly twice as fast, and better at not making things up. Those are not consolation categories. A model that costs $0.29 per task against $0.36, streams at about 129 output tokens per second against 55, and scores 33 against 31 on a benchmark that explicitly penalizes hallucination is winning on three axes that many teams weight above raw reasoning power.

What tips the decision is the third column, the one neither vendor prints as a number. Gemini 3.1 Pro Preview is a pre-general-availability offering. Google's Service Specific Terms state that Pre-GA Offerings "are not covered by any SLA or Google indemnity" and "may be changed, suspended or discontinued at any time without prior notice to Customer." Google has already exercised that on this exact product line: the previous Gemini 3 Pro Preview was shut down on March 9, 2026, 111 days after its release. Anyone building on the current one is building under the same terms.

Is Gemini 3.1 Pro Preview still in preview?

Yes, as of July 28, 2026. Google's Gemini Enterprise Agent Platform documentation lists gemini-3.1-pro-preview with "Launch stage: Public preview" and a release date of February 19, 2026 — five months and nine days in preview. On the Gemini API models page the model sits under the heading Preview, while Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.1 Flash-Lite all sit under Stable.

That split inside Google's own catalog is the fact worth pausing on. Google is perfectly capable of promoting Gemini models to general availability, and it did so seven days before this page was written. The changelog entry for July 21, 2026 reads: "Gemini 3.6 Flash and Gemini 3.5 Flash-Lite generally available (GA): Released stable, production-ready versions of our latest 3.x Flash models." The Flash tier has now gone GA four separate times in 2026. The reasoning Pro tier has not gone GA once.

One precision matters here, because a careless version of this claim is easy to refute. A Google model with "Pro" in its name is generally available: gemini-3-pro-image, announced on May 28, 2026 as one of "the generally available (GA) versions of our native visual models." That is an image generation model. The claim on this page is narrower and holds: Google ships no generally available Gemini 3.x reasoning Pro model. Every occurrence of gemini-3.1-pro in Google's model catalog is followed by -preview.

Google has published no general availability date for this model, and that is an absence we checked rather than assumed. The Gemini API deprecations table lists gemini-3.1-pro-preview with a release date of February 19, 2026 and, in the shutdown column, "No shutdown date announced." There is no published GA target, no stable-version date, and no end-of-support date anywhere in Google's documentation for it. The model has now been in preview for 159 days.

We covered the shape of this problem when it first became visible, in our reporting on why the promised Gemini 3.5 Pro is late and on the three Gemini models Google shipped without the flagship. The comparison here is the downstream consequence: buyers evaluating Google's best reasoning model in July 2026 are evaluating a preview.

What does preview status actually cost a buyer?

Three concrete things: no service level agreement, no indemnity, and a shorter and less certain support runway. Google's Service Specific Terms classify anything labeled "Preview" as a Pre-GA Offering, and state that Pre-GA Offerings "are provided 'as is' without any express or implied warranties or representations of any kind," that they "may be changed, suspended or discontinued at any time without prior notice to Customer," and that they "are not covered by any SLA or Google indemnity." A separate clause caps Google's liability for Pre-GA Offerings at $25,000.

Set that against what Google's developer documentation promises, because the two do not say the same thing. The Gemini API models page states that preview models "will be deprecated with at least 2 weeks notice." The contract reserves the right to discontinue with no notice at all. The two weeks is a documentation courtesy, not a contractual guarantee, and a procurement review will read the contract.

One widespread misreading needs correcting before anyone repeats it, because it is wrong and Google would rightly push back. The generic Pre-GA terms say customers "should not use Pre-GA Offerings to process personal data," but they also allow Google to state otherwise in its documentation, and for this model Google does. The Generative AI Preview banner on the Gemini 3.1 Pro page says that "Customers may elect to use it for production or commercial purposes, or disclose Generated Output to third-parties, and may process personal data as outlined in the Cloud Data Processing Addendum." Production use and personal data are both explicitly permitted. What the banner does not restore is the SLA or the indemnity.

The risk is therefore not a licensing prohibition, it is continuity — and Google's own deprecations table quantifies it. Preview models on this API have been given runways of roughly three to four months, while generally available models get a year or more. That table is published by Google and needs no interpretation.

Gemini modelStageRelease dateShutdown datePublished runway
gemini-3-pro-previewPreviewNovember 18, 2025March 9, 2026111 days
gemini-3.1-flash-lite-previewPreviewMarch 3, 2026May 25, 202683 days
gemini-3.1-flash-image-previewPreviewFebruary 26, 2026June 25, 2026119 days
gemini-3.1-flash-liteGenerally availableMay 7, 2026May 7, 2027365 days
gemini-2.5-proStableJune 17, 2025October 16, 2026about 16 months
gemini-3.1-pro-previewPreviewFebruary 19, 2026none announcednot published

The precedent is worth reading closely, because the failure mode was quieter than a shutdown usually is. Google announced the retirement of Gemini 3 Pro Preview on February 26, 2026 and shut it down on March 9, eleven days later — shorter than the two weeks its own documentation describes, though we cannot rule out an earlier private notice to customers. The changelog then records: "The Gemini 3 Pro Preview model has been shut down. The gemini-3-pro-preview now points to gemini-3.1-pro-preview." The identifier kept resolving. Calls kept returning results. The model underneath had changed.

Then there is the fact that reframes the whole question, and it comes from the same Google table. Gemini 2.5 Pro, the only Pro-tier model Google currently lists as stable, has a published shutdown date of October 16, 2026. The recommended replacement Google names for it is gemini-3.1-pro-preview. A buyer following Google's own migration guidance moves from an SLA-covered stable model to a preview offering that is contractually "as is." At the time of writing there is no generally available successor in the Pro tier to move to instead.

Rate limits carry the same caveat. Google's rate-limits page declines to publish requests-per-minute or tokens-per-minute figures for this model, deferring to the console, and states plainly that "rate limits are more restricted for experimental and preview models." The only hard quota Google does publish for gemini-3.1-pro-preview is batch enqueued tokens: 5,000,000 at Tier 1, 500,000,000 at Tier 2, and 1,000,000,000 at Tier 3.

Anthropic's position is different in kind rather than degree, and it is worth quoting accurately rather than overstating. Anthropic does not use the phrase "generally available" for Claude Opus 5; the launch announcement says "Claude Opus 5 is available today on all platforms." What matters is the absence of qualifiers: Anthropic explicitly labels Claude Mythos Preview "a research preview available to invited customers" and says Claude Mythos 5 "is not generally available." Opus 5 carries none of those labels, needs no beta header, and sits in the main current-models table.

Opus 5 is not qualifier-free either, and pretending otherwise would be dishonest. Several of its features are pre-release: Fast mode is "in research preview" at $10.00 and $50.00 per million tokens, the 300,000-token output limit sits under a heading reading "Extended output (beta)" and requires a beta header, and server-side fallback and compaction are both in beta. The distinction is that Anthropic ships a production model with some preview features attached, while Google ships the flagship model itself as a preview.

What does each model's reasoning ladder look like?

Claude Opus 5 appears on the Artificial Analysis leaderboard five times, once per effort level, spanning index 51 to 61 and $0.36 to $2.03 per task. Gemini 3.1 Pro Preview appears exactly once, at index 46 and $0.29 per task, with no effort suffix on the row. Reading only the headline row for Opus 5 hides four other price-performance points sold under the same product name.

The two ladders are controlled by different parameters and the levels are not translatable. Anthropic exposes effort nested inside output_config, with five values: low, medium, high, xhigh, and max, defaulting to high on the Claude API and Claude Code. Google exposes thinking_level, and for this specific model the documented set is three values: low, medium, and high, defaulting to high.

That three-value set is narrower than the Gemini family default and is a common source of error. Gemini 3.6 Flash, Gemini 3.5 Flash, and Gemini 3 Flash Preview all accept minimal. Gemini 3.1 Pro Preview does not: Google's per-model support table lists only low, medium, and high for it. Anyone carrying a four-level assumption over from the Flash tier will send an argument this model rejects.

Configuration (measured July 28, 2026)Claude Opus 5Gemini 3.1 Pro Preview
Top of the vendor's laddermax — 61, $2.03 per taskhigh — no per-level score published
Second rungxhigh — 60, $1.56 per taskmedium — no per-level score published
Vendor defaulthigh — 59, $1.06 per taskhigh — the top of its own ladder
Middle rungmedium — 56, $0.62 per tasklow — no per-level score published
Bottom of the ladderlow — 51, $0.36 per taskno level below low exists
Single published leaderboard rowfive rows, one per levelone row, unlabeled — 46, $0.29 per task

Three different kinds of blank appear in that table and they must not be read the same way. A level below low does not exist for Gemini 3.1 Pro Preview, because Google's parameter does not define one. The low and medium levels exist and are documented, but Artificial Analysis has not measured them separately, so no per-level score is published. And the single Gemini row carries no effort suffix at all, so we could not verify which level the 46 was measured at.

That last blank deserves care rather than a guess. Artificial Analysis does label Google thinking levels when it tests them — the same leaderboard carries Gemini 3.5 Flash (medium) and Gemini 3.5 Flash (minimal) as distinct rows. The absence of a suffix on Gemini 3.1 Pro Preview therefore indicates a single default-configuration run, not an evaluator that never labels. Since Google documents the default for this model as high, the most likely reading is that 46 is a score at or near the top of Gemini's ladder. We report that as an inference, not as a measurement, because Artificial Analysis does not state it.

The structural consequence survives the uncertainty either way. Opus 5 has four measured rungs above 51, and Gemini's ladder ends at high. Whatever level produced the 46, there is no higher setting on this model to buy more capability with.

Table comparing Claude Opus 5 flat rates of $5.00 input and $25.00 output against Gemini 3.1 Pro Preview at $2.00 input below 200,000 tokens and $4.00 input with $18.00 output above the threshold
The long-context mechanic side by side: Anthropic bills one flat rate across the full window, Google reprices the entire request above 200,000 input tokens.

Which model charges more for long context?

Gemini 3.1 Pro Preview does, and the mechanic is harsher than a simple higher rate. Google charges $2.00 per million input tokens for prompts at or below 200,000 tokens and $4.00 above it, with output rising from $12.00 to $18.00. Claude Opus 5 charges $5.00 and $25.00 per million tokens at every prompt length. Anthropic's documentation states that "a 900k-token request is billed at the same per-token rate as a 9k-token request."

The critical detail is which tokens the higher rate applies to, and Google answers it in a footnote on the Vertex AI pricing page rather than on the developer pricing page. The wording is unambiguous: "If a query input context is longer than 200K tokens, all tokens (input and output) are charged at long context rates."

Two consequences follow, and both are worse than the intuitive reading. First, this is not marginal pricing. A 250,000-token prompt does not bill 200,000 tokens at $2.00 and 50,000 at $4.00; it bills all 250,000 at $4.00. Second, the penalty reaches output tokens even though only input length triggers it, so the same request bills its output at $18.00 rather than $12.00 regardless of how short the answer is.

The threshold behaves as a cliff, not a slope. A request with exactly 200,000 input tokens and 10,000 output tokens costs about $0.52. Add one input token and the same request costs about $0.98, an increase of roughly 88 percent for a single token. Google's column header reads "<= 200K input tokens" and the footnote says "longer than 200K," so exactly 200,000 stays on the low rate.

500,000-token prompt, 20,000-token answerClaude Opus 5Gemini 3.1 Pro Preview
Input rate applied$5.00 per million$4.00 per million, long-context rate
Output rate applied$25.00 per million$18.00 per million, long-context rate
Total cost of the request$3.00$2.36
Same request at the short-prompt rate$3.00, unchanged$1.24
Effect of crossing the thresholdnoneabout 90 percent more expensive

The honest conclusion is narrower than "Google is more expensive," and it is worth stating precisely because the opposite overstatement is tempting. Gemini remains the cheaper model even after the surcharge: $2.36 against $3.00 on that request. What the surcharge destroys is the size of the advantage. Below the threshold Opus 5 costs about 2.5 times as much per input token; above it, about 1.25 times. A workload that routinely crosses 200,000 tokens loses roughly half of Google's price advantage, and it loses it in a step rather than gradually.

Two smaller asymmetries belong here. Google's context caching tiers the same way, from $0.20 to $0.40 per million tokens, with cache storage at $4.50 per million tokens per hour. Anthropic states that "prompt caching and batch processing discounts apply at standard rates across the full context window." And Anthropic's flat-rate claim is scoped to a model family rather than to Opus 5 by name — the documentation says "Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing," which includes Opus 5. If token-level billing is unfamiliar territory, our explainer on how input, output, and cached tokens are billed covers the mechanics.

How do the two token prices compare?

Gemini 3.1 Pro Preview is substantially cheaper per token below the 200,000-token threshold: $2.00 and $12.00 per million input and output tokens against $5.00 and $25.00 for Claude Opus 5. That is 2.5 times more for input and about 2.08 times more for output. Artificial Analysis publishes blended list prices of $1.74 per million tokens for Gemini and $3.85 for Opus 5.

Price component, per million tokensClaude Opus 5Gemini 3.1 Pro Preview
Input, standard tier$5.00 at any length$2.00 at or below 200k, $4.00 above
Output, standard tier$25.00 at any length$12.00 at or below 200k, $18.00 above
Cached input read$0.50$0.20 at or below 200k, $0.40 above
Cache write or storage$6.25 for five minutes, $10.00 for one hour$4.50 per million tokens per hour
Batch tier$2.50 and $12.50$1.00 and $6.00 at or below 200k
Priority tiernot supported on Opus 5$3.60 and $21.60 at or below 200k
Blended rate published by Artificial Analysis$3.85$1.74
Free tiernone"Not available" on all four serving tiers

The blended figures come from Artificial Analysis, which weights them seven parts cached input, two parts fresh input, and one part output. They are a convention for collapsing list prices into one number, not a measurement of any workload. They are therefore not comparable to the $2.03 and $0.29 cost-per-task figures, which count tokens actually consumed running the index. Mixing the two is the easiest error to make on this page.

Two vendor-specific wrinkles cut in opposite directions. Anthropic applies a 1.1 times multiplier when a request pins inference to United States geography, with the same 10 percent premium on partner-cloud regional endpoints. Google charges nothing extra for its global endpoint but bills grounding with Google Search at $14.00 per 1,000 queries after a free allowance of 5,000 search requests per month that is shared across the whole Gemini 3 family rather than granted per model.

Which model is more factually reliable?

Gemini 3.1 Pro Preview, by two points. Artificial Analysis scores it 33 on AA-Omniscience against 31 for Claude Opus 5 at max effort. This is the one head-to-head measurement on this page that goes against Opus 5, and it goes against it despite Opus 5 leading the overall intelligence index by fifteen points. Both trail Claude Fable 5, which leads that benchmark at 40.

AA-Omniscience is a different measurement from the intelligence index and must not be blended into it. It spans 6,000 questions across 42 topics in six domains, and Artificial Analysis describes it as follows: "It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct." Abstention is free; confident wrong answers are not.

AA-Omniscience (measured July 28, 2026)Claude Opus 5, max effortGemini 3.1 Pro Preview
Index score3133
Accuracy54.20 percent55.25 percent
Attempt rate86.65 percent79.37 percent
Correct answers, of 6,0003,2523,315
Incorrect answers, of 6,0001,3761,339
Not attempted, of 6,0008011,238

The underlying counts explain the gap and make it less flattering to Google than the headline implies. Gemini answers more questions correctly and fewer incorrectly, but it also declines 1,238 questions against Opus 5's 801. Part of its lead is earned by better recall and part by abstaining more often on a scale that rewards abstention. Both mechanisms are legitimate under the benchmark's stated rules; a buyer who needs an answer rather than a refusal should read the attempt rate alongside the score.

Opus 5's own ladder shows the same axis behaving independently of reasoning effort. Its AA-Omniscience score rises from 23.18 at low to 31.27 at max — real improvement, but not enough at any setting to pass Gemini's 33. Spending more on thinking does not buy past this particular gap.

One caveat on reading that leaderboard: the default chart on the page renders a subset of models and Gemini 3.1 Pro Preview is not among the bars, which makes Opus 5 look better placed than it is. The 33 comes from Artificial Analysis's own summary text and underlying data rather than from the visible chart. We also declined to publish the raw hallucination-rate field on that page, because its denominator is not documented and does not reconcile with the answer counts.

What does the intelligence index fail to measure?

Throughput and latency, and Gemini 3.1 Pro Preview wins both by a wide margin. Artificial Analysis measures it at about 129 output tokens per second against 55 for Claude Opus 5 at max effort, and at 42.32 seconds to first token against 78.91 seconds. These are sliding measurements read from live endpoints on July 28, 2026 and they move as providers tune capacity.

Configuration (measured July 28, 2026)Output tokens per secondTime to first tokenTotal response time
Claude Opus 5, low543.32 seconds12.53 seconds
Claude Opus 5, medium605.98 seconds14.37 seconds
Claude Opus 5, high5720.90 seconds29.66 seconds
Claude Opus 5, xhigh5842.33 seconds50.99 seconds
Claude Opus 5, max5578.91 seconds87.93 seconds
Gemini 3.1 Pro Preview12942.32 seconds46.18 seconds

The comparison that decides interactive workloads is not the headline pairing. Gemini's 42.32 seconds to first token matches Opus 5 at xhigh almost exactly and is far slower than Opus 5 at low, which answers in 3.32 seconds. If first-response latency is the constraint, Opus 5 at a low effort setting is the faster model of the two, and it scores 51 while doing it. Gemini's advantage is in how fast tokens arrive once they start, not in how quickly they start.

Artificial Analysis also flags Opus 5 at max effort as "very verbose," consuming about 100 million output tokens across the index run. Verbosity is invisible in the index score and fully visible on an invoice, and it is one reason Opus 5 costs $2.03 per task at max against Gemini's $0.29.

How do the published specifications compare?

The two are matched on input context at roughly one million tokens and separated on nearly everything else. Claude Opus 5 doubles Gemini's maximum output, 128,000 tokens against 65,536, and carries a knowledge cutoff about sixteen months fresher: May 2026 against January 2025. Gemini accepts audio and video, which Opus 5 does not accept at any price.

SpecificationClaude Opus 5Gemini 3.1 Pro Preview
Model identifierclaude-opus-5gemini-3.1-pro-preview
Launch stageProduction model, no preview label"Public preview"
Release dateJuly 24, 2026February 19, 2026
Input context window1,000,000 tokens1,048,576 tokens
Maximum output128,000 tokens synchronously65,536 tokens
Knowledge cutoffMay 2026January 2025
Input modalitiesText, image, PDF and plain-text documentsText, image, video, audio, PDF
Reasoning parameteroutput_config.effort, five levelsthinking_level, three levels
Default reasoning settinghigh on the Claude API and Claude Codehigh
Long-context surchargenone at any lengthabove 200,000 input tokens
Realtime bidirectional APInot offeredLive API not supported on this model
Priority service tiernot supported on Opus 5supported

The knowledge cutoff is the one specification where the gap is large, one-directional, and unaffected by configuration. Google publishes "Jan 2025" for this model in its Gemini 3 developer guide, and states in the same document that "Gemini 3 models have a knowledge cutoff of January 2025." A trap sits next to it: the model card shows "Latest update: February 2026," which is the release date, not the cutoff. The card publishes no cutoff row at all.

Two capability gaps run in opposite directions and are easy to miss. On Google's side, the Live API is not supported on Gemini 3.1 Pro Preview, so bidirectional realtime streaming requires dropping to the Flash tier. On Anthropic's side, web fetch is not available on Claude Opus 5 and the Priority Tier is not supported on it either, both of which Opus 4.8 offers.

The 300,000-token output figure sometimes quoted for Opus 5 needs its conditions attached. It requires the output-300k-2026-03-24 beta header, works only on the Message Batches API rather than synchronous calls, and is available on the Claude API and Claude Platform on AWS but not on Amazon Bedrock, Google Cloud, or Microsoft Foundry.

Which Gemini model is this, exactly?

This comparison is about gemini-3.1-pro-preview, Google's top reasoning model, scored at 46 on Intelligence Index v4.1. Google's catalog contains several similarly named entries carrying different numbers, and one widely cited name does not exist at all. Mixing them up is the most common factual error in comparisons of this kind.

There is no Gemini 3.5 Pro. Google shipped Gemini 3.6 Flash and Gemini 3.5 Flash-Lite in July 2026 without ever releasing a 3.5 Pro flagship. Any benchmark figure attributed to "Gemini 3.5 Pro" belongs to a model that was never released and should be discarded rather than reassigned to a nearby variant.

  • Gemini 3.1 Pro Preview — the subject of this comparison. Index 46, $0.29 per task, $2.00 and $12.00 per million tokens below 200,000, "Public preview."
  • Gemini 3 Pro Preview — the predecessor, shut down March 9, 2026. Its identifier gemini-3-pro-preview now resolves to gemini-3.1-pro-preview as an alias.
  • Gemini 3 Pro Image — a generally available image generation model, not a reasoning model. Its GA status does not transfer to the Pro reasoning tier.
  • Gemini 3.5 Flash and Gemini 3.6 Flash — cheaper, faster, generally available Flash-tier models with different scores and a four-level thinking_level ladder that includes minimal.
  • Gemini 2.5 Pro — the older Pro-tier entry listed as stable, the only non-preview Pro option Google currently offers, and scheduled for shutdown on October 16, 2026 with gemini-3.1-pro-preview named as its replacement.

One documentation inconsistency is worth knowing about before someone cites it against this page. Google's Gemini 3 developer guide still says "All Gemini 3 models are currently in preview," which stopped being true when the Flash models went GA. The current authorities are the models index and the Gemini Enterprise Agent Platform launch-stage block, both cited below.

Who wins each category?

CategoryWinnerMargin
Peak measured intelligenceClaude Opus 561 against 46, fifteen points on Index v4.1
Intelligence at the cheapest configurationClaude Opus 551 at low against a single published 46
Reasoning headroomClaude Opus 5four measured rungs above its own floor; Gemini's ladder ends at high
Cost per taskGemini 3.1 Pro Preview$0.29 against $0.36, about 19 percent cheaper
Price per million tokensGemini 3.1 Pro Preview$2.00 and $12.00 against $5.00 and $25.00 below 200k
Long-context price predictabilityClaude Opus 5flat at any length; Gemini reprices the whole request above 200k
Output throughputGemini 3.1 Pro Previewabout 129 tokens per second against 55
Fastest first tokenClaude Opus 53.32 seconds at low against 42.32 seconds
Factual reliabilityGemini 3.1 Pro Preview33 against 31 on AA-Omniscience
Input modality breadthGemini 3.1 Pro Previewadds video and audio, which Opus 5 does not accept
Maximum output lengthClaude Opus 5128,000 tokens against 65,536
Knowledge freshnessClaude Opus 5May 2026 against January 2025, about sixteen months
Release stabilityClaude Opus 5production model against "Public preview" with two weeks' deprecation notice

That is eight categories to five, which is a real but narrower margin than fifteen index points implies. The split is not random. Opus 5 wins the categories describing how capable the model is, how fresh it is, and how predictable it is to depend on. Gemini wins the categories describing what it costs, how fast tokens arrive, what data it can read, and how often it avoids inventing an answer.

What are the strengths and weaknesses of each?

Claude Opus 5

Strengths. Highest measured intelligence of the pair at 61, fifteen points clear, with a floor of 51 that still sits five points above Gemini's only published score. Five reasoning levels spanning a 5.6-fold cost range, so one integration serves both cheap and expensive work. One million tokens of context at a flat rate with no length surcharge and no beta header. A May 2026 knowledge cutoff, about sixteen months fresher. Maximum output of 128,000 tokens, roughly double Gemini's. First-token latency of 3.32 seconds at low, an order of magnitude faster to respond than Gemini. A production model with no preview label attached to it.

Weaknesses. Loses the factual-reliability head-to-head at 31 against 33 on AA-Omniscience, and cannot close it at any effort level. Roughly 2.5 times Gemini's input price and about 2.08 times its output price below 200,000 tokens. About 19 percent more expensive per task at its cheapest rung. Under half Gemini's throughput. No audio or video input at all. Flagged "very verbose" by the evaluator, consuming about 100 million output tokens across the index run. Web fetch and the Priority Tier are both unsupported, though Opus 4.8 offers them. Several capabilities remain pre-release, including Fast mode and 300,000-token output.

Gemini 3.1 Pro Preview

Strengths. Cheaper on every published axis below the threshold: $0.29 per task, $2.00 and $12.00 per million tokens, $1.74 blended. Better factual reliability at 33 on AA-Omniscience, ahead of a model that leads it by fifteen intelligence points. About 2.3 times the output throughput. Native text, image, video, audio, and PDF input. A 1,048,576-token context window. Broad feature support including code execution, Search and Maps grounding, URL context, structured outputs, and Batch, Flex, and Priority service tiers. Production and commercial use explicitly permitted under the Pre-GA Offerings Terms.

Weaknesses. Still labeled "Public preview" 159 days after release, contractually a Pre-GA Offering provided "as is" with no SLA and no indemnity, and discontinuable without prior notice — a risk Google has already realized once on this product line, retiring the predecessor after 111 days. A single published index score of 46, fifteen points behind, with no higher thinking_level to buy more capability. A long-context surcharge that reprices the entire request, output included, above 200,000 input tokens. A January 2025 knowledge cutoff, about sixteen months behind. Maximum output of 65,536 tokens. Time to first token of 42.32 seconds. No free tier on any serving tier. Live API unsupported. Rate limits unpublished and explicitly "more restricted for experimental and preview models."

When should you pick each one?

When to pick Claude Opus 5

Choose Opus 5 when the work is reasoning-shaped and a wrong answer costs more than the inference: agentic coding loops, complex refactors, multi-step research. Choose it when you need capability above 46, because Gemini has no configuration that buys more. Choose it for document-heavy pipelines that routinely exceed 200,000 tokens, where flat pricing beats a rate that steps up and drags output with it. Choose it when the answer must reflect the last eighteen months, since May 2026 against January 2025 is the difference between knowing a library, a regulation, or a competitor exists and not knowing. And choose it when an SLA and a predictable deprecation path are procurement requirements rather than preferences.

When to pick Gemini 3.1 Pro Preview

Choose Gemini when the workload is high-volume and cost-sensitive and index 46 clears your bar, because it is cheaper per task and roughly 2.5 times cheaper per input token below the threshold. Choose it when factual recall matters more than reasoning depth, since it wins that measurement outright. Choose it when you need audio or video input, which Opus 5 cannot accept, or when sustained throughput decides the experience and 129 tokens per second against 55 is the deciding number. Choose it for retrieval-shaped work where the January 2025 cutoff is irrelevant because facts arrive in the prompt, and where Google's Search grounding covers the rest.

When the comparison does not apply

If your governance rules forbid building on pre-GA services, this is a one-model comparison and Gemini is out regardless of its numbers. If you need a stable Google Pro-tier model specifically, the only option is Gemini 2.5 Pro, which is older, not covered here, and carries a published shutdown date of October 16, 2026 — so it is a decision you will have to make again. If your workload runs acceptably well below index 46, the Flash tier is dramatically cheaper than either model and both are overspecified. And if factual reliability is the single requirement, neither wins: Claude Fable 5 leads AA-Omniscience at 40, ahead of both.

Verdict panel showing Claude Opus 5 winning on index score, flat pricing and knowledge cutoff, Gemini 3.1 Pro Preview winning on cost per task, throughput and omniscience, and the preview risk
The verdict in one view: Opus 5 for capability, freshness and pricing predictability; Gemini 3.1 Pro Preview for cost, speed and factual recall.

Final verdict

Claude Opus 5 wins this comparison on capability and on dependability, and the second half of that sentence carries more weight than the first. Fifteen index points is a wide margin, and Opus 5's cheapest configuration still scores five points above the only figure Artificial Analysis publishes for Gemini 3.1 Pro Preview. But the argument that should decide a production commitment is the one printed in Google's own contract: the flagship is a Pre-GA Offering, provided "as is," not covered by any SLA or indemnity, and open to being "changed, suspended or discontinued at any time without prior notice," 159 days into a preview with no announced general availability date.

Gemini's case is genuinely strong and this page has not tried to weaken it. It costs less per task, less per token, and streams tokens more than twice as fast. It beats Opus 5 on the one benchmark here that explicitly penalizes hallucination, which is an uncomfortable result for a model leading the intelligence index by fifteen points and one we are not going to bury. It reads audio and video that Opus 5 cannot open. For a high-volume pipeline that clears its bar at index 46, Gemini is the rational choice on economics alone.

What that pipeline is buying, though, is a model that cannot go higher and might not stay. Gemini's thinking_level ladder ends at high, so there is no setting that buys past 46 at any price, and Google has already shut down this model's direct predecessor after 111 days and repointed its identifier. Opus 5 has four measured rungs above Gemini's score and no preview label.

The uncomfortable part is that Google currently offers no way out of this within its own Pro tier. Gemini 2.5 Pro, the stable alternative, is scheduled for shutdown on October 16, 2026, and the replacement Google names for it is the preview model. Teams that can absorb a migration on short notice, and whose work fits under index 46, should take Gemini and the savings — they are real and substantial. Teams whose governance requires an SLA on the model serving production traffic do not have a Google option in this tier right now, and that, more than fifteen index points, is what decides the comparison.

Frequently asked questions

Is Gemini 3.1 Pro Preview still in preview in July 2026?

Yes. Google's Gemini Enterprise Agent Platform documentation lists gemini-3.1-pro-preview with "Launch stage: Public preview" and a release date of February 19, 2026, which is five months and nine days before July 28, 2026. On the Gemini API models page it appears under the Preview heading while Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.1 Flash-Lite appear under Stable. Google has published no general availability date.

Can Gemini 3.1 Pro Preview be used in production?

Yes, Google explicitly permits it, but without an SLA. The Generative AI Preview banner places the model under the Pre-GA Offerings Terms and states that customers "may elect to use it for production or commercial purposes, or disclose Generated Output to third-parties, and may process personal data as outlined in the Cloud Data Processing Addendum." Google's Service Specific Terms separately state that Pre-GA Offerings "are not covered by any SLA or Google indemnity." The restriction is on the guarantees, not on the permission.

How much notice does Google give before retiring a preview model?

Google's documentation promises "at least 2 weeks notice," but its Service Specific Terms reserve the right to change, suspend or discontinue a Pre-GA Offering "at any time without prior notice to Customer." In practice, Gemini 3 Pro Preview was announced for deprecation on February 26, 2026 and shut down on March 9, eleven days later. Its identifier was then repointed to gemini-3.1-pro-preview, so calls kept succeeding against a different model.

Does Gemini 3.1 Pro Preview charge more for long prompts?

Yes, and the surcharge applies to the whole request. Input rises from $2.00 to $4.00 per million tokens and output from $12.00 to $18.00 above 200,000 input tokens. Google's Vertex AI pricing footnote states: "If a query input context is longer than 200K tokens, all tokens (input and output) are charged at long context rates." This is not marginal pricing — a 250,000-token prompt bills all 250,000 tokens at the higher rate.

Does Claude Opus 5 charge more for long context?

No. Anthropic states that "Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing," adding that "a 900k-token request is billed at the same per-token rate as a 9k-token request." Prompt caching and batch discounts apply at standard rates across the full window, and the one-million-token context requires no beta header. Claude Opus 5 falls inside that model family.

Which model scores higher on the Artificial Analysis Intelligence Index?

Claude Opus 5, by fifteen points. It scores 61 at max effort on Intelligence Index v4.1 against 46 for Gemini 3.1 Pro Preview. Opus 5 also scores 51 at its cheapest effort level, five points above Gemini's only published score. Artificial Analysis lists five rows for Opus 5, one per effort level, and a single unlabeled row for Gemini 3.1 Pro Preview, so no per-thinking-level breakdown is published for Google's model.

At which thinking level was Gemini 3.1 Pro Preview's score of 46 measured?

Artificial Analysis does not disclose it. The leaderboard row carries no effort suffix, unlike its Gemini 3.5 Flash rows which are labeled medium and minimal. Google documents the default for this model as high, which is also the top of its three-level ladder, so 46 most likely reflects a default-configuration run at or near Gemini's ceiling. We report that as an inference rather than a measurement, because the evaluator does not state it.

Which model hallucinates less?

Gemini 3.1 Pro Preview, by two points. It scores 33 on AA-Omniscience against 31 for Claude Opus 5 at max effort. Artificial Analysis describes the benchmark as one that "rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer," on a scale from -100 to 100. Gemini answers 3,315 of 6,000 questions correctly against 3,252, but also declines 1,238 against 801, so part of its lead comes from abstaining more often.

Which reasoning levels does each model support?

Claude Opus 5 exposes five effort levels through output_config.effort: low, medium, high, xhigh, and max, defaulting to high on the Claude API and Claude Code. Gemini 3.1 Pro Preview exposes three thinking levels through thinking_level: low, medium, and high, defaulting to high. This model does not accept minimal, unlike Gemini 3.6 Flash and Gemini 3.5 Flash. The legacy numeric thinking_budget remains supported for backward compatibility but cannot be combined with thinking_level.

Which model is faster?

It depends on which kind of speed. Gemini 3.1 Pro Preview streams about 129 output tokens per second against 55 for Claude Opus 5 at max effort, roughly 2.3 times faster. But its time to first token is 42.32 seconds, while Opus 5 at low effort answers in 3.32 seconds. Gemini wins sustained throughput; Opus 5 at a low effort setting wins first-response latency by a wide margin. Both figures were measured on July 28, 2026.

Is there a Gemini 3.5 Pro?

No. Google has not released a model called Gemini 3.5 Pro. The current top reasoning entry is Gemini 3.1 Pro Preview, and the nearby models that do exist are Gemini 3.5 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, Gemini 3 Pro Image, and Gemini 2.5 Pro. Any benchmark figure attributed to "Gemini 3.5 Pro" belongs to a model that was never released and should be discarded rather than reassigned to a nearby variant.

How do the context windows, output limits, and knowledge cutoffs compare?

Claude Opus 5 has a one-million-token context window, a 128,000-token synchronous output limit, and a May 2026 knowledge cutoff. Gemini 3.1 Pro Preview has a 1,048,576-token input limit, a 65,536-token output limit, and a January 2025 knowledge cutoff, about sixteen months older. Note that Gemini's model card shows "Latest update: February 2026," which is its release date rather than its cutoff; the card publishes no cutoff row at all.

Sources and references

Every figure on this page comes from a vendor's own documentation or from an independent evaluator, and the two are labeled separately throughout rather than stacked into a single ranking. Google's published benchmark scores for Gemini 3.1 Pro Preview are vendor self-reported and are not used for any head-to-head claim here; neither are Anthropic's, which compare Claude Opus 5 only to other Anthropic models. Cost-per-task, throughput, and latency figures are sliding measurements read on July 28, 2026.

Related comparisons on ThePlanetTools: Claude Opus 4.8 against Gemini 3.1 Pro, Claude Sonnet 5 against Gemini 3.1 Pro, Claude Fable 5 against Gemini 3.1 Pro, and GPT-5.6 Sol against Gemini 3.1 Pro. Background reading: our report on the Claude Opus 5 launch.

Our Verdict

Claude Opus 5 wins, on capability and on dependability. It scores 61 on Artificial Analysis Intelligence Index v4.1 at max effort against 46 for Gemini 3.1 Pro Preview, and its cheapest configuration still scores 51, five points above the only figure published for Google’s model. Gemini’s case is real and this page does not weaken it: it costs $0.29 per task against $0.36, streams about 129 output tokens per second against 55, beats Opus 5 on factual reliability by 33 to 31 on AA-Omniscience, and reads audio and video that Opus 5 cannot accept. What decides it is the label. Gemini 3.1 Pro Preview has been a "Public preview" offering for 159 days with no announced general availability date; Google’s Service Specific Terms exclude Pre-GA offerings from any SLA or indemnity and allow discontinuation without prior notice, and Google retired this model’s direct predecessor after 111 days. Google also charges a long-context surcharge that reprices every token in a request, output included, above 200,000 input tokens, while Anthropic bills its full one-million-token window flat. Teams that can absorb a short-notice migration and fit under index 46 should take Gemini and the savings; teams that need an SLA on the model serving production traffic have no Google option in this tier today.

Winner:Claude Opus 5

Choose Claude Opus 5

Anthropic's frontier reasoning model — top of the independent index at half the price of Fable 5.

Try Claude Opus 5

Choose Gemini 3.1 Pro Preview

Google DeepMind's flagship Gemini 3.1 Pro Preview — 94.3% GPQA Diamond, 77.1% ARC-AGI-2, 1M-token context, multimodal in/text out, vibe coding plus agentic tool use. Preview status as of April 2026.

Try Gemini 3.1 Pro Preview

Frequently Asked Questions

Is Claude Opus 5 better than Gemini 3.1 Pro Preview?

Claude Opus 5 wins, on capability and on dependability. It scores 61 on Artificial Analysis Intelligence Index v4.1 at max effort against 46 for Gemini 3.1 Pro Preview, and its cheapest configuration still scores 51, five points above the only figure published for Google’s model. Gemini’s case is real and this page does not weaken it: it costs $0.29 per task against $0.36, streams about 129 output tokens per second against 55, beats Opus 5 on factual reliability by 33 to 31 on AA-Omniscience, and reads audio and video that Opus 5 cannot accept. What decides it is the label. Gemini 3.1 Pro Preview has been a "Public preview" offering for 159 days with no announced general availability date; Google’s Service Specific Terms exclude Pre-GA offerings from any SLA or indemnity and allow discontinuation without prior notice, and Google retired this model’s direct predecessor after 111 days. Google also charges a long-context surcharge that reprices every token in a request, output included, above 200,000 input tokens, while Anthropic bills its full one-million-token window flat. Teams that can absorb a short-notice migration and fit under index 46 should take Gemini and the savings; teams that need an SLA on the model serving production traffic have no Google option in this tier today.

Which is cheaper, Claude Opus 5 or Gemini 3.1 Pro Preview?

Claude Opus 5 is priced at $5 in / $25 out per M tokens. Gemini 3.1 Pro Preview is priced at $2 in / $12 out per M tokens. Check the pricing comparison section above for a full breakdown.

What are the main differences between Claude Opus 5 and Gemini 3.1 Pro Preview?

The key differences span across 13 features we compared. For Artificial Analysis Intelligence Index v4.1, top measured configuration, Claude Opus 5 offers 61 (max effort) while Gemini 3.1 Pro Preview offers 46 (default configuration, level not disclosed). For Index at the cheapest measured configuration, Claude Opus 5 offers 51 (low effort), $0.36 per task while Gemini 3.1 Pro Preview offers 46, $0.29 per task — only row published. For Cost per task, measured July 28, 2026, Claude Opus 5 offers $0.36 at low effort, $2.03 at max while Gemini 3.1 Pro Preview offers $0.29. See the full feature comparison table above for all details.

Related Comparisons