Claude Opus 5 vs Gemini 3.1 Pro Preview: Fifteen Points and a Preview Label (2026)
Claude Opus 5 scores 61 to Gemini 3.1 Pro Preview’s 46. Gemini is cheaper and more accurate, but 159 days into preview with no SLA. July 28, 2026.
Feature Comparison
| Feature | Claude Opus 5 | Gemini 3.1 Pro Preview |
|---|---|---|
| Artificial Analysis Intelligence Index v4.1, top measured configuration | 61 (max effort) | 46 (default configuration, level not disclosed) |
| Index at the cheapest measured configuration | 51 (low effort), $0.36 per task | 46, $0.29 per task — only row published |
| Cost per task, measured July 28, 2026 | $0.36 at low effort, $2.03 at max | $0.29 |
| Release stage | Production model, no preview label | "Public preview" since February 19, 2026 |
| Long-context pricing above 200,000 input tokens | Flat, no surcharge at any length | All tokens repriced: input $2.00 to $4.00, output $12.00 to $18.00 |
| Input price per million tokens, standard tier | $5.00 | $2.00 up to 200k, $4.00 above |
| AA-Omniscience factual reliability | 31 (max effort) | 33 |
| Output throughput | About 55 tokens per second at max effort | About 129 tokens per second |
| Time to first token | 3.32 seconds at low effort | 42.32 seconds |
| Reasoning ladder | Five levels via output_config.effort | Three levels via thinking_level, no minimal |
| Input modalities | Text, image, PDF and plain-text documents | Text, image, video, audio, PDF |
| Maximum output tokens | 128,000 synchronously | 65,536 |
| Knowledge cutoff | May 2026 | January 2025 |
Pricing Comparison
Claude Opus 5
Gemini 3.1 Pro Preview
Detailed Comparison
Claude Opus 5 scores 61 on the Artificial Analysis Intelligence Index v4.1 at max effort against 46 for Gemini 3.1 Pro Preview, a gap of fifteen points. The label matters more than the gap: Google's flagship reasoning model has carried the launch stage "Public preview" since February 19, 2026, and Google's contract terms exclude Pre-GA offerings from any SLA and reserve the right to discontinue them without prior notice, while Claude Opus 5 shipped as an ordinary production model on July 24, 2026. Gemini is the cheaper and faster model, at $0.29 per task against $0.36 for Opus 5's lowest rung and about 129 output tokens per second against 55, and it beats Opus 5 on factual reliability by 33 to 31 on AA-Omniscience. We researched both from vendor documentation and independent evaluations; index figures are v4.1 and sliding measurements were read on July 28, 2026.
Which model should you choose?
Pick Claude Opus 5 if you need capability above index 46, a knowledge cutoff from 2026 rather than January 2025, flat pricing across a one-million-token window, or a model covered by an ordinary support commitment. Pick Gemini 3.1 Pro Preview if you need audio or video input, roughly 2.3 times the throughput, a lower bill per task, or measurably better factual recall.
Claude Opus 5 wins this comparison, but not the way the headline numbers suggest and not without real concessions. Fifteen index points is the widest measured gap on the page, and Opus 5's cheapest configuration still scores five points above the only score Artificial Analysis publishes for Gemini 3.1 Pro Preview. That part is a clean win.
The part that is not clean: Gemini is cheaper per task, roughly twice as fast, and better at not making things up. Those are not consolation categories. A model that costs $0.29 per task against $0.36, streams at about 129 output tokens per second against 55, and scores 33 against 31 on a benchmark that explicitly penalizes hallucination is winning on three axes that many teams weight above raw reasoning power.
What tips the decision is the third column, the one neither vendor prints as a number. Gemini 3.1 Pro Preview is a pre-general-availability offering. Google's Service Specific Terms state that Pre-GA Offerings "are not covered by any SLA or Google indemnity" and "may be changed, suspended or discontinued at any time without prior notice to Customer." Google has already exercised that on this exact product line: the previous Gemini 3 Pro Preview was shut down on March 9, 2026, 111 days after its release. Anyone building on the current one is building under the same terms.
Is Gemini 3.1 Pro Preview still in preview?
Yes, as of July 28, 2026. Google's Gemini Enterprise Agent Platform documentation lists gemini-3.1-pro-preview with "Launch stage: Public preview" and a release date of February 19, 2026 — five months and nine days in preview. On the Gemini API models page the model sits under the heading Preview, while Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.1 Flash-Lite all sit under Stable.
That split inside Google's own catalog is the fact worth pausing on. Google is perfectly capable of promoting Gemini models to general availability, and it did so seven days before this page was written. The changelog entry for July 21, 2026 reads: "Gemini 3.6 Flash and Gemini 3.5 Flash-Lite generally available (GA): Released stable, production-ready versions of our latest 3.x Flash models." The Flash tier has now gone GA four separate times in 2026. The reasoning Pro tier has not gone GA once.
One precision matters here, because a careless version of this claim is easy to refute. A Google model with "Pro" in its name is generally available: gemini-3-pro-image, announced on May 28, 2026 as one of "the generally available (GA) versions of our native visual models." That is an image generation model. The claim on this page is narrower and holds: Google ships no generally available Gemini 3.x reasoning Pro model. Every occurrence of gemini-3.1-pro in Google's model catalog is followed by -preview.
Google has published no general availability date for this model, and that is an absence we checked rather than assumed. The Gemini API deprecations table lists gemini-3.1-pro-preview with a release date of February 19, 2026 and, in the shutdown column, "No shutdown date announced." There is no published GA target, no stable-version date, and no end-of-support date anywhere in Google's documentation for it. The model has now been in preview for 159 days.
We covered the shape of this problem when it first became visible, in our reporting on why the promised Gemini 3.5 Pro is late and on the three Gemini models Google shipped without the flagship. The comparison here is the downstream consequence: buyers evaluating Google's best reasoning model in July 2026 are evaluating a preview.
What does preview status actually cost a buyer?
Three concrete things: no service level agreement, no indemnity, and a shorter and less certain support runway. Google's Service Specific Terms classify anything labeled "Preview" as a Pre-GA Offering, and state that Pre-GA Offerings "are provided 'as is' without any express or implied warranties or representations of any kind," that they "may be changed, suspended or discontinued at any time without prior notice to Customer," and that they "are not covered by any SLA or Google indemnity." A separate clause caps Google's liability for Pre-GA Offerings at $25,000.
Set that against what Google's developer documentation promises, because the two do not say the same thing. The Gemini API models page states that preview models "will be deprecated with at least 2 weeks notice." The contract reserves the right to discontinue with no notice at all. The two weeks is a documentation courtesy, not a contractual guarantee, and a procurement review will read the contract.
One widespread misreading needs correcting before anyone repeats it, because it is wrong and Google would rightly push back. The generic Pre-GA terms say customers "should not use Pre-GA Offerings to process personal data," but they also allow Google to state otherwise in its documentation, and for this model Google does. The Generative AI Preview banner on the Gemini 3.1 Pro page says that "Customers may elect to use it for production or commercial purposes, or disclose Generated Output to third-parties, and may process personal data as outlined in the Cloud Data Processing Addendum." Production use and personal data are both explicitly permitted. What the banner does not restore is the SLA or the indemnity.
The risk is therefore not a licensing prohibition, it is continuity — and Google's own deprecations table quantifies it. Preview models on this API have been given runways of roughly three to four months, while generally available models get a year or more. That table is published by Google and needs no interpretation.
| Gemini model | Stage | Release date | Shutdown date | Published runway |
|---|---|---|---|---|
gemini-3-pro-preview | Preview | November 18, 2025 | March 9, 2026 | 111 days |
gemini-3.1-flash-lite-preview | Preview | March 3, 2026 | May 25, 2026 | 83 days |
gemini-3.1-flash-image-preview | Preview | February 26, 2026 | June 25, 2026 | 119 days |
gemini-3.1-flash-lite | Generally available | May 7, 2026 | May 7, 2027 | 365 days |
gemini-2.5-pro | Stable | June 17, 2025 | October 16, 2026 | about 16 months |
gemini-3.1-pro-preview | Preview | February 19, 2026 | none announced | not published |
The precedent is worth reading closely, because the failure mode was quieter than a shutdown usually is. Google announced the retirement of Gemini 3 Pro Preview on February 26, 2026 and shut it down on March 9, eleven days later — shorter than the two weeks its own documentation describes, though we cannot rule out an earlier private notice to customers. The changelog then records: "The Gemini 3 Pro Preview model has been shut down. The gemini-3-pro-preview now points to gemini-3.1-pro-preview." The identifier kept resolving. Calls kept returning results. The model underneath had changed.
Then there is the fact that reframes the whole question, and it comes from the same Google table. Gemini 2.5 Pro, the only Pro-tier model Google currently lists as stable, has a published shutdown date of October 16, 2026. The recommended replacement Google names for it is gemini-3.1-pro-preview. A buyer following Google's own migration guidance moves from an SLA-covered stable model to a preview offering that is contractually "as is." At the time of writing there is no generally available successor in the Pro tier to move to instead.
Rate limits carry the same caveat. Google's rate-limits page declines to publish requests-per-minute or tokens-per-minute figures for this model, deferring to the console, and states plainly that "rate limits are more restricted for experimental and preview models." The only hard quota Google does publish for gemini-3.1-pro-preview is batch enqueued tokens: 5,000,000 at Tier 1, 500,000,000 at Tier 2, and 1,000,000,000 at Tier 3.
Anthropic's position is different in kind rather than degree, and it is worth quoting accurately rather than overstating. Anthropic does not use the phrase "generally available" for Claude Opus 5; the launch announcement says "Claude Opus 5 is available today on all platforms." What matters is the absence of qualifiers: Anthropic explicitly labels Claude Mythos Preview "a research preview available to invited customers" and says Claude Mythos 5 "is not generally available." Opus 5 carries none of those labels, needs no beta header, and sits in the main current-models table.
Opus 5 is not qualifier-free either, and pretending otherwise would be dishonest. Several of its features are pre-release: Fast mode is "in research preview" at $10.00 and $50.00 per million tokens, the 300,000-token output limit sits under a heading reading "Extended output (beta)" and requires a beta header, and server-side fallback and compaction are both in beta. The distinction is that Anthropic ships a production model with some preview features attached, while Google ships the flagship model itself as a preview.
What does each model's reasoning ladder look like?
Claude Opus 5 appears on the Artificial Analysis leaderboard five times, once per effort level, spanning index 51 to 61 and $0.36 to $2.03 per task. Gemini 3.1 Pro Preview appears exactly once, at index 46 and $0.29 per task, with no effort suffix on the row. Reading only the headline row for Opus 5 hides four other price-performance points sold under the same product name.
The two ladders are controlled by different parameters and the levels are not translatable. Anthropic exposes effort nested inside output_config, with five values: low, medium, high, xhigh, and max, defaulting to high on the Claude API and Claude Code. Google exposes thinking_level, and for this specific model the documented set is three values: low, medium, and high, defaulting to high.
That three-value set is narrower than the Gemini family default and is a common source of error. Gemini 3.6 Flash, Gemini 3.5 Flash, and Gemini 3 Flash Preview all accept minimal. Gemini 3.1 Pro Preview does not: Google's per-model support table lists only low, medium, and high for it. Anyone carrying a four-level assumption over from the Flash tier will send an argument this model rejects.
| Configuration (measured July 28, 2026) | Claude Opus 5 | Gemini 3.1 Pro Preview |
|---|---|---|
| Top of the vendor's ladder | max — 61, $2.03 per task | high — no per-level score published |
| Second rung | xhigh — 60, $1.56 per task | medium — no per-level score published |
| Vendor default | high — 59, $1.06 per task | high — the top of its own ladder |
| Middle rung | medium — 56, $0.62 per task | low — no per-level score published |
| Bottom of the ladder | low — 51, $0.36 per task | no level below low exists |
| Single published leaderboard row | five rows, one per level | one row, unlabeled — 46, $0.29 per task |
Three different kinds of blank appear in that table and they must not be read the same way. A level below low does not exist for Gemini 3.1 Pro Preview, because Google's parameter does not define one. The low and medium levels exist and are documented, but Artificial Analysis has not measured them separately, so no per-level score is published. And the single Gemini row carries no effort suffix at all, so we could not verify which level the 46 was measured at.
That last blank deserves care rather than a guess. Artificial Analysis does label Google thinking levels when it tests them — the same leaderboard carries Gemini 3.5 Flash (medium) and Gemini 3.5 Flash (minimal) as distinct rows. The absence of a suffix on Gemini 3.1 Pro Preview therefore indicates a single default-configuration run, not an evaluator that never labels. Since Google documents the default for this model as high, the most likely reading is that 46 is a score at or near the top of Gemini's ladder. We report that as an inference, not as a measurement, because Artificial Analysis does not state it.
The structural consequence survives the uncertainty either way. Opus 5 has four measured rungs above 51, and Gemini's ladder ends at high. Whatever level produced the 46, there is no higher setting on this model to buy more capability with.
Which model charges more for long context?
Gemini 3.1 Pro Preview does, and the mechanic is harsher than a simple higher rate. Google charges $2.00 per million input tokens for prompts at or below 200,000 tokens and $4.00 above it, with output rising from $12.00 to $18.00. Claude Opus 5 charges $5.00 and $25.00 per million tokens at every prompt length. Anthropic's documentation states that "a 900k-token request is billed at the same per-token rate as a 9k-token request."
The critical detail is which tokens the higher rate applies to, and Google answers it in a footnote on the Vertex AI pricing page rather than on the developer pricing page. The wording is unambiguous: "If a query input context is longer than 200K tokens, all tokens (input and output) are charged at long context rates."
Two consequences follow, and both are worse than the intuitive reading. First, this is not marginal pricing. A 250,000-token prompt does not bill 200,000 tokens at $2.00 and 50,000 at $4.00; it bills all 250,000 at $4.00. Second, the penalty reaches output tokens even though only input length triggers it, so the same request bills its output at $18.00 rather than $12.00 regardless of how short the answer is.
The threshold behaves as a cliff, not a slope. A request with exactly 200,000 input tokens and 10,000 output tokens costs about $0.52. Add one input token and the same request costs about $0.98, an increase of roughly 88 percent for a single token. Google's column header reads "<= 200K input tokens" and the footnote says "longer than 200K," so exactly 200,000 stays on the low rate.
| 500,000-token prompt, 20,000-token answer | Claude Opus 5 | Gemini 3.1 Pro Preview |
|---|---|---|
| Input rate applied | $5.00 per million | $4.00 per million, long-context rate |
| Output rate applied | $25.00 per million | $18.00 per million, long-context rate |
| Total cost of the request | $3.00 | $2.36 |
| Same request at the short-prompt rate | $3.00, unchanged | $1.24 |
| Effect of crossing the threshold | none | about 90 percent more expensive |
The honest conclusion is narrower than "Google is more expensive," and it is worth stating precisely because the opposite overstatement is tempting. Gemini remains the cheaper model even after the surcharge: $2.36 against $3.00 on that request. What the surcharge destroys is the size of the advantage. Below the threshold Opus 5 costs about 2.5 times as much per input token; above it, about 1.25 times. A workload that routinely crosses 200,000 tokens loses roughly half of Google's price advantage, and it loses it in a step rather than gradually.
Two smaller asymmetries belong here. Google's context caching tiers the same way, from $0.20 to $0.40 per million tokens, with cache storage at $4.50 per million tokens per hour. Anthropic states that "prompt caching and batch processing discounts apply at standard rates across the full context window." And Anthropic's flat-rate claim is scoped to a model family rather than to Opus 5 by name — the documentation says "Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing," which includes Opus 5. If token-level billing is unfamiliar territory, our explainer on how input, output, and cached tokens are billed covers the mechanics.
How do the two token prices compare?
Gemini 3.1 Pro Preview is substantially cheaper per token below the 200,000-token threshold: $2.00 and $12.00 per million input and output tokens against $5.00 and $25.00 for Claude Opus 5. That is 2.5 times more for input and about 2.08 times more for output. Artificial Analysis publishes blended list prices of $1.74 per million tokens for Gemini and $3.85 for Opus 5.
| Price component, per million tokens | Claude Opus 5 | Gemini 3.1 Pro Preview |
|---|---|---|
| Input, standard tier | $5.00 at any length | $2.00 at or below 200k, $4.00 above |
| Output, standard tier | $25.00 at any length | $12.00 at or below 200k, $18.00 above |
| Cached input read | $0.50 | $0.20 at or below 200k, $0.40 above |
| Cache write or storage | $6.25 for five minutes, $10.00 for one hour | $4.50 per million tokens per hour |
| Batch tier | $2.50 and $12.50 | $1.00 and $6.00 at or below 200k |
| Priority tier | not supported on Opus 5 | $3.60 and $21.60 at or below 200k |
| Blended rate published by Artificial Analysis | $3.85 | $1.74 |
| Free tier | none | "Not available" on all four serving tiers |
The blended figures come from Artificial Analysis, which weights them seven parts cached input, two parts fresh input, and one part output. They are a convention for collapsing list prices into one number, not a measurement of any workload. They are therefore not comparable to the $2.03 and $0.29 cost-per-task figures, which count tokens actually consumed running the index. Mixing the two is the easiest error to make on this page.
Two vendor-specific wrinkles cut in opposite directions. Anthropic applies a 1.1 times multiplier when a request pins inference to United States geography, with the same 10 percent premium on partner-cloud regional endpoints. Google charges nothing extra for its global endpoint but bills grounding with Google Search at $14.00 per 1,000 queries after a free allowance of 5,000 search requests per month that is shared across the whole Gemini 3 family rather than granted per model.
Which model is more factually reliable?
Gemini 3.1 Pro Preview, by two points. Artificial Analysis scores it 33 on AA-Omniscience against 31 for Claude Opus 5 at max effort. This is the one head-to-head measurement on this page that goes against Opus 5, and it goes against it despite Opus 5 leading the overall intelligence index by fifteen points. Both trail Claude Fable 5, which leads that benchmark at 40.
AA-Omniscience is a different measurement from the intelligence index and must not be blended into it. It spans 6,000 questions across 42 topics in six domains, and Artificial Analysis describes it as follows: "It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct." Abstention is free; confident wrong answers are not.
| AA-Omniscience (measured July 28, 2026) | Claude Opus 5, max effort | Gemini 3.1 Pro Preview |
|---|---|---|
| Index score | 31 | 33 |
| Accuracy | 54.20 percent | 55.25 percent |
| Attempt rate | 86.65 percent | 79.37 percent |
| Correct answers, of 6,000 | 3,252 | 3,315 |
| Incorrect answers, of 6,000 | 1,376 | 1,339 |
| Not attempted, of 6,000 | 801 | 1,238 |
The underlying counts explain the gap and make it less flattering to Google than the headline implies. Gemini answers more questions correctly and fewer incorrectly, but it also declines 1,238 questions against Opus 5's 801. Part of its lead is earned by better recall and part by abstaining more often on a scale that rewards abstention. Both mechanisms are legitimate under the benchmark's stated rules; a buyer who needs an answer rather than a refusal should read the attempt rate alongside the score.
Opus 5's own ladder shows the same axis behaving independently of reasoning effort. Its AA-Omniscience score rises from 23.18 at low to 31.27 at max — real improvement, but not enough at any setting to pass Gemini's 33. Spending more on thinking does not buy past this particular gap.
One caveat on reading that leaderboard: the default chart on the page renders a subset of models and Gemini 3.1 Pro Preview is not among the bars, which makes Opus 5 look better placed than it is. The 33 comes from Artificial Analysis's own summary text and underlying data rather than from the visible chart. We also declined to publish the raw hallucination-rate field on that page, because its denominator is not documented and does not reconcile with the answer counts.
What does the intelligence index fail to measure?
Throughput and latency, and Gemini 3.1 Pro Preview wins both by a wide margin. Artificial Analysis measures it at about 129 output tokens per second against 55 for Claude Opus 5 at max effort, and at 42.32 seconds to first token against 78.91 seconds. These are sliding measurements read from live endpoints on July 28, 2026 and they move as providers tune capacity.
| Configuration (measured July 28, 2026) | Output tokens per second | Time to first token | Total response time |
|---|---|---|---|
Claude Opus 5, low | 54 | 3.32 seconds | 12.53 seconds |
Claude Opus 5, medium | 60 | 5.98 seconds | 14.37 seconds |
Claude Opus 5, high | 57 | 20.90 seconds | 29.66 seconds |
Claude Opus 5, xhigh | 58 | 42.33 seconds | 50.99 seconds |
Claude Opus 5, max | 55 | 78.91 seconds | 87.93 seconds |
| Gemini 3.1 Pro Preview | 129 | 42.32 seconds | 46.18 seconds |
The comparison that decides interactive workloads is not the headline pairing. Gemini's 42.32 seconds to first token matches Opus 5 at xhigh almost exactly and is far slower than Opus 5 at low, which answers in 3.32 seconds. If first-response latency is the constraint, Opus 5 at a low effort setting is the faster model of the two, and it scores 51 while doing it. Gemini's advantage is in how fast tokens arrive once they start, not in how quickly they start.
Artificial Analysis also flags Opus 5 at max effort as "very verbose," consuming about 100 million output tokens across the index run. Verbosity is invisible in the index score and fully visible on an invoice, and it is one reason Opus 5 costs $2.03 per task at max against Gemini's $0.29.
How do the published specifications compare?
The two are matched on input context at roughly one million tokens and separated on nearly everything else. Claude Opus 5 doubles Gemini's maximum output, 128,000 tokens against 65,536, and carries a knowledge cutoff about sixteen months fresher: May 2026 against January 2025. Gemini accepts audio and video, which Opus 5 does not accept at any price.
| Specification | Claude Opus 5 | Gemini 3.1 Pro Preview |
|---|---|---|
| Model identifier | claude-opus-5 | gemini-3.1-pro-preview |
| Launch stage | Production model, no preview label | "Public preview" |
| Release date | July 24, 2026 | February 19, 2026 |
| Input context window | 1,000,000 tokens | 1,048,576 tokens |
| Maximum output | 128,000 tokens synchronously | 65,536 tokens |
| Knowledge cutoff | May 2026 | January 2025 |
| Input modalities | Text, image, PDF and plain-text documents | Text, image, video, audio, PDF |
| Reasoning parameter | output_config.effort, five levels | thinking_level, three levels |
| Default reasoning setting | high on the Claude API and Claude Code | high |
| Long-context surcharge | none at any length | above 200,000 input tokens |
| Realtime bidirectional API | not offered | Live API not supported on this model |
| Priority service tier | not supported on Opus 5 | supported |
The knowledge cutoff is the one specification where the gap is large, one-directional, and unaffected by configuration. Google publishes "Jan 2025" for this model in its Gemini 3 developer guide, and states in the same document that "Gemini 3 models have a knowledge cutoff of January 2025." A trap sits next to it: the model card shows "Latest update: February 2026," which is the release date, not the cutoff. The card publishes no cutoff row at all.
Two capability gaps run in opposite directions and are easy to miss. On Google's side, the Live API is not supported on Gemini 3.1 Pro Preview, so bidirectional realtime streaming requires dropping to the Flash tier. On Anthropic's side, web fetch is not available on Claude Opus 5 and the Priority Tier is not supported on it either, both of which Opus 4.8 offers.
The 300,000-token output figure sometimes quoted for Opus 5 needs its conditions attached. It requires the output-300k-2026-03-24 beta header, works only on the Message Batches API rather than synchronous calls, and is available on the Claude API and Claude Platform on AWS but not on Amazon Bedrock, Google Cloud, or Microsoft Foundry.
Which Gemini model is this, exactly?
This comparison is about gemini-3.1-pro-preview, Google's top reasoning model, scored at 46 on Intelligence Index v4.1. Google's catalog contains several similarly named entries carrying different numbers, and one widely cited name does not exist at all. Mixing them up is the most common factual error in comparisons of this kind.
There is no Gemini 3.5 Pro. Google shipped Gemini 3.6 Flash and Gemini 3.5 Flash-Lite in July 2026 without ever releasing a 3.5 Pro flagship. Any benchmark figure attributed to "Gemini 3.5 Pro" belongs to a model that was never released and should be discarded rather than reassigned to a nearby variant.
- Gemini 3.1 Pro Preview — the subject of this comparison. Index 46, $0.29 per task, $2.00 and $12.00 per million tokens below 200,000, "Public preview."
- Gemini 3 Pro Preview — the predecessor, shut down March 9, 2026. Its identifier
gemini-3-pro-previewnow resolves togemini-3.1-pro-previewas an alias. - Gemini 3 Pro Image — a generally available image generation model, not a reasoning model. Its GA status does not transfer to the Pro reasoning tier.
- Gemini 3.5 Flash and Gemini 3.6 Flash — cheaper, faster, generally available Flash-tier models with different scores and a four-level
thinking_levelladder that includesminimal. - Gemini 2.5 Pro — the older Pro-tier entry listed as stable, the only non-preview Pro option Google currently offers, and scheduled for shutdown on October 16, 2026 with
gemini-3.1-pro-previewnamed as its replacement.
One documentation inconsistency is worth knowing about before someone cites it against this page. Google's Gemini 3 developer guide still says "All Gemini 3 models are currently in preview," which stopped being true when the Flash models went GA. The current authorities are the models index and the Gemini Enterprise Agent Platform launch-stage block, both cited below.
Who wins each category?
| Category | Winner | Margin |
|---|---|---|
| Peak measured intelligence | Claude Opus 5 | 61 against 46, fifteen points on Index v4.1 |
| Intelligence at the cheapest configuration | Claude Opus 5 | 51 at low against a single published 46 |
| Reasoning headroom | Claude Opus 5 | four measured rungs above its own floor; Gemini's ladder ends at high |
| Cost per task | Gemini 3.1 Pro Preview | $0.29 against $0.36, about 19 percent cheaper |
| Price per million tokens | Gemini 3.1 Pro Preview | $2.00 and $12.00 against $5.00 and $25.00 below 200k |
| Long-context price predictability | Claude Opus 5 | flat at any length; Gemini reprices the whole request above 200k |
| Output throughput | Gemini 3.1 Pro Preview | about 129 tokens per second against 55 |
| Fastest first token | Claude Opus 5 | 3.32 seconds at low against 42.32 seconds |
| Factual reliability | Gemini 3.1 Pro Preview | 33 against 31 on AA-Omniscience |
| Input modality breadth | Gemini 3.1 Pro Preview | adds video and audio, which Opus 5 does not accept |
| Maximum output length | Claude Opus 5 | 128,000 tokens against 65,536 |
| Knowledge freshness | Claude Opus 5 | May 2026 against January 2025, about sixteen months |
| Release stability | Claude Opus 5 | production model against "Public preview" with two weeks' deprecation notice |
That is eight categories to five, which is a real but narrower margin than fifteen index points implies. The split is not random. Opus 5 wins the categories describing how capable the model is, how fresh it is, and how predictable it is to depend on. Gemini wins the categories describing what it costs, how fast tokens arrive, what data it can read, and how often it avoids inventing an answer.
What are the strengths and weaknesses of each?
Claude Opus 5
Strengths. Highest measured intelligence of the pair at 61, fifteen points clear, with a floor of 51 that still sits five points above Gemini's only published score. Five reasoning levels spanning a 5.6-fold cost range, so one integration serves both cheap and expensive work. One million tokens of context at a flat rate with no length surcharge and no beta header. A May 2026 knowledge cutoff, about sixteen months fresher. Maximum output of 128,000 tokens, roughly double Gemini's. First-token latency of 3.32 seconds at low, an order of magnitude faster to respond than Gemini. A production model with no preview label attached to it.
Weaknesses. Loses the factual-reliability head-to-head at 31 against 33 on AA-Omniscience, and cannot close it at any effort level. Roughly 2.5 times Gemini's input price and about 2.08 times its output price below 200,000 tokens. About 19 percent more expensive per task at its cheapest rung. Under half Gemini's throughput. No audio or video input at all. Flagged "very verbose" by the evaluator, consuming about 100 million output tokens across the index run. Web fetch and the Priority Tier are both unsupported, though Opus 4.8 offers them. Several capabilities remain pre-release, including Fast mode and 300,000-token output.
Gemini 3.1 Pro Preview
Strengths. Cheaper on every published axis below the threshold: $0.29 per task, $2.00 and $12.00 per million tokens, $1.74 blended. Better factual reliability at 33 on AA-Omniscience, ahead of a model that leads it by fifteen intelligence points. About 2.3 times the output throughput. Native text, image, video, audio, and PDF input. A 1,048,576-token context window. Broad feature support including code execution, Search and Maps grounding, URL context, structured outputs, and Batch, Flex, and Priority service tiers. Production and commercial use explicitly permitted under the Pre-GA Offerings Terms.
Weaknesses. Still labeled "Public preview" 159 days after release, contractually a Pre-GA Offering provided "as is" with no SLA and no indemnity, and discontinuable without prior notice — a risk Google has already realized once on this product line, retiring the predecessor after 111 days. A single published index score of 46, fifteen points behind, with no higher thinking_level to buy more capability. A long-context surcharge that reprices the entire request, output included, above 200,000 input tokens. A January 2025 knowledge cutoff, about sixteen months behind. Maximum output of 65,536 tokens. Time to first token of 42.32 seconds. No free tier on any serving tier. Live API unsupported. Rate limits unpublished and explicitly "more restricted for experimental and preview models."
When should you pick each one?
When to pick Claude Opus 5
Choose Opus 5 when the work is reasoning-shaped and a wrong answer costs more than the inference: agentic coding loops, complex refactors, multi-step research. Choose it when you need capability above 46, because Gemini has no configuration that buys more. Choose it for document-heavy pipelines that routinely exceed 200,000 tokens, where flat pricing beats a rate that steps up and drags output with it. Choose it when the answer must reflect the last eighteen months, since May 2026 against January 2025 is the difference between knowing a library, a regulation, or a competitor exists and not knowing. And choose it when an SLA and a predictable deprecation path are procurement requirements rather than preferences.
When to pick Gemini 3.1 Pro Preview
Choose Gemini when the workload is high-volume and cost-sensitive and index 46 clears your bar, because it is cheaper per task and roughly 2.5 times cheaper per input token below the threshold. Choose it when factual recall matters more than reasoning depth, since it wins that measurement outright. Choose it when you need audio or video input, which Opus 5 cannot accept, or when sustained throughput decides the experience and 129 tokens per second against 55 is the deciding number. Choose it for retrieval-shaped work where the January 2025 cutoff is irrelevant because facts arrive in the prompt, and where Google's Search grounding covers the rest.
When the comparison does not apply
If your governance rules forbid building on pre-GA services, this is a one-model comparison and Gemini is out regardless of its numbers. If you need a stable Google Pro-tier model specifically, the only option is Gemini 2.5 Pro, which is older, not covered here, and carries a published shutdown date of October 16, 2026 — so it is a decision you will have to make again. If your workload runs acceptably well below index 46, the Flash tier is dramatically cheaper than either model and both are overspecified. And if factual reliability is the single requirement, neither wins: Claude Fable 5 leads AA-Omniscience at 40, ahead of both.
Final verdict
Claude Opus 5 wins this comparison on capability and on dependability, and the second half of that sentence carries more weight than the first. Fifteen index points is a wide margin, and Opus 5's cheapest configuration still scores five points above the only figure Artificial Analysis publishes for Gemini 3.1 Pro Preview. But the argument that should decide a production commitment is the one printed in Google's own contract: the flagship is a Pre-GA Offering, provided "as is," not covered by any SLA or indemnity, and open to being "changed, suspended or discontinued at any time without prior notice," 159 days into a preview with no announced general availability date.
Gemini's case is genuinely strong and this page has not tried to weaken it. It costs less per task, less per token, and streams tokens more than twice as fast. It beats Opus 5 on the one benchmark here that explicitly penalizes hallucination, which is an uncomfortable result for a model leading the intelligence index by fifteen points and one we are not going to bury. It reads audio and video that Opus 5 cannot open. For a high-volume pipeline that clears its bar at index 46, Gemini is the rational choice on economics alone.
What that pipeline is buying, though, is a model that cannot go higher and might not stay. Gemini's thinking_level ladder ends at high, so there is no setting that buys past 46 at any price, and Google has already shut down this model's direct predecessor after 111 days and repointed its identifier. Opus 5 has four measured rungs above Gemini's score and no preview label.
The uncomfortable part is that Google currently offers no way out of this within its own Pro tier. Gemini 2.5 Pro, the stable alternative, is scheduled for shutdown on October 16, 2026, and the replacement Google names for it is the preview model. Teams that can absorb a migration on short notice, and whose work fits under index 46, should take Gemini and the savings — they are real and substantial. Teams whose governance requires an SLA on the model serving production traffic do not have a Google option in this tier right now, and that, more than fifteen index points, is what decides the comparison.
Frequently asked questions
Is Gemini 3.1 Pro Preview still in preview in July 2026?
Yes. Google's Gemini Enterprise Agent Platform documentation lists gemini-3.1-pro-preview with "Launch stage: Public preview" and a release date of February 19, 2026, which is five months and nine days before July 28, 2026. On the Gemini API models page it appears under the Preview heading while Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.1 Flash-Lite appear under Stable. Google has published no general availability date.
Can Gemini 3.1 Pro Preview be used in production?
Yes, Google explicitly permits it, but without an SLA. The Generative AI Preview banner places the model under the Pre-GA Offerings Terms and states that customers "may elect to use it for production or commercial purposes, or disclose Generated Output to third-parties, and may process personal data as outlined in the Cloud Data Processing Addendum." Google's Service Specific Terms separately state that Pre-GA Offerings "are not covered by any SLA or Google indemnity." The restriction is on the guarantees, not on the permission.
How much notice does Google give before retiring a preview model?
Google's documentation promises "at least 2 weeks notice," but its Service Specific Terms reserve the right to change, suspend or discontinue a Pre-GA Offering "at any time without prior notice to Customer." In practice, Gemini 3 Pro Preview was announced for deprecation on February 26, 2026 and shut down on March 9, eleven days later. Its identifier was then repointed to gemini-3.1-pro-preview, so calls kept succeeding against a different model.
Does Gemini 3.1 Pro Preview charge more for long prompts?
Yes, and the surcharge applies to the whole request. Input rises from $2.00 to $4.00 per million tokens and output from $12.00 to $18.00 above 200,000 input tokens. Google's Vertex AI pricing footnote states: "If a query input context is longer than 200K tokens, all tokens (input and output) are charged at long context rates." This is not marginal pricing — a 250,000-token prompt bills all 250,000 tokens at the higher rate.
Does Claude Opus 5 charge more for long context?
No. Anthropic states that "Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing," adding that "a 900k-token request is billed at the same per-token rate as a 9k-token request." Prompt caching and batch discounts apply at standard rates across the full window, and the one-million-token context requires no beta header. Claude Opus 5 falls inside that model family.
Which model scores higher on the Artificial Analysis Intelligence Index?
Claude Opus 5, by fifteen points. It scores 61 at max effort on Intelligence Index v4.1 against 46 for Gemini 3.1 Pro Preview. Opus 5 also scores 51 at its cheapest effort level, five points above Gemini's only published score. Artificial Analysis lists five rows for Opus 5, one per effort level, and a single unlabeled row for Gemini 3.1 Pro Preview, so no per-thinking-level breakdown is published for Google's model.
At which thinking level was Gemini 3.1 Pro Preview's score of 46 measured?
Artificial Analysis does not disclose it. The leaderboard row carries no effort suffix, unlike its Gemini 3.5 Flash rows which are labeled medium and minimal. Google documents the default for this model as high, which is also the top of its three-level ladder, so 46 most likely reflects a default-configuration run at or near Gemini's ceiling. We report that as an inference rather than a measurement, because the evaluator does not state it.
Which model hallucinates less?
Gemini 3.1 Pro Preview, by two points. It scores 33 on AA-Omniscience against 31 for Claude Opus 5 at max effort. Artificial Analysis describes the benchmark as one that "rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer," on a scale from -100 to 100. Gemini answers 3,315 of 6,000 questions correctly against 3,252, but also declines 1,238 against 801, so part of its lead comes from abstaining more often.
Which reasoning levels does each model support?
Claude Opus 5 exposes five effort levels through output_config.effort: low, medium, high, xhigh, and max, defaulting to high on the Claude API and Claude Code. Gemini 3.1 Pro Preview exposes three thinking levels through thinking_level: low, medium, and high, defaulting to high. This model does not accept minimal, unlike Gemini 3.6 Flash and Gemini 3.5 Flash. The legacy numeric thinking_budget remains supported for backward compatibility but cannot be combined with thinking_level.
Which model is faster?
It depends on which kind of speed. Gemini 3.1 Pro Preview streams about 129 output tokens per second against 55 for Claude Opus 5 at max effort, roughly 2.3 times faster. But its time to first token is 42.32 seconds, while Opus 5 at low effort answers in 3.32 seconds. Gemini wins sustained throughput; Opus 5 at a low effort setting wins first-response latency by a wide margin. Both figures were measured on July 28, 2026.
Is there a Gemini 3.5 Pro?
No. Google has not released a model called Gemini 3.5 Pro. The current top reasoning entry is Gemini 3.1 Pro Preview, and the nearby models that do exist are Gemini 3.5 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, Gemini 3 Pro Image, and Gemini 2.5 Pro. Any benchmark figure attributed to "Gemini 3.5 Pro" belongs to a model that was never released and should be discarded rather than reassigned to a nearby variant.
How do the context windows, output limits, and knowledge cutoffs compare?
Claude Opus 5 has a one-million-token context window, a 128,000-token synchronous output limit, and a May 2026 knowledge cutoff. Gemini 3.1 Pro Preview has a 1,048,576-token input limit, a 65,536-token output limit, and a January 2025 knowledge cutoff, about sixteen months older. Note that Gemini's model card shows "Latest update: February 2026," which is its release date rather than its cutoff; the card publishes no cutoff row at all.
Sources and references
Every figure on this page comes from a vendor's own documentation or from an independent evaluator, and the two are labeled separately throughout rather than stacked into a single ranking. Google's published benchmark scores for Gemini 3.1 Pro Preview are vendor self-reported and are not used for any head-to-head claim here; neither are Anthropic's, which compare Claude Opus 5 only to other Anthropic models. Cost-per-task, throughput, and latency figures are sliding measurements read on July 28, 2026.
- Anthropic — Introducing Claude Opus 5 (vendor: release date, token pricing, availability wording)
- Anthropic — Pricing (vendor: token rates, caching, batch, long-context pricing, geography multiplier)
- Anthropic — Models overview (vendor: context window, maximum output, knowledge cutoffs, modalities, availability labels)
- Anthropic — Effort (vendor: five effort levels, default high, thinking constraints)
- Anthropic — Context windows (vendor: 1M default, no beta header, standard pricing)
- Anthropic — Batch processing (vendor: 300,000-token extended output beta and its conditions)
- Anthropic — Model migration guide (vendor: model identifier scheme, feature parity with Opus 4.8)
- Google — Gemini API models (vendor: Preview and Stable groupings, two-week deprecation notice)
- Google — Gemini 3.1 Pro Preview model card (vendor: model ID, token limits, modalities, feature support)
- Google — Gemini 3 developer guide (vendor: knowledge cutoff, thinking level defaults, thinking_budget back-compatibility)
- Google — Gemini thinking (vendor: per-model thinking_level support table)
- Google — Gemini API pricing (vendor: standard, batch, flex and priority rates, free tier availability, grounding)
- Google Cloud — Generative AI pricing (vendor: long-context threshold mechanic applying to all tokens)
- Google Cloud — Gemini 3.1 Pro model page (vendor: launch stage, release date, Pre-GA Offerings Terms banner)
- Google Cloud — Service Specific Terms (vendor: Pre-GA Offerings Terms, SLA and indemnity exclusion, liability cap, discontinuation without notice)
- Google — Gemini API deprecations (vendor: release and shutdown dates, recommended replacements, preview and stable runways)
- Google — Gemini API changelog (vendor: Gemini 3 Pro Preview shutdown, Flash general availability entries)
- Google — Gemini API rate limits (vendor: preview rate-limit caveat, batch enqueued token quotas)
- Google DeepMind — Gemini Pro (vendor self-reported benchmark scores, measured at thinking level high)
- Artificial Analysis — Model leaderboard (independent: index scores, cost per task, throughput, time to first token)
- Artificial Analysis — Claude Opus 5 (independent: index version, blended price, per-configuration measurements)
- Artificial Analysis — Gemini 3.1 Pro Preview (independent: index version, blended price, single measured configuration)
- Artificial Analysis — Intelligence benchmarking methodology (independent: Index v4.1 composition, cost-per-task definition)
- Artificial Analysis — AA-Omniscience (independent: scale definition, scores, accuracy and attempt rates)
Related comparisons on ThePlanetTools: Claude Opus 4.8 against Gemini 3.1 Pro, Claude Sonnet 5 against Gemini 3.1 Pro, Claude Fable 5 against Gemini 3.1 Pro, and GPT-5.6 Sol against Gemini 3.1 Pro. Background reading: our report on the Claude Opus 5 launch.
Our Verdict
Claude Opus 5 wins, on capability and on dependability. It scores 61 on Artificial Analysis Intelligence Index v4.1 at max effort against 46 for Gemini 3.1 Pro Preview, and its cheapest configuration still scores 51, five points above the only figure published for Google’s model. Gemini’s case is real and this page does not weaken it: it costs $0.29 per task against $0.36, streams about 129 output tokens per second against 55, beats Opus 5 on factual reliability by 33 to 31 on AA-Omniscience, and reads audio and video that Opus 5 cannot accept. What decides it is the label. Gemini 3.1 Pro Preview has been a "Public preview" offering for 159 days with no announced general availability date; Google’s Service Specific Terms exclude Pre-GA offerings from any SLA or indemnity and allow discontinuation without prior notice, and Google retired this model’s direct predecessor after 111 days. Google also charges a long-context surcharge that reprices every token in a request, output included, above 200,000 input tokens, while Anthropic bills its full one-million-token window flat. Teams that can absorb a short-notice migration and fit under index 46 should take Gemini and the savings; teams that need an SLA on the model serving production traffic have no Google option in this tier today.
Choose Claude Opus 5
Anthropic's frontier reasoning model — top of the independent index at half the price of Fable 5.
Try Claude Opus 5 →Choose Gemini 3.1 Pro Preview
Google DeepMind's flagship Gemini 3.1 Pro Preview — 94.3% GPQA Diamond, 77.1% ARC-AGI-2, 1M-token context, multimodal in/text out, vibe coding plus agentic tool use. Preview status as of April 2026.
Try Gemini 3.1 Pro Preview →Frequently Asked Questions
Is Claude Opus 5 better than Gemini 3.1 Pro Preview?
Claude Opus 5 wins, on capability and on dependability. It scores 61 on Artificial Analysis Intelligence Index v4.1 at max effort against 46 for Gemini 3.1 Pro Preview, and its cheapest configuration still scores 51, five points above the only figure published for Google’s model. Gemini’s case is real and this page does not weaken it: it costs $0.29 per task against $0.36, streams about 129 output tokens per second against 55, beats Opus 5 on factual reliability by 33 to 31 on AA-Omniscience, and reads audio and video that Opus 5 cannot accept. What decides it is the label. Gemini 3.1 Pro Preview has been a "Public preview" offering for 159 days with no announced general availability date; Google’s Service Specific Terms exclude Pre-GA offerings from any SLA or indemnity and allow discontinuation without prior notice, and Google retired this model’s direct predecessor after 111 days. Google also charges a long-context surcharge that reprices every token in a request, output included, above 200,000 input tokens, while Anthropic bills its full one-million-token window flat. Teams that can absorb a short-notice migration and fit under index 46 should take Gemini and the savings; teams that need an SLA on the model serving production traffic have no Google option in this tier today.
Which is cheaper, Claude Opus 5 or Gemini 3.1 Pro Preview?
Claude Opus 5 is priced at $5 in / $25 out per M tokens. Gemini 3.1 Pro Preview is priced at $2 in / $12 out per M tokens. Check the pricing comparison section above for a full breakdown.
What are the main differences between Claude Opus 5 and Gemini 3.1 Pro Preview?
The key differences span across 13 features we compared. For Artificial Analysis Intelligence Index v4.1, top measured configuration, Claude Opus 5 offers 61 (max effort) while Gemini 3.1 Pro Preview offers 46 (default configuration, level not disclosed). For Index at the cheapest measured configuration, Claude Opus 5 offers 51 (low effort), $0.36 per task while Gemini 3.1 Pro Preview offers 46, $0.29 per task — only row published. For Cost per task, measured July 28, 2026, Claude Opus 5 offers $0.36 at low effort, $2.03 at max while Gemini 3.1 Pro Preview offers $0.29. See the full feature comparison table above for all details.

