Claude Opus 5 vs Gemini 3.5 Flash: Cheaper Per Token, Costlier Per Task (2026)
Gemini 3.5 Flash scores 50 at its top setting for $0.59 per task. Claude Opus 5 scores 51 at its cheapest for $0.36. Measured July 28, 2026.
Feature Comparison
| Feature | Claude Opus 5 | Gemini 3.5 Flash |
|---|---|---|
| Artificial Analysis Intelligence Index v4.1, top measured configuration | 61 (max effort) | 50 (high thinking level) |
| Index at the vendor default configuration | 59 (high effort) | 45 (medium thinking level) |
| Index at the bottom of the vendor ladder | 51 (low effort) | 35 (minimal thinking level) |
| Cost per task near index 50 (measured July 28, 2026) | $0.36 (low effort, index 51) | $0.59 (high, index 50) |
| Price per million input and output tokens | $5.00 and $25.00 | $1.50 and $9.00 |
| Blended price per million tokens (Artificial Analysis) | $3.85 | $1.31 |
| Output throughput (measured July 28, 2026) | About 55 tokens per second | About 199 tokens per second |
| Fastest time to first token (measured July 28, 2026) | 3.32 seconds at low effort | 0.92 seconds at minimal |
| Input modalities | Text, image, PDF and plain-text documents | Text, image, video, audio, PDF |
| Input context window | 1,000,000 tokens, flat pricing | 1,048,576 tokens, flat pricing |
| Maximum output tokens | 128,000 synchronously; 300,000 in batch behind a beta header | 65,536 |
| Knowledge cutoff | May 2026 | January 2025 |
| AA-Omniscience score | 31 (max effort) | Not published |
Pricing Comparison
Claude Opus 5
Gemini 3.5 Flash
Detailed Comparison
Claude Opus 5 scores 61 on the Artificial Analysis Intelligence Index v4.1 at max effort, eleven points above Gemini 3.5 Flash at 50. The more useful number is the cheap end: Opus 5 also scores 51 at its lowest effort setting for $0.36 per task, while Gemini 3.5 Flash reaches 50 for $0.59 per task. Google's model is therefore one point lower and roughly 64 percent more expensive per task than Anthropic's floor. Per token the picture inverts completely: Flash costs $1.50 and $9.00 per million input and output tokens against $5.00 and $25.00 for Opus 5. We researched both from vendor documentation and independent evaluations; all index figures are v4.1 and all cost-per-task figures were read on July 28, 2026.
Which model should you choose?
Pick Claude Opus 5 if your work is reasoning-heavy, if you need anything above index 51, or if you want the cheapest route to roughly index 50 on a per-task basis. Pick Gemini 3.5 Flash if you need audio or video input, roughly three and a half times the throughput, or if your workload is dominated by large inputs and short outputs rather than by reasoning.
The headline gap is eleven points, and that gap is real. But the number that decides most procurement questions is narrower and points the same way. Artificial Analysis measures Gemini 3.5 Flash at 50 at high, the top of the four-level ladder Google documents. It measures Claude Opus 5 at 51 at low, the bottom of Anthropic's five-level ladder. The same neighborhood on the index means opposite things: for Google it is a ceiling, for Anthropic it is a floor.
That asymmetry is not a rhetorical framing, it is the whole comparison. Above 51, Gemini 3.5 Flash has no measured configuration at any price, and Opus 5 has four more rungs to climb. At roughly 50, Opus 5 is the cheaper of the two per task. Below 50, Flash has two measured configurations and Opus 5 has none, and that is the region where Google's model earns its place.
Where Flash wins, it wins on axes the intelligence index does not score at all: throughput, time to first token at its lowest setting, native audio and video input, and raw token price on input-heavy jobs. Those are not consolation prizes. They decide a large class of production workloads, and we treat them as their own categories below rather than folding them into a single score.
What does each model's full effort scale look like?
Artificial Analysis publishes one leaderboard row per reasoning setting, not one row per model. Claude Opus 5 appears five times, spanning 51 to 61 and $0.36 to $2.03 per task. Gemini 3.5 Flash appears three times, spanning 35 to 50. Reading only each model's headline row hides the fact that both vendors sell several different price-performance points under one product name.
The two ladders are controlled by different parameters and the levels are not translatable. Anthropic exposes effort nested inside output_config, with a closed set of five values: low, medium, high, xhigh, and max. There is no minimal level on Opus 5. Google exposes thinking_level on Gemini 3.5 Flash, with four values: minimal, low, medium, and high. The word high appears on both ladders and means nothing comparable across them.
Two clarifications on Google's side, because this parameter changed shape recently. Named levels are current: earlier Gemini models used a numeric thinking_budget, which Google says is "still supported for backward compatibility" while recommending migration to thinking_level, and which must not be sent alongside it in the same request. And Artificial Analysis labels its Gemini rows high, medium, and minimal, which are Google's own documented values rather than labels the evaluator invented.
| Configuration | Claude Opus 5 | Gemini 3.5 Flash |
|---|---|---|
| Top of the vendor's ladder | max — 61, $2.03 per task | high — 50, $0.59 per task |
| Second rung | xhigh — 60, $1.56 per task | none above high |
| Vendor default | high — 59, $1.06 per task | medium — 45, no cost published |
| Middle rung | medium — 56, $0.62 per task | low — not measured |
| Bottom of the ladder | low — 51, $0.36 per task | minimal — 35, no cost published |
Three different kinds of blank appear in that table and they should not be read the same way. Gemini 3.5 Flash has nothing above high because Google's parameter does not define a higher level: that level does not exist. Google's low level exists and is documented, but Artificial Analysis has not measured it, so no score is published. And the scores at medium and minimal carry an asterisk on the leaderboard with no cost-per-task figure beside them; Artificial Analysis does not publish a legend defining that asterisk on the pages we checked, so we report the marking and decline to interpret it.
One consequence deserves its own sentence, because default-to-default is how most teams will actually experience these models. Anthropic's default effort on the Claude API and Claude Code is high, which scores 59. Google's default thinking level on Gemini 3.5 Flash is medium, which scores 45. Out of the box, without either team touching a reasoning parameter, the measured gap is fourteen points, not eleven.
Is Gemini 3.5 Flash strictly dominated by Claude Opus 5?
On the two axes Artificial Analysis publishes together, yes. Gemini 3.5 Flash at high scores 50 at $0.59 per task. Claude Opus 5 at low scores 51 at $0.36 per task. Opus 5 is one point higher and roughly 39 percent cheaper per task at that comparison. Flash costs about 64 percent more to reach a score one point lower. That is strict domination on intelligence per unit of task cost.
This is worth stating plainly because it is unusual. In the comparable matchups we have researched recently, the competing model bought its position with price: it scored below Opus 5's floor but undercut it per task, which is a defensible trade. Gemini 3.5 Flash does not make that trade on this measurement. It lands just under the floor and above the price.
The mechanism is visible in Google's own pricing note. Gemini 3.5 Flash bills output at $9.00 per million tokens, and Google states that this rate is charged "including thinking tokens." Cost per task is not a list price, it is a measurement of tokens actually consumed across a workload multiplied by the prices that apply to them. Reaching index 50 evidently requires enough thinking for Flash's cheaper per-token rate not to translate into a cheaper task. Opus 5 at low thinks less and finishes cheaper despite list prices roughly three times higher.
Two honest qualifications. First, domination holds on Artificial Analysis's workload, which is reasoning-weighted by construction: version 4.1 of the index aggregates nine evaluations including Humanity's Last Exam, GPQA Diamond, Terminal-Bench v2.1, SciCode, and AA-LCR. A workload with a different shape can invert the result, and the next two sections describe exactly when it does. Second, domination on this axis says nothing about throughput, latency, or input modalities, where Flash is straightforwardly the better model.
How do the two token prices compare?
Per token, Gemini 3.5 Flash is substantially cheaper, and this is the fact that makes the cost-per-task result counterintuitive. Flash costs $1.50 per million input tokens and $9.00 per million output tokens. Claude Opus 5 costs $5.00 per million input tokens and $25.00 per million output tokens. That is roughly 3.3 times more for input and 2.8 times more for output.
| Price component | Claude Opus 5 | Gemini 3.5 Flash |
|---|---|---|
| Input, per million tokens | $5.00 | $1.50 |
| Output, per million tokens | $25.00 | $9.00, stated as including thinking tokens |
| Cached input read, per million tokens | $0.50 | $0.15 |
| Cache storage | Write at $6.25 for five minutes or $10.00 for one hour | $1.00 per million tokens per hour |
| Batch discount | 50 percent, giving $2.50 and $12.50 | Batch tier published on Google's pricing page |
| Blended rate published by Artificial Analysis | $3.85 per million tokens | $1.31 per million tokens |
The blended figures come from Artificial Analysis, which weights them seven parts cached input, two parts fresh input, and one part output. They are a pricing convention for comparing list prices on a single number, not a measurement of any workload. The $3.85 for Opus 5 and the $1.31 for Flash are therefore not comparable to the $2.03 and $0.59 cost-per-task figures, which measure real token consumption on the index. Confusing the two is the single easiest mistake to make on this page.
One further wrinkle cuts against Opus 5 and belongs in any per-token comparison. Anthropic notes that models from Claude 4.7 onward use a newer tokenizer that "produces approximately 30% more tokens for the same text," with the exact increase depending on content and workload shape. The same document therefore costs more than the headline ratio suggests when it is sent to Opus 5. Cost per task already absorbs this effect, because it counts tokens actually consumed; the list-price ratio does not.
Both vendors also charge a premium for pinning inference to a geography, and the premium is the same size on each. Anthropic applies a 1.1 times multiplier when a request sets US-only inference, and the same 10 percent premium applies to regional and multi-region endpoints on partner clouds. Google's Vertex AI pricing lists Gemini 3.5 Flash at $1.65 and $9.90 per million input and output tokens for non-global endpoints against $1.50 and $9.00 for global.
Does either model charge more for long context?
Neither does. Both models carry a context window of roughly one million tokens and both bill it at a flat rate with no length-based surcharge. This is worth checking rather than assuming, because Google does apply a long-context surcharge to other models in the same family, and because a surcharge at this scale would dominate the economics of any document-heavy workload.
Anthropic states the position directly: "Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing," adding that "a 900k-token request is billed at the same per-token rate as a 9k-token request," and that caching and batch discounts apply at standard rates across the full window. A separate page confirms that one million tokens is the default and that "you don't need a beta header."
Google's position for this specific model is the same, and the Vertex AI pricing table makes it explicit by publishing both tiers with identical values: Gemini 3.5 Flash is listed at $1.50 per million tokens for prompts at or below 200,000 input tokens and $1.50 per million tokens for prompts above 200,000. The surcharge column exists and the model does not use it.
Other Gemini models on the same page do use it. Gemini 3.1 Pro Preview is listed at $2.00 per million input tokens at or below 200,000 tokens and "$4 per 1M tokens" above 200,000. Gemini 2.5 Pro is listed at $1.25 below the threshold and "$2.50 per 1M tokens" above it. So the doubling above 200,000 tokens is a real and current Google practice; it simply does not apply to Gemini 3.5 Flash. Anyone carrying an assumption over from Gemini 2.5 Pro or 3.1 Pro Preview should drop it here.
The practical effect is that long-context economics do not separate these two models structurally. They separate them by rate: a 500,000-token prompt costs about $0.75 sent to Flash and about $2.50 sent to Opus 5, before the tokenizer difference, and neither figure carries a penalty for length.
What does the intelligence index fail to measure?
Throughput and latency, and on both Gemini 3.5 Flash is the faster model by a wide margin. Artificial Analysis measures Flash at high at roughly 199 output tokens per second against roughly 54 for Opus 5 at low and 55 at max, a gap of about three and a half times. These are sliding measurements taken from live endpoints and we read them on July 28, 2026; they move as providers tune capacity.
| Configuration (measured July 28, 2026) | Output tokens per second | Time to first token |
|---|---|---|
Claude Opus 5, low | 54 | 3.32 seconds |
Claude Opus 5, high | 57 | 20.90 seconds |
Claude Opus 5, max | 55 | 78.91 seconds |
Gemini 3.5 Flash, minimal | 170 | 0.92 seconds |
Gemini 3.5 Flash, medium | 207 | 23.05 seconds |
Gemini 3.5 Flash, high | 199 | 28.40 seconds |
Time to first token complicates the simple story, and the complication favors Opus 5 in one place. Flash at minimal answers in under a second, which nothing in Anthropic's lineup approaches. But Flash at high takes 28.40 seconds before its first token, longer than Opus 5 at high at 20.90 seconds. If you run Flash at the setting that produces its 50, you are not getting a fast first response; you are getting fast tokens after a long wait.
Artificial Analysis also notes that Opus 5 is verbose, consuming about 100 million output tokens across the index against a median of about 63 million for comparable models, and describes its time to first token as at the higher end among reasoning models. Opus 5 at max ranks 115th of 191 models on output speed. None of this is captured by the number 61.
For interactive products, this matters more than eleven index points. A support assistant that must respond in under two seconds cannot run Opus 5 at any effort level and cannot run Flash at high either; it runs Flash at minimal, and the relevant score is 35, not 50.
Which model accepts audio and video?
Only Gemini 3.5 Flash. Google's model card lists input modalities as text, image, video, audio, and PDF, with text output. Anthropic's Messages API accepts text, image, and document blocks, where documents are PDF and plain text; there is no audio or video content block in the API schema. For workloads that ingest recordings or footage, this is not a score difference, it is a capability the comparison cannot bridge.
This is the clearest case of the fast model being not merely sufficient but necessary. A meeting-transcription pipeline, a video-moderation queue, or a call-center analytics job can send its media directly to Flash. The same job routed to Opus 5 needs a separate transcription or captioning stage first, which adds a vendor, a failure mode, and a cost line that no index score reflects.
The reverse asymmetry is narrower. Opus 5 accepts up to 600 images or PDF pages in a single request, with a 32 MB request size limit for PDFs, which is a substantial document-ingestion capability. But on modality breadth Google is ahead, and nothing on Anthropic's side offsets audio and video.
How do the published specifications compare?
The two models are closely matched on context and far apart on output length and training freshness. Both offer roughly one million tokens of input. Opus 5 allows nearly double Flash's maximum output synchronously and can go far higher in batch mode. The widest specification gap is the knowledge cutoff: Opus 5 is current to May 2026, Gemini 3.5 Flash to January 2025, a difference of about sixteen months.
| Specification | Claude Opus 5 | Gemini 3.5 Flash |
|---|---|---|
| Model identifier | claude-opus-5 | gemini-3.5-flash |
| Input context window | 1,000,000 tokens | 1,048,576 tokens |
| Maximum output | 128,000 tokens synchronously; 300,000 in batch behind a beta header | 65,536 tokens |
| Knowledge cutoff | May 2026, for both reliable knowledge and training data | January 2025 |
| Input modalities | Text, image, PDF and plain-text documents | Text, image, video, audio, PDF |
| Reasoning parameter | output_config.effort, five levels | thinking_level, four levels |
| Reasoning off by default | No — adaptive thinking is on | No — thinking is on at medium |
| Release | July 24, 2026 | May 19, 2026 |
| Status label | Current; listed outside the legacy group | "generally available (GA), stable" |
The knowledge cutoff deserves emphasis because it is the one specification where the gap is large, one-directional, and unaffected by how either model is configured. Google states in its Gemini 3.5 guide that "Gemini 3.5 Flash has a knowledge cutoff of January 2025," and recommends the Search Grounding tool "for more recent information." Anthropic lists May 2026 for Opus 5 as both the reliable knowledge cutoff and the training data cutoff. Note that Google publishes this date in its release guide rather than on the model card itself, which is where most readers will look for it and not find it.
Two details behind the table are easy to overstate. The 300,000-token output on Opus 5 requires the output-300k-2026-03-24 beta header, works only on the Message Batches API rather than synchronous calls, and is available on the Claude API and Claude Platform on AWS but not on Amazon Bedrock, Google Cloud, or Microsoft Foundry. It is a real capability with three conditions attached, not a plain headline number.
Neither model lets you turn reasoning fully off, but the constraints differ. Opus 5 accepts thinking: {type: "disabled"} at effort high or below; combining it with xhigh or max returns a 400 error. Google's documentation shows thinking on at medium for Gemini 3.5 Flash with minimal as the floor rather than an off switch. Anthropic also hard-blocks sampling parameters on Opus 5: non-default temperature, top_p, or top_k values "return a 400 error on every request."
One classification is worth quoting rather than paraphrasing, because it is the kind of word that carries consequences. Anthropic's documentation places Claude Opus 4.8 and its predecessors inside a group introduced as "The following models are still available. Consider migrating to current models for improved performance," under the heading "Legacy models." Opus 5 sits outside that group. Separately, Claude Opus 4.1 was retired on August 5, 2026: Anthropic's documentation now lists its state as "Retired," requests to it on the Claude API return an error, and Opus 5 is the named migration target.
Which model hallucinates less?
This question cannot be answered as a head-to-head, and the reason is a genuine gap in the public record rather than a close call. Artificial Analysis publishes an AA-Omniscience score of 31 for Claude Opus 5, measured on the adaptive-reasoning configuration at max effort. Gemini 3.5 Flash is absent from the AA-Omniscience leaderboard entirely, and the corresponding sections on its model page render without a value.
AA-Omniscience is described as "a bounded metric (-100 to 100) measuring factual recall that jointly penalizes hallucinations and rewards abstention when uncertain, with 0 equating to a model that answers questions correctly as much as it does incorrectly." It spans 6,000 questions across six domains and 42 topics, and it is one of the nine evaluations aggregated into Intelligence Index v4.1. So it contributes to Flash's 50 without being separately published for it.
Opus 5's 31 is a positive score on that scale, meaning it gets substantially more right than wrong and abstains usefully. It is not, however, a leading score. On the same leaderboard Claude Fable 5 posts 40 and Gemini 3.1 Pro Preview posts 33, both above Opus 5, despite Opus 5 leading the overall intelligence index. Peak intelligence and factual reliability are not the same axis, and Opus 5 illustrates the difference within Anthropic's own lineup.
Anyone needing a hallucination comparison between these two specific models will have to run it themselves. We will not estimate Flash's Omniscience score from its overall index position, and no figure published by either vendor substitutes for the missing measurement.
Which Gemini model is this, exactly?
Gemini 3.5 Flash is gemini-3.5-flash, Google's stable fast-tier model, scored at 50 on Intelligence Index v4.1. Google's lineup contains several similarly named entries that carry different numbers, and mixing them up is the most common factual error in comparisons of this kind. Four distinct models sit close enough in name to be confused, and one widely referenced name does not exist at all.
There is no Gemini 3.5 Pro. Google shipped Gemini 3.6 Flash and other models without releasing a 3.5 Pro flagship, which we covered when it happened in our report on the three new Gemini models that did not include the expected flagship. Any figure attributed to "Gemini 3.5 Pro" is attached to a model that was never released and should be discarded rather than reassigned.
- Gemini 3.5 Flash — the subject of this comparison. Index 50 at
high, $0.59 per task, $1.50 and $9.00 per million tokens. - Gemini 3.5 Flash-Lite — a separate model, not an effort level of Flash. Index 36 at $0.09 per task.
- Gemini 3.6 Flash — a later release in the Flash line, and a different model from 3.5 Flash.
- Gemini 3.1 Pro Preview — the Pro-tier entry, which does carry the long-context surcharge that 3.5 Flash does not.
The Flash-Lite confusion is the one most likely to distort a price argument, because $0.09 per task is genuinely dramatic and belongs to a model scoring 36, fourteen points below Flash. If a workload can live at 36, Flash-Lite is far cheaper than anything in this comparison, and neither model here is the right answer. We covered Flash's own arrival in our write-up of the Gemini 3.5 Flash launch at Google I/O 2026.
Who wins each category?
| Category | Winner | Margin |
|---|---|---|
| Peak measured intelligence | Claude Opus 5 | 61 against 50, eleven points on Index v4.1 |
| Intelligence at the vendor default | Claude Opus 5 | 59 at high against 45 at medium |
| Cost per task near index 50 | Claude Opus 5 | $0.36 against $0.59, about 39 percent cheaper |
| Price per million tokens | Gemini 3.5 Flash | $1.50 and $9.00 against $5.00 and $25.00 |
| Blended list price | Gemini 3.5 Flash | $1.31 against $3.85 per million tokens |
| Output throughput | Gemini 3.5 Flash | About 199 tokens per second against about 55 |
| Fastest first token | Gemini 3.5 Flash | 0.92 seconds at minimal against 3.32 seconds at low |
| Input modality breadth | Gemini 3.5 Flash | Adds video and audio, which Opus 5 does not accept |
| Maximum output length | Claude Opus 5 | 128,000 tokens against 65,536 |
| Knowledge freshness | Claude Opus 5 | May 2026 against January 2025, about sixteen months |
| Reasoning headroom | Claude Opus 5 | Four measured rungs above Flash's ceiling |
| Hallucination resistance | Not determinable | Opus 5 posts 31; Flash has no published score |
Counted crudely that is six categories to five with one undecided, which is a far closer split than eleven index points implies. The count is not the point, though. Opus 5 wins the categories that describe how well the model reasons and what that reasoning costs to complete a task. Flash wins the categories that describe how fast tokens arrive, what they list for, and what kinds of data the model can accept at all.
What are the strengths and weaknesses of each?
Claude Opus 5
Strengths. Highest measured intelligence of the pair at every comparable rung, topping out at 61. Cheaper per task than Flash at comparable measured intelligence, at $0.36 against $0.59. Five reasoning levels spanning a 5.6-fold cost range, so one integration serves both cheap and expensive work. One million tokens of context at flat pricing with no beta header. A knowledge cutoff of May 2026 for both reliable knowledge and training data, about sixteen months fresher than Flash. Maximum output of 128,000 tokens synchronously, nearly double Flash's 65,536.
Weaknesses. Roughly 3.3 times Flash's input price and 2.8 times its output price. Under one third of Flash's throughput. No audio or video input at all. A newer tokenizer that produces roughly 30 percent more tokens for the same text, raising effective cost beyond the headline ratio. Verbose on the index relative to peers. An AA-Omniscience score of 31 that trails both Claude Fable 5 and Gemini 3.1 Pro Preview. Not even Anthropic's most capable model — its own documentation points to Claude Fable 5 for "the highest available capability."
Gemini 3.5 Flash
Strengths. Substantially cheaper per token, at $1.50 and $9.00 per million against $5.00 and $25.00. Roughly three and a half times the output throughput. Sub-second time to first token at minimal. Native text, image, video, audio, and PDF input. A 1,048,576-token context window at flat pricing, with no long-context surcharge despite Google applying one to its Pro-tier models. Cached input at $0.15 per million tokens. Broad distribution across the Gemini API, AI Studio, Vertex AI, and Google's consumer surfaces.
Weaknesses. Its ceiling of 50 sits one point below Opus 5's floor of 51, and it costs about 64 percent more per task to get there. No configuration above high at any price. Its default level scores 45 with no published cost figure. Time to first token of 28.40 seconds at high, worse than Opus 5 at high. Maximum output of 65,536 tokens against Opus 5's 128,000. A January 2025 knowledge cutoff, roughly sixteen months behind Opus 5, published in Google's release guide rather than on the model card. No published AA-Omniscience score, so factual reliability cannot be compared directly.
When should you pick each one?
When to pick Claude Opus 5
Choose Opus 5 when the task is reasoning-shaped and the output is the product: agentic coding loops, multi-step research, complex refactors, or analysis where a wrong answer costs more than the inference. Choose it when you need any measured capability above 51, because Flash has none at any price. Choose it when you want a single integration that can be dialed from $0.36 to $2.03 per task without changing vendors. And choose it when recency matters: its May 2026 cutoff is about sixteen months ahead of Flash's January 2025, which is the difference between knowing about a library, a regulation, or a competitor and not knowing it exists.
When to pick Gemini 3.5 Flash
Choose Flash when the job is input-heavy and output-light, because that is where per-token price dominates and cost per task on a reasoning benchmark stops predicting your bill. Summarizing a 500,000-token corpus, classifying a large document set, or extracting fields from long records all cost roughly a third as much on Flash, and both models bill long context flat. Choose it when you need audio or video input, which Opus 5 cannot accept. Choose it for interactive products where sub-second first tokens or high throughput decide the experience, and where minimal at index 35 is enough. The January 2025 cutoff is less limiting than it sounds for retrieval-shaped work, since the facts arrive in the prompt rather than from training, and Google points to its Search Grounding tool "for more recent information."
When the comparison does not apply
If your workload lives comfortably below index 36, neither model is the economical choice: Gemini 3.5 Flash-Lite runs the index at $0.09 per task, a fraction of both. If you need the highest capability Anthropic sells rather than the best cost-to-intelligence ratio, Anthropic's own guidance points past Opus 5 to Claude Fable 5. And if your requirement is factual reliability specifically rather than reasoning power, neither model's published record settles it, since Flash has no AA-Omniscience score and Opus 5's 31 trails two models we could name.
Final verdict
Claude Opus 5 wins this comparison, and it wins it in an unusual way. The eleven-point headline gap is the least interesting evidence. The decisive finding is that Gemini 3.5 Flash's best measured configuration lands one point below Opus 5's cheapest one while costing about 64 percent more per task, which makes it strictly dominated on the two axes Artificial Analysis publishes side by side. When a model is both lower and more expensive at its own ceiling, the argument for it has to come from somewhere other than the leaderboard.
That argument exists and it is not weak. Flash moves tokens about three and a half times faster, starts answering in under a second at its lowest setting, lists at roughly a third of Opus 5's token price, and accepts audio and video that Opus 5 cannot read at any price. For input-heavy, latency-sensitive, or multimodal work, those properties decide the choice and the index score is close to irrelevant. The fast model does not merely suffice there; it is the only one of the two that can do the job.
What Flash cannot do is go higher. Its ladder stops at high, and Opus 5 has four measured rungs above that point. Teams whose workloads occasionally need more than index 50 will end up integrating a second model anyway, and at that point the cost-per-task advantage of running Opus 5 at low for the easy work makes the two-vendor split hard to justify. Teams whose workloads never exceed 50, and who care about speed, price per token, or media input, should ignore the eleven points and pick Flash without apology.
Frequently asked questions
Is Claude Opus 5 better than Gemini 3.5 Flash?
On measured intelligence, yes. Claude Opus 5 scores 61 on Artificial Analysis Intelligence Index v4.1 at max effort against 50 for Gemini 3.5 Flash at high thinking level. Unusually, Opus 5 also wins on cost at comparable intelligence: it scores 51 at low effort for $0.36 per task, while Flash costs $0.59 to reach 50. Flash remains better on throughput, latency at its lowest setting, token price, and audio and video input.
Is Gemini 3.5 Flash cheaper than Claude Opus 5?
It depends entirely on the unit. Per token Flash is much cheaper, at $1.50 and $9.00 per million input and output tokens against $5.00 and $25.00 for Opus 5. Per task on the Artificial Analysis index, Flash is more expensive at $0.59 against $0.36 for Opus 5 at low effort, because reaching index 50 consumes enough thinking tokens to overwhelm the cheaper rate. Input-heavy work favors Flash; reasoning-heavy work favors Opus 5.
What is the difference between cost per task and blended price per million tokens?
Blended price is a list-price convention. Artificial Analysis weights it seven parts cached input, two parts fresh input, and one part output, giving $3.85 per million tokens for Opus 5 and $1.31 for Gemini 3.5 Flash. Cost per task measures tokens actually consumed running the index workload, giving $2.03 for Opus 5 at max effort and $0.59 for Flash. The two metrics rank the models in opposite orders, which is why they must never be mixed.
Which reasoning levels does each model support?
Claude Opus 5 exposes five effort levels through output_config.effort: low, medium, high, xhigh, and max. There is no minimal level. Gemini 3.5 Flash exposes four thinking levels through thinking_level: minimal, low, medium, and high. The ladders are controlled by different parameters and the shared word "high" does not mean the same thing on both. Opus 5 defaults to high on the Claude API and Claude Code; Flash defaults to medium.
Why does Gemini 3.5 Flash have fewer scores on the leaderboard?
Artificial Analysis publishes three of Flash's four levels: high at 50, medium at 45, and minimal at 35. Google's low level is documented but has not been measured, so no score exists for it. The medium and minimal scores carry an asterisk on the leaderboard and have no cost-per-task figure beside them. Artificial Analysis does not publish a legend defining that asterisk on the pages we checked, so we report the marking without interpreting it.
Does either model charge extra for long context?
No. Anthropic states that Claude 4.6 and later models include the full one-million-token context window at standard pricing, and that a 900,000-token request bills at the same per-token rate as a 9,000-token request. Google's Vertex AI pricing lists Gemini 3.5 Flash at $1.50 per million tokens both at or below 200,000 input tokens and above it. Other Gemini models do carry the surcharge: Gemini 3.1 Pro Preview doubles from $2.00 to $4.00 above 200,000 tokens.
How do the context windows, output limits, and knowledge cutoffs compare?
Claude Opus 5 has a one-million-token context window with no beta header, a 128,000-token synchronous output limit, and a May 2026 knowledge cutoff. Gemini 3.5 Flash has a 1,048,576-token input limit, a 65,536-token output limit, and a January 2025 knowledge cutoff, about sixteen months older. Opus 5 batch requests can reach 300,000 output tokens using the output-300k-2026-03-24 beta header, on the Claude API and Claude Platform on AWS only.
Which model can process audio and video?
Only Gemini 3.5 Flash. Google lists its input modalities as text, image, video, audio, and PDF, with text output. Anthropic's Messages API accepts text, image, and document blocks, where documents are PDF and plain text, and defines no audio or video content block. Workloads that ingest recordings or footage must either use Flash or add a separate transcription stage before sending anything to Opus 5.
Which model hallucinates less?
This cannot be answered from published data. Artificial Analysis lists an AA-Omniscience score of 31 for Claude Opus 5 at max effort, but Gemini 3.5 Flash is absent from that leaderboard and the corresponding fields on its model page are blank. AA-Omniscience is one of nine evaluations inside Intelligence Index v4.1, so it contributes to Flash's 50 without being published separately. For context, Claude Fable 5 scores 40 and Gemini 3.1 Pro Preview scores 33, both above Opus 5.
Is there a Gemini 3.5 Pro?
No. Google has not released a model called Gemini 3.5 Pro. The models that exist nearby are Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, and Gemini 3.1 Pro Preview, all of which are distinct models with different scores and prices. Any benchmark figure attributed to "Gemini 3.5 Pro" refers to a model that was never released and should be discarded rather than reassigned to another variant.
Is Gemini 3.5 Flash-Lite the same model as Gemini 3.5 Flash?
No, it is a separate model rather than a reasoning level of Flash. Artificial Analysis scores Gemini 3.5 Flash-Lite at 36 on Intelligence Index v4.1 at $0.09 per task, against 50 at $0.59 for Gemini 3.5 Flash. That is fourteen index points lower for roughly one sixth of the task cost. If a workload runs acceptably at 36, Flash-Lite is cheaper than either model discussed in this comparison.
Can reasoning be turned off on either model?
Not fully on either. Claude Opus 5 has adaptive thinking on by default and accepts thinking: {type: "disabled"} at effort high or below; combining that with xhigh or max effort returns a 400 error. Gemini 3.5 Flash has thinking on at medium by default, and Google warns that "minimal does not guarantee that thinking is off." Google also keeps the older numeric thinking_budget "supported for backward compatibility" but instructs developers not to send both parameters in one request.
Sources and references
Every figure on this page comes from a vendor's own documentation or from an independent evaluator. Vendor claims and independent measurements are labeled separately throughout and are never stacked into a single ranking. Anthropic's published benchmarks compare Claude Opus 5 to other Anthropic models only and never to Google models, so no head-to-head figure on this page is sourced from either vendor.
- Anthropic — Introducing Claude Opus 5 (vendor: release date, token pricing, fast mode)
- Anthropic — Pricing (vendor: token rates, caching, batch, long-context pricing, tokenizer note, data residency multiplier)
- Anthropic — Models overview (vendor: context window, maximum output, knowledge cutoffs, legacy classification, modalities)
- Anthropic — Effort (vendor: five effort levels, default high)
- Anthropic — Context windows (vendor: 1M default, no beta header)
- Anthropic — Batch processing (vendor: 300,000-token extended output beta)
- Anthropic — Messages API reference (vendor: effort enum, content block types, sampling parameter restrictions)
- Google — Gemini API pricing (vendor: token rates, caching, thinking tokens included in output price)
- Google — Gemini 3.5 Flash model card (vendor: model identifier, token limits, modalities, status)
- Google — What's new in Gemini 3.5 (vendor: knowledge cutoff, general availability, thinking levels and default)
- Google — Gemini thinking (vendor: thinking_level values, default, backward compatibility with thinking_budget)
- Google DeepMind — Gemini 3.5 Flash model card (vendor: output length, model lineage)
- Google Cloud — Vertex AI generative AI pricing (vendor: long-context tiers, global and non-global rates)
- Artificial Analysis — Model leaderboard (independent: index scores, cost per task, throughput, time to first token)
- Artificial Analysis — Claude Opus 5 (independent: blended price, per-configuration measurements)
- Artificial Analysis — Gemini 3.5 Flash (independent: blended price, per-configuration measurements)
- Artificial Analysis — Intelligence benchmarking methodology (independent: Index v4.1 composition, cost-per-task definition)
- Artificial Analysis — AA-Omniscience (independent: Omniscience scale and scores)
Related comparisons on ThePlanetTools: Gemini 3.5 Flash against Claude Haiku 4.5, Gemini 3.5 Flash against MiniMax M3, and Claude Opus 4.8 against Gemini 3.1 Pro. Background reading: our report on the Claude Opus 5 launch.
Our Verdict
Claude Opus 5 wins, and it wins in an unusual way: Gemini 3.5 Flash's best measured configuration scores 50 at $0.59 per task, one point below Opus 5's cheapest configuration at 51 for $0.36, which makes Google's model strictly dominated on intelligence per unit of task cost. Opus 5 also holds four measured rungs above Flash's ceiling and a knowledge cutoff sixteen months fresher. The concessions are real and matter: Gemini 3.5 Flash moves tokens about three and a half times faster, lists at roughly a third of Opus 5's token price, and accepts audio and video that Opus 5 cannot read at any price. For input-heavy, latency-sensitive or multimodal work, Flash remains the correct choice and the index gap is close to irrelevant.
Choose Claude Opus 5
Anthropic's frontier reasoning model — top of the independent index at half the price of Fable 5.
Try Claude Opus 5 →Choose Gemini 3.5 Flash
Google DeepMind's generally available fast tier — frontier-adjacent intelligence at roughly four times the speed, with a 1M-token context window and native multimodal input.
Try Gemini 3.5 Flash →Frequently Asked Questions
Is Claude Opus 5 better than Gemini 3.5 Flash?
Claude Opus 5 wins, and it wins in an unusual way: Gemini 3.5 Flash's best measured configuration scores 50 at $0.59 per task, one point below Opus 5's cheapest configuration at 51 for $0.36, which makes Google's model strictly dominated on intelligence per unit of task cost. Opus 5 also holds four measured rungs above Flash's ceiling and a knowledge cutoff sixteen months fresher. The concessions are real and matter: Gemini 3.5 Flash moves tokens about three and a half times faster, lists at roughly a third of Opus 5's token price, and accepts audio and video that Opus 5 cannot read at any price. For input-heavy, latency-sensitive or multimodal work, Flash remains the correct choice and the index gap is close to irrelevant.
Which is cheaper, Claude Opus 5 or Gemini 3.5 Flash?
Claude Opus 5 is priced at $5 in / $25 out per M tokens. Gemini 3.5 Flash is priced at $1.5 in / $9 out per M tokens (free plan available). Check the pricing comparison section above for a full breakdown.
What are the main differences between Claude Opus 5 and Gemini 3.5 Flash?
The key differences span across 13 features we compared. For Artificial Analysis Intelligence Index v4.1, top measured configuration, Claude Opus 5 offers 61 (max effort) while Gemini 3.5 Flash offers 50 (high thinking level). For Index at the vendor default configuration, Claude Opus 5 offers 59 (high effort) while Gemini 3.5 Flash offers 45 (medium thinking level). For Index at the bottom of the vendor ladder, Claude Opus 5 offers 51 (low effort) while Gemini 3.5 Flash offers 35 (minimal thinking level). See the full feature comparison table above for all details.

