Claude Opus 5 vs DeepSeek V4: 17 Points, 29x the Price
Opus 5 scores 61, DeepSeek V4 Pro 44 (measured July 28, 2026). Opus 5 costs 29x more per output token. V4 ships MIT weights. Which gap is worth it.
Feature Comparison
| Feature | Claude Opus 5 | DeepSeek V4 |
|---|---|---|
| Intelligence Index v4.1, matched max effort | 61 | 44 |
| Intelligence Index v4.1, each model default | 59 at high | 43 at high |
| AA-Omniscience, max effort | 31 | -10 |
| Cost to run index suite, max effort | $2.028 | $0.045 |
| Published output price per million tokens | $25.00 | $0.87 |
| Published input price per million tokens | $5.00 | $0.435 |
| Context window | 1M tokens | 1M tokens |
| Long-context surcharge | None | None |
| Distinct reasoning effort levels | 5 | 2 |
| Open weights | No | Yes, MIT license |
| Self-hosting permitted | No | Yes, unrestricted |
| Published knowledge cutoff | May 2026 | Not published |
| Release status wording | Released | Preview |
Pricing Comparison
Claude Opus 5
DeepSeek V4
Detailed Comparison
Claude Opus 5 scores 61 on the Artificial Analysis Intelligence Index v4.1 at max reasoning effort. DeepSeek V4 Pro scores 44 at its own max effort. Both readings were taken on July 28, 2026. That is a 17-point gap. On published API rates, Opus 5 charges $25 per million output tokens against DeepSeek V4 Pro's $0.87 — about 29 times more — and that ratio holds at every effort setting, because reasoning effort changes how many tokens a model emits, never the rate per token. Measured on completed work, Artificial Analysis spent $2.03 running its index suite on Opus 5 at max against $0.045 on DeepSeek V4 Pro at max, a ratio of about 45 to 1; even Opus 5's cheapest configuration cost roughly 8 times DeepSeek V4 Pro's most expensive. No Opus 5 setting reaches DeepSeek's cost band on either measure. DeepSeek V4 also ships its weights under a plain, unmodified MIT license, which no index measures at all. What the index does measure, and what cost comparisons routinely omit, is reliability: on AA-Omniscience, Opus 5 scores 31 while DeepSeek V4 Pro scores −10, meaning it asserts more wrong answers than correct ones.
Quick verdict
Claude Opus 5 wins on capability and on knowledge reliability. DeepSeek V4 wins on cost, on deployability, and on license freedom. The two are not substitutes at the same job. Opus 5 returns 17 index points plus a 41-point swing on hallucination resistance, for roughly 29 times the published output-token rate. DeepSeek V4 returns weights you can host yourself under MIT.
Winner on raw capability: Claude Opus 5. It leads 61 to 44 at matched max effort, and 59 to 43 at each model's default effort.
Winner on cost: DeepSeek V4, decisively and unconditionally, on published rates and on measured spend alike. No Opus 5 effort setting reaches DeepSeek V4 Pro's cost band.
Winner on knowledge reliability: Claude Opus 5, by the widest margin on this page. AA-Omniscience places Opus 5 at 31 and DeepSeek V4 Pro at −10 on a scale running from −100 to 100.
Winner on license and control: DeepSeek V4. Plain MIT weights on Hugging Face, no revenue threshold, no user cap, no acceptable-use appendix.
Winner on long-context economics: Tie. Both bill a one-million-token context at a flat per-token rate, with no threshold surcharge. This is genuinely unusual in July 2026.
Our overall pick is Claude Opus 5, but on a narrower basis than the headline gap suggests. The 17 points get paid only when a wrong answer costs more than the price difference — and the price difference is large enough that for a great many tasks, it does not. We set out exactly where that line falls in the section on difficulty.
What DeepSeek V4 actually is
DeepSeek V4 is not one model. It is a two-model family released on April 24, 2026: deepseek-v4-pro at 1.6 trillion total parameters with 49 billion activated, and deepseek-v4-flash at 284 billion total with 13 billion activated. Both are Mixture-of-Experts models supporting a one-million-token context. DeepSeek's own word for the release is "Preview", and it still is today.
This distinction matters more than usual, and we want to be precise about it because the number most often quoted for "DeepSeek V4" — the 44 on the Artificial Analysis Intelligence Index — belongs specifically to DeepSeek V4 Pro at max reasoning effort. It is not a family-wide score. DeepSeek V4 Flash at max effort scores 40, and both models score materially lower with reasoning switched off.
On availability, we are quoting rather than paraphrasing, because the classifying word is the vendor's and not ours. The DeepSeek API documentation titles the release announcement "DeepSeek V4 Preview Release", dated 2026/04/24, and the body reads: "DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length." The Hugging Face model card for both repositories opens: "We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models". As of today, the deepseek.com homepage still carries: "DeepSeek-V4 Preview is here with stronger Agent capabilities and top-tier reasoning. Now available on web, app, and API."
So the accurate description is preview, live in production. DeepSeek has never applied the phrase "generally available" to V4, and we will not apply it on their behalf. Equally, "preview" here does not mean vapour: the weights are physically present on Hugging Face (64 safetensors shards for Pro, 46 for Flash, with config, tokenizer and an inference folder), the production API serves both model identifiers, and the pricing page bills them. Three months after launch, the label has not been lifted and DeepSeek has published no timeline for lifting it.
We flag this carefully because we have been burned in the opposite direction. DeepSeek R2 was reported across the technology press as a shipped model and never actually shipped; we published on it and had to retract. That experience is why every claim on this page traces to deepseek.com, the API documentation, Hugging Face, or an independent evaluator — and to no media outlet whatsoever. DeepSeek V4 passes the test that R2 failed: the weights exist, the API answers, and the invoice is real.
One consequence of the April release deserves a note for anyone maintaining older integration code. The API documentation states: "deepseek-chat & deepseek-reasoner will be fully retired and inaccessible after Jul 24th, 2026, 15:59 (UTC Time)." That deadline has now passed. Those two identifiers no longer resolve, and the only valid strings are deepseek-v4-flash and deepseek-v4-pro.
How we compared the two
We researched both models rather than running our own private benchmark suite. Every performance number on this page comes from Artificial Analysis, an independent evaluator, read directly from its leaderboard on July 28, 2026. Every price and specification comes from the vendor's own documentation, fetched directly rather than summarized from search results.
We want to be explicit about method, because a comparison of this kind can go wrong in several quiet ways and we have designed around each of them.
We did not run these benchmarks ourselves. We are not an evaluation lab. Publishing our own scores for two frontier models would mean asking you to trust a methodology we could not properly document. Instead we take an independent third party's published measurements and are precise about their provenance.
Every score carries its configuration. A reasoning model does not have "a score" — it has a score at an effort level. Opus 5's 61 is at max; at its default high it is 59. Quoting the ceiling of one model against the default of the other is the single most common way these comparisons mislead, and we have matched configurations throughout.
Index readings are dated because they move. Artificial Analysis re-runs evaluations and revises index versions; a score is a reading, not a constant. Everything here was read on July 28, 2026, against Intelligence Index v4.1, whose methodology page describes it as incorporating nine evaluations weighted Agents 34 percent, Coding 24 percent, Scientific Reasoning 24 percent, and General 18 percent.
We never stack vendor claims against independent measurements. DeepSeek's model card publishes strong headline figures — LiveCodeBench 93.5, Codeforces 3206, GPQA Diamond 90.1 among them — all at max reasoning effort. Anthropic publishes its own. Those are self-reported and are not directly comparable to each other, so we keep them out of the head-to-head tables entirely and rely on the one source that measured both models under the same harness.
We compare the two models to each other only. Anthropic's published comparisons for Opus 5 are against Claude Opus 4.8 and Claude Fable 5. They never mention DeepSeek. Nothing on this page infers an Opus 5 versus DeepSeek V4 result by chaining through a third model.
The effort scales, side by side
Opus 5 exposes five reasoning effort levels and Artificial Analysis measures all five, from 51 at low to 61 at max. DeepSeek V4 exposes only two, high and max — its API maps low and medium onto high, and xhigh onto max. Comparing the two at matched max effort gives 61 against 44; at each model's own default, 59 against 43.
This is the table that most comparisons get wrong, so here it is in full. All values read from the Artificial Analysis leaderboard on July 28, 2026, Intelligence Index v4.1.
| Entry as labelled by Artificial Analysis | Index score | Cost to run the index suite |
|---|---|---|
| Claude Opus 5 (max) | 61 | $2.028 |
| Claude Opus 5 (xhigh) | 60 | $1.561 |
| Claude Opus 5 (high) — default | 59 | $1.057 |
| Claude Opus 5 (medium) | 56 | $0.618 |
| Claude Opus 5 (low) | 51 | $0.361 |
| DeepSeek V4 Pro (max) | 44 | $0.045 |
| DeepSeek V4 Pro (high) — default | 43 | $0.041 |
| DeepSeek V4 Flash (max) | 40 | $0.022 |
| DeepSeek V4 Flash (high) — default | 37 | $0.041 |
| DeepSeek V4 Pro, non-reasoning mode | 31, estimated | not measured |
| DeepSeek V4 Flash, non-reasoning mode | 29, estimated | not measured |
Three points of precision about that table.
DeepSeek V4's scale is genuinely shorter, not merely unmeasured. This is a real architectural difference rather than a gap in the evaluator's coverage. The DeepSeek thinking-mode documentation states: "In thinking mode, the default effort is high for regular requests; for some complex agent requests (such as Claude Code, OpenCode), effort is automatically set to max", and then: "In thinking mode, for compatibility, low and medium are mapped to high, and xhigh is mapped to max". You may send low, but you will not receive a low-effort run — you will receive a high-effort run and a high-effort bill. DeepSeek V4 has two effort levels wearing five names.
The two non-reasoning rows are not an effort level. They are a separate mode with reasoning disabled, and their scores carry Artificial Analysis's estimation marker, defined on the leaderboard as: "Figure is estimated as Artificial Analysis has not independently conducted all required benchmarks. Estimate is based on the subset of benchmarks we have conducted or based on claims of the AI lab behind the model." Their cost per task is not published at all. We include them for completeness and exclude them from every comparison on this page.
One figure in that table is counter-intuitive, and it matters enough that we treat it separately. DeepSeek V4 Flash cost more to run at high effort than at max. That is not a transcription error on our part, and it is the clearest available demonstration that this column measures spend rather than a rate. We unpack it in the next section.
The shape of the comparison, then: Opus 5's floor of 51 sits well above DeepSeek V4 Pro's ceiling of 44. There is no overlap in capability between the two ranges. Opus 5's worst measured configuration beats DeepSeek V4's best by seven points.
What it cost to run the index, and why that is not a price
Artificial Analysis publishes what it spent running its own benchmark suite on each configuration. That figure is not a vendor rate. It combines the rate with how many tokens the model emitted, reasoning tokens included, so it moves with verbosity as well as with price. It is the most decision-useful number on this page precisely because it measures completed work — but calling it a price would be wrong.
The dataset contains its own proof of this. DeepSeek V4 Flash cost $0.0411 to run at high effort and $0.0223 at max — the same model, from the same vendor, on the same per-token rate, costing 1.84 times more at the lower effort setting. A price ratio cannot invert like that. What can invert is total spend: a higher-effort run that reaches an answer in fewer total tokens ends up cheaper than a lower-effort run that flounders and emits more. That is a fact about token economics, not about DeepSeek's price list.
So this page keeps the two ideas apart. Published rates come from each vendor's own pricing page and are set out in the pricing section. The figures below are measured spend on a fixed benchmark suite, which is a different and complementary thing.
| Opus 5 setting | Index lead over DeepSeek V4 Pro | Multiple of DeepSeek V4 Pro's measured spend |
|---|---|---|
| max (61, $2.03) | +17 points | about 45 times |
| xhigh (60, $1.56) | +16 points | about 35 times |
| high (59, $1.06) | +15 points | about 24 times |
| medium (56, $0.62) | +12 points | about 14 times |
| low (51, $0.36) | +7 points | about 8 times |
Read the bottom row carefully, because it is the least intuitive result on this page. Dropping Opus 5 from max to low cuts measured spend by roughly 82 percent and surrenders 10 of the 17 points. You give up most of the capability advantage and the run still costs about eight times what DeepSeek V4 Pro's most expensive setting costs. The economics do not reward meeting in the middle; they reward deciding which of the two jobs you are actually doing.
Multiples above are computed from the unrounded values in the leaderboard's underlying data, not from the two-decimal figures displayed on screen. The displayed $0.04 for DeepSeek V4 Pro is really $0.0448, which is why an eyeball calculation from the visible table returns about 51 to 1 at max instead of the correct 45 to 1.
Note also that this metric is measured only where Artificial Analysis ran the full suite. The non-reasoning modes of both DeepSeek models have no published cost figure at all, so they cannot be placed on this scale.
API pricing, and a tokenizer detail that moves the numbers
These are the actual prices, taken from each vendor's own pricing page. Claude Opus 5 charges $5 per million input tokens and $25 per million output tokens. DeepSeek V4 Pro charges $0.435 and $0.87; DeepSeek V4 Flash charges $0.14 and $0.28. Opus 5 therefore costs about 29 times more per output token than V4 Pro and about 89 times more than V4 Flash. Neither vendor applies a long-context surcharge.
One structural point deserves emphasis before the table, because it settles the central question of this comparison more cleanly than any benchmark figure. Reasoning effort does not change the rate you are charged. Opus 5 bills $5 and $25 per million tokens whether you run it at low or at max; effort changes how many tokens it emits, not their price. The same is true of DeepSeek. So the question "can I turn Opus 5 down until it undercuts DeepSeek V4?" has a definitive answer on rates alone: no, and not by any amount. Every Opus 5 configuration sits at 11.5 times DeepSeek V4 Pro's input rate and 28.7 times its output rate. Turning effort down reduces your bill by emitting fewer tokens; it never moves you into DeepSeek's price band.
| Claude Opus 5 | DeepSeek V4 Pro | DeepSeek V4 Flash | |
|---|---|---|---|
| Input, standard | $5.00 | $0.435 | $0.14 |
| Input, cache hit | $0.50 | $0.003625 | $0.0028 |
| Output | $25.00 | $0.87 | $0.28 |
| Cache write, 5 minute | $6.25 | not published | not published |
| Cache write, 1 hour | $10.00 | not published | not published |
| Batch input and output | $2.50 and $12.50 | not published | not published |
| Context window | 1M tokens | 1M tokens | 1M tokens |
| Maximum output | not restated here | 384K tokens | 384K tokens |
| Long-context surcharge | none | none | none |
| Concurrency limit | tier-dependent | 500 | 2500 |
The flat long-context rate is a genuine convergence, and it is unusual. Several competitors step their prices above a token threshold, sometimes charging the higher rate on the entire request rather than the excess. Neither of these two does. Anthropic's pricing page states: "Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.)" DeepSeek's deduction rule is a single sentence with no tier function in it: "The expense = number of tokens × price." For long-document work, both bill honestly at scale.
DeepSeek's off-peak discount is gone. DeepSeek historically ran time-windowed discounts, and older write-ups still assume them. The V4 pricing page carries no discount table, no time windows and no percentages. Anyone modelling costs on a remembered off-peak rate is modelling a price that no longer exists.
Now the detail that quietly moves every figure above. Anthropic's pricing page carries this note: "Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape."
Opus 5 is a Claude 4.7-or-later model, so this applies. The practical consequence is that comparing sticker prices per million tokens understates Opus 5's real cost on identical text, because the same document becomes roughly 30 percent more tokens before either model has done any thinking. The per-token gap on output is about 29 times; adjusted for tokenization on comparable text, the effective gap is larger still. Anthropic's own caveat — that the increase depends on content and workload shape — is why we do not publish a single adjusted multiplier. Measure it on your own corpus.
This is also why the cost-per-task figures in the previous section are the more trustworthy comparison. They are measured on completed work rather than derived from list prices, so tokenizer differences are already inside them.
One further Opus 5 pricing mode is worth knowing about even though it does not affect the head-to-head: a fast mode in research preview, billed at $10 and $50 per million input and output tokens, which Anthropic notes "applies across the full context window, including requests over 200k input tokens". It is not available with the Batch API.
What the index does not fully price: knowledge reliability
On AA-Omniscience, which rewards correct answers, penalizes hallucinations and applies no penalty for declining to answer, Claude Opus 5 scores 31 and DeepSeek V4 Pro scores −10. The scale runs from −100 to 100, where zero means a model answers correctly as often as incorrectly. A negative score means DeepSeek V4 Pro asserts more wrong answers than correct ones on this benchmark.
This is the finding that most changes the shape of the recommendation, and it runs against the direction people usually expect when they hear "the cheap model is closer than you think".
| Entry as labelled by Artificial Analysis | AA-Omniscience Index |
|---|---|
| Claude Opus 5 (max) | 31 |
| DeepSeek V4 Pro (Reasoning, Max Effort) | −10 |
| DeepSeek V4 Flash (Reasoning, Max Effort) | −23 |
The benchmark's own description, as published: "AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer." The scoring design is the important part. A model is never punished for saying it does not know. A negative result therefore cannot be blamed on excessive caution — it can only come from confidently asserting things that are false.
The 41-point spread between Opus 5 and DeepSeek V4 Pro is much wider than the 17-point spread on general intelligence, and it measures a different property. General capability tells you what a model can do when it knows the answer. This tells you what it does when it does not. On that second question the two models behave in genuinely different ways: Opus 5 is substantially more likely to decline, DeepSeek V4 substantially more likely to produce a confident fabrication.
Two honest caveats. This evaluation covers only the max-effort variants of each model — Opus 5's low through xhigh settings and DeepSeek's high setting are not measured on it, so we cannot tell you how the gap moves across the effort ladder. And it measures factual knowledge reliability specifically, not reasoning correctness or code quality. A model can be poor at recalling obscure facts and still competent at transforming text you have handed it. That distinction turns out to be the practical key to using DeepSeek V4 well, and we return to it below.
What the index does not measure at all: MIT weights
DeepSeek publishes V4 Pro and V4 Flash weights on Hugging Face under a plain, unmodified MIT license. We read the LICENSE files as raw bytes rather than trusting the model card or the platform tag: both are 1084 bytes and byte-identical to each other, with no revenue threshold, no user cap, no acceptable-use appendix and no distillation clause. Claude Opus 5 is API-only.
We check licenses this way because the summary and the file disagree often enough to matter. Among models we have covered: one was widely reported as "Modified MIT" but carried a custom license with a revenue threshold; another was genuinely plain MIT; a third shipped a "Community License" that was non-commercial with internally contradictory clauses. The only reliable method is to open the file.
For DeepSeek V4, the file says MIT License on its first line, with the copyright line Copyright (c) 2023 DeepSeek, followed by the canonical MIT text granting rights "without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software". Both repositories carry exactly one LICENSE file each; the LICENSE-MODEL file DeepSeek has used in the past for restrictive model terms does not exist in either repository. This is the real thing, not a permissive label on restrictive terms.
What that buys, none of which appears anywhere in an intelligence index:
- You can run it on your own hardware. Data never leaves your infrastructure, which for regulated, classified or contractually restricted workloads is not a preference but a precondition. Opus 5 offers data-residency controls, at a 1.1 times pricing multiplier for US-only inference, but the request still leaves your building.
- You can fine-tune it on proprietary data and keep the resulting weights.
- You cannot be deprecated. Weights on your disk keep working when a vendor retires an endpoint. DeepSeek's own retirement of
deepseek-chatanddeepseek-reasoneron July 24, 2026 is a live illustration of what API-only dependency exposes you to. - Your unit economics become hardware economics. At sufficient sustained volume, self-hosting changes the cost structure rather than merely reducing the rate.
The counterweight is honest and substantial: V4 Pro is 1.6 trillion parameters. Self-hosting it is a serious infrastructure undertaking, not a weekend exercise, and for most teams the practical open-weights option is V4 Flash at 284 billion total and 13 billion activated — which scores 40 rather than 44. The MIT license is real and unrestricted; the hardware bill for exercising it is real too.
| DeepSeek V4 Pro | DeepSeek V4 Flash | |
|---|---|---|
| Total parameters | 1.6T | 284B |
| Activated parameters | 49B | 13B |
| Context length | 1M | 1M |
| Precision | FP4 and FP8 mixed | FP4 and FP8 mixed |
| License | MIT | MIT |
Anthropic publishes no parameter count for Opus 5 and no weights, so the architecture row simply has no counterpart. That is itself part of the comparison: one of these models is an inspectable artifact and the other is a service.
Feature comparison
Opus 5 leads on capability, reliability, effort granularity and ecosystem depth. DeepSeek V4 leads on price, license, self-hosting and raw concurrency. Both match on context window and on flat long-context billing. DeepSeek publishes no knowledge cutoff for V4, which is a genuine gap rather than an omission on our part.
| Feature | Claude Opus 5 | DeepSeek V4 Pro | Advantage |
|---|---|---|---|
| Intelligence Index v4.1, matched max effort | 61 | 44 | Opus 5 |
| Intelligence Index v4.1, each default | 59 at high | 43 at high | Opus 5 |
| AA-Omniscience, max effort | 31 | −10 | Opus 5 |
| Measured spend running the index suite, max effort | $2.03 | $0.045 | DeepSeek V4 |
| Output price per million tokens | $25.00 | $0.87 | DeepSeek V4 |
| Context window | 1M | 1M | Tie |
| Long-context surcharge | None | None | Tie |
| Effort levels actually distinct | 5 | 2 | Opus 5 |
| Weights available | No | Yes, MIT | DeepSeek V4 |
| Self-hosting permitted | No | Yes, unrestricted | DeepSeek V4 |
| Published knowledge cutoff | May 2026 | Not published | Opus 5 |
| Prompt caching discount published | Yes, $0.50 read | Yes, $0.003625 read | Tie |
| Batch pricing published | Yes, 50 percent | Not published | Opus 5 |
| Release status wording | Released | "Preview" | Opus 5 |
On the knowledge cutoff row, we searched DeepSeek's pricing page, FAQ, change log, release announcement, thinking-mode guide, full chat-completions reference and both model cards. No cutoff is stated for either V4 model. We report that as not published rather than estimating one.
At what difficulty the 17 points get paid
The 17-point gap is worth paying for when a wrong answer is expensive to detect. On measured spend per completed task, $0.045 against $2.03, you can run DeepSeek V4 Pro about 45 times for what one Opus 5 run costs at max effort. So on any task where correctness is cheaply verifiable — tests pass, code compiles, output validates against a schema — repeated cheap attempts plus a verifier frequently beat one expensive attempt. Where no cheap verifier exists, that arithmetic collapses.
This is the question the brief for this page was really asking, and it has a reasonably crisp answer. The deciding variable is not task difficulty in the abstract. It is the cost of detecting a wrong answer.
Where cheap verification exists, DeepSeek V4 is very hard to beat. If a unit test suite, a compiler, a JSON schema, a type checker or a deterministic reference can tell you whether an output is correct, then a wrong answer costs you one retry. With a roughly 45-to-1 ratio on measured spend you can afford a great many retries. Sampling a cheaper model several times and keeping the attempt that passes is a well-understood pattern, and a ratio of this size makes it comfortable rather than marginal. Under these conditions, paying roughly 45 times more to raise first-attempt quality is usually poor value.
Where verification is as expensive as the work, the gap gets paid immediately. If checking the answer means a domain expert reading it carefully — legal analysis, medical summarization, financial reconciliation, architectural review, anything where a plausible wrong answer looks exactly like a right one — then the error does not cost a retry. It costs expert time, or it escapes into production. Here the AA-Omniscience result matters more than the intelligence gap: a model at −10 produces confident falsehoods, and confident falsehoods are precisely the failure mode that cheap verification cannot catch. Against an expert-hour rate, the difference between $0.045 and $2.03 is a rounding error.
Where the task is long-horizon and agentic, the gap compounds. Index v4.1 weights agentic evaluations most heavily, at 34 percent. In a multi-step agent run, step-level errors do not stay local — a wrong tool call or a misread intermediate result propagates, and the run either fails late or produces something subtly wrong after consuming its whole budget. A per-step advantage compounds across steps, so a 17-point gap on single-turn measurement understates the difference over a 40-step trajectory. It is also the regime where DeepSeek's own documentation escalates effort automatically to max for complex agent requests, which is a signal about where it needs its ceiling.
Where volume is extreme, price wins by arithmetic. At ten million tasks of comparable shape, the difference between $0.045 and $2.03 of measured spend is roughly twenty million dollars. No quality argument survives contact with that number unless each task carries real downside risk. Classification, extraction, routing, summarization at scale, first-pass triage: these belong on the cheap model, and it is not close.
The practical synthesis most teams land on is not choosing one model but routing between them: DeepSeek V4 for high-volume, verifiable, transformation-shaped work, Opus 5 reserved for the fraction of traffic that is genuinely hard or genuinely consequential. The 17 points are worth buying for that fraction and wasteful everywhere else. The two flat long-context billing models make such routing unusually clean, since neither will surprise you with a threshold charge on long documents.
Pros and cons of each
Opus 5's strengths are capability, reliability and control granularity; its weaknesses are price, API-only access and a tokenizer that inflates effective cost. DeepSeek V4's strengths are price, MIT weights and a full one-million-token context; its weaknesses are a negative hallucination score, only two real effort levels and unpublished basic specifications.
Claude Opus 5 — strengths
- Highest measured score of the two by 17 points at matched max effort, and its floor still beats DeepSeek's ceiling by seven.
- AA-Omniscience of 31 against −10 — it declines rather than fabricates.
- Five genuinely distinct effort levels, so cost and quality are tunable across a real range.
- Full one-million-token context at standard per-token pricing, with no threshold surcharge.
- Prompt caching at $0.50 per million read tokens and a published 50 percent batch discount.
- Published knowledge cutoff of May 2026.
Claude Opus 5 — weaknesses
- About 11.5 times DeepSeek V4 Pro's input rate and 28.7 times its output rate, at every effort setting.
- No configuration reaches DeepSeek's cost band; even at low effort a completed run cost about eight times DeepSeek V4 Pro's most expensive setting.
- API-only. No weights, no self-hosting, no fine-tuning on your own infrastructure.
- The newer tokenizer produces approximately 30 percent more tokens for the same text, so sticker prices understate real cost on identical documents.
- US-only inference carries a 1.1 times pricing multiplier.
- No published parameter count or architecture detail.
DeepSeek V4 — strengths
- Roughly 29 times cheaper per output token than Opus 5 on published rates, and about 45 times cheaper on measured spend running the index suite at max effort.
- Plain, unmodified MIT weights, verified byte-level, with no revenue threshold or use restriction.
- Full one-million-token context with a flat billing rule and no long-context surcharge.
- Two model sizes, letting you trade 44 points at 1.6 trillion parameters against 40 at 284 billion.
- Cache-hit input pricing of $0.003625 per million tokens on Pro.
- High concurrency limits, 500 on Pro and 2500 on Flash.
- Maximum output of 384,000 tokens.
DeepSeek V4 — weaknesses
- AA-Omniscience of −10 for Pro and −23 for Flash: more confident wrong answers than correct ones on knowledge questions.
- Only two genuinely distinct effort levels; low, medium and xhigh are silently remapped.
- Still labelled "Preview" by the vendor three months after release, with no announced timeline for lifting it.
- No published knowledge cutoff for either model.
- No published batch pricing, and the former off-peak discounts no longer exist.
- Self-hosting the 1.6 trillion parameter Pro model is a heavy infrastructure commitment.
- Thinking mode ignores temperature, top_p, presence_penalty and frequency_penalty entirely.
When to pick each
Pick Claude Opus 5 when errors are expensive to detect, when the work is long-horizon and agentic, or when factual reliability is the product. Pick DeepSeek V4 when outputs are cheaply verifiable, when volume is high, when weights must live on your own hardware, or when the license is a procurement requirement.
Pick Claude Opus 5 when
- A wrong answer is expensive to catch. Legal, medical, financial and compliance work where a plausible fabrication passes casual review. The AA-Omniscience gap of 41 points speaks directly to this case.
- The task is a long agentic trajectory. Multi-step tool use where per-step errors compound, and where a failed run wastes its entire budget rather than one call.
- You need fine cost control. Five distinct effort levels, spanning measured spend from $0.361 to $2.028 on the same suite, let you tune per workload rather than per model.
- Recall accuracy is the product. Research assistance and knowledge work where "I do not know" is a correct and valuable answer.
- Volume is modest and stakes are high. When you run thousands rather than millions of tasks, the absolute price difference stops being a business consideration.
Pick DeepSeek V4 when
- Outputs are machine-verifiable. Code with tests, structured extraction against a schema, anything a validator can grade. Fifty cheap attempts with a verifier beat one expensive attempt.
- Volume is the dominant cost. Classification, routing, tagging, bulk summarization, first-pass triage at millions of calls.
- Data cannot leave your infrastructure. MIT weights make on-premise and air-gapped deployment legally straightforward, which no amount of API-side residency configuration matches.
- You need to fine-tune and keep the result. Domain adaptation on proprietary data, with the resulting weights yours.
- Vendor lock-in is an identified risk. Weights on disk cannot be deprecated out from under you.
- The work is transformation rather than recall. Rewriting, translating, reformatting and summarizing text you supply — the regime where a weak knowledge-reliability score matters least, because the facts are in the prompt rather than in the weights.
That final bullet is the most useful distinction we can offer. AA-Omniscience measures what a model knows unaided. It says comparatively little about how well a model manipulates material you have put in front of it. Since both models carry a one-million-token context at a flat rate, feeding DeepSeek V4 the facts instead of relying on its recall is both technically easy and economically painless — and it neutralizes its single worst measured weakness. Retrieval-grounded architectures suit this model unusually well.
Frequently asked questions
Is DeepSeek V4 a real, released model?
Yes. DeepSeek V4 was announced on April 24, 2026, and both models are live on the production API and downloadable from Hugging Face. The weights are physically present as 64 safetensors shards for Pro and 46 for Flash. The vendor's own classifying word is "Preview" rather than generally available, and that label is still on the deepseek.com homepage today. It is materially shipped, unlike DeepSeek R2, which was widely reported but never released.
Which DeepSeek model does the score of 44 refer to?
It refers specifically to DeepSeek V4 Pro at max reasoning effort, as measured by Artificial Analysis on Intelligence Index v4.1 and read on July 28, 2026. It is not a family-wide figure. DeepSeek V4 Pro at high effort scores 43, DeepSeek V4 Flash at max scores 40, and Flash at high scores 37.
Can lowering Claude Opus 5's effort make it cheaper than DeepSeek V4?
No, and not by any amount. Reasoning effort changes how many tokens a model emits, never the rate charged per token, so Opus 5 bills $5 and $25 per million tokens at low effort exactly as it does at max. That leaves it at about 11.5 times DeepSeek V4 Pro's input rate and 28.7 times its output rate at every setting. On measured spend, Opus 5's cheapest configuration still cost roughly eight times DeepSeek V4 Pro's most expensive. Lowering effort reduces your bill by emitting fewer tokens; it does not move you into DeepSeek's cost band.
Is DeepSeek V4 really MIT licensed?
Yes, and we verified it by reading the raw LICENSE files rather than the model card or platform tag. Both the V4 Pro and V4 Flash files are 1084 bytes and byte-identical to each other, beginning with "MIT License" and carrying the canonical MIT text. There is no revenue threshold, no user cap, no acceptable-use appendix, no distillation clause and no non-commercial restriction. Neither repository contains a separate LICENSE-MODEL file.
What does DeepSeek V4's negative AA-Omniscience score mean?
AA-Omniscience runs from −100 to 100 and rewards correct answers, penalizes hallucinations, and applies no penalty for refusing to answer. A score of zero means a model answers correctly as often as incorrectly. DeepSeek V4 Pro's −10 means it asserts more wrong answers than correct ones on knowledge questions. Because refusing costs nothing, the negative result cannot be explained by excessive caution. Claude Opus 5 scores 31 on the same evaluation.
Does either model charge more for long context?
Neither does. Anthropic states that Claude 4.6 and later models include the full one-million-token context window at standard pricing, and that a 900,000-token request is billed at the same per-token rate as a 9,000-token request. DeepSeek's deduction rule is that the expense equals the number of tokens multiplied by the price, with no tier. This makes both unusual in July 2026, as several competitors apply a surcharge above a threshold.
How many reasoning effort levels does DeepSeek V4 have?
Two that are genuinely distinct: high and max. The API accepts five names, but its documentation states that low and medium are mapped to high, and xhigh is mapped to max. Sending low therefore produces a high-effort run at high-effort cost. Claude Opus 5 has five distinct levels: low, medium, high, xhigh and max, with high as the default.
What is the actual price difference between the two?
On published rates per million tokens, Opus 5 charges $5.00 input and $25.00 output against DeepSeek V4 Pro's $0.435 and $0.87 — about 11.5 times more on input and 28.7 times more on output. DeepSeek V4 Flash is cheaper still at $0.14 and $0.28. Separately, on measured spend running the Artificial Analysis index suite at matched max effort, Opus 5 cost $2.028 against V4 Pro's $0.045, a ratio of about 45 to 1. Those two ratios differ because measured spend also reflects how many tokens each model emits, not only what each token costs.
Does Claude Opus 5's tokenizer affect the price comparison?
Yes, and it works against Opus 5. Anthropic notes that Claude 4.7 and later models use a newer tokenizer that produces approximately 30 percent more tokens for the same text. The same document therefore becomes more tokens before any work is done, so per-token list price comparisons understate Opus 5's real cost on identical text. Cost-per-task figures already account for this, since they measure completed work.
What is DeepSeek V4's knowledge cutoff?
DeepSeek does not publish one. We searched the pricing page, FAQ, change log, release announcement, thinking-mode guide, the full chat completions reference and both model cards, and found no cutoff stated for either V4 Pro or V4 Flash. Claude Opus 5's published cutoff is May 2026. We report DeepSeek's as not published rather than estimating it.
Can I self-host DeepSeek V4 in practice?
Legally, yes, without restriction under MIT. Practically, it depends which model. V4 Pro is 1.6 trillion total parameters with 49 billion activated, which is a serious infrastructure commitment. V4 Flash is 284 billion total with 13 billion activated and is far more approachable, at the cost of scoring 40 rather than 44 on the index. Claude Opus 5 cannot be self-hosted at all.
Are the old deepseek-chat and deepseek-reasoner model names still usable?
No. DeepSeek's documentation states that deepseek-chat and deepseek-reasoner were fully retired and inaccessible after July 24, 2026 at 15:59 UTC. That date has passed. The only valid model identifiers are deepseek-v4-flash and deepseek-v4-pro, on base URLs api.deepseek.com for the OpenAI-compatible format and api.deepseek.com/anthropic for the Anthropic format.
Final verdict
Claude Opus 5 is the better model and DeepSeek V4 is the better deal, and both statements are true by wide margins rather than narrow ones. Opus 5 leads by 17 index points and by 41 points on hallucination resistance. DeepSeek V4 charges about 29 times less per output token, cost about 45 times less to run the index suite, and ships MIT weights. Our overall pick is Opus 5, on the condition that you are buying it for work where a wrong answer is expensive to detect.
We give the overall verdict to Claude Opus 5 because the two gaps point the same way and reinforce each other. A 17-point capability lead alone would be an interesting trade against a cost ratio of this size, and for a large share of production traffic the cheap model would simply win that trade. What tips it is the AA-Omniscience result. A model scoring −10 on an evaluation that never penalizes refusal is telling you something specific about its failure mode: when it does not know, it does not say so. That failure mode is the one that survives casual review, reaches production, and costs money somewhere other than the inference bill.
But we want to be exact about the scope of that verdict, because a general recommendation would be the wrong output from this analysis. Opus 5 is the correct default for consequential work. It is the wrong default for high-volume work with cheap verification, and the price ratio there is not close enough to argue about. If your outputs are graded by a test suite or a schema validator, DeepSeek V4 Pro is very likely the right choice and Opus 5 is an expensive habit.
DeepSeek V4 also wins outright, regardless of any score, on one axis the index cannot express: if your data cannot leave your infrastructure, or your procurement requires a permissive license, Opus 5 is not a candidate at any price. That is not a close call either — it is a different question, and MIT weights answer it.
What we would tell a team choosing today: do not pick one. Route. Send verifiable, high-volume, transformation-shaped work to DeepSeek V4, keep Opus 5 for the fraction of traffic that is genuinely hard or genuinely consequential, and ground DeepSeek's context with retrieved facts rather than trusting its recall. Both models bill a one-million-token context at a flat rate, which makes that architecture unusually clean to build. The 17 points are worth buying for the hard fraction and wasteful for everything else.
For adjacent comparisons, we have covered Claude Opus 4.8 against DeepSeek V4, Claude Fable 5 against DeepSeek V4, and GPT-5.6 Sol against DeepSeek V4. Readers weighing MIT-licensed options specifically may also want GLM-5.2 against DeepSeek V4, since GLM-5.2 is the other genuinely plain-MIT model we have verified at the file level. Full specifications live on our Claude Opus 5 and DeepSeek V4 pages.
Sources and references
Every figure on this page comes from a vendor primary source or an independent evaluator. We cite no media coverage, no aggregators and no secondary write-ups. Benchmark readings were taken on July 28, 2026 and are point-in-time measurements that the evaluator revises over time.
Vendor primary sources
- Anthropic — Introducing Claude Opus 5: launch announcement, July 24, 2026.
- Anthropic — Claude Opus 5 System Card: the vendor-reported evaluations referenced on this page.
- Anthropic — Models overview: model IDs, context window, knowledge cutoff.
- Anthropic — Effort parameter: the five effort levels and per-model defaults.
- Anthropic — Claude pricing documentation: Opus 5 token prices, cache and batch rates, long-context pricing statement, tokenizer note, fast mode, data residency multiplier.
- DeepSeek — official homepage: current "Preview" availability wording.
- DeepSeek — API pricing: V4 Pro and V4 Flash token prices, context window, maximum output, concurrency limits, deduction rule.
- DeepSeek — V4 Preview release announcement, April 24, 2026: release date, open-source statement, legacy model retirement deadline.
- DeepSeek — thinking mode guide: effort levels, default effort, remapping of low, medium and xhigh, unsupported sampling parameters.
- DeepSeek — chat completions API reference: valid model identifiers.
- Hugging Face — DeepSeek-V4-Pro model card: parameter counts, precision, architecture, context length.
- Hugging Face — DeepSeek-V4-Pro LICENSE file, raw: the MIT license text we verified byte for byte.
- Hugging Face — DeepSeek-V4-Flash model card: Flash parameter counts and precision.
- Hugging Face — DeepSeek-V4-Flash LICENSE file, raw: byte-identical to the Pro license.
Independent evaluators
- Artificial Analysis — model leaderboard: Intelligence Index v4.1 scores and cost per task for every effort level quoted here.
- Artificial Analysis — AA-Omniscience: knowledge reliability and hallucination scores, and the index definition.
- Artificial Analysis — intelligence benchmarking methodology: Index v4.1 composition and category weights.
- Artificial Analysis — Claude Opus 5: index score by effort level, cost per task, speed and latency.
- Artificial Analysis — DeepSeek V4 Pro: independent index score and cost to run the suite.
Last compared: July 28, 2026. We researched both models using vendor documentation and independent evaluations rather than running our own benchmark suite. Benchmark scores are point-in-time readings against Artificial Analysis Intelligence Index v4.1 and change as the evaluator re-runs and revises its index. We hold no affiliate relationship with Anthropic or DeepSeek.
Our Verdict
Claude Opus 5 wins overall, but conditionally. It leads DeepSeek V4 Pro by 17 index points at matched max effort (61 to 44) and by 41 points on AA-Omniscience (31 to -10), where DeepSeek V4 Pro's negative score means it asserts more wrong answers than correct ones. DeepSeek V4 wins decisively on cost, charging about 29 times less per output token, and wins outright on licensing with plain MIT weights verified at file level. Reasoning effort never changes the per-token rate, so no Opus 5 setting reaches DeepSeek's cost band. Pick Opus 5 when a wrong answer is expensive to detect; pick DeepSeek V4 when outputs are cheaply verifiable, volume dominates, or weights must run on your own hardware.
Choose Claude Opus 5
Anthropic's frontier reasoning model — top of the independent index at half the price of Fable 5.
Try Claude Opus 5 →Choose DeepSeek V4
Chinese open-source flagship: 1.6T MoE (49B active), 1M context, 80.6% SWE-bench Verified, MIT license — V4-Pro input costs about one-eleventh of Claude Opus 4.7
Try DeepSeek V4 →Frequently Asked Questions
Is Claude Opus 5 better than DeepSeek V4?
Claude Opus 5 wins overall, but conditionally. It leads DeepSeek V4 Pro by 17 index points at matched max effort (61 to 44) and by 41 points on AA-Omniscience (31 to -10), where DeepSeek V4 Pro's negative score means it asserts more wrong answers than correct ones. DeepSeek V4 wins decisively on cost, charging about 29 times less per output token, and wins outright on licensing with plain MIT weights verified at file level. Reasoning effort never changes the per-token rate, so no Opus 5 setting reaches DeepSeek's cost band. Pick Opus 5 when a wrong answer is expensive to detect; pick DeepSeek V4 when outputs are cheaply verifiable, volume dominates, or weights must run on your own hardware.
Which is cheaper, Claude Opus 5 or DeepSeek V4?
Claude Opus 5 is priced at $5 in / $25 out per M tokens. DeepSeek V4 is priced at $0.14 in / $0.28 out per M tokens (free plan available). Check the pricing comparison section above for a full breakdown.
What are the main differences between Claude Opus 5 and DeepSeek V4?
The key differences span across 13 features we compared. For Intelligence Index v4.1, matched max effort, Claude Opus 5 offers 61 while DeepSeek V4 offers 44. For Intelligence Index v4.1, each model default, Claude Opus 5 offers 59 at high while DeepSeek V4 offers 43 at high. For AA-Omniscience, max effort, Claude Opus 5 offers 31 while DeepSeek V4 offers -10. See the full feature comparison table above for all details.

