Claude Opus 4.8 vs DeepSeek V4: Frontier Capability vs Open-Weight Cost (2026)
Opus 4.8 leads the Artificial Analysis index 56 vs 44 and SWE-bench 88.6% vs 80.6%; DeepSeek V4 is ~90x cheaper and open-weight. Split verdict by use case.
Feature Comparison
| Feature | Claude Opus 4.8 | DeepSeek V4 |
|---|---|---|
| AA Intelligence Index (Artificial Analysis, same evaluator) | 61, ranked number one | 52 (V4-Pro, max reasoning) |
| SWE-bench Verified | 88.6% (Anthropic reports) | 80.6% (DeepSeek reports) |
| Output price (per million tokens) | $25.00 (verified) | Flash $0.28 / Pro $0.87 (verified) |
| Input price (per million tokens) | $5.00 (verified) | Flash $0.14 / Pro $0.435 (verified) |
| License / openness | Closed, API and cloud only | Open weights, MIT license |
| Self-hostable | No | Yes, incl. Huawei Ascend |
| Context window | 1,000,000 tokens (verified) | 1,000,000 tokens (verified) |
| Western data residency / compliance | US, plus AWS/GCP/Microsoft Foundry | China-hosted API (or self-host) |
| SWE-bench Pro (agentic coding) | 69.2% (Anthropic reports) | Not reported |
| Computer use — Online-Mind2Web | 84% (Anthropic reports) | Not reported |
Pricing Comparison
Claude Opus 4.8
DeepSeek V4
Detailed Comparison
Claude Opus 4.8 and DeepSeek V4 are the two flagship large language models compared here, and they sit at opposite ends of the same frontier. Claude Opus 4.8 is Anthropic's closed, US-hosted top model, announced May 28, 2026, priced at 5 dollars per million input tokens and 25 dollars per million output tokens. DeepSeek V4 is the open-weight Chinese flagship from DeepSeek, shipped under an MIT license on Hugging Face, with a hosted API tier (V4-Flash at 0.14 dollars input and 0.28 dollars output per million tokens, V4-Pro at 0.435 dollars input and 0.87 dollars output). On the one independent evaluator that scores both the same way, Artificial Analysis, Opus 4.8 leads the Intelligence Index — 56 versus 44 for DeepSeek V4-Pro as of June 2026 — and on SWE-bench Verified Anthropic reports 88.6 percent against DeepSeek's reported 80.6 percent. DeepSeek V4 is dramatically cheaper and can be downloaded and self-hosted; Opus 4.8 is the stronger model on capability, agentic coding, computer use, and Western data residency. Best for raw frontier capability, agentic coding, and US-hosted compliance: Claude Opus 4.8. Best for cost, open weights, and self-hosting: DeepSeek V4.
Quick Verdict
This is a split verdict by use case, not a single overall winner. We compared the published benchmarks and specifications of Claude Opus 4.8 and DeepSeek V4 side-by-side, pulled the pricing directly from each vendor's own pages, and added our own hands-on observations on Opus 4.8, which we use daily in production. We have not run weeks of controlled side-by-side benchmarking of both models on identical tasks, so where we lean on numbers we attribute them. The honest summary is that these two models are not really fighting for the same buyer. Here is the short version.
- Best for raw frontier capability: Claude Opus 4.8. On the Artificial Analysis Intelligence Index — the one composite both models are scored on by the same evaluator — Opus 4.8 sits among the very top models at 56 as of June 2026, while DeepSeek V4-Pro in its maximum reasoning mode scores 44.
- Best for agentic coding: Claude Opus 4.8. Anthropic reports 88.6 percent on SWE-bench Verified and 69.2 percent on SWE-bench Pro; DeepSeek reports 80.6 percent on SWE-bench Verified and 55.4 percent on SWE-bench Pro. Opus 4.8 leads on both comparable coding benchmarks — about 8 points on Verified and about 14 on Pro.
- Best for cost: DeepSeek V4, and it is not close. V4-Flash output at 0.28 dollars per million tokens is roughly 90 times cheaper than Opus 4.8 output at 25 dollars. Even V4-Pro at 0.87 dollars output is about 29 times cheaper.
- Best for open weights and self-hosting: DeepSeek V4. The weights ship under an MIT license and run on your own hardware, including Huawei Ascend chips. Opus 4.8 is closed and API-only.
- Best for Western data residency and compliance: Claude Opus 4.8. It is hosted by Anthropic in the US and on AWS, Google Cloud, and Microsoft Foundry. DeepSeek's API is hosted in China, which is a non-starter for many regulated buyers unless self-hosted.
Bottom line: if you need the strongest model and you can pay for it — or you need US-hosted compliance — pick Claude Opus 4.8. If you are cost-constrained, want to own your weights, or need to self-host for sovereignty reasons, DeepSeek V4 gives you frontier-adjacent quality at a fraction of the price. We did not crown a single winner because the two models optimize for different things.
At a Glance
Before the detail, here is the side-by-side that frames everything below. All pricing in this table was fetched directly from each vendor's pricing page in June 2026. All benchmark figures are attributed to their source.
| Dimension | Claude Opus 4.8 | DeepSeek V4 |
|---|---|---|
| Vendor / origin | Anthropic (US) | DeepSeek (China) |
| License | Closed, API and cloud only | Open weights, MIT license |
| Released | May 28, 2026 | April 2026 |
| Input price (per million tokens) | 5 dollars (verified) | Flash 0.14 dollars, Pro 0.435 dollars (verified) |
| Output price (per million tokens) | 25 dollars (verified) | Flash 0.28 dollars, Pro 0.87 dollars (verified) |
| Cache-read input (per million tokens) | 0.50 dollars (verified) | Flash 0.0028 dollars, Pro 0.0036 dollars (verified) |
| Context window | 1,000,000 tokens (verified) | 1,000,000 tokens (verified) |
| AA Intelligence Index (June 2026) | 56, among the top models (Artificial Analysis) | 44 for V4-Pro max reasoning (Artificial Analysis) |
| SWE-bench Verified | 88.6 percent (Anthropic reports) | 80.6 percent (DeepSeek reports) |
| Self-hostable | No | Yes, including Huawei Ascend chips |
| Data residency | US, plus AWS, Google Cloud, Microsoft Foundry | China-hosted API, or self-host anywhere |
Overview of Each Model
Claude Opus 4.8
Claude Opus 4.8 is Anthropic's flagship model, announced May 28, 2026, aimed squarely at agentic coding, computer use, and multi-agent orchestration. It kept the same headline price as Opus 4.7 — 5 dollars per million input tokens and 25 dollars per million output tokens — while improving on the metrics Anthropic cares most about. It ranks among the very top of the Artificial Analysis Intelligence Index at 56 as of June 2026, second only to Anthropic's own Claude Fable 5 across the models that evaluator ranks. Anthropic reports 88.6 percent on SWE-bench Verified, 69.2 percent on SWE-bench Pro, and 74.6 percent on Terminal-Bench 2.1, alongside an 84 percent score on Online-Mind2Web that Anthropic calls its best-tested computer-use result. It carries a 1,000,000-token context window at standard pricing, and adds Dynamic Workflows for orchestrating large numbers of parallel subagents plus explicit effort controls that trade latency against reasoning depth. In our daily production use, the real day-to-day win over Opus 4.7 is speed and reliability rather than a leap in raw capability: it reaches correct results faster and verifies its own edits instead of declaring a task done without checking. It is closed and available only through Anthropic's API, claude.ai, Claude Code, and the major cloud marketplaces. For the full breakdown, see our Claude Opus 4.8 review.
DeepSeek V4
DeepSeek V4 is the Chinese open-weight flagship, shipped in two sizes: V4-Pro, a 1.6-trillion-parameter mixture-of-experts model with about 49 billion parameters active per token, and V4-Flash, a 284-billion-parameter model with about 13 billion active. Both carry a 1,000,000-token context window with up to 384K tokens of output, and both ship under an MIT license that permits free commercial use, redistribution, and modification of the weights — although the training code and data recipe are not released, so this is open weights rather than fully open source. DeepSeek reports 80.6 percent on SWE-bench Verified for V4-Pro, and Artificial Analysis scores V4-Pro at 44 on its Intelligence Index in maximum reasoning mode as of June 2026, among the leading open-weight models. The architecture is genuinely novel rather than just bigger: a Hybrid Attention design combining Compressed Sparse Attention and Heavily Compressed Attention cuts inference compute and KV-cache footprint sharply versus the previous generation, and three built-in thinking modes let you dial cost against quality per request. It is the first major Chinese frontier model with day-one inference on Huawei Ascend hardware, and the hosted API is OpenAI-compatible. The headline, though, is price: V4-Flash output sits at 0.28 dollars per million tokens. Our full DeepSeek V4 review covers the architecture and licensing in more depth.
Pricing Compared
This is where the two models diverge most violently, and it is the single most important thing to understand about this matchup. We fetched every number below directly from each vendor's pricing page in June 2026.
| Tier | Input (per million tokens) | Output (per million tokens) | Cache read (per million tokens) |
|---|---|---|---|
| Claude Opus 4.8 (standard) | 5.00 dollars | 25.00 dollars | 0.50 dollars |
| Claude Opus 4.8 (Batch API, 50 percent off) | 2.50 dollars | 12.50 dollars | — |
| Claude Opus 4.8 (Fast Mode) | 10.00 dollars | 50.00 dollars | — |
| DeepSeek V4-Flash | 0.14 dollars | 0.28 dollars | 0.0028 dollars |
| DeepSeek V4-Pro | 0.435 dollars | 0.87 dollars | 0.0036 dollars |
Run the arithmetic and the gap is staggering. On output tokens, the comparison most people care about because output dominates real agentic spend, Opus 4.8 at 25 dollars is roughly 90 times the cost of V4-Flash at 0.28 dollars, and roughly 29 times the cost of V4-Pro at 0.87 dollars. On input tokens, Opus 4.8 at 5 dollars is about 36 times V4-Flash and about 11 times V4-Pro. Even Opus 4.8's Batch API discount and aggressive prompt caching — a cache read at 0.50 dollars is genuinely cheap by frontier standards — cannot close a gap of that magnitude.
One scheduling change to price in: DeepSeek has announced that from mid-July 2026 its API moves to peak and off-peak pricing, with peak hours — 9 a.m. to noon and 2 p.m. to 6 p.m. Beijing time — billed at roughly twice the off-peak rate. The flat rates quoted above become the off-peak tariff, so batch workloads that can run outside China's business hours keep today's economics, while always-on traffic pays more during those windows. DeepSeek is also deprecating the legacy deepseek-chat and deepseek-reasoner API names on July 24, 2026, so integrations should move to the current versioned model names before then.
One nuance worth flagging honestly: a self-hosted DeepSeek deployment is not free. The API prices above are the cheap path; running V4-Pro yourself in full precision requires enterprise GPU clusters, and even V4-Flash needs INT4 or INT8 quantization to fit on a single high-end consumer card. The open weights buy you control and remove per-token billing, but they shift cost into hardware and operations. For most teams, the hosted DeepSeek API is the relevant comparison, and there Opus 4.8 simply costs an order of magnitude more per token.
Benchmarks Compared
Benchmarks across two different labs are a minefield, because vendors pick favorable evaluations and report them their own way. We discipline this by leaning on the one independent evaluator that scores both models the same way — Artificial Analysis — and treating vendor-reported figures as attributed claims, not verified facts.
| Benchmark | Claude Opus 4.8 | DeepSeek V4 | Like-for-like? |
|---|---|---|---|
| AA Intelligence Index (Artificial Analysis, June 2026) | 56 (top tier) | 44 (V4-Pro, max reasoning) | Yes — same evaluator |
| SWE-bench Verified | 88.6 percent (Anthropic reports) | 80.6 percent (DeepSeek reports) | Same benchmark, different labs |
| SWE-bench Pro | 69.2 percent (Anthropic reports) | 55.4 percent (DeepSeek reports) | Same benchmark, different labs |
| Terminal-Bench | 74.6 percent on 2.1 (Anthropic reports) | 67.9 percent on 2.0 (DeepSeek reports) | Different benchmark versions |
| Online-Mind2Web (computer use) | 84 percent (Anthropic reports) | Not reported | No counterpart |
| Output speed (Artificial Analysis) | 60.6 tokens per second | Not listed comparably | No clean counterpart |
The cleanest signal is the Artificial Analysis Intelligence Index, because it is one evaluator running the same battery on both: Opus 4.8 at 56 versus V4-Pro at 44 as of June 2026, a roughly 12-point gap, with Opus 4.8 sitting in the top tier overall. On SWE-bench Verified — the same benchmark, but self-reported by each lab — Anthropic's 88.6 percent leads DeepSeek's 80.6 percent by 8 points. We will not pretend that 8-point gap is exact, because it comes from two separate evaluation harnesses, but the direction is consistent with the independent Intelligence Index.
The one place we have no clean head-to-head is computer use: Anthropic reports 84 percent on Online-Mind2Web, while DeepSeek does not publish a comparable computer-use benchmark, so we leave that DeepSeek cell blank rather than invent a number. On SWE-bench Pro and Terminal-Bench, both labs do report — Opus 4.8 leads SWE-bench Pro 69.2 percent to 55.4 percent, and its Terminal-Bench 2.1 result of 74.6 percent sits above DeepSeek's 67.9 percent on Terminal-Bench 2.0, though the versions differ. What the numbers say clearly is that Opus 4.8 is the stronger model on capability, and that DeepSeek V4 is remarkably close given its open weights and its price.
Architecture and What Is Actually Different
It is tempting to treat two frontier models as interchangeable black boxes that you poke through an API, but the engineering underneath shapes how they behave, what they cost to run, and where they can be deployed. The two could hardly be more different in philosophy.
Claude Opus 4.8 is a closed model, so Anthropic discloses behavior rather than internals. What it does surface is a set of product-level capabilities built for agentic work: Dynamic Workflows, a research-preview feature in Claude Code that orchestrates large numbers of parallel subagents on a single task, and explicit effort controls that let you trade response latency and token spend against reasoning depth. Anthropic also notes that Opus 4.7 and later use a new tokenizer that can consume more tokens for the same text, which matters when you are modeling cost. In practice, the model's defining trait is a reliability personality: it verifies its own edits and flags problems rather than declaring a task fixed without checking, and Anthropic claims it is roughly four times less likely than Opus 4.7 to let code flaws pass unremarked. We cannot independently verify that multiplier, but the self-checking behavior is real and observable in our daily use.
DeepSeek V4 is the opposite — fully transparent at the architecture level because the weights and a technical report ship publicly. It is a mixture-of-experts model: V4-Pro carries 1.6 trillion total parameters with about 49 billion active per token, V4-Flash carries 284 billion total with about 13 billion active, both trained on more than 32 trillion tokens. The headline innovation is a Hybrid Attention design that combines Compressed Sparse Attention, at four-times compression, with Heavily Compressed Attention, at 128-times compression, to make a 1,000,000-token context affordable to serve. DeepSeek reports this cuts inference compute to a small fraction of the previous generation and shrinks the KV cache dramatically. It also swaps the usual AdamW optimizer for Muon for faster, more stable training, serves experts in FP4 with most other parameters in FP8 to save memory, and bakes three reasoning modes — Non-Think, Think High, and Think Max — directly into the model rather than bolting them on as a separate API. This is why DeepSeek V4 can be both frontier-adjacent in quality and an order of magnitude cheaper: the efficiency is engineered in, not just priced in.
The practical upshot is that Opus 4.8 gives you a polished, self-verifying agent you cannot inspect or move, while DeepSeek V4 gives you an inspectable, movable model that you operate yourself. Neither philosophy is wrong; they serve different risk and cost profiles.
Total Cost of Ownership
Per-token price is the headline, but the real economics depend on volume, caching, and whether you self-host. Here is how to think about it without overstating the case in either direction.
For the hosted-API path, the gap is so large that for high-volume workloads it changes what is buildable. A pipeline that processes, say, a billion output tokens a month costs about 25,000 dollars on Opus 4.8 at standard pricing, around 12,500 dollars with the Batch API discount, roughly 870 dollars on DeepSeek V4-Pro, and about 280 dollars on V4-Flash. Those are not small percentage differences; they are different orders of magnitude, and they decide whether an idea is economically viable at all. Prompt caching narrows the input side meaningfully — Opus 4.8 cache reads at 0.50 dollars per million tokens are genuinely cheap — but output dominates agentic spend, and there Opus 4.8 has no answer to DeepSeek's pricing.
For the self-hosted path, the calculus flips from per-token billing to capital and operations. DeepSeek's open weights remove the API meter entirely, but you pay in hardware: full-precision V4-Pro requires enterprise GPU clusters, and even V4-Flash needs INT4 or INT8 quantization to fit on a single high-end consumer card. For a team with steady, predictable, very high volume and the operational maturity to run model infrastructure, self-hosting V4 can be the cheapest option of all and the only one that guarantees data never leaves your premises. For a team with spiky or modest volume, the hosted DeepSeek API is the sensible comparison — and it is still an order of magnitude cheaper than Opus 4.8. The honest conclusion is that DeepSeek wins on cost in every scenario; the only question is by how much and at what operational price.
How We Tested
Honesty about methodology matters more in a cross-lab, cross-country comparison than almost anywhere else. Here is exactly what is hands-on and what is research.
We use Claude Opus 4.8 daily in our own production workflow — agentic coding in Claude Code, multi-file refactors, and content pipelines — so our observations on its behavior (speed over Opus 4.7, tighter instruction-following, self-verification of edits) are first-hand. We have run DeepSeek V4 through its hosted API on coding and reasoning prompts to confirm it behaves as documented, including its three thinking modes, but we have not stood up a self-hosted V4-Pro cluster, and we have not run weeks of controlled, identical-task benchmarking of both models against each other. For that reason, every capability claim that rests on numbers is attributed to its source — Artificial Analysis for the independent index, each vendor for their own SWE-bench figures — and we pulled all pricing by fetching each vendor's pricing page directly rather than trusting secondhand summaries. Where we could not verify a like-for-like number, we said so and left the cell blank. That is the standard we hold ourselves to.
Winner by Category
A single overall winner would be dishonest here, because these models are tuned for different buyers. Here is who wins what.
- Best for raw frontier capability: Claude Opus 4.8. Top tier on the Artificial Analysis Intelligence Index at 56, ahead of V4-Pro at 44 as of June 2026.
- Best for agentic coding: Claude Opus 4.8. Leads SWE-bench Verified 88.6 percent to 80.6 percent and SWE-bench Pro 69.2 percent to 55.4 percent, and adds a computer-use result (84 percent on Online-Mind2Web) DeepSeek does not report.
- Best for cost: DeepSeek V4. An order of magnitude cheaper per token on the hosted API, and roughly 90 times cheaper on V4-Flash output specifically.
- Best for open weights and self-hosting: DeepSeek V4. MIT-licensed downloadable weights, with native Huawei Ascend support; Opus 4.8 cannot be self-hosted at all.
- Best for Western data residency and compliance: Claude Opus 4.8. US-hosted with major cloud options; DeepSeek's hosted API runs in China.
- Best for long-context work: Tie. Both ship a 1,000,000-token context window at standard pricing, and DeepSeek allows up to 384K output tokens.
- Best for reliability and safety posture: Claude Opus 4.8. Anthropic's self-verification behavior and safety tuning are a real production advantage, though the underlying claims are vendor-reported.
Pros and Cons
Claude Opus 4.8 — Pros
- Top tier on the Artificial Analysis Intelligence Index at 56 as of June 2026, second only to Anthropic's own Claude Fable 5.
- Leads agentic coding on the comparable benchmark: 88.6 percent SWE-bench Verified versus 80.6 percent reported for DeepSeek V4-Pro.
- Best-tested computer-use model in Anthropic's lineup at 84 percent on Online-Mind2Web.
- US-hosted with AWS, Google Cloud, and Microsoft Foundry options — clears Western data-residency requirements DeepSeek cannot.
- Cautious, self-verifying reliability personality that flags problems instead of declaring tasks done unchecked.
- 1,000,000-token context at standard pricing, plus Dynamic Workflows and effort controls for large agentic jobs.
Claude Opus 4.8 — Cons
- Costs an order of magnitude more per token than DeepSeek's hosted API — 25 dollars output per million versus 0.28 to 0.87 dollars.
- Closed model: no self-hosting, no weights, no sovereignty option.
- Coding benchmarks are vendor-reported and not yet independently verified outside the Artificial Analysis composite.
- Fast Mode doubles per-token cost to 10 dollars input and 50 dollars output per million for its speed-up.
- Time to first token is on the higher end for its price tier at around 31 seconds in Artificial Analysis testing.
DeepSeek V4 — Pros
- Frontier-adjacent capability at open weights: 44 on the Artificial Analysis Intelligence Index (June 2026) and a reported 80.6 percent SWE-bench Verified.
- Dramatically cheaper hosted API — V4-Flash output at 0.28 dollars per million tokens is roughly 90 times cheaper than Opus 4.8 output.
- MIT-licensed weights downloadable from Hugging Face for free commercial use, redistribution, and modification.
- Self-hostable for full data sovereignty, with day-one support on Huawei Ascend chips that removes NVIDIA dependency.
- 1,000,000-token context with up to 384K output tokens, plus three built-in reasoning modes to tune cost against quality.
- Near-free cache-hit input pricing at 0.0028 dollars per million tokens for V4-Flash, which makes stable-prompt RAG and tool loops almost cost-free.
DeepSeek V4 — Cons
- Trails Opus 4.8 on the independent Intelligence Index, 44 versus 56 as of June 2026, and on reported SWE-bench Verified, 80.6 versus 88.6 percent.
- Hosted API runs in China, a non-starter for US Federal, EU healthcare, and many regulated buyers without a Western reseller.
- Open weights, not open source: training code and data recipe are not released, so the run cannot be fully reproduced.
- Self-hosting requires serious hardware — full-precision V4-Pro needs enterprise GPU clusters, and V4-Flash needs quantization to fit a single high-end card.
- Open-source inference frameworks take time to land first-class V4 support, and the Think-mode protocol is more complex to integrate than the previous generation.
When to Pick Each
When to pick Claude Opus 4.8
Pick Opus 4.8 when capability and trust matter more than per-token cost. If you are doing serious agentic coding, multi-file refactors, or computer-use automation, it is the stronger model on every comparable benchmark, and in our own daily use it is the more reliable one — it verifies its own work and stays on the brief. Pick it if you are a Western enterprise with data-residency or compliance obligations, because US hosting and the AWS, Google Cloud, and Microsoft Foundry options clear bars DeepSeek's China-hosted API cannot. And pick it if your workload is moderate in volume but high in value, where paying 25 dollars per million output tokens for the best result is a rounding error against engineer time.
When to pick DeepSeek V4
Pick DeepSeek V4 when cost, control, or sovereignty dominate. If you are running high-volume inference where token spend is the binding constraint, an order-of-magnitude cheaper API changes what is economically viable — and V4-Flash at 0.28 dollars output makes use cases that are simply unaffordable on Opus 4.8 routine. Pick it if you need to own your weights: the MIT license lets you self-host, fine-tune, and redistribute, and the Huawei Ascend support means you are not locked to a single chip vendor. Pick it if you are operating where Chinese hosting is acceptable or where self-hosting is mandatory for data sovereignty. You give up a measurable slice of frontier capability and the Western compliance story, but you get most of the quality at a fraction of the price.
Final Verdict
This is a split verdict by use case, tilted toward Claude Opus 4.8 on capability and toward DeepSeek V4 on cost and openness. On the one independent evaluator that scores both — Artificial Analysis — Opus 4.8 sits in the top tier at 56 versus 44 for DeepSeek V4-Pro as of June 2026, and it leads the comparable SWE-bench Verified figure 88.6 percent to 80.6 percent. It is the stronger model, the more reliable one in our hands-on production use, and the only one that clears Western data-residency requirements. DeepSeek V4, in return, costs roughly an order of magnitude less per token on its hosted API, ships MIT-licensed open weights you can self-host, and lands within striking distance of the frontier — a genuinely remarkable result for an open model.
We did not crown a single overall winner because the two models are not really competing for the same buyer. If you need the best model, the strongest agentic coding, or US-hosted compliance, the answer is Claude Opus 4.8. If you are cost-constrained, want to own your weights, or need to self-host for sovereignty, the answer is DeepSeek V4. Both answers are correct — for different people. All benchmark numbers here are vendor-reported or drawn from the Artificial Analysis index and press-relayed; only the pricing is fetch-verified directly from each vendor.
If you are weighing Opus 4.8 against other frontier models, we also ran it head-to-head with OpenAI's flagship in Claude Opus 4.8 vs GPT-5.5, with Google's in Claude Opus 4.8 vs Gemini 3.1 Pro, and against its own predecessor in Claude Opus 4.8 vs Claude Opus 4.7.
Frequently Asked Questions
Is Claude Opus 4.8 better than DeepSeek V4?
On capability, yes. Claude Opus 4.8 sits in the top tier of the Artificial Analysis Intelligence Index at 56 versus 44 for DeepSeek V4-Pro as of June 2026, and it leads SWE-bench Verified 88.6 percent to 80.6 percent. But DeepSeek V4 is roughly an order of magnitude cheaper and is open-weight and self-hostable, so the better choice depends on whether you are optimizing for capability or for cost and control.
How much cheaper is DeepSeek V4 than Claude Opus 4.8?
Dramatically. On output tokens, DeepSeek V4-Flash at 0.28 dollars per million is roughly 90 times cheaper than Claude Opus 4.8 at 25 dollars per million, and V4-Pro at 0.87 dollars is about 29 times cheaper. On input tokens, V4-Flash at 0.14 dollars is about 36 times cheaper than Opus 4.8 at 5 dollars. All prices were fetched directly from each vendor's pricing page in June 2026. From mid-July 2026, DeepSeek moves to peak and off-peak pricing — these flat rates become the off-peak tariff, and peak hours (9 a.m. to noon and 2 p.m. to 6 p.m. Beijing time) are billed at roughly twice that rate.
Is DeepSeek V4 open source?
It is open weights, not fully open source. DeepSeek V4 ships its model weights under an MIT license on Hugging Face, allowing free commercial use, redistribution, and modification. However, the training code and data recipe are not released, so the community cannot fully reproduce the training run. You can self-host and fine-tune the model, but you cannot rebuild it from scratch.
Can I self-host DeepSeek V4 or Claude Opus 4.8?
You can self-host DeepSeek V4 because its weights are MIT-licensed and downloadable, including native support for Huawei Ascend chips. You cannot self-host Claude Opus 4.8 — it is a closed model available only through Anthropic's API, claude.ai, Claude Code, and the major cloud marketplaces. Self-hosting V4 requires serious hardware: full-precision V4-Pro needs enterprise GPU clusters, and V4-Flash needs quantization to fit a single high-end card.
What is the context window for each model?
Both ship a 1,000,000-token context window. Claude Opus 4.8 offers the full 1M-token window at standard pricing. DeepSeek V4 also provides 1,000,000 tokens of context on both V4-Pro and V4-Flash, with up to 384K tokens of output. On context length specifically, the two are effectively tied.
Which model is better for coding?
Claude Opus 4.8 leads on every comparable coding signal. Anthropic reports 88.6 percent on SWE-bench Verified and 69.2 percent on SWE-bench Pro, against DeepSeek's reported 80.6 percent on SWE-bench Verified and 55.4 percent on SWE-bench Pro, and Opus 4.8 also adds a computer-use result (84 percent on Online-Mind2Web) DeepSeek does not report. In our daily hands-on use, Opus 4.8 is also the more reliable agentic coder, verifying its own edits. DeepSeek V4 is still strong and far cheaper, which makes it attractive for high-volume coding where cost dominates.
Is DeepSeek V4 safe to use for a Western company?
It depends on your data-residency rules. DeepSeek's hosted API runs in China, which keeps many regulated buyers — US Federal, EU healthcare — from adopting it without a Western reseller. The MIT-licensed open weights let you sidestep this by self-hosting the model on your own infrastructure anywhere in the world. If compliance is the concern and you cannot self-host, Claude Opus 4.8's US hosting and major-cloud options are the safer default.
How do the two models score on independent benchmarks?
The cleanest independent signal is the Artificial Analysis Intelligence Index, which scores both with the same battery: Claude Opus 4.8 sits in the top tier at 56, while DeepSeek V4-Pro in maximum reasoning mode scores 44 as of June 2026. That gap is consistent with the vendor-reported SWE-bench Verified figures, where Anthropic's 88.6 percent leads DeepSeek's 80.6 percent.
What are the different DeepSeek V4 tiers?
DeepSeek V4 ships in two sizes. V4-Pro is a 1.6-trillion-parameter mixture-of-experts model with about 49 billion parameters active per token, priced at 0.435 dollars input and 0.87 dollars output per million tokens. V4-Flash is a 284-billion-parameter model with about 13 billion active, priced at 0.14 dollars input and 0.28 dollars output. Both carry a 1,000,000-token context window, and both support three reasoning modes — Non-Think, Think High, and Think Max.
Does Claude Opus 4.8 have a cheaper mode?
Yes, two cost levers. The Batch API offers a 50 percent discount for asynchronous workloads, bringing Opus 4.8 to 2.50 dollars input and 12.50 dollars output per million tokens. Prompt caching drops repeated input to 0.50 dollars per million tokens on cache reads. Even with both applied, however, Opus 4.8 remains far more expensive per token than DeepSeek V4's hosted API.
Which should I choose for a high-volume production workload?
For pure high volume where token cost is the binding constraint, DeepSeek V4 is usually the rational choice — an order-of-magnitude cheaper API makes workloads viable that Opus 4.8 cannot support economically. Choose Claude Opus 4.8 instead when each call is high-value, when you need the strongest agentic coding, or when Western data residency is mandatory. Many teams run both: Opus 4.8 for the hardest tasks, DeepSeek V4 for everything bulk.
When were these models released and is this comparison current?
Claude Opus 4.8 was announced May 28, 2026, and DeepSeek V4 shipped in April 2026. This comparison was last updated in July 2026, with all pricing verified directly against each vendor's own pages and all benchmark figures attributed to Artificial Analysis or to each vendor's own reports.
Our Verdict
Split verdict by use case, tilted toward Claude Opus 4.8 on capability and toward DeepSeek V4 on cost and openness. On the one independent evaluator that scores both the same way, Artificial Analysis, Opus 4.8 leads the Intelligence Index at 56 versus 44 for DeepSeek V4-Pro as of June 2026 — second only to Anthropic's own Claude Fable 5 among ranked models — and it leads the comparable SWE-bench Verified figure 88.6% to 80.6%. It is the stronger, more reliable model in our hands-on use and the only one clearing Western data-residency requirements. DeepSeek V4 costs roughly an order of magnitude less per token on its hosted API (V4-Flash output $0.28 versus $25 per million), ships MIT-licensed open weights you can self-host, and lands within striking distance of the frontier. No single overall winner: the two models target different buyers. Best for raw capability, agentic coding, and US-hosted compliance: Claude Opus 4.8. Best for cost, open weights, and self-hosting: DeepSeek V4. All benchmark numbers are vendor-reported or from the Artificial Analysis index; only pricing is fetch-verified.
Choose Claude Opus 4.8
Anthropic's flagship model for agentic coding, computer use, and multi-agent orchestration.
Try Claude Opus 4.8 →Choose DeepSeek V4
Chinese open-source flagship: 1.6T MoE (49B active), 1M context, 80.6% SWE-bench Verified, MIT license — V4-Pro input costs about one-eleventh of Claude Opus 4.7
Try DeepSeek V4 →Frequently Asked Questions
Is Claude Opus 4.8 better than DeepSeek V4?
Split verdict by use case, tilted toward Claude Opus 4.8 on capability and toward DeepSeek V4 on cost and openness. On the one independent evaluator that scores both the same way, Artificial Analysis, Opus 4.8 leads the Intelligence Index at 56 versus 44 for DeepSeek V4-Pro as of June 2026 — second only to Anthropic's own Claude Fable 5 among ranked models — and it leads the comparable SWE-bench Verified figure 88.6% to 80.6%. It is the stronger, more reliable model in our hands-on use and the only one clearing Western data-residency requirements. DeepSeek V4 costs roughly an order of magnitude less per token on its hosted API (V4-Flash output $0.28 versus $25 per million), ships MIT-licensed open weights you can self-host, and lands within striking distance of the frontier. No single overall winner: the two models target different buyers. Best for raw capability, agentic coding, and US-hosted compliance: Claude Opus 4.8. Best for cost, open weights, and self-hosting: DeepSeek V4. All benchmark numbers are vendor-reported or from the Artificial Analysis index; only pricing is fetch-verified.
Which is cheaper, Claude Opus 4.8 or DeepSeek V4?
Claude Opus 4.8 is priced at $5 in / $25 out per M tokens. DeepSeek V4 is priced at $0.14 in / $0.28 out per M tokens (free plan available). Check the pricing comparison section above for a full breakdown.
What are the main differences between Claude Opus 4.8 and DeepSeek V4?
The key differences span across 10 features we compared. For AA Intelligence Index (Artificial Analysis, same evaluator), Claude Opus 4.8 offers 61, ranked number one while DeepSeek V4 offers 52 (V4-Pro, max reasoning). For SWE-bench Verified, Claude Opus 4.8 offers 88.6% (Anthropic reports) while DeepSeek V4 offers 80.6% (DeepSeek reports). For Output price (per million tokens), Claude Opus 4.8 offers $25.00 (verified) while DeepSeek V4 offers Flash $0.28 / Pro $0.87 (verified). See the full feature comparison table above for all details.

