Claude Opus 5 vs Kimi K3: 4 Points, 2.8x the Bill (2026)
We researched both: Opus 5 scores 61 to Kimi K3’s 57 on the independent index but costs 2.8 times more per task. A threshold, not an average.
Feature Comparison
| Feature | Claude Opus 5 | Kimi K3 |
|---|---|---|
| Artificial Analysis Intelligence Index (v4.1, independent), highest effort | 61 (raw 60.69) | 57 (raw 57.11) |
| Same index at each model’s default effort | 59 (high) | 57 (max) |
| Measured cost per index task, highest effort | USD 2.03 | USD 0.72 |
| Measured cost per index task, default effort | USD 1.06 | USD 0.72 |
| Input price per million tokens | USD 5.00 | USD 3.00 |
| Output price per million tokens | USD 25.00 | USD 15.00 |
| Cache-hit input per million tokens | USD 0.50 | USD 0.30 |
| Context window | 1,000,000 tokens | 1,048,576 tokens |
| Long-context price premium | None at any length | None published |
| Maximum output tokens | 128,000 (300,000 via batch) | 131,072 default, up to 1,048,576 |
| Training cutoff | May 2026 | Not published |
| Reasoning effort levels | Five: low, medium, high, xhigh, max (default high) | Three: low, high, max (default max) |
| AA-Omniscience accuracy (independent) | 54.20 percent | 45.95 percent |
| AA-Omniscience hallucination rate (independent) | 50.07 percent | 50.94 percent |
| Median output speed, long prompts (independent) | 52.8 tokens per second | 33.0 tokens per second |
| Median time to first token, 100K-token context | 43.3 seconds | 9.0 seconds |
| Native vision | Yes | Yes |
| Model weights | Closed, API only | Published July 27, 2026 — 96 weight files, ungated, under the custom Kimi K3 License (not MIT) |
| Architecture disclosure | Undisclosed | Mixture-of-experts, 2.8T total, 16 of 896 experts active |
| Data-retention requirement | None; zero-retention eligible | None published |
Pricing Comparison
Claude Opus 5
Kimi K3
Detailed Comparison
Claude Opus 5 vs Kimi K3 in 2026: Claude Opus 5 is Anthropic’s closed frontier model, released July 24, 2026, priced at USD 5 per million input tokens and USD 25 per million output tokens, with a 1,000,000-token context window carrying no long-context premium, a 128,000-token maximum synchronous output, and a May 2026 training cutoff. It scores 61 on version 4.1 of the independent Artificial Analysis Intelligence Index at its highest effort setting, and 59 at its default setting. Kimi K3 is Moonshot AI’s 2.8-trillion-parameter mixture-of-experts model, available by API since July 16, 2026, priced at USD 3 per million input tokens, USD 0.30 per million cache-hit input tokens, and USD 15 per million output tokens, with a 1,048,576-token context window and native vision. It scores 57 on the same version of the same independent index. Measured cost per index task is USD 2.03 for Opus 5 at max effort against USD 0.72 for Kimi K3 — about 2.8 times. We researched both from vendor documentation and independent measurement rather than running our own benchmarks. There is no overall winner: Opus 5 owns the top of the capability curve and a fresher cutoff, while Kimi K3 delivers near-frontier capability, a matching context window, and a route to self-hosting at a fraction of the measured cost.
Our verdict at a glance
This page exists because of one number, and that number is smaller than it looks. On version 4.1 of the Artificial Analysis Intelligence Index — the same index, the same version, measured by the same third party — Claude Opus 5 scores 61 and Kimi K3 scores 57. Four points. The raw figures behind those rounded scores are 60.69 and 57.11, so the true gap is 3.58 points, and the rounding flatters Opus 5 slightly.
Against those four points sits a cost difference of about 2.8 times on the metric that matters most, which is not the price list but the measured spend to finish the work. Artificial Analysis reports USD 2.03 to run its index task set on Opus 5 at max effort, against USD 0.72 on Kimi K3.
The wrong way to read this is as an average. Nobody buys four points of index score spread evenly across their workload. The right way to read it is as a threshold. Somewhere in your work there is a hardest task, and the question is whether that task sits above or below what a model scoring 57 can finish. If it sits above, the four points are the entire reason Opus 5 exists and no amount of Kimi K3’s value substitutes for a ceiling you actually hit. If it sits below — and for most real workloads it does — you are paying 2.8 times for headroom you never touch.
There is a second finding that surprised us, and it is the reason this page does not declare a winner. The 2.8-times multiplier is not a fixed property of these two models. It is a property of a dial. Opus 5 has five effort levels, and at its default setting the comparison changes shape entirely: 59 against 57, at USD 1.06 against USD 0.72, a gap of about 1.5 times rather than 2.8. Drop one more level to medium effort and Opus 5 costs USD 0.62 per task, which is less than Kimi K3, while scoring 56 against Kimi K3’s 57. At that point the two models sit on essentially the same cost-to-capability curve, and Kimi K3 is marginally ahead of it.
So we are naming no overall winner, deliberately. Opus 5 takes the ceiling, the freshness, and long-prompt speed. Kimi K3 takes the price, the openness, and the speed at very long context. Neither of those is a consolation prize.
How we ran this comparison
We researched both models rather than benchmarking them ourselves, so here is exactly where every figure comes from.
Pricing, context windows, output limits, effort parameters, and cutoff dates were read directly from each vendor’s own live documentation on July 27, 2026 — Anthropic’s pricing and model pages, and Moonshot’s API reference and pricing pages. Not from summaries or secondary reporting, because token prices and parameter defaults change quietly and secondary sources go stale within days.
Capability figures come from Artificial Analysis, a third-party evaluator that runs both models on the same task set and publishes the cost of doing so. We use their Intelligence Index version 4.1 throughout. This matters more than it sounds: index versions are not comparable to each other, and a score quoted without its version number is not a score. Artificial Analysis also re-runs models, so these numbers move — earlier reporting quoted Kimi K3 at USD 0.94 per index task, and the live figure today is USD 0.72.
Where a vendor publishes its own benchmark results, we keep them in a separate section and never place them in the same table, sentence, or chart as an independent measurement. And where neither vendor has published something, we say so rather than inferring it: Kimi K3 has no published knowledge cutoff, which is a gap in the record rather than a zero.
Claude Opus 5 in brief
Anthropic released Claude Opus 5 on July 24, 2026, and the headline of the launch was the price that did not change: USD 5 per million input tokens and USD 25 per million output tokens, identical to Claude Opus 4.8 before it. Cache reads are USD 0.50 per million tokens, and the Batch API halves the base rates to USD 2.50 and USD 12.50.
The context window is 1,000,000 tokens with no length-based surcharge anywhere in it. Anthropic’s pricing documentation states that a 900,000-token request bills at the same per-token rate as a 9,000-token request, and that caching and batch discounts apply at standard rates across the full window. Several competing frontier models charge a premium above 200,000 tokens; Opus 5 does not. Maximum output is 128,000 tokens synchronously, rising to 300,000 through the Batch API with a beta header.
Training data cutoff is May 2026, and unusually Anthropic lists the reliable knowledge cutoff as the same month rather than an earlier one. That is the single most concrete advantage Opus 5 holds that has nothing to do with benchmark scores.
Opus 5 exposes five effort levels — low, medium, high, xhigh, and max — with high as the API default. Thinking cannot be disabled at xhigh or max. Weights are closed and the architecture is undisclosed. Opus 5 carries no data-retention requirement and remains eligible for zero-retention agreements; that restriction applies to Anthropic’s Claude Fable 5 and Mythos tiers, not to this one. Teams for whom the price is the obstacle rather than the capability should look at Claude Sonnet 5 instead. Our full write-up of the launch is in our Opus 5 launch analysis.
Kimi K3 in brief
Moonshot AI unveiled Kimi K3 on July 16, 2026, and made it available the same day through its hosted API and consumer platform. It is a mixture-of-experts model with 2.8 trillion total parameters, activating 16 of 896 experts per token, built on Moonshot’s Kimi Delta Attention architecture. Moonshot has said the full architectural and training details will arrive with a technical report that has not yet been published.
Pricing is USD 3 per million input tokens on a cache miss, USD 0.30 per million on a cache hit — a 90 percent discount — and USD 15 per million output tokens. The API is pay-as-you-go with no subscription tier; Moonshot’s documentation draws that line explicitly between the developer platform and its consumer Kimi Membership products. We checked the product plans page on July 27, 2026 and found no notice of suspended or closed signups on any tier.
The context window is 1,048,576 tokens. Maximum output defaults to 131,072 tokens and can be raised to the full window through the max_completion_tokens field; the older max_tokens field is deprecated. The model has native visual understanding, as does Opus 5, and it always reasons — Moonshot calls this Preserved Thinking and it cannot be switched off.
No knowledge cutoff has been published for Kimi K3. Not by Moonshot, not in the API documentation, and Artificial Analysis records the field as empty. Against a model whose cutoff is May 2026 and stated plainly, that absence is itself a fact worth weighing. Our launch coverage sits in our Kimi K3 launch analysis.
The numbers side by side
Every figure in this table is either a vendor-published specification or an independent measurement, read on July 27, 2026. Nothing here is vendor-reported benchmark performance — that is kept separate further down.
| Feature | Claude Opus 5 | Kimi K3 | Edge |
|---|---|---|---|
| Artificial Analysis Intelligence Index (v4.1, independent), highest effort | 61 (raw 60.69) | 57 (raw 57.11) | Opus 5 |
| Same index, at each model’s default effort | 59 (high) | 57 (max) | Opus 5 |
| Measured cost per index task, highest effort | USD 2.03 | USD 0.72 | Kimi K3 |
| Measured cost per index task, default effort | USD 1.06 | USD 0.72 | Kimi K3 |
| Input price per million tokens | USD 5.00 | USD 3.00 | Kimi K3 |
| Output price per million tokens | USD 25.00 | USD 15.00 | Kimi K3 |
| Cache-hit input per million tokens | USD 0.50 | USD 0.30 | Kimi K3 |
| Context window | 1,000,000 tokens | 1,048,576 tokens | Kimi K3 |
| Long-context price premium | None at any length | None published | Tie |
| Maximum output tokens | 128,000 (300,000 via batch) | 131,072 default, up to 1,048,576 | Kimi K3 |
| Training cutoff | May 2026 | Not published | Opus 5 |
| Reasoning effort levels | Five: low, medium, high, xhigh, max (default high) | Three: low, high, max (default max) | Opus 5 |
| AA-Omniscience accuracy (independent) | 54.20 percent | 45.95 percent | Opus 5 |
| AA-Omniscience hallucination rate (independent) | 50.07 percent | 50.94 percent | Tie |
| Median output speed, long prompts (independent) | 52.8 tokens per second | 33.0 tokens per second | Opus 5 |
| Median time to first token, 100K-token context | 43.3 seconds | 9.0 seconds | Kimi K3 |
| Native vision | Yes | Yes | Tie |
| Model weights | Closed, API only | Published July 27, 2026 — 96 weight files, ungated, under the custom Kimi K3 License | Kimi K3 |
| Architecture disclosure | Undisclosed | Mixture-of-experts, 2.8T total, 16 of 896 experts active | Kimi K3 |
| Data-retention requirement | None; zero-retention eligible | None published | Tie |
Pricing: list rates against measured spend
There are two different price comparisons here, and conflating them is the most common mistake made about these two models.
The first is the list rate. Opus 5 charges USD 5 per million input tokens against Kimi K3’s USD 3, and USD 25 per million output tokens against USD 15 — identical ratios, so Opus 5 is almost exactly 1.7 times the sticker price on both sides of the meter. The same ratio holds on cache hits, USD 0.50 against USD 0.30, and both vendors offer a 90 percent cache discount.
The second is what you actually spend to finish a unit of work. Artificial Analysis publishes the cost of running its full index task set on each model: USD 2.03 per task on Opus 5 at max effort, USD 0.72 on Kimi K3. That is about 2.8 times, not 1.7.
The difference between 1.7 and 2.8 is reasoning tokens. Opus 5 at max effort thinks a great deal before it answers, and every one of those thinking tokens is billed at the output rate. The list price tells you what a token costs; the cost per task tells you how many tokens the model chooses to spend. For a reasoning model the second number is the one that lands on your invoice, and it is the one we use throughout this page. It also means the effort dial moves your bill far more than any negotiation on rates ever will: dropping from max to high cuts the cost per task roughly in half, and dropping to medium cuts it by about 70 percent.
The four-point question: where the threshold sits
Four index points for 2.8 times the bill is the trade as it is usually stated, and stated that way it sounds like a bad deal. It is worth taking apart, because the shape of the trade is more interesting than the headline.
Start with what you are buying at the margin. Moving from Kimi K3 to Opus 5 at max effort buys 3.58 raw index points for an extra USD 1.31 per task — roughly 37 cents per point. Moving to Opus 5 at its default high effort instead buys 1.75 points for an extra 34 cents — roughly 19 cents per point. Climbing the last step from high to max buys 1.83 points for an extra 97 cents — roughly 53 cents per point. Marginal points get steeply more expensive as you climb, which is worth knowing before you set effort to max out of habit. The cheapest route to a score above Kimi K3 is not Opus 5 at max; it is Opus 5 at its default.
Now the threshold argument. Index scores are an average over a task set, and averages do not describe how capability failures actually feel. A model does not deliver 93 percent of a working refactor; it either completes the task or it does not. What the four points buy is a higher probability of completion on the hardest tasks in the distribution — and no improvement at all on the tasks both models already finish.
That gives you a decision rule that does not require trusting any benchmark very far. Identify the hardest task you routinely need done and run it on Kimi K3. If it completes, the four points are headroom you are not using. If it fails, try Opus 5 at high effort before reaching for max — the 19-cents-per-point step clears the bar more often than its two-point margin suggests. Reach for max only when high fails too.
The inverse case is worth naming. Where a single wrong output is expensive — production code shipping without review, legal or financial drafting, anything where a confident error propagates — the arithmetic reverses. A task that has to be redone costs many times the entire per-task difference. In that regime the four points are cheap and the 2.8 times is noise.
Effort levels: comparing like with like
Both of these models let you spend more compute per request, and any comparison that ignores this is comparing two arbitrary configurations.
Opus 5 exposes five levels — low, medium, high, xhigh, max — with high as the API default. Artificial Analysis publishes a separate index score for every one of them, which makes Opus 5 unusually easy to reason about. Kimi K3 exposes three levels — low, high, max — with max as the default. Moonshot’s documentation is explicit that the model always reasons and that reasoning_effort takes those three values with max as the default. When Kimi K3 launched, Moonshot stated that it would ship using max effort by default with low and high modes to follow in later updates; those modes are now documented, so all three are live.
This is where the two models are not symmetrical, and it needs saying rather than glossing. Artificial Analysis publishes five separately-labeled entries for Opus 5, each tagged with its effort level. It publishes exactly one entry for Kimi K3, with no effort label attached to it at all. So we can tell you with certainty that Opus 5’s 61 is a max-effort figure and its 59 is a high-effort figure, because Artificial Analysis labels them. We cannot tell you the same about Kimi K3’s 57 on the same authority, because Artificial Analysis never stated it.
What we can say is bounded and it is this: Kimi K3’s API default is max effort, the index run was published on July 17, 2026, and the low and high modes did not exist at launch. On those two facts the 57 was almost certainly measured at max. But that is our reasoning from the record, not a published label, and we are flagging it as such rather than presenting it as Artificial Analysis’s statement.
The practical consequence: the 61-against-57 comparison is a max-against-max comparison, which is the fair one to headline. The 59-against-57 comparison is default-against-default, which is the fair one for anyone who never touches the parameter — and that is most people. Both are in the table above, and they tell different stories about the same two models.
Accuracy and hallucination, measured on one bench
This section exists because the obvious claim is the wrong one, and we nearly made it.
Anthropic’s Opus 5 system card reports hallucination results, and they are measured against Claude Opus 4.8. Not against Kimi K3. Nothing in that document supports any statement about how Opus 5 compares to Kimi K3 on factual reliability, and a comparison assembled by chaining two models together through a third would be worthless. So we set the system card aside for this purpose — though it is worth noting in passing that its finding runs against expectation: Anthropic reports Opus 5 as more accurate than Opus 4.8 overall while hallucinating slightly more factual claims, because it abstains less and asserts more.
There is, however, one measurement that covers both models on the same bench under the same conditions. Artificial Analysis runs AA-Omniscience, a closed-book knowledge benchmark that rewards correct answers, penalizes hallucinations, and applies no penalty for declining to answer. Here is what it records for these two.
| AA-Omniscience (independent, closed-book) | Claude Opus 5 (max) | Kimi K3 |
|---|---|---|
| Accuracy | 54.20 percent | 45.95 percent |
| Hallucination rate | 50.07 percent | 50.94 percent |
| Omniscience Index | 31.27 | 18.42 |
Read that carefully, because it does not say what a quick glance suggests. Opus 5 is well ahead on the composite index — 31.27 against 18.42 — but the two hallucination rates are 0.87 percentage points apart, which is a tie in any practical sense. Opus 5’s advantage here is entirely an accuracy advantage: it knows more, by about eight percentage points. It does not make things up meaningfully less often.
So if you are choosing between these two because you need fewer fabricated citations, fewer invented API methods, fewer confident wrong dates, the independent evidence does not favor either one. Neither is a reliable narrator without retrieval, tool access, or verification in the loop, and building for that is a better use of your effort than choosing between them on this axis. One caveat on the evidence itself: we found no second independent hallucination measurement covering both models, so treat this as one measurement rather than a consensus.
Throughput and latency
Both models are measured by Artificial Analysis against each vendor’s own first-party API, so this is like-for-like — with one caveat: Opus 5 is served by three hosts and Kimi K3 by one, so Kimi K3 has no alternative provider to route around a slow day.
On long prompts Opus 5 is clearly faster: a median 52.8 output tokens per second against Kimi K3’s 33.0, and a first token in a median 69.7 seconds against 177.2. Both are large absolute latencies, which is what reasoning at maximum effort costs in wall-clock terms.
At very long context the ordering flips hard. On 100,000-token prompts Kimi K3 sustains 62.1 output tokens per second against Opus 5’s 53.4, and returns its first token in a median 9.0 seconds against 43.3. If your work is long-document work — feed in a large codebase or corpus and ask a question about it — Kimi K3 is the more responsive of the two by a wide margin, and the cheaper one.
One warning about headline latency figures, including the ones we just quoted. The single advertised time-to-first-token number for a model is the long-prompt bucket, and the spread across prompt types is enormous: Opus 5’s median ranges from 30.0 seconds on medium prompts to 69.7 on long ones, and Kimi K3’s from 9.0 seconds at 100K context to 177.2 on long prompts. Match the bucket to your own prompt shape before you plan around it. Neither vendor publishes absolute speed figures of its own, so every number in this section is third-party measurement.
What open-weight actually buys you, and what it does not
This is the axis where the two models are not competing on the same dimension at all, and it deserves precision rather than enthusiasm.
First, the state of play. The weights are out. Moonshot published them on July 27, 2026, and we verified the repository on July 28: 96 weight shards across 118 files, ungated, no access request in front of the download. This is the advantage Opus 5 structurally cannot match, and it is now delivered rather than promised.
Second, the license — and here the widely repeated shorthand is wrong. Press coverage described the release as “Modified MIT.” What Moonshot shipped is a custom license it calls the Kimi K3 License, published as license: other on the model card. The permission grant is MIT-style and genuinely broad: use, copy, modify, publish, distribute, sublicense, sell, fine-tune, and build derivative works. But two conditions are attached that no MIT license carries. Model-as-a-service operators whose aggregate revenue exceeds USD 20 million over any consecutive twelve months must sign a separate agreement with Moonshot before commercial use. And products with more than 100 million monthly active users, or more than USD 20 million in monthly revenue, must display “Kimi K3” prominently in the interface. Internal use is exempt, as is access through Moonshot’s own products and certified inference partners. For a startup or a research team none of this binds; for a scaled AI provider it is a commercial negotiation, and calling it MIT would have set the wrong expectation.
Third, what the term still does not mean. The repository carries the parameters, the inference and modeling code, the tokenizer, and the vision processing — enough to run and fine-tune the model. It does not carry the training data, the training code, or the data-curation pipeline, so you cannot reproduce the model or audit what went into it. The model card discloses more architecture than the launch post did, including quantization-aware training from the supervised fine-tuning stage onward using MXFP4 weights with MXFP8 activations, but links no technical report. Downloadable is not reproducible, and the Kimi K3 License is not an OSI-approved open-source license.
Fourth, what self-hosting requires, because “downloadable” and “runnable” are not the same claim. In MXFP4 four-bit precision the weights need roughly 1.4 terabytes of fast memory, and Moonshot recommends at least 64 accelerators. This is not a workstation project. For most teams the practical value of open weights is not that they will run it, but that they could — which removes the single-vendor dependency and makes the model survivable if the vendor changes terms or disappears. For organizations with a hard data-residency or sovereignty requirement, that is the only thing on this page that matters, and Opus 5 cannot offer it at any price.
Against that, Opus 5 carries no data-retention requirement and remains eligible for zero-retention agreements, so the closed option is not automatically the one that keeps your data. These are two different trust models — verify-by-possession against contract-and-audit — and which one you need is an organizational question, not a technical one.
What Moonshot reports, kept separate
Moonshot published its own benchmark results for Kimi K3 on July 16, 2026, and some are striking. They are also vendor self-reported, produced on Moonshot’s own harness, and not reproduced by any independent party — so they get their own section, with no independent figure anywhere near them, because placing them alongside third-party numbers would create a false equivalence. Moonshot’s stated conditions, as published on the model card that shipped with the weights: all results were obtained with reasoning effort set to max and temperature 1.0, with top-p at 0.95 for single-step tasks such as GPQA-Diamond and at 1.0 for agentic tasks, and each benchmark run under one of three agentic harnesses — KimiCode, which is Moonshot’s own, or Claude Code, or Codex.
| Vendor self-reported by Moonshot — Kimi K3 only | Result | Conditions |
|---|---|---|
| Terminal-Bench 2.1 | 88.3 | KimiCode harness (Moonshot’s own) |
| BrowseComp | 91.2 | Context compaction at 300K |
| GPQA-Diamond | 93.5 | Reasoning effort max, top-p 0.95 |
There is one instructive case of what happens when someone else runs the same test. Artificial Analysis independently measured Kimi K3 on Terminal-Bench 2.1 and got 85.0 percent against Moonshot’s self-reported 88.3 — a 3.3-point difference attributable to harness rather than model, since Artificial Analysis uses Terminus 2 and Moonshot uses KimiCode. On GPQA-Diamond the two agree exactly at 93.5. So Moonshot’s numbers are not inflated across the board; they are harness-dependent in the way vendor benchmarks usually are, and the size of the gap tells you how much slack to leave. Moonshot published no SWE-bench Verified figure, and no hallucination or truthfulness benchmark at all.
The same discipline applies in the other direction. Anthropic published a detailed system card alongside Opus 5, and its evaluations compare Opus 5 to earlier Claude models. Kimi K3 is not in that document, so nothing in it can rank these two against each other, and we have not used it that way.
Winner by category
Highest measured capability — Claude Opus 5. 61 against 57 at max effort, 59 against 57 at default. There is no reading of the independent data in which Kimi K3 is the more capable model.
Cost per unit of work — Kimi K3. USD 0.72 per index task against USD 2.03 at Opus 5’s max effort, and against USD 1.06 at its default. Kimi K3 is only beaten on cost by dropping Opus 5 to medium effort, where it scores below Kimi K3.
Knowledge freshness — Claude Opus 5. A stated May 2026 cutoff against no published cutoff at all. For work touching recent events, libraries, or APIs, this is decisive and it is not a benchmark artifact.
Long-document work — Kimi K3. Nearly five times faster to first token at 100K context, higher throughput at that length, a marginally larger window, a far higher output ceiling, and cheaper.
Interactive and long-prompt latency — Claude Opus 5. Roughly 1.6 times the output throughput on long prompts, less than half the time to first token, and three serving hosts against one.
Control over the model — Kimi K3. Open weights once published, a disclosed architecture, self-hosting, fine-tuning, and no single-vendor dependency. Opus 5 offers none of this and cannot.
Factual reliability — no winner. Hallucination rates of 50.07 and 50.94 percent on the only independent bench covering both. Opus 5 knows more; it does not fabricate less.
Configurability — Claude Opus 5. Five effort levels against three, each independently scored and priced, which makes the cost-capability trade explicit rather than something you discover on your invoice.
Pros and cons
Claude Opus 5
Strengths
- Highest independently measured intelligence of the two — 61 on version 4.1 of the Artificial Analysis Intelligence Index at max effort, 59 at default
- A stated May 2026 training cutoff, the freshest of any current frontier model, against no published cutoff for Kimi K3
- A 1,000,000-token context window with no length-based price premium at any length, confirmed in Anthropic’s own pricing documentation
- Five effort levels, each separately scored and priced by an independent evaluator, so the cost-capability trade is explicit
- Roughly 1.6 times the output throughput on long prompts, less than half the time to first token, and three serving hosts
- No data-retention requirement and eligible for zero-retention agreements
Weaknesses
- About 2.8 times the measured cost per task at max effort, and about 1.5 times at default
- Closed weights and an undisclosed architecture — no self-hosting, no inspection, no fine-tuning, no route to data sovereignty
- Marginal index points get steeply more expensive as effort climbs, reaching about 53 cents per point on the last step from high to max
- Hallucinates at essentially the same rate as Kimi K3 on the one independent bench covering both, despite the accuracy lead
- Slower to first token at 100K context — a median 43.3 seconds against 9.0
Kimi K3
Strengths
- Near-frontier capability at a measured USD 0.72 per index task — about a third of Opus 5 at max effort
- Within four rounded index points of the highest-scoring model available, on the same version of the same independent index
- Weights published and ungated since July 27, 2026, with a disclosed mixture-of-experts architecture — a delivered route to self-hosting, fine-tuning, and vendor independence
- Markedly faster at very long context — a median 9.0 seconds to first token at 100K against 43.3, and higher throughput at that length
- A 1,048,576-token window and an output ceiling that can be raised to the full window, far above Opus 5’s 128,000 synchronous limit
- Cheaper cache-hit input at USD 0.30 per million tokens, with the same 90 percent cache discount
Weaknesses
- The license is not MIT despite the reporting: model-as-a-service operators above USD 20 million in revenue over any twelve months must negotiate a separate agreement with Moonshot before commercial use
- No published knowledge cutoff from any source, which makes recency of knowledge impossible to plan around
- Four rounded index points behind, and the gap is concentrated exactly where it hurts: the hardest tasks in a workload
- Lower independent accuracy — 45.95 percent against 54.20 on AA-Omniscience — with no hallucination-rate advantage to offset it
- A single serving host, so no provider to route around during degradation, and much slower on long prompts
- Self-hosting needs roughly 1.4 terabytes of fast memory in four-bit precision and at least 64 accelerators — out of reach for most teams
- Moonshot’s headline coding results are self-reported on its own harness and unreproduced; the one independent re-run came in 3.3 points lower
When to pick each one
Pick Claude Opus 5 when the hardest task in your workload decides the tool. If you have watched a mid-fifties model stall on genuinely difficult work — multi-hour agentic runs, unfamiliar large codebases, reasoning chains where an early error poisons everything downstream — the four points are the product and the price is not the deciding factor. Same answer when an error is expensive: production changes that ship without human review, regulated drafting, financial or legal work. One avoided rework pays for a great deal of per-task premium.
Pick Claude Opus 5 when knowledge recency matters. A stated May 2026 cutoff against no published cutoff is not a small difference for anyone working with libraries, frameworks, or events from the last year. This is the cleanest, least contestable advantage on the page.
Pick Kimi K3 for the bulk of real work. Most production workloads do not sit at the capability frontier. Summarizing, extraction, classification, routine code, first drafts, retrieval-augmented answering with the facts supplied in context — a model at 57 finishes these as reliably as a model at 61, and you keep about two thirds of the bill. At high volume, this is where the money is.
Pick Kimi K3 for long-document work. Nearly five times faster to first token at 100K context, higher throughput at that length, a marginally larger window, an output ceiling eight times higher, and cheaper. There is no argument for Opus 5 here unless the reasoning difficulty is genuinely at the frontier.
Pick Kimi K3 when you need the model itself, not access to it. Data residency, sovereignty requirements, air-gapped deployment, fine-tuning on proprietary data, or refusing to depend on one vendor’s continued goodwill. Opus 5 has no answer to this, and since July 27, 2026 the weights are published and ungated. Two caveats before you plan on it: running them needs roughly 1.4 terabytes of fast memory, and the Kimi K3 License requires a separate agreement with Moonshot if you resell inference above USD 20 million in revenue.
Consider running both. For a high-volume team the most defensible setup is not a choice at all: route the bulk to Kimi K3 and escalate to Opus 5 at high effort when the cheaper model fails or flags low confidence. Both take a 1M-token context, so prompts port across with little rework, and the escalation costs 34 cents per task rather than USD 1.31.
Frequently Asked Questions
Which is better overall, Claude Opus 5 or Kimi K3?
Neither, and that is a conclusion rather than a hedge. On version 4.1 of the independent Artificial Analysis Intelligence Index, Claude Opus 5 scores 61 at max effort and Kimi K3 scores 57 — a four-point gap on rounded figures, 3.58 on the raw ones. Kimi K3 costs a measured USD 0.72 per index task against USD 2.03. Opus 5 wins the capability ceiling, the May 2026 cutoff, and long-prompt speed; Kimi K3 wins the cost, the openness, and the speed at very long context. The rule is a threshold rather than an average: if the hardest task in your workload sits above what a model scoring 57 can finish, the four points are the whole reason to pay. If it sits below, and for most workloads it does, they are headroom you never touch.
Is Kimi K3 really 2.8 times cheaper than Claude Opus 5?
On measured cost per task at Opus 5’s maximum effort, yes: USD 0.72 against USD 2.03, published by Artificial Analysis and read July 27, 2026. On list token rates the gap is smaller and exactly 1.7 times — USD 3 against USD 5 per million input tokens, USD 15 against USD 25 per million output. The difference between the ratios is reasoning volume: Opus 5 at max effort spends far more tokens thinking, and those bill at the output rate. At Opus 5’s default high effort the measured gap narrows to about 1.5 times, and at medium effort Opus 5 is cheaper per task than Kimi K3 while scoring one point lower.
Are the 61 and the 57 measured at comparable settings?
Both are version 4.1 of the same independent index, but the effort labeling is not symmetrical. Artificial Analysis publishes five separate entries for Claude Opus 5, each tagged with its effort level, so its 61 is confirmed as max-effort and its 59 as high-effort. It publishes exactly one entry for Kimi K3, with no effort label. Kimi K3’s API default is max, and its low and high modes did not exist when the index run was published on July 17, 2026, so the 57 was almost certainly measured at max — but that is our reasoning from the record, not something Artificial Analysis stated. Treat 61 against 57 as max-against-max with that caveat attached.
Does Kimi K3 have reasoning effort levels like Claude Opus 5?
Yes. Moonshot’s API documentation specifies a reasoning_effort field accepting low, high, and max, with max as the default, and states that Kimi K3 always reasons — Moonshot calls this Preserved Thinking and it cannot be disabled. Claude Opus 5 has five levels rather than three, with high as the default. The meaningful difference is not that one has effort control and the other does not; it is that Artificial Analysis publishes a separate score and cost for each of Opus 5’s five levels, while Kimi K3 has a single published figure, so its cost-capability curve is not visible in the same way.
Are Kimi K3’s open weights actually available?
Yes. Moonshot published them on July 27, 2026, and we verified the repository directly on July 28: 96 weight shards across 118 files, ungated, with no access request in front of the download. The repository also ships the inference and modeling code, the tokenizer, and the vision processing — enough to run and fine-tune the model on your own hardware. Two things to know first. The license is not MIT despite widespread reporting to the contrary; it is a custom Kimi K3 License with revenue-triggered commercial conditions. And the weights need roughly 1.4 terabytes of fast memory in four-bit precision, with at least 64 accelerators recommended, so this is a data-center deployment rather than a workstation one.
Is Kimi K3 open source?
No, and now that the weights are out the distinction is easier to state precisely. The repository carries the parameters, the inference and modeling code, the tokenizer, and the vision processing — but not the training data, the training code, or the data-curation pipeline, so you can run and modify the model without being able to reproduce it or audit what went into it. The license is a custom Kimi K3 License, not the Modified MIT that circulated in the press and not an OSI-approved open-source license: it attaches a separate-agreement requirement for model-as-a-service operators above USD 20 million in revenue, and an attribution requirement for products above 100 million monthly active users. Open-weight, generously licensed, genuinely useful — but not open source.
Which model hallucinates less?
On the only independent measurement covering both, neither. Artificial Analysis runs AA-Omniscience, a closed-book benchmark that rewards correct answers, penalizes hallucinations, and does not penalize declining to answer. It records a hallucination rate of 50.07 percent for Claude Opus 5 at max effort and 50.94 percent for Kimi K3 — 0.87 percentage points apart, which is a tie in practice. Opus 5 leads the composite Omniscience Index 31.27 to 18.42, but that lead comes entirely from accuracy, 54.20 percent against 45.95. It knows more; it does not fabricate less. Anthropic’s own system card measures Opus 5 against Claude Opus 4.8 and never against Kimi K3, so nothing in it speaks to this question.
Can I compare Moonshot’s 88.3 on Terminal-Bench to Claude Opus 5?
Not directly, no. Moonshot’s 88.3 is self-reported on KimiCode, Moonshot’s own agentic harness, at reasoning effort max with temperature 1.0 and top-p 1.0. Nobody outside Moonshot has reproduced it. When Artificial Analysis ran the same benchmark on Kimi K3 with its own Terminus 2 harness it measured 85.0 percent — 3.3 points lower, a gap attributable to harness rather than model. Independent figures on the same benchmark are 89.1 percent for Opus 5 at max effort against 85.0 for Kimi K3. Use those if you want a comparison, and keep the vendor number in its own column.
Which one is faster?
It depends entirely on prompt length, and the ordering reverses. On long prompts Claude Opus 5 sustains a median 52.8 output tokens per second against Kimi K3’s 33.0, and reaches its first token in 69.7 seconds against 177.2. At 100,000-token context Kimi K3 sustains 62.1 tokens per second against 53.4 and returns its first token in a median 9.0 seconds against 43.3 — nearly five times faster. Opus 5 is also served by three hosts against Kimi K3’s one, so it has somewhere to fail over. All of these are third-party measurements against each vendor’s own API; neither vendor publishes absolute speed figures.
Do both models have a one-million-token context window?
Effectively yes, with a small edge to Kimi K3: 1,048,576 tokens against 1,000,000. More useful than the window size is what happens at the edges. Anthropic states explicitly that Opus 5 carries no length-based premium at any length — a 900,000-token request bills at the same per-token rate as a 9,000-token one — and Moonshot publishes no long-context surcharge either. On output the gap is larger: Opus 5 caps at 128,000 tokens synchronously and 300,000 through the Batch API, while Kimi K3 defaults to 131,072 and can be raised to the full window.
What is Kimi K3’s knowledge cutoff?
There isn’t a published one. Not from Moonshot’s announcement, not in the API documentation, and Artificial Analysis records the field as empty. That is an absence in the public record rather than a claim about the model being outdated, and we are not going to guess at it. Claude Opus 5, by contrast, has a stated training cutoff of May 2026, with Anthropic listing the reliable knowledge cutoff as the same month rather than an earlier one. If knowledge recency matters to your work, that asymmetry is one of the strongest arguments on this page for Opus 5.
What would change this verdict?
One trigger has already fired: the weights were published on July 27, 2026, which is reflected throughout this page. Two remain open. Independent reproduction of Moonshot’s coding figures on a neutral harness would make the capability gap look narrower than the index alone suggests, and a published knowledge cutoff for Kimi K3 would remove the freshness asymmetry that currently favors Opus 5. A third could move things the other way: now that the weights are downloadable, evaluators can measure the self-hosted model rather than Moonshot’s API, and any divergence would matter. We will revisit this page as each resolves.
The verdict
We are not naming a winner, and the reason is not indecision. These two models answer different questions and both answers are good.
Claude Opus 5 answers “what is the most capable model I can buy right now,” and the answer is independently verified: 61 at max effort, the highest figure recorded, with a May 2026 cutoff, a million-token window that costs no more at the top than at the bottom, and five effort levels. Kimi K3 answers “how close to the frontier can I get for a third of the money, with a route to owning the model,” and the answer is: within four rounded points, at a measured USD 0.72 per task, with a matching window, native vision, and weights you can now download and run yourself.
The number that decides it is not 61 or 57. It is the difficulty of the hardest task you actually need done. Run it on Kimi K3 first; if it completes, you have your answer and it costs a third as much. If it fails, escalate to Opus 5 at its default high effort before reaching for max — that step buys 1.75 index points for 34 cents per task, the best value on the curve. Only when high effort fails does max earn its 2.8 times, and where a wrong answer is expensive it earns it easily.
Two things we will not tell you, because the evidence does not support them. We will not tell you Opus 5 hallucinates less — independently measured, the two are 0.87 percentage points apart and that is a tie. And we will not tell you Kimi K3 is open source: the weights did land on July 27, 2026 and they are ungated, but the repository ships no training data and no training code, and the Kimi K3 License is a custom license with revenue-triggered commercial conditions rather than the Modified MIT the press reported. What we will tell you is that a 2.8-trillion-parameter model from a Chinese lab landing four points behind the best model in the world, at a third of the cost, with weights now downloadable by anyone, is the most interesting thing to happen to this market in 2026 — and that Anthropic answered it not by cutting the price but by moving the ceiling.
If you want the wider field, we have also compared Kimi K3 against Claude Fable 5, against Claude Opus 4.8, and against GPT-5.6 Sol, and you can read our full write-ups of GPT-5.6 Sol and Claude Opus 4.8 for the models sitting between these two on the index.
Last compared July 27, 2026, updated July 28. Pricing, context limits, effort parameters, and cutoff dates come from Anthropic’s and Moonshot’s own live documentation. Capability, cost, accuracy, hallucination, and speed figures are Artificial Analysis Intelligence Index version 4.1, read July 27; those move as models are re-run. Kimi K3’s weight availability and license terms were verified on the published repository and its LICENSE file on July 28, after Moonshot released the weights on July 27 — an earlier version of this page reported them unpublished, which was accurate when written. Moonshot’s self-reported results are labeled as such and kept separate from independent measurement throughout.
Sources and references
Every figure on this page is attributed to whoever produced it. Independent measurement, vendor-reported results and third-party-verified scores are listed separately and never merged into a single ranking.
- Anthropic — Introducing Claude Opus 5 (launch announcement, July 24, 2026)
- Anthropic — Claude Opus 5 System Card (source for every Anthropic-reported benchmark quoted here)
- Anthropic — Models overview (model IDs, context windows, knowledge cutoffs)
- Anthropic — Pricing (per-token rates, cache pricing, Batch API)
- Anthropic — Effort parameter (the five effort levels and per-model defaults)
- Artificial Analysis — Model leaderboard (Intelligence Index v4.1 scores and cost per task, consulted July 27, 2026)
- Artificial Analysis — Claude Opus 5 (index score by effort, speed, latency and verbosity)
- Artificial Analysis — Kimi K3 (independent Intelligence Index v4.1 score and cost per task)
- Moonshot AI — Chat API reference (reasoning_effort levels and defaults)
- Hugging Face — moonshotai/Kimi-K3 (weight release status, checked July 27, 2026)
Our Verdict
There is no overall winner here, and that is a conclusion rather than a hedge. On version 4.1 of the independent Artificial Analysis Intelligence Index, Claude Opus 5 scores 61 at max effort against Kimi K3’s 57 — four points on the rounded figures, 3.58 on the raw ones — while Kimi K3 costs a measured USD 0.72 per index task against USD 2.03, about 2.8 times less. Read that as a threshold rather than an average: if the hardest task in your workload sits above what a model scoring 57 can finish, the four points are the entire reason Opus 5 exists; if it sits below, and for most workloads it does, you are paying 2.8 times for headroom you never touch. The multiplier is also a property of a dial rather than of the two models — at Opus 5’s default high effort the comparison is 59 against 57 at about 1.5 times the cost, and at medium effort Opus 5 is cheaper than Kimi K3 while scoring one point lower. Opus 5 takes the capability ceiling, the stated May 2026 cutoff and long-prompt speed; Kimi K3 takes the cost, the responsiveness at very long context, and openness that is now delivered rather than promised: Moonshot published the weights on July 27, 2026, ungated, though under a custom Kimi K3 License with revenue-triggered commercial conditions rather than the Modified MIT widely reported. On the one independent bench covering both, hallucination rates are effectively tied at 50.07 and 50.94 percent, so Opus 5’s advantage there is accuracy, not fewer fabrications.
Choose Claude Opus 5
Anthropic's frontier reasoning model — top of the independent index at half the price of Fable 5.
Try Claude Opus 5 →Choose Kimi K3
Moonshot AI's ~2.8T-parameter open Mixture-of-Experts flagship — ~50B active, 1M context, native vision. Testable via API today at $3 in / $15 out per million tokens; open weights expected July 27, 2026.
Try Kimi K3 →Frequently Asked Questions
Is Claude Opus 5 better than Kimi K3?
There is no overall winner here, and that is a conclusion rather than a hedge. On version 4.1 of the independent Artificial Analysis Intelligence Index, Claude Opus 5 scores 61 at max effort against Kimi K3’s 57 — four points on the rounded figures, 3.58 on the raw ones — while Kimi K3 costs a measured USD 0.72 per index task against USD 2.03, about 2.8 times less. Read that as a threshold rather than an average: if the hardest task in your workload sits above what a model scoring 57 can finish, the four points are the entire reason Opus 5 exists; if it sits below, and for most workloads it does, you are paying 2.8 times for headroom you never touch. The multiplier is also a property of a dial rather than of the two models — at Opus 5’s default high effort the comparison is 59 against 57 at about 1.5 times the cost, and at medium effort Opus 5 is cheaper than Kimi K3 while scoring one point lower. Opus 5 takes the capability ceiling, the stated May 2026 cutoff and long-prompt speed; Kimi K3 takes the cost, the responsiveness at very long context, and openness that is now delivered rather than promised: Moonshot published the weights on July 27, 2026, ungated, though under a custom Kimi K3 License with revenue-triggered commercial conditions rather than the Modified MIT widely reported. On the one independent bench covering both, hallucination rates are effectively tied at 50.07 and 50.94 percent, so Opus 5’s advantage there is accuracy, not fewer fabrications.
Which is cheaper, Claude Opus 5 or Kimi K3?
Claude Opus 5 is priced at $5 in / $25 out per M tokens. Kimi K3 is priced at $3 in / $15 out per M tokens (free plan available). Check the pricing comparison section above for a full breakdown.
What are the main differences between Claude Opus 5 and Kimi K3?
The key differences span across 20 features we compared. For Artificial Analysis Intelligence Index (v4.1, independent), highest effort, Claude Opus 5 offers 61 (raw 60.69) while Kimi K3 offers 57 (raw 57.11). For Same index at each model’s default effort, Claude Opus 5 offers 59 (high) while Kimi K3 offers 57 (max). For Measured cost per index task, highest effort, Claude Opus 5 offers USD 2.03 while Kimi K3 offers USD 0.72. See the full feature comparison table above for all details.

