Claude Opus 5 vs MiniMax M3: Seventeen Points, Seventeen Times the Cost
Claude Opus 5 scores 61 on Artificial Analysis v4.1 against MiniMax M3's 44 — at USD 2.03 per task against USD 0.12. Verdict, pricing and the real license.
Feature Comparison
| Feature | Claude Opus 5 | MiniMax M3 |
|---|---|---|
| Artificial Analysis Intelligence Index v4.1 | 61 at max effort, 59 at default high effort | 44, effort level not labeled |
| Cost per index task (measured July 28, 2026) | USD 2.03 at max, USD 1.06 at default high, USD 0.36 at low | USD 0.12 |
| Input price per million tokens | USD 5.00 | USD 0.30 at or below 512K input, USD 0.60 above |
| Output price per million tokens | USD 25.00 | USD 1.20 at or below 512K input, USD 2.40 above |
| Context window | 1,000,000 tokens | 1,000,000 tokens |
| Long-context surcharge | None at any length | Rates double above 512,000 input tokens |
| Maximum output | 128,000 synchronous, 300,000 via Batch API beta | 131,072 recommended, 524,288 maximum |
| Training cutoff | May 2026 | Not published |
| Reasoning control | Five effort levels, default high | Binary thinking parameter, default differs by endpoint |
| Input modalities | Text and image | Text, image, and video |
| Output speed (measured July 28, 2026) | 55 tokens per second | 77 tokens per second |
| Time to first token (measured July 28, 2026) | 78.91 seconds | 1.64 seconds |
| Weights | Not published | Published on Hugging Face, ungated |
| License | Proprietary (Artificial Analysis classification) | MINIMAX COMMUNITY LICENSE, non-commercial default grant |
Pricing Comparison
Claude Opus 5
MiniMax M3
Detailed Comparison
Claude Opus 5 vs MiniMax M3 in 2026: Claude Opus 5 is Anthropic's closed frontier model, released July 24, 2026, priced at USD 5 per million input tokens and USD 25 per million output tokens, with a 1,000,000-token context window carrying no long-context premium, a 128,000-token maximum synchronous output, and a May 2026 training cutoff. It scores 61 on version 4.1 of the independent Artificial Analysis Intelligence Index at its highest effort setting and 59 at its default setting. MiniMax M3 is Shanghai-based MiniMax's open-weight multimodal model, released June 1, 2026, built on a 428-billion-parameter mixture-of-experts architecture with roughly 23 billion active parameters, priced at USD 0.30 per million input tokens and USD 1.20 per million output tokens below 512,000 tokens of input and double that above it, with a 1,000,000-token context window and text, image, and video input. It scores 44 on the same version of the same independent index. Measured cost per index task is USD 2.03 for Opus 5 at maximum effort against USD 0.12 for MiniMax M3 — about seventeen times. We researched both from vendor documentation and independent measurement rather than running our own benchmarks. Claude Opus 5 wins the capability comparison decisively and it is not close; MiniMax M3 wins on price, on downloadable weights, and on the ability to run the model on your own hardware, none of which the index measures.
Our verdict at a glance
Seventeen points is not a duel. On version 4.1 of the Artificial Analysis Intelligence Index — the same index, the same version, measured by the same third party — Claude Opus 5 scores 61 and MiniMax M3 scores 44. That is one of the widest gaps we have written up in this series, and no amount of framing turns it into a photo finish. If your question is which model is more capable, the answer is Claude Opus 5 and it is not close.
The question worth several thousand more words is a different one: what does MiniMax M3 give you that a capability index cannot see, and at what task difficulty do those seventeen points start costing you real money?
Three answers, all sourced below. Price: M3 costs USD 0.12 per index task against USD 2.03 for Opus 5 at maximum effort — roughly seventeen times less. Weights: MiniMax publishes M3's weights on Hugging Face, ungated, with 154,969 downloads recorded at the time of writing, so you can run the model inside your own network and never send a token to anyone's API. Latency: M3 returns its first token in a median 1.64 seconds against 78.91 seconds for Opus 5 at maximum effort — roughly forty-eight times, which matters enormously for anything interactive and not at all for batch work.
And one finding that runs against the pattern we found in earlier comparisons on this site. Against Kimi K3, Opus 5 could be dialed down to medium effort and land below its rival on cost per task. That does not happen here. Opus 5's cheapest rung — low effort, USD 0.36 per task — is still three times M3's cost, and it still scores 51 against 44. No configuration of Claude Opus 5 undercuts MiniMax M3 on price. What the effort ladder buys instead is a much better exchange rate: seven points of lead for three times the cost, rather than seventeen for seventeen.
So the verdict is decisive and narrow at once. Claude Opus 5 wins on capability, knowledge recency, and flat long-context pricing. MiniMax M3 wins on cost, latency, and the one thing no API-only model can offer — weights you can download. It is not, however, the winner on licensing freedom, and the reason is the single most surprising thing we found on this page.
How we ran this comparison
We researched both models rather than benchmarking them ourselves, so here is exactly where every figure on this page comes from and how we treated it.
Pricing, context windows, output limits, reasoning parameters, modalities, and cutoff dates were read directly from each vendor's own live documentation on July 28, 2026 — Anthropic's pricing, models, and effort pages, and MiniMax's pay-as-you-go pricing page, API references, and release notes. Not from summaries and not from secondary reporting, because token prices and parameter defaults change quietly and a page that cites a summary of a price is citing a price that may no longer exist.
Capability figures come from Artificial Analysis, a third-party evaluator that runs both models on the same task set and publishes the cost of doing so. We use their Intelligence Index version 4.1 throughout. This matters more than it sounds: index versions are not comparable to one another, and a score quoted without its version number is not a score.
Some of these figures slide and some do not. Index scores are stable within a given index version. Cost per task, output speed, and time to first token are live measurements against each vendor's API that move as models are re-run and as serving capacity changes; every one of them on this page carries the date we read it, July 28, 2026, and should be treated as a reading rather than a constant.
Where a vendor publishes its own benchmark results, we keep them separate and never place them in the same table, sentence, or chart as an independent measurement. MiniMax publishes self-reported scores for M3; Anthropic publishes safety figures for Opus 5 that compare it only to other Anthropic models. Neither is used for a head-to-head claim here. And where neither vendor has published something — MiniMax has never stated a knowledge cutoff for M3 — we say so, rather than inferring a date from the release month.
Claude Opus 5 in brief
Anthropic released Claude Opus 5 on July 24, 2026 at USD 5 per million input tokens and USD 25 per million output tokens, identical to Claude Opus 4.8 before it. Cache hits are USD 0.50 per million tokens, five-minute cache writes USD 6.25 and one-hour writes USD 10, and the Batch API halves the base rates to USD 2.50 and USD 12.50. The announcement positions it against Claude Fable 5, which Anthropic describes it as coming close to "at half the price."
The context window is 1,000,000 tokens with no length-based surcharge anywhere in it. Anthropic's pricing documentation states plainly that "a 900k-token request is billed at the same per-token rate as a 9k-token request," and that caching and batch processing discounts apply at standard rates across the full context window. Maximum output is 128,000 tokens synchronously, rising to 300,000 through the Batch API behind a beta header — and that extended output is Batch-only, explicitly not available on the synchronous Messages API.
Training data cutoff is May 2026, and Anthropic lists the reliable knowledge cutoff as the same month rather than an earlier one — the single most concrete advantage Opus 5 holds that has nothing to do with benchmark scores. Two constraints are worth knowing before you build on it: Anthropic's migration guide lists web fetch as not available on Claude Opus 5, and Priority Tier as not supported. Neither is a capability claim, but both are the kind of detail that surfaces late in an integration.
MiniMax M3 in brief
MiniMax released M3 on June 1, 2026 according to the company's own release notes, describing it as "the latest M-series language model for agentic reasoning, tool use, coding, multimodal chat input, and long-context tasks." It is a mixture-of-experts model with roughly 428 billion total parameters and about 23 billion active per token, built on a sparse attention scheme MiniMax calls MSA. It succeeds M2.7, not M2 — the M-series shipped M2 in October 2025, then M2.1, M2.5, and M2.7 before M3 arrived.
Pricing is tiered by context length, which is the detail most summaries drop. At or below 512,000 input tokens the published rates are USD 0.30 per million input tokens, USD 1.20 per million output tokens, and USD 0.06 per million cache-read tokens; above that threshold every one of those doubles. MiniMax describes the current rates as a permanent 50 percent reduction from list with no expiry published — a vendor claim rather than a contractual commitment. A priority tier sits at 1.5 times standard.
The context window is 1,000,000 tokens, and the published weights carry a maximum position embedding of 1,048,576. Maximum output is unusually generous: MiniMax's API reference gives a recommended value of 131,072 tokens and a maximum of 524,288, four times what Opus 5 will emit in a synchronous call. Input modalities are text, image, and video; output is text only. It is a model that watches video and writes about it, not a video generator — MiniMax's Hailuo line handles video generation and is a separate product family with separate pricing.
MiniMax publishes no training data cutoff for M3 — not in the model card, not on the model page, not in the API documentation, and not in the MSA architecture paper. That is an absence in the public record rather than a statement about the model being stale, and we are not going to invent a date for it.
What the seventeen-point gap actually measures
Claude Opus 5 scores 61 and MiniMax M3 scores 44 on Artificial Analysis Intelligence Index version 4.1. That is a blend of nine evaluations: GDPval-AA v2 and Tau-cubed Banking under agents at 34 percent, Terminal-Bench v2.1 and SciCode under coding at 24 percent, Humanity's Last Exam, GPQA Diamond and CritPt under scientific reasoning at 24 percent, and AA-LCR and AA-Omniscience under general at 18 percent.
Knowing the composition matters because it tells you what the seventeen points are made of. They are weighted heavily toward agentic task completion and hard scientific reasoning — those two categories together account for 58 percent of the index. They are not a measure of everyday instruction-following, summarization, or extraction, and there is no reason to assume the gap on those tasks resembles the gap on the index.
One component is worth isolating. AA-Omniscience measures knowledge reliability and hallucination on a scale from -100 to 100, rewarding correct answers, penalizing hallucinations, and applying no penalty for refusing to answer. Claude Opus 5, at adaptive reasoning and maximum effort, scores 31 — third among the models Artificial Analysis displays, behind Claude Fable 5 at 40 and Gemini 3.1 Pro Preview at 33. MiniMax M3 does not appear among the models displayed on that page. Precision matters here: AA-Omniscience is a 12 percent component of the index M3 scores 44 on, so a score exists. It is simply not shown in the excerpt we could read, and we are not going to back it out of the index total. This is "we could not verify it," not "it does not exist."
The honest way to use the seventeen points is as a threshold rather than an average. Nobody consumes seventeen index points spread evenly across a workload. Somewhere in your work sits a hardest task, and the only question that matters is whether it lands above or below what a model scoring 44 can finish.
An effort ladder against an on-off switch
Comparing two models means comparing two configurations, and these two do not expose the same dial at all. Claude Opus 5 has five graded effort levels; MiniMax M3 has a binary. Quoting a single number for each without saying which configuration produced it is the most common way these comparisons go wrong.
Opus 5's effort levels are low, medium, high, xhigh, and max, set through the output_config.effort field. Anthropic's documentation states that "the API default is high" and that setting effort to high "produces exactly the same behavior as omitting the effort parameter entirely." Artificial Analysis measures every rung: 61 at max for USD 2.03 per task, 60 at xhigh for USD 1.56, 59 at high for USD 1.06, 56 at medium for USD 0.62, and 51 at low for USD 0.36. One constraint is worth knowing: on Opus 5, thinking cannot be disabled at xhigh or max, and a request that tries returns a 400 error.
MiniMax M3 has no effort ladder. The parameter is thinking, and the documented values are adaptive and disabled — on or off, with nothing in between. And here MiniMax's own documentation contradicts itself, which we flag rather than resolve. The OpenAI-compatible API reference states that when the parameter is omitted, adaptive thinking is enabled by default. The Anthropic-compatible API reference states the opposite: "If thinking is omitted, thinking is off by default and the response does not include thinking blocks." The default flips depending on which endpoint you call, and the Hugging Face model card adds a third value, enabled, that neither API reference lists. If you are building on M3, set the parameter explicitly.
Artificial Analysis lists MiniMax M3 exactly once, as "MiniMax-M3," with no effort or reasoning suffix in the displayed name, and notes on the model page that it is showing the reasoning version. So the 44 is measured with reasoning on. Whether that constitutes M3's ceiling is a question the leaderboard does not answer, because M3 has no graded ceiling to be at — the level is simply not labeled, and we are not going to guess at one.
What this does change is the comparison at defaults. Left alone, a Claude Opus 5 request runs at high effort and scores 59, not 61. The default-against-measured comparison is therefore 59 against 44, fifteen points rather than seventeen — and USD 1.06 against USD 0.12, about nine times rather than seventeen. Both framings are true; a page that quotes one without naming the configuration is quietly answering neither.
Head to head: the full specification table
Every row below is sourced from a vendor's live documentation or from Artificial Analysis, and the two are never mixed inside a row. Figures marked as measured on July 28, 2026 are live readings that move.
| Specification | Claude Opus 5 | MiniMax M3 |
|---|---|---|
| Vendor | Anthropic (United States) | MiniMax (Shanghai, China) |
| Released | July 24, 2026 | June 1, 2026 |
| Intelligence Index v4.1 | 61 at max effort, 59 at default high effort | 44, effort level not labeled |
| Cost per index task (measured July 28, 2026) | USD 2.03 at max, USD 1.06 at default high, USD 0.36 at low | USD 0.12 |
| Input price per million tokens | USD 5.00 | USD 0.30 at or below 512K input, USD 0.60 above |
| Output price per million tokens | USD 25.00 | USD 1.20 at or below 512K input, USD 2.40 above |
| Cache read per million tokens | USD 0.50 | USD 0.06 at or below 512K input, USD 0.12 above |
| Batch discount | 50 percent — USD 2.50 input, USD 12.50 output | Not published |
| Context window | 1,000,000 tokens | 1,000,000 tokens |
| Long-context surcharge | None at any length | Rates double above 512,000 input tokens |
| Maximum output | 128,000 synchronous, 300,000 via Batch API beta | 131,072 recommended, 524,288 maximum |
| Training cutoff | May 2026 | Not published |
| Reasoning control | Five effort levels, default high | Binary thinking parameter, default differs by endpoint |
| Input modalities | Text and image | Text, image, and video |
| Output speed (measured July 28, 2026) | 55 tokens per second | 77 tokens per second |
| Time to first token (measured July 28, 2026) | 78.91 seconds | 1.64 seconds |
| Weights | Not published | Published on Hugging Face, ungated |
| License | Proprietary (Artificial Analysis classification) | MINIMAX COMMUNITY LICENSE |
One row deserves a footnote: the "Proprietary" and "Open Weights" labels are Artificial Analysis's third-party classification, not a term either vendor uses. Anthropic does not describe Opus 5 as proprietary; it simply does not publish weights.
What each model actually costs to run
List prices and real spend are different numbers, and the gap between them is where most model-selection decisions go wrong. On list, Opus 5 costs about 16.7 times more per million input tokens and about 20.8 times more per million output tokens; on Artificial Analysis's blended figure, which weights input, cached input, and output at seven to two to one, it is USD 3.85 per million against USD 0.22, about 17.5 times.
The number we trust most is neither of those. Cost per index task measures what it actually cost to run the same nine evaluations on each model, including every reasoning token the model chose to emit. Read on July 28, 2026, that is USD 2.03 for Opus 5 at maximum effort against USD 0.12 for MiniMax M3 — about seventeen times. A reasoning model that thinks longer spends more even at an identical token price, and only a cost-per-task figure captures that.
Now walk down the effort ladder, because this is where the decision actually lives. At default high effort Opus 5 costs USD 1.06 per task and scores 59, so you give up two index points and save 48 percent against max. At medium it is USD 0.62 and 56 points. At low it is USD 0.36 and 51 points — still seven points clear of M3, at three times the cost.
That last figure is what separates this comparison from the others in this series. Against Kimi K3, Opus 5 at medium effort landed below its rival on cost per task while scoring within a point of it. Nothing like that happens here: every rung of Claude Opus 5 costs more than MiniMax M3, and the cheapest costs three times more. There is no setting at which Opus 5 is the budget option.
For high-volume, low-difficulty work the arithmetic is brutal and it favors M3. A million tasks at USD 0.12 comes to USD 120,000; the same million at USD 2.03 comes to USD 2,030,000. No plausible quality difference on a routing task justifies the USD 1.9 million between them.
Two million-token windows, priced differently
Both models advertise a 1,000,000-token context window, but only one charges the same rate at the top of that window as at the bottom — the clearest example on this page of something a capability index cannot see. Anthropic's position is explicit and quotable: models with a 1M-token window include it "at standard pricing," a 900,000-token request "is billed at the same per-token rate as a 9k-token request," and caching and batch discounts apply at standard rates across the full window. A separate page adds that 1M is the default and needs no beta header. There is no long-context tier to fall into.
MiniMax M3 has one. At or below 512,000 input tokens you pay USD 0.30 and USD 1.20 per million; above it, USD 0.60 and USD 2.40, with cache reads doubling from USD 0.06 to USD 0.12. MiniMax's own model page describes the window as "up to 1M tokens context window with a guaranteed minimum of 512K tokens," a candid way of saying the second half is priced as a different product.
Work the numbers on a genuinely long prompt and the picture stays lopsided but narrows. A single 800,000-token input costs USD 4.00 on Opus 5 and USD 0.48 on M3 at the above-512K rate — about eight times, down from about seventeen at short context. If your workload lives above 512,000 tokens, price it at the upper tier and not at the first row of the pricing page.
On output length the advantage flips to M3 outright. Opus 5 emits at most 128,000 tokens in a synchronous call, reaching 300,000 only through the Batch API behind a beta header. M3's API reference recommends 131,072 and permits up to 524,288 in a normal call — roughly four times the ceiling, without batching.
What the MiniMax M3 license actually says
We opened the LICENSE file in the Hugging Face repository and read all 3,339 bytes of it, because the difference between "open weights" and a permissive license has cost people real money before. MiniMax M3 is not released under MIT, not under Apache 2.0, and not under any license the Open Source Initiative has approved. The repository metadata records the license as other, named minimax-community; the file itself is titled MINIMAX COMMUNITY LICENSE.
The grant reads like MIT with one phrase changed, and that phrase changes everything. Permission is granted "to deal in the Software for non-commercial purposes, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or provide copies of the Software." MIT's grant carries no such qualifier.
Clause 2 then sets terms for commercial use anyway, which is where a careful reader raises an eyebrow — the grant says non-commercial, and the next clause describes what you must do if you go commercial. We report that tension rather than resolve it, because resolving it is a lawyer's job. What clause 2 requires is concrete: you "shall prominently display 'Built with MiniMax M3' on a related website, user interface, blogpost, about page or product documentation," and you must send a one-time notice to MiniMax. Above a revenue threshold the requirement hardens — if your products and services "generate more than 20 million US dollars (or equivalent in other currencies) in yearly revenue," you "shall obtain a separate, prior written authorization from MiniMax."
Clause 3 defines commercial use expansively enough that most business deployments land inside it: offering products or services to third parties for a fee that rely on the software, commercial use of APIs built on it, and deployment of any fine-tuned or post-trained derivative for a commercial purpose. Serving a fine-tune of M3 to your own paying customers is commercial use under this license. An appendix adds prohibited uses with no analogue in MIT or Apache, including no assistance with "any military purpose" — and use restrictions are precisely what disqualifies a license from the open source definition.
None of this makes M3 a bad deal — an attribution line and an email are a trivial price for a near-frontier model you can run yourself. But "MiniMax released M3 as open source" is wrong, and if your legal team hears "MIT-like" and later reads clause 2, the conversation will be unpleasant. The same mistake was made about Kimi K3, where coverage reported a modified MIT license and the actual file carried revenue-triggered conditions. Read the file.
Open weight is not open source
Even setting the license aside, the word "open" is doing two different jobs in most write-ups of Chinese model releases. What MiniMax publishes for M3 is substantial: 59 safetensors shards holding the weights, the model and generation configurations, the full tokenizer set, the chat template, and inference-side code for image, video, and text preprocessing. The repository is ungated — no access request, no acceptance form — and recorded 154,969 downloads when we checked. You can pull it and serve it through SGLang, vLLM, or Transformers on your own hardware today.
What MiniMax does not publish is everything upstream of the weights: no training data, no dataset description, no data composition, no training scripts, and no reproduction recipe. The MSA repository on GitHub is MIT-licensed but contains inference attention kernels targeting NVIDIA SM100 hardware — not a training codebase. The arXiv paper numbered 2606.13392 is an architecture paper about MiniMax Sparse Attention, submitted June 11, 2026, ten days after the model shipped; it is not an M3 technical report, and we could not verify that a distinct one exists.
So the accurate label is open weight, under a restricted community license. Not open source. You can run M3 privately, audit its outputs, and fine-tune it — but you cannot reproduce it, cannot audit what went into it, and cannot verify any claim about its training data. For a sovereignty argument that is usually enough: weights running inside your own network are the point, not the recipe. For a reproducibility argument, it is not.
Claude Opus 5 offers none of this and does not pretend to. If your requirement is that inference happens on hardware you control, Opus 5 cannot meet it at any price, and no configuration changes that.
Speed, latency, and what the measurements mean
All figures in this section are third-party measurements against each vendor's API read on July 28, 2026; neither vendor publishes absolute speed figures, and these readings move as serving capacity changes. MiniMax M3 sustains 77 output tokens per second against 55 for Claude Opus 5 — about 40 percent faster, meaningful but not decisive. Time to first token is a different story entirely: a median 1.64 seconds for M3 against 78.91 seconds for Opus 5, roughly forty-eight times. That figure is not a serving deficiency. It is what maximum effort looks like from the outside — Opus 5 is thinking before it speaks, and the score of 61 is the product of exactly that thinking. Dial effort down and the wait falls with it, along with the score.
These two models therefore suit different interaction shapes. Anything a human waits on in real time is close to unusable at a 79-second first token. Anything queued, batched, or run overnight does not care about first-token latency at all, and there the 55-against-77 throughput difference is the only speed number that matters.
Winner by category
Claude Opus 5 takes capability, knowledge recency, long-context pricing, and reasoning control. MiniMax M3 takes cost, latency, throughput, output length, input modalities, and self-hosting. Neither wins licensing freedom outright. These categories are not evenly weighted, and we are not going to pretend they are.
- Raw capability — Claude Opus 5, decisively. 61 against 44, and 59 against 44 even at Opus 5's default setting.
- Cost — MiniMax M3, decisively. USD 0.12 per index task against USD 2.03, and no Opus 5 effort level gets within three times of it.
- Latency — MiniMax M3. A median 1.64 seconds to first token against 78.91 at maximum effort.
- Throughput — MiniMax M3. 77 output tokens per second against 55.
- Long-context pricing — Claude Opus 5. Flat rates across the full million tokens against M3's doubling above 512,000.
- Maximum output length — MiniMax M3. Up to 524,288 tokens in a normal call against 128,000 synchronous.
- Knowledge recency — Claude Opus 5. A stated May 2026 cutoff against no published cutoff at all.
- Input modalities — MiniMax M3. Text, image, and video in, against text and image.
- Reasoning control — Claude Opus 5. Five graded effort levels against a binary switch whose documented default contradicts itself between endpoints.
- Self-hosting and data sovereignty — MiniMax M3. Downloadable ungated weights against none.
- Licensing freedom — neither, and this surprised us. Opus 5 is API-only. M3's weights arrive under a license whose default grant is non-commercial, with a mandatory attribution line and a USD 20 million revenue threshold above which you need written authorization. Downloadable is not the same as unencumbered.
When to pick each model
Pick Claude Opus 5 when the hardest task in your workload is genuinely hard and a wrong answer is expensive: complex multi-step agentic work, research-grade scientific reasoning, large refactors across an unfamiliar codebase. Pick it when your prompts routinely exceed 512,000 tokens, because flat pricing across the full window is worth more at that length than any headline rate. Pick it when knowledge recency matters, because a stated May 2026 cutoff beats an unstated one. And pick it when you want to tune spend against quality without switching vendors — the five-rung effort ladder has no equivalent on M3.
Pick MiniMax M3 when volume dominates difficulty. Classification, extraction, routing, tagging, first-pass summarization, and any pipeline where you run millions of similar tasks and the quality bar is "correct," not "brilliant." Pick it when a human is waiting on the response, because 1.64 seconds against 79 is not a preference, it is a different product. Pick it when the data cannot leave your network — that is the argument no index can score. Pick it when you need video input, which Opus 5 does not accept, or output longer than 128,000 tokens in a single call.
Run both if the workload has a shape. The most economical pattern we can defend from these numbers is a difficulty router: send everything to M3 first and escalate only what fails to Opus 5, at default high effort before reaching for max. If most of your traffic completes on M3, the blended cost lands far closer to USD 0.12 than to USD 2.03, and the hard tail still gets a frontier model. Before committing to M3 commercially, read the license and budget for the attribution line.
Frequently asked questions
Which model is better, Claude Opus 5 or MiniMax M3?
Claude Opus 5, on capability, and it is not close. It scores 61 against 44 on version 4.1 of the independent Artificial Analysis Intelligence Index, a seventeen-point gap that is among the widest in this comparison series. MiniMax M3 wins on price, latency, maximum output length, video input, and the availability of downloadable weights. If your question is purely which model is smarter, the answer is Opus 5, and no configuration of M3 changes it.
How much cheaper is MiniMax M3 than Claude Opus 5?
About seventeen times cheaper on measured cost per index task — USD 0.12 against USD 2.03 for Opus 5 at maximum effort, read July 28, 2026. On list prices it is USD 0.30 against USD 5.00 per million input tokens, and USD 1.20 against USD 25.00 per million output tokens. Even at Opus 5's cheapest effort level, low, which costs USD 0.36 per task, M3 is still three times cheaper. There is no Opus 5 setting that undercuts it.
Is MiniMax M3 open source?
No. It is open weight under a restricted license, which is a different thing. The weights are published ungated on Hugging Face, but the repository contains no training data, no training code, and no reproduction recipe. The license file is titled MINIMAX COMMUNITY LICENSE, is recorded in the repository metadata as "other," and is not approved by the Open Source Initiative. Calling it open source is wrong; calling it open weight is accurate.
What does the MiniMax M3 license actually allow?
Its grant covers dealing in the software "for non-commercial purposes." A separate clause then sets conditions for commercial use: you must prominently display "Built with MiniMax M3" on a website, user interface, blog post, about page, or product documentation, and send a one-time notice to MiniMax. If your products generate more than 20 million US dollars in yearly revenue, you must obtain prior written authorization instead. An appendix also prohibits military use and several other categories.
At what effort level is each model's index score measured?
Claude Opus 5's 61 is measured at maximum effort. Its default setting, high, scores 59. Artificial Analysis publishes all five rungs: 61 at max, 60 at xhigh, 59 at high, 56 at medium, and 51 at low. MiniMax M3 appears exactly once, as "MiniMax-M3," with no effort suffix in the displayed name — the level is not labeled, because M3 exposes no graded effort ladder. Artificial Analysis notes that it is showing M3's reasoning version.
Do both models really have a one-million-token context window?
Both are documented at 1,000,000 tokens, but they price it differently. Anthropic states that Opus 5 includes the full window at standard pricing and that a 900,000-token request bills at the same per-token rate as a 9,000-token one. MiniMax doubles M3's rates above 512,000 input tokens, from USD 0.30 and USD 1.20 to USD 0.60 and USD 2.40 per million. Both windows are real; only one costs the same throughout.
Which model can generate longer outputs?
MiniMax M3, by a wide margin. Its API reference gives a recommended maximum of 131,072 output tokens and permits up to 524,288 in a normal call. Claude Opus 5 caps at 128,000 tokens synchronously and reaches 300,000 only through the Batch API behind a beta header, which is not available on the synchronous Messages API. For long-form generation in a single request, M3 has roughly four times the ceiling.
What is MiniMax M3's knowledge cutoff?
There is no published one. Not in the Hugging Face model card, not on MiniMax's model page, not in the API documentation, and not in the MSA architecture paper. That is an absence in the public record rather than a claim about the model being outdated, and we are not going to infer a date from the June 1, 2026 release. Claude Opus 5, by contrast, has a stated training data cutoff and reliable knowledge cutoff of May 2026.
Why is Claude Opus 5 so slow to respond?
Because at maximum effort it thinks before it answers. Independent measurement on July 28, 2026 puts its median time to first token at 78.91 seconds against 1.64 seconds for MiniMax M3. That wait is not a serving problem; it is the reasoning that produces the score of 61. Lower effort levels return faster and score lower. For throughput once generation starts, the gap narrows sharply: 55 output tokens per second against 77.
Can MiniMax M3 process video?
Yes, as input. M3 accepts text, image, and video and outputs text only, so it can watch and describe a video but cannot generate one. Claude Opus 5 accepts text and image. Video generation at MiniMax is handled by the separate Hailuo product line, which has its own models and its own pricing and should not be confused with the M-series language models.
Can I self-host either model?
MiniMax M3, yes. The weights are on Hugging Face without a gate and can be served through SGLang, vLLM, or Transformers on your own hardware, which is the strongest argument for M3 in any environment where data cannot leave the network. Claude Opus 5, no. Anthropic publishes no weights and the model is available only through APIs, so no budget or configuration makes on-premises inference possible.
What would change this verdict?
Three things. A published knowledge cutoff for MiniMax M3 would close the freshness asymmetry that currently favors Opus 5. A second Artificial Analysis entry for M3 at a labeled reasoning configuration would tell us whether 44 is a ceiling or a midpoint, which we currently cannot say. And a license revision by MiniMax would materially improve its case, since the current terms are the main obstacle to recommending M3 for commercial products above the revenue threshold. We will revisit this page as each resolves.
The verdict
Claude Opus 5 is the winner, and the scope of that win needs stating precisely: capability, knowledge recency, flat long-context pricing, and granular control of the quality-against-cost trade. Seventeen index points is among the widest gaps we have documented between two models on this site, and there is no reading of the evidence in which MiniMax M3 is the more capable model.
But the index is not the whole product, and this is where M3 earns its page. It is about seventeen times cheaper to run, returns a first token roughly forty-eight times faster, emits four times more tokens in a single call, accepts video, and — uniquely between these two — can be downloaded and run on hardware you control. Where difficulty is modest and volume is large, or where data cannot leave the building, those properties beat seventeen index points comfortably. That is not a consolation prize; it is a different purchase.
The number that decides between them is not 61 or 44. It is the difficulty of the hardest task you actually need done, and the honest way to find it is empirical: run that task on M3 first. If it completes, you have your answer at a seventeenth of the cost. If it fails, escalate to Opus 5 at its default high effort before reaching for max — high delivers 59 of the available 61 points at roughly half the cost per task, the best value on the entire curve.
Two things we will not tell you, because the evidence does not support them. We will not tell you that MiniMax M3 is open source — the weights are genuinely downloadable and ungated, but the repository ships no training data and no training code, and the MINIMAX COMMUNITY LICENSE grants rights for non-commercial purposes with attribution duties and a USD 20 million revenue threshold above it. And we will not tell you where M3's 44 sits within its own range, because MiniMax exposes no effort ladder and Artificial Analysis lists the model exactly once, unlabeled. What we will tell you is that a model seventeen points below the best in the world, at a seventeenth of the price, with weights anyone can download, is a rational default for most production traffic.
If you want the wider field, we have compared MiniMax M3 against Kimi K3, against Claude Opus 4.8, and against Claude Fable 5, and you can read our full write-ups of GLM-5.2 and Claude Opus 4.8 for the open-weight and frontier models sitting between these two on the index.
Sources and references
Every figure on this page comes from a vendor's own documentation or from an independent evaluator, labeled separately throughout rather than stacked into a single ranking. MiniMax's self-reported benchmarks and Anthropic's own safety figures are not used for any head-to-head claim. Cost per task, output speed, and latency are sliding measurements read July 28, 2026; index scores are stable within Intelligence Index version 4.1.
- Anthropic — Introducing Claude Opus 5 (vendor: release date)
- Anthropic — Pricing (vendor: token rates, cache, batch, long-context pricing)
- Anthropic — Models overview (vendor: context window, maximum output, cutoffs, modalities)
- Anthropic — Effort (vendor: five effort levels, default high)
- Anthropic — Context windows (vendor: 1M default, standard pricing)
- Anthropic — Batch processing (vendor: 300,000-token extended output beta)
- Anthropic — Model migration guide (vendor: model identifier, web fetch and Priority Tier exceptions)
- MiniMax — Model release notes (vendor: M3 release date, M-series lineage)
- MiniMax — Pay-as-you-go pricing (vendor: token rates, 512K tiering, cache reads)
- MiniMax — OpenAI-compatible chat API (vendor: context, maximum output, modalities, thinking default)
- MiniMax — Anthropic-compatible chat API (vendor: thinking values and conflicting default)
- MiniMax — Models introduction (vendor: line-up, Hailuo video family)
- MiniMax — M3 model page (vendor: architecture, guaranteed-minimum context wording)
- Hugging Face — MiniMaxAI/MiniMax-M3 (vendor: weights, repository contents, downloads, gating)
- Hugging Face — MiniMax M3 LICENSE file (vendor: full license text)
- arXiv — MiniMax Sparse Attention (vendor: architecture paper)
- Artificial Analysis — Model leaderboard (independent: index scores per effort level, cost per task)
- Artificial Analysis — MiniMax M3 (independent: index score, open-weights label, reasoning-version note)
- Artificial Analysis — Claude Opus 5 vs MiniMax M3 (independent: blended price, speed, latency, license labels)
- Artificial Analysis — AA-Omniscience (independent: knowledge reliability scoring)
- Artificial Analysis — Intelligence benchmarking methodology (independent: index v4.1 composition)
Last compared July 28, 2026. Pricing, context limits, output limits, reasoning parameters, modalities, and cutoff dates come from Anthropic's and MiniMax's own live documentation read that day. Capability, cost, speed, and latency figures are Artificial Analysis Intelligence Index version 4.1, also read July 28; the cost and speed figures move as models are re-run. The MINIMAX COMMUNITY LICENSE terms were read directly from the LICENSE file in the published Hugging Face repository. MiniMax's self-reported benchmark results are labeled as such and kept separate from independent measurement throughout.
Our Verdict
Claude Opus 5 wins on capability and it is not close: 61 against 44 on version 4.1 of the Artificial Analysis Intelligence Index, among the widest gaps in this comparison series, plus a stated May 2026 cutoff and flat pricing across its full million-token window. MiniMax M3 wins everything the index does not measure — about seventeen times cheaper per task at USD 0.12 against USD 2.03, a first token in 1.64 seconds against 78.91, four times the single-call output ceiling, video input, and downloadable ungated weights. Unlike earlier comparisons in this series, no Claude Opus 5 effort level undercuts M3 on cost; even its cheapest rung costs three times more while scoring 51 against 44. One caveat that changes the open-weight case: M3's LICENSE file is the MINIMAX COMMUNITY LICENSE, not MIT — its grant covers non-commercial purposes, commercial use requires displaying “Built with MiniMax M3”, and products above USD 20 million in yearly revenue need prior written authorization.
Choose Claude Opus 5
Anthropic's frontier reasoning model — top of the independent index at half the price of Fable 5.
Try Claude Opus 5 →Choose MiniMax M3
Open-weight frontier model from MiniMax combining near-frontier coding, a 1M token context window, and native multimodality — from $0.30 per million input tokens.
Try MiniMax M3 →Frequently Asked Questions
Is Claude Opus 5 better than MiniMax M3?
Claude Opus 5 wins on capability and it is not close: 61 against 44 on version 4.1 of the Artificial Analysis Intelligence Index, among the widest gaps in this comparison series, plus a stated May 2026 cutoff and flat pricing across its full million-token window. MiniMax M3 wins everything the index does not measure — about seventeen times cheaper per task at USD 0.12 against USD 2.03, a first token in 1.64 seconds against 78.91, four times the single-call output ceiling, video input, and downloadable ungated weights. Unlike earlier comparisons in this series, no Claude Opus 5 effort level undercuts M3 on cost; even its cheapest rung costs three times more while scoring 51 against 44. One caveat that changes the open-weight case: M3's LICENSE file is the MINIMAX COMMUNITY LICENSE, not MIT — its grant covers non-commercial purposes, commercial use requires displaying “Built with MiniMax M3”, and products above USD 20 million in yearly revenue need prior written authorization.
Which is cheaper, Claude Opus 5 or MiniMax M3?
Claude Opus 5 is priced at $5 in / $25 out per M tokens. MiniMax M3 is priced at $0.3 in / $1.2 out per M tokens. Check the pricing comparison section above for a full breakdown.
What are the main differences between Claude Opus 5 and MiniMax M3?
The key differences span across 14 features we compared. For Artificial Analysis Intelligence Index v4.1, Claude Opus 5 offers 61 at max effort, 59 at default high effort while MiniMax M3 offers 44, effort level not labeled. For Cost per index task (measured July 28, 2026), Claude Opus 5 offers USD 2.03 at max, USD 1.06 at default high, USD 0.36 at low while MiniMax M3 offers USD 0.12. For Input price per million tokens, Claude Opus 5 offers USD 5.00 while MiniMax M3 offers USD 0.30 at or below 512K input, USD 0.60 above. See the full feature comparison table above for all details.

