Skip to content
C

Claude Opus 5

Anthropic's frontier reasoning model — top of the independent index at half the price of Fable 5.

9.7/10
Last updated July 27, 2026
Author
Anthony M.
57 min readVerified July 27, 2026Tested hands-on

Quick Summary

Claude Opus 5 is Anthropic's frontier reasoning model (API claude-opus-5), released July 24, 2026. It runs a 1M-token context window with a May 2026 knowledge cutoff and costs 5 dollars per million input tokens and 25 dollars per million output tokens — unchanged since Opus 4.5. It scores 61 on the independent Artificial Analysis Intelligence Index v4.1, first of 190 models at max effort, one point ahead of Claude Fable 5 at half the price. Research-led review. Our rating: 9.7 out of 10.

Claude Opus 5 — 1M context, 5 and 25 dollars per million tokens
Claude Opus 5 — released July 24, 2026, at the same list price as Opus 4.8

Claude Opus 5 (API model claude-opus-5) is Anthropic's frontier reasoning model, released July 24, 2026. It costs 5 dollars per million input tokens and 25 dollars per million output tokens — the same list price Opus has carried since 4.5 — runs a 1M-token context window, and ships with a May 2026 knowledge cutoff, four months fresher than Fable 5, Opus 4.8, and Sonnet 5. On the independent Artificial Analysis Intelligence Index v4.1 it scores 61 at max effort, ranking first of 190 models, one point ahead of Claude Fable 5 at half Fable 5's price. The catch is latency: 62.68 seconds to first token at max effort.

Quick verdict

Claude Opus 5 is the strongest reasoning model you can currently rent, and it is priced like a mid-generation refresh rather than a frontier launch. Anthropic held the Opus line at 5 and 25 dollars per million tokens for the fifth release running, which means the model that now tops the independent Artificial Analysis leaderboard costs half of Claude Fable 5, the model it just displaced.

Our rating is 9.7 out of 10, and we want to be exact about what that number represents: it is a research-led assessment built from Anthropic's documentation, the 194-page Claude Opus 5 system card, independent third-party measurement, and reported user feedback. It is not the output of a metered hands-on test protocol of our own. We set out precisely what we did and did not do in the methodology section below.

Three things make this release matter. First, the price hold: frontier-tier index performance at half the cost of the previous leader changes the economics of hard, low-volume reasoning work. Second, the May 2026 knowledge cutoff, the most recent in the Claude family by four months, which reduces how much recent context you have to supply yourself. Third, a genuine five-level effort ladder that lets one model span a range of cost and latency profiles which previously required switching models.

Three things should give you pause. The latency is severe — 62.68 seconds to first token at max effort. The model is measurably more verbose than its price-tier peers, and lowering the effort level does not reliably shorten what it says. And Anthropic's own system card documents a hallucination rate that rose alongside accuracy, an unusual and important admission we quote in full below.

How we researched this review

We have not run our own benchmark suite against Claude Opus 5, and this review makes no hands-on claims. The model was released on July 24, 2026, and this page was compiled the same day from primary and independent sources.

What we did: we read Anthropic's launch announcement and the full Claude Opus 5 system card directly, and we pulled every hard specification — model identifiers, list pricing, cache rates, context and output limits, rate-limit tiers, effort levels, thinking behavior, and platform availability — from the Anthropic model documentation and the official pricing tables rather than from press summaries. For independent performance data we used the Artificial Analysis evaluation of Opus 5, which runs its own harness rather than republishing vendor figures.

What we did not do: we have not reproduced any benchmark, measured cost per task on our own workloads, tracked latency across a working day, or characterized failure modes that only surface over weeks of use. Every capability number on this page is attributed to whoever produced it, and we are careful about three distinct provenance buckets — evaluations Anthropic ran itself, evaluations run or scored by third parties such as Cognition, Zapier, and the ARC Prize Foundation, and independent measurement by Artificial Analysis. We never merge them into a single ranking, because they are measured on different harnesses and are not comparable.

One sourcing note worth stating plainly: the benchmark charts on Anthropic's launch page are images with no text layer, so no number on this page is cited from the announcement. Every vendor figure below comes from the system card, with the page number given.

Why publish on day one anyway: the questions teams actually have at launch — what does it cost, what breaks if I migrate, how fresh is its knowledge, is it fast enough for my product — are all answerable from primary documentation without testing. Those answers are the substance of this page. We will revise it as independent evaluations accumulate.

What is Claude Opus 5?

Claude Opus 5 is the fifth generation of Anthropic's Opus tier, the line reserved for its hardest reasoning work. Anthropic announced it on July 24, 2026 in a post titled Introducing Claude Opus 5, describing a model "priced at $5 per million input tokens and $25 per million output tokens" and available "today on all platforms."

The API identifier is claude-opus-5, with no date suffix. That format is deliberate and worth understanding: since the Claude 4.6 generation, Anthropic's model IDs are dateless but still pinned snapshots rather than evergreen pointers, so claude-opus-5 will not silently become a different set of weights. On Amazon Bedrock the identifier is anthropic.claude-opus-5; on Google Cloud it is claude-opus-5, matching the first-party string.

Where it sits in the family is the interesting part. Claude Fable 5, generally available since June 9, 2026, remains Anthropic's most capable widely released model by the company's own designation, and it is the top capacity tier at 10 and 50 dollars per million tokens. Opus 5 slots in beneath it on paper at half the price — and then edges past it on the leading independent index. Below both sit Claude Sonnet 5 at 3 and 15 dollars and Claude Haiku 4.5 at 1 and 5 dollars. Claude Opus 4.8, the previous Opus flagship, has moved into Anthropic's collapsed Legacy models table.

On the consumer side, Opus 5 is the new default model on Claude Max and the strongest model available on Claude Pro. Note that wording carefully: strongest on Pro, not most capable overall — that designation still belongs to Fable 5. It is available through the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry, and it powers agentic sessions in Claude Code and Claude Cowork.

One migration deadline attaches to this release. Claude Opus 4.1 (claude-opus-4-1-20250805) is deprecated and, per the Anthropic migration guide, will be retired on August 5, 2026, with Opus 5 named as the recommended target. If you still have Opus 4.1 pinned anywhere in production, that is a dated task, not a someday task.

Key features

The headline capabilities, all drawn from Anthropic's model documentation:

  • 1M-token context window. One million tokens is both the default and the maximum. There is no beta header to enable it and no premium long-context tier — Anthropic's documentation states that for every model with a 1M-token window, "1M is the default: you don't need a beta header, and long-context requests are billed at standard pricing."
  • 128k maximum output, or 300k in batch. The synchronous Messages API caps output at 128,000 tokens. The Message Batches API supports up to 300,000 output tokens with the output-300k-2026-03-24 beta header.
  • May 2026 knowledge cutoff. Both the reliable knowledge cutoff and the training data cutoff are May 2026 — unusual, since those two dates normally diverge. Fable 5, Opus 4.8, and Sonnet 5 all sit at January 2026, making Opus 5 four months fresher than every sibling.
  • Five effort levels. Opus 5 supports the full ladder: low, medium, high, xhigh, and max. The default is high on the Claude API and in Claude Code.
  • Adaptive thinking. Opus 5 supports adaptive thinking and rejects the older extended-thinking flag: a request with thinking.type: "enabled" returns a 400 error. Thinking output is hidden by default, since thinking.display defaults to omitted; pass display: "summarized" to surface it.
  • Fast mode. An opt-in tier delivering up to 2.5 times higher output tokens per second, at twice the base price. Covered in detail under pricing, including its several restrictions.
  • Cheaper prompt caching floor. The minimum cacheable prompt length drops to 512 tokens, down from 1,024 on Opus 4.8, so shorter system prompts and tool definitions now qualify for cache discounts.
  • Lower tool-use overhead. Tool definitions consume 286 tokens with tool_choice set to auto or none, and 406 tokens for any or tool. Opus 4.8 costs 290 and 410; Opus 4.7 cost 675 and 804.
  • Quadrupled rate limits in a separate bucket. Opus 5 does not share the combined Opus 4.x rate-limit pool. Detailed below.
  • No data retention requirement. Anthropic states that "consistent with prior Opus models, Opus 5 does not have data retention requirements for general access" — a meaningful contrast with Fable 5 and Mythos 5, which are designated Covered Models carrying 30-day retention and are unavailable under zero data retention.
  • Lighter-touch cyber classifiers. Anthropic expects its cyber classifiers "to intervene around 85% less often than they do for Fable 5." When a request is flagged in Claude.ai, Claude Code, or Claude Cowork, it falls back to Opus 4.8 by default.
Claude Opus 5 specifications: 1M context, 128k output, May 2026 cutoff, five effort levels
Claude Opus 5 headline specifications, per Anthropic's model documentation

The rate-limit change deserves its own paragraph because it is easy to miss and it directly affects throughput planning. Anthropic's rate-limit documentation notes that the Opus limit is "a total limit that applies to combined traffic across Claude Opus 4.8, Opus 4.7, Opus 4.6, and Opus 4.5" and that "Claude Opus 5 has a separate rate limit and is not part of this combined bucket." On the Start tier — Anthropic's tiers are named Start, Build, Scale, and Custom rather than numbered — Opus 5 gets 1,000 requests per minute, 2,000,000 input tokens per minute, and 400,000 output tokens per minute. Fable 5 on the same tier gets 500,000 input and 100,000 output tokens per minute, so Opus 5 has four times the input and output throughput allowance of the model above it in the lineup.

One capability is explicitly absent. Priority Tier is not supported on Claude Opus 5, stated plainly in both the migration guide and the service-tiers documentation, which lists Opus 5 among the exceptions. The broader context softens this: Anthropic notes that Priority Tier capacity commitments are no longer available for purchase at all.

Pricing

Standard rates, from Anthropic's pricing table:

  • Input: 5 dollars per million tokens.
  • Output: 25 dollars per million tokens.
  • Prompt cache write, 5-minute: 6.25 dollars per million tokens.
  • Prompt cache write, 1-hour: 10 dollars per million tokens.
  • Prompt cache hits and refreshes: 0.50 dollars per million tokens.
  • Batch API: 2.50 dollars input and 12.50 dollars output per million tokens, a 50 percent discount.

The number that matters most is the one that did not change. Opus has been 5 and 25 dollars per million tokens since Opus 4.5, and Anthropic held that line through 4.6, 4.7, 4.8, and now 5. A workload that runs on Opus 4.8 today costs exactly the same to run on Opus 5, which removes the usual frontier-upgrade budgeting exercise entirely.

Set against the rest of the family, the positioning is aggressive. Fable 5 costs 10 and 50 dollars per million tokens, double Opus 5, and Fable 5 is the model Opus 5 narrowly outscores on the independent index. Sonnet 5 lists at 3 and 15 dollars, currently discounted to 2 and 10 dollars per million tokens through August 31, 2026. Haiku 4.5 sits at 1 and 5 dollars. Opus 5 is therefore the second-most-expensive model in the lineup while topping the independent leaderboard — an inversion of the usual price-capability ordering.

Fast mode is a separate product and it is easy to describe imprecisely, so here is what Anthropic's fast-mode documentation actually says. It "delivers up to 2.5x higher output tokens per second from Claude Opus 5 and Claude Opus 4.8 at premium pricing," enabled by setting speed: "fast" with the fast-mode-2026-02-01 beta header. Pricing is 10 dollars input and 50 dollars output per million tokens, twice the base rate. The critical nuance: "Speed benefits are focused on output tokens per second (OTPS), not time to first token (TTFT)." Fast mode makes the model type faster once it starts; it does not make it start sooner. Given that Opus 5's weak point is precisely time to first token, fast mode does not address the latency problem most teams will hit.

Fast mode also carries real restrictions. It is a research preview available on the Claude API only, including Claude Managed Agents — not on Amazon Bedrock, Google Cloud, Microsoft Foundry, or Claude Platform on AWS. It cannot be combined with the Batch API. It has a dedicated rate limit separate from standard Opus limits, it is not available with a Priority Tier commitment, and access is gated: Anthropic directs teams to contact an account manager or join a waitlist.

Because API billing is per token, there is no monthly figure to quote — your bill is a function of token volume, cache hit rate, and how much the model chooses to say. That last variable is unusually significant here, and we return to it under limitations. On the consumer side, Opus 5 access comes bundled with paid Claude plans rather than metered per token, as the new default on Max and the strongest option on Pro.

Independent benchmarks: Artificial Analysis Intelligence Index v4.1

This is the section to read if you want a number Anthropic did not produce. Artificial Analysis runs its own harness across models and publishes a composite Intelligence Index; version 4.1 combines nine evaluations, including GDPval-AA v2, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, and CritPt. Every score in this section is on index v4.1, retrieved July 24, 2026, and is only comparable to other v4.1 scores — index versions are not interchangeable, and comparing across them produces nonsense.

Claude Opus 5 scores 61 on Intelligence Index v4.1, ranking first of 190 models, in the configuration Artificial Analysis labels "Adaptive Reasoning, Max Effort."

That headline needs two immediate qualifications. The 61 is achieved at max effort, which is not the default — the default high setting scores 59. And the margin over second place is a single point. Here is the full ladder Artificial Analysis published across Opus 5's effort settings, alongside the v4.1 field:

  • Opus 5 at max effort: 61 — first of 190.
  • Opus 5 at xhigh effort: 60.
  • Opus 5 at high effort (the default): 59.
  • Opus 5 at medium effort: 56.
  • Opus 5 at low effort: 51.
  • Claude Fable 5: 60.
  • GPT-5.6 Sol: 59.
  • Kimi K3: 57.
  • Claude Opus 4.8: 56.
  • Median among comparably priced reasoning models: 32.

That median deserves a footnote of its own, because it is easy to overstate. Artificial Analysis reports 32 as the median "among other reasoning models in a similar price tier" — it is a peer-group median, not the midpoint of all 190 models on the board. The same qualification applies to every median we quote from this evaluation.

Read the ladder honestly and the picture is narrower than "new number one." Opus 5 at its default effort setting scores 59, level with GPT-5.6 Sol and one point below Fable 5. You have to explicitly request max effort — accepting the latency and token cost that come with it — to claim the top spot, and it wins by one point. Two of those comparisons are close enough that rounding is doing real work: on the unrounded figures Opus 5 at high effort and GPT-5.6 Sol at max effort are separated by roughly three hundredths of a point, which is a tie in any meaningful sense. A one-point separation at the top sits well inside the range where a harness change or a re-run could reorder the leaders, so "narrowly ahead" is the accurate phrasing, not "ahead."

There is one more caveat that cuts against over-reading the Fable 5 comparison. Artificial Analysis does not evaluate Fable 5 as a bare model: its official entry is named "Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)," shortened on the leaderboard to "Claude Fable 5 (with fallback)." That fallback is baked into the measured configuration rather than noted as an asterisk. Artificial Analysis never defines what the fallback does, so we quote the label and decline to explain the mechanism. What it means practically is that the 60 is a score for a configuration, not for a single set of weights — one more reason to treat a one-point gap as a tie in practice. Worth noting too that Opus 5 at xhigh effort also rounds to 60 but is fractionally ahead on the raw figure, which places Fable 5 third of 190 rather than second.

Artificial Analysis Intelligence Index v4.1: Opus 5 at max effort 61, Fable 5 60, GPT-5.6 Sol 59, Kimi K3 57, Opus 4.8 56
Artificial Analysis Intelligence Index v4.1, independent measurement, retrieved July 24, 2026 — the top five sit within five points

The same evaluation produced cost and speed measurements that are less flattering and, for most production decisions, more consequential:

  • Output speed: 52.8 tokens per second, ranking 115th of 190 models, against a peer median of 70.9 tokens per second.
  • Time to first token: 62.68 seconds at max effort, against a 2.87-second median among comparably priced reasoning models.
  • Verbosity: roughly 100 million output tokens to complete the index run, against a 63 million peer median — about 59 percent more output than the typical comparable model.
  • Cost per Intelligence Index task: 2.03 dollars, versus 2.75 dollars for Fable 5, 1.04 dollars for GPT-5.6 Sol, and 0.95 dollars for Kimi K3.

Those figures reframe the leaderboard win. Opus 5 buys four index points over Kimi K3 for more than double the cost per task, and it is slower than roughly 60 percent of the field on output throughput. The time-to-first-token figure is the one to sit with: 62.68 seconds is more than twenty times the peer median. That is a batch-processing latency profile, not an interactive one.

The effort ladder does give you a real dial on both cost and latency, and the numbers are worth having in front of you when you choose a setting. Artificial Analysis measured cost per task and time to first token at every level: 2.03 dollars and 62.68 seconds at max, 1.56 dollars and 33.84 seconds at xhigh, 1.06 dollars and 22.29 seconds at high, 0.62 dollars and 5.04 seconds at medium, and 0.36 dollars and 3.86 seconds at low. Dropping from max to the default high setting costs two index points and returns nearly two-thirds of the cost and two-thirds of the latency. Dropping to medium costs five points and brings time to first token down to roughly five seconds. For most workloads that trade is worth making deliberately rather than inheriting.

Vendor and third-party benchmarks: what the system card reports

Everything in this section comes from the 194-page Claude Opus 5 system card, with page numbers given. We keep it separate from the Artificial Analysis data above so the two are never blended, and we distinguish within it between evaluations Anthropic ran itself and evaluations run or scored by outside parties — a distinction the system card makes and which is easy to lose in summary.

One methodological caution applies throughout. The system card notes that unless stated otherwise, Opus 5 results use adaptive thinking at max effort averaged over five trials, while "competitor figures are drawn from the respective developers' published system cards or benchmark leaderboards." Competitor numbers in the summary table therefore have a different provenance from Anthropic's own runs.

Anthropic-run evaluations:

  • FrontierBench v0.1 (pages 152 to 153): Opus 5 achieved "a 44.4% mean reward, averaged over 5 attempts for each one of the 74 unique tasks, using xhigh effort" — note xhigh, not max, which "scored similarly and within noise, landing at 43%." On the same harness and infrastructure, GPT-5.6 Sol reached 37.5 percent, Fable 5 33.7 percent, and Opus 4.8 18.7 percent. A separate table earlier in the card reports a different FrontierBench set sourced from Harbor's evaluations, with Opus 5 at 43.3 and Opus 4.8 at 21.1; we quote the internal run and do not mix the two.
  • OSWorld 2.0 (page 173): 70.57 percent, "first-attempt success rate, averaged over five runs," against Opus 4.8 at 55.7 and Fable 5 at 66.1 — all three run by Anthropic on the same harness. The GPT-5.6 Sol figure of 62.6 sits on a different basis and is not comparable like-for-like: the card states that Sol and Muse Spark 1.1 scores were "sourced from their respective release posts."
  • IMO (page 153): "Opus 5's final score of 42/42 corresponds to gold-medal performance, well above the 2026 gold-medal cutoff score of 29/42 points." Run without agent harness or tools, with four independent solutions per problem judged by a panel of three frontier models requiring unanimity, plus human expert grading of one pre-specified solution per problem. The card also discloses that attempts exhausting the output limit were resampled at lower thinking efforts, and that one solution required this.
  • AA-Omniscience (page 107): a net score of 0.49, placing Opus 5 between Opus 4.8 and the two Mythos models. Discussed under limitations, because the interesting part is the breakdown.
  • ExploitBench (page 37): on the Full ACEs measure, Opus 5 records 99 against Opus 4.8's 2, with Mythos 5 at 132 and Sonnet 5 at 0. Two caveats belong with that number: Full ACEs pools results across both the plain and AutoNudge arms, and the card states that "as elsewhere in this card, production safety interventions are disabled during evaluation." We deliberately do not express this as a multiple of Opus 4.8's score — a ratio built on a base of 2 is arithmetically true and practically meaningless. Anthropic's own framing is that Opus 5 "is substantially behind Mythos 5 in its ability to exploit" vulnerabilities. This is a capability measurement, not a safety score, and it is the clearest quantitative reason cyber classifiers ship alongside this model.

Third-party run or scored evaluations:

  • FrontierCode (pages 150 to 152), built and scored by Cognition across 150 tasks: Opus 5 "ranks 2nd on FrontierCode (Main) with a 53.4% score," improving on Opus 4.8 at 46.5 and leading GPT-5.6 Sol at 47.5 — second because Fable 5 scores 53.5. On the extended set Opus 5 also ranks second at 63.6 percent, against Opus 4.8 at 59.6 and GPT-5.6 Sol at 60.6. Both best scores are reached at medium effort, which we return to under limitations.
  • AutomationBench (page 180), Zapier's leaderboard on a private held-out set: Opus 5 at max effort scored 26.0 percent, "a substantial gain over Claude Opus 4.8 (max effort) at 17.0% and Claude Fable 5 at 17.4%," with GPT-5.6 Sol at 18.1. At medium effort Opus 5 scored 24 percent at 0.89 dollars cost per task.
  • ARC-AGI (pages 181 to 182), verified by the ARC Prize Foundation on semi-private datasets — these are third-party verified figures, not Anthropic self-reported: 97.50 percent on ARC-AGI-1 and 90.42 percent on ARC-AGI-2 at max effort, and 30.16 percent on ARC-AGI-3 at high effort, the latter a Relative Human Action Efficiency score, with max-effort results unavailable at release. On ARC-AGI-3 the comparison is lopsided in Opus 5's favor, with GPT-5.6 Sol at 7.78 percent and Opus 4.8 at 1.52 percent. On ARC-AGI-2, however, Opus 5 loses: the card's summary table puts GPT-5.6 Sol at 92.5 against Opus 5's 90.4.
  • Indirect prompt injection (pages 72 to 73), built with Gray Swan, the UK AI Security Institute, and the US Center for AI Standards and Innovation: Opus 5 "reduc[ed] the probability of an attacker succeeding within 15 attempts from 5.5% to 2.0%" versus Opus 4.8, and from 0.5 percent to 0.2 percent at a single attempt. Comparisons on the same page put Sonnet 5 at 5.9 percent, Mythos 5 at 2.6 percent, and GPT-5.6 Sol at 20.0 percent at k=15. Anthropic attaches a caveat we are repeating rather than burying: "We evaluated Claude models without additional safeguards; other frontier models are evaluated on their publicly available endpoints, which may or may not include additional safeguards." That is not an apples-to-apples comparison, and the card says so.

One further disclosure from the FrontierBench section is worth surfacing, because it quantifies how often safety classifiers intercede in practice. During that evaluation, "Opus 5 safety classifiers flagged and refused 5% of the API calls, in 4% of the total trials, falling back to Opus 4.8," while "Fable safety classifiers flagged 42% API calls on 26% of trials." That gap — 5 percent of calls against 42 percent — is the concrete, measured version of Anthropic's claim that Opus 5's classifiers intervene far less often than Fable 5's.

Where Claude Opus 5 falls short

This section is the reason to read a review rather than a launch post. None of what follows is speculation — all of it comes from Anthropic's own documentation or from independent measurement.

It hallucinates more than Opus 4.8, and Anthropic says so. The system card states on page 3: "We found a surprising number of cases in which Opus 5 confidently stated an answer about which it was in fact unsure. The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall." Page 107 quantifies the trade-off on the public split of AA-Omniscience: "the Claude Opus 5's accuracy is 11% higher than Opus 4.8, but its rate of hallucinations is also 6% higher." Those are relative changes, not percentage points. The model is simultaneously more correct and more confidently wrong. For any workflow where an unflagged fabrication is expensive, that argues for keeping verification in the loop rather than relaxing it because headline capability went up.

The card corroborates this from user feedback as well, listing "overconfident and unsupported claims, sometimes from model-fabricated data, often followed by theatrical retractions" and "self-correction loops where the model continually attempted to reconsider its answer, especially at higher effort levels" among reported issues in pilot use.

The latency rules it out for interactive products at high effort. 62.68 seconds to first token at max effort, against a 2.87-second peer median, is not a tuning problem you can prompt your way out of. Lower effort levels help substantially — 22.29 seconds at the default high setting, 5.04 seconds at medium — but the configuration that earns the 61 is the slowest one. And fast mode, as covered above, improves output tokens per second and explicitly does not improve time to first token. If your product has a user waiting on a response, Opus 5 above medium effort is the wrong choice.

It is verbose, and the effort dial does not fix it. Independent measurement puts Opus 5 at roughly 100 million output tokens across the index run against a 63 million peer median. Anthropic's Opus 5 prompting guide is unusually direct about the mechanism: "Claude Opus 5's default user-facing responses run longer than prior Opus models'. The effort parameter controls how much the model thinks rather than how much it says: lowering effort can reduce thinking volume without reliably shortening the visible response. To control response length, prompt for it explicitly." The same page notes that files it writes to disk are often longer than on prior models, and that it can expand the scope of a task. Turning the effort dial down is a cost lever on reasoning, not a concision lever on output.

The cost per task is over double a strong open-weight competitor. 2.03 dollars per Intelligence Index task against 0.95 dollars for Kimi K3, for a four-point index difference. Whether that is good value depends entirely on how much those four points are worth on your specific work — on genuinely frontier problems the gap may be decisive, and on routine work it is very hard to justify.

Higher effort is not monotonically better — but read the full explanation. The system card notes on page 150 a "decline in FrontierCode score above high effort," and attributes it to "a tendency for Opus 5 at these effort levels to make more changes than the task requires (e.g., refactoring or making other edits to improve the codebase)," which the evaluation's grader penalizes as out-of-scope. Best results land at medium effort. Crucially, Anthropic does not leave it there: "We found that adding a brief instruction to the prompt telling the model to stay within the scope of the task recovered performance on most of these tasks, showing this is not primarily a model limitation. However, we report the Opus 5 scores on this benchmark without this prompt change." So the honest reading is that this is a scope-discipline and scoring artifact rather than a capability ceiling — and that Anthropic reported the unimproved numbers anyway, which is to its credit. The practical lesson stands: on scoped coding work, tell the model to stay in scope, and do not assume max effort is better.

No Priority Tier. If your capacity planning assumed a Priority Tier commitment for predictable throughput, Opus 5 does not support it.

Claude Opus 5 trade-offs: first on intelligence index, 62.68 seconds to first token, 2.03 dollars per task
The Opus 5 trade-off: top independent index score, bought with latency and output volume

What changes for developers migrating from Opus 4.8

Opus 5 is not a drop-in replacement. The migration guide documents several behavior changes that will surface as errors or as unexpected bills rather than as gradual degradation, so audit before you switch the model string.

  • Thinking is on by default now. On Opus 4.8, a request with no thinking field ran without thinking. On Opus 5, the same request runs with adaptive thinking. max_tokens remains a hard limit on total output — thinking plus response text together — so a request that previously spent its full budget on the answer now shares that budget with reasoning. Anthropic explicitly advises revisiting max_tokens for workloads that ran without thinking on Opus 4.8.
  • Disabling thinking now has a ceiling. thinking: {type: "disabled"} still works, but only at effort high or below. Combining it with xhigh or max returns a 400 error. Opus 4.8 accepted that combination, so any code path that disables thinking needs auditing before migration.
  • The extended-thinking flag is gone. thinking.type: "enabled" returns a 400 error on Opus 5, as it does on Opus 4.7, Opus 4.8, Sonnet 5, Fable 5, and Mythos 5. Use adaptive thinking instead.
  • Sampling parameters are hard-blocked. Non-default temperature, top_p, or top_k values return a 400 error on every request, whether or not thinking is used. Any wrapper that sets a house temperature by default will fail on every call.
  • Prompt caching starts paying off sooner. The minimum cacheable prompt length is 512 tokens on Opus 5, down from 1,024 on Opus 4.8, so system prompts and tool blocks that previously fell below the threshold now qualify.
  • Reasoning traces are hidden unless requested. thinking.display defaults to omitted; pass display: "summarized" if your application surfaces reasoning to users or logs.
  • Rate limits are a fresh pool. Opus 5 traffic does not draw from the combined Opus 4.x bucket, so migrating shifts load off that shared limit and onto a separate, considerably larger one.
  • Prompt hygiene changes too. Anthropic's Opus 5 prompting guide advises removing explicit verification instructions that helped on earlier models, because they now cause over-verification, and prompting explicitly for length if you want shorter answers.

Pros and cons

What we rate highly: the price hold at 5 and 25 dollars per million tokens while taking the top independent index position; the May 2026 knowledge cutoff, four months ahead of every sibling model; a genuine five-level effort ladder that lets one model cover several cost and latency profiles; 1M context as the default with no header and no premium; quadrupled rate limits in a dedicated bucket; the 512-token caching floor; safety classifiers that flagged 5 percent of calls during Anthropic's own FrontierBench run against Fable 5's 42 percent; and no data retention requirement, which matters for teams that cannot use Fable 5 for compliance reasons.

What gives us pause: 62.68 seconds to first token at max effort; a documented hallucination-rate increase of roughly 6 percent alongside the accuracy gain; verbosity around 59 percent above the peer median that the effort dial does not control; cost per task more than double Kimi K3; second place rather than first on the FrontierCode coding evaluations and on ARC-AGI-2; no Priority Tier support; and several breaking changes that make migration from Opus 4.8 an audit rather than a string swap.

Best use cases

Opus 5's profile — highest independent index score, severe latency at high effort, high output volume, moderate price — points at a specific class of work.

  • Hard problems where a wrong answer is expensive. Complex refactors, architectural decisions, and analysis where correctness dominates and a minute of waiting is irrelevant. This is the case the model was built for.
  • Asynchronous and batch pipelines. Overnight jobs, bulk document analysis, and offline evaluation, where the Batch API halves the price to 2.50 and 12.50 dollars per million tokens and the 300k output ceiling is available. Latency costs nothing here.
  • Long-context work that previously needed chunking. A 1M-token default window at standard pricing, paired with a May 2026 cutoff, means large codebases and document sets fit in one pass with less recent context to supply by hand.
  • Agentic coding with the effort dial tuned deliberately. In Claude Code, note that FrontierCode results peak at medium effort, and that a short in-scope instruction recovers most of the loss at higher settings. Both are worth testing on your own tasks rather than defaulting to max.
  • Compliance-constrained deployments. Teams that need frontier capability but cannot accept the 30-day retention attached to Fable 5 and Mythos 5 as Covered Models.
  • Security research within policy. The ExploitBench figures indicate a step change in offensive-security capability over Opus 4.8, which matters for defensive tooling — with the classifier behavior and fallback to Opus 4.8 factored into your design.
  • Migrating off deprecated Opus versions. Anyone still pinned to Opus 4.1 has until August 5, 2026, with Opus 5 as the named target.

Where we would not reach for it: anything with a user waiting on the first token at high effort or above, high-volume routine classification or extraction where Haiku 4.5 or Sonnet 5 costs a fraction, and cost-sensitive bulk work where Kimi K3 delivers 57 index points at 0.95 dollars per task.

Alternatives to Claude Opus 5

Claude Fable 5 is the direct in-family comparison and the strange one. It remains Anthropic's most capable widely released model by designation and costs double, at 10 and 50 dollars per million tokens, yet scores 60 on Intelligence Index v4.1 against Opus 5's 61 — and its cost per task on the same evaluation is higher, at 2.75 dollars. Fable 5 does still lead Opus 5 on the FrontierCode main set, 53.5 against 53.4, so the coding picture is not one-sided. Its other trade-off is compliance: Fable 5 is a Covered Model carrying 30-day data retention, which Opus 5 is not. Our Fable 5 versus Opus 4.8 comparison and the Fable 5 launch coverage have the fuller picture on that tier.

Claude Opus 4.8 is now a Legacy model at the identical 5 and 25 dollars per million tokens, scoring 56 on v4.1 against Opus 5's 61. At the same price for five fewer points there is little reason to start new work on it — though it is worth remembering that Opus 4.8 is what Claude.ai, Claude Code, and Cowork fall back to when a cyber classifier flags an Opus 5 request, so it remains in your dependency chain whether you choose it or not.

Claude Sonnet 5 at 3 and 15 dollars, currently 2 and 10 dollars through August 31, 2026, is the pragmatic choice for the large middle of real workloads, and it is the model most products default to. Our Sonnet 5 launch analysis covers why it landed as a default. Claude Haiku 4.5 at 1 and 5 dollars handles the high-volume tail.

GPT-5.6 Sol is the closest non-Anthropic comparison, scoring 59 on Intelligence Index v4.1 — level with Opus 5 at its default high effort setting, and two points behind at max. It also beats Opus 5 on ARC-AGI-2, at 92.5 against 90.4, and costs about half as much per index task at 1.04 dollars. Our GPT-5.6 Sol versus Opus 4.8 comparison covers that matchup against the previous generation.

Kimi K3 is the value argument and the hardest one to dismiss: 57 on v4.1 at 0.95 dollars per task, against Opus 5's 61 at 2.03 dollars. Four index points for more than twice the money is a trade many workloads should decline. See our Fable 5 versus Kimi K3 comparison for how that plays out against Anthropic's top tier.

Final verdict

Claude Opus 5 earns 9.7 out of 10 in our research-led assessment. The core achievement is economic rather than purely technical: Anthropic put the top score on the leading independent index at the same 5 and 25 dollars per million tokens that Opus has cost since 4.5, which makes the previous leader in its own lineup look mispriced at double the rate.

The honest framing of the capability win is narrow. Opus 5 leads Intelligence Index v4.1 by one point over Fable 5, and only at max effort — at its default high setting it scores 59 and sits level with GPT-5.6 Sol. Fable 5's measured configuration includes an Opus 4.8 fallback, and on the coding evaluations Opus 5 places second rather than first. Treat the top of that leaderboard as a cluster, not a ranking, and choose on the secondary characteristics instead: price, context, cutoff freshness, retention policy, and latency.

On those secondary characteristics Opus 5 is genuinely strong — May 2026 knowledge, 1M context at standard pricing, four times the rate-limit headroom of Fable 5, no retention requirement — and genuinely weak in one dimension that will disqualify it outright for a whole category of products. 62.68 seconds to first token at max effort is not a number you design an interactive experience around, and fast mode does not address it.

What would change our assessment: independent confirmation of how the documented hallucination increase behaves on real workloads rather than in system-card aggregates, a second independent index run to see whether the one-point lead over Fable 5 survives, and measured cost-per-task figures from production workloads rather than benchmark harnesses. We will update this page as those arrive.

If you are choosing today: use Opus 5 for hard asynchronous reasoning where correctness dominates and latency does not, prompt it explicitly for concision, test medium effort before assuming max is better, tell it to stay in scope on coding tasks, and audit your thinking and sampling parameters before you migrate anything from Opus 4.8.

Frequently asked questions

What is Claude Opus 5?

Claude Opus 5 is Anthropic's frontier reasoning model, released July 24, 2026, under the API model string claude-opus-5. It runs a 1M-token context window with up to 128,000 output tokens synchronously, carries a May 2026 knowledge cutoff, and costs 5 dollars per million input tokens and 25 dollars per million output tokens. On the independent Artificial Analysis Intelligence Index v4.1 it scores 61 at max effort, ranking first of 190 models. It is the default model on Claude Max and the strongest model available on Claude Pro.

How much does Claude Opus 5 cost?

Standard API pricing is 5 dollars per million input tokens and 25 dollars per million output tokens — unchanged from Opus 4.8, 4.7, 4.6, and 4.5. Prompt cache writes cost 6.25 dollars per million tokens for the 5-minute tier and 10 dollars for the 1-hour tier, while cache hits cost 0.50 dollars per million tokens. The Batch API halves the base rates to 2.50 and 12.50 dollars per million tokens. Because billing is per token, your actual cost depends on volume, cache hit rate, and how verbose the responses are.

Is Claude Opus 5 better than Claude Fable 5?

On the independent Artificial Analysis Intelligence Index v4.1, Opus 5 scores 61 at max effort against Fable 5's 60 — a one-point lead, narrow enough to treat as a tie. Opus 5 also costs half as much, at 5 and 25 dollars per million tokens versus 10 and 50, and its measured cost per index task is lower at 2.03 dollars against 2.75. Fable 5 still leads narrowly on the FrontierCode main set, 53.5 against 53.4, and retains Anthropic's designation as its most capable widely released model. One practical difference: Fable 5 is a Covered Model carrying 30-day data retention, while Opus 5 has no retention requirement.

What is Claude Opus 5's knowledge cutoff?

May 2026, for both the reliable knowledge cutoff and the training data cutoff — unusual, as those two dates normally differ. That makes Opus 5 four months fresher than Claude Fable 5, Claude Opus 4.8, and Claude Sonnet 5, which all sit at January 2026. In practice it means less recent context to supply yourself when asking about developments from early 2026.

How slow is Claude Opus 5?

Slow enough to rule out interactive use at high effort. Artificial Analysis measured time to first token at 62.68 seconds at max effort, against a 2.87-second median among comparably priced reasoning models, and output speed at 52.8 tokens per second, ranking 115th of 190. Lower effort levels help considerably: 22.29 seconds at the default high setting, 5.04 seconds at medium, and 3.86 seconds at low. Fast mode does not solve this, because Anthropic's documentation states its speed benefits are "focused on output tokens per second (OTPS), not time to first token (TTFT)."

Does Claude Opus 5 hallucinate more than Opus 4.8?

Yes, and Anthropic documents it. The Opus 5 system card states on page 3: "We found a surprising number of cases in which Opus 5 confidently stated an answer about which it was in fact unsure. The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall." Page 107 quantifies it on the AA-Omniscience benchmark: accuracy 11 percent higher than Opus 4.8, with a hallucination rate also 6 percent higher. The model is both more accurate and more confidently wrong, so verification steps should stay in place rather than be relaxed on the strength of the capability gain.

What are Claude Opus 5's effort levels?

Five: low, medium, high, xhigh, and max. The default is high on the Claude API and in Claude Code. On the Artificial Analysis Intelligence Index v4.1 those settings score 51, 56, 59, 60, and 61 respectively, at 0.36, 0.62, 1.06, 1.56, and 2.03 dollars per task. Higher is not always better: Anthropic's system card records a decline in FrontierCode scores above high effort, with the best results at medium, because the model makes more out-of-scope changes at higher settings.

What breaks when migrating from Opus 4.8 to Opus 5?

Four things to audit. Thinking runs by default on Opus 5 whereas omitting the field on Opus 4.8 meant no thinking, and max_tokens remains a hard cap on thinking plus response together, so budgets need revisiting. Combining thinking: {type: "disabled"} with effort xhigh or max returns a 400 error, though Opus 4.8 accepted it. Non-default temperature, top_p, or top_k values return a 400 error on every request. And the minimum cacheable prompt length drops to 512 tokens from 1,024, which changes which prompts qualify for cache pricing.

What is Claude Opus 5 fast mode and is it worth it?

Fast mode delivers up to 2.5 times higher output tokens per second at twice the base price — 10 dollars input and 50 dollars output per million tokens — enabled with speed: "fast" and the fast-mode-2026-02-01 beta header. It is worth it only for workloads bottlenecked on generation throughput, because it explicitly does not improve time to first token, which is Opus 5's real latency problem. It is also a research preview on the Claude API only, unavailable on Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS, incompatible with the Batch API and Priority Tier, and access-gated through an account manager or waitlist.

What are Claude Opus 5's rate limits?

Opus 5 has its own rate-limit bucket and does not share the pool covering combined traffic across Opus 4.8, 4.7, 4.6, and 4.5. On the Start tier it allows 1,000 requests per minute, 2,000,000 input tokens per minute, and 400,000 output tokens per minute. That is four times the input and output throughput Claude Fable 5 gets on the same tier, which allows 500,000 input and 100,000 output tokens per minute. Priority Tier is not supported on Opus 5, and Anthropic notes that Priority Tier capacity commitments are no longer available for purchase at all.

Can Claude Opus 5 output more than 128,000 tokens?

Yes, through the Message Batches API. The synchronous Messages API caps output at 128,000 tokens, but batch requests support up to 300,000 output tokens using the output-300k-2026-03-24 beta header. Batch pricing is also 50 percent lower, at 2.50 dollars input and 12.50 dollars output per million tokens, which makes batch the sensible route for long-form generation where latency does not matter.

Should I use Claude Opus 5 or a cheaper model?

It depends on how much four index points are worth on your work. Kimi K3 scores 57 on Intelligence Index v4.1 at a measured 0.95 dollars per task, against Opus 5's 61 at 2.03 dollars — more than double the cost for a modest gap. For routine extraction, classification, and high-volume work, Claude Haiku 4.5 at 1 and 5 dollars per million tokens or Claude Sonnet 5 at 3 and 15 dollars will be far more economical. Reserve Opus 5 for genuinely hard problems where a wrong answer is expensive and you can absorb the latency.

Sources and references

Methodology note: this review is research-led. Capability figures are attributed to their source throughout — evaluations Anthropic ran itself, evaluations run or scored by third parties such as Cognition, Zapier, and the ARC Prize Foundation, and independent measurement by Artificial Analysis are labeled distinctly and never merged. We have not independently reproduced any benchmark, and we do not present hands-on findings for this model.

Key Features

1M-token context window — default and maximum, no beta header, billed at standard pricing
128,000 max output tokens synchronously, up to 300,000 via the Message Batches API with a beta header
May 2026 knowledge cutoff — four months fresher than Fable 5, Opus 4.8, and Sonnet 5
Five effort levels: low, medium, high, xhigh, max — default high on the Claude API and Claude Code
Adaptive thinking supported; the older extended-thinking flag returns a 400 error
Fast mode research preview: up to 2.5 times higher output tokens per second at twice the base price
Prompt caching minimum lowered to 512 tokens, down from 1,024 on Opus 4.8
Tool-use overhead reduced to 286 tokens for auto or none, 406 for any or tool
Separate rate-limit bucket: 1,000 requests, 2,000,000 input and 400,000 output tokens per minute on the Start tier
No data retention requirement, unlike Fable 5 and Mythos 5 which are Covered Models
Cyber classifiers expected to intervene around 85 percent less often than for Fable 5
Available on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, Microsoft Foundry, Claude Code, and Claude Cowork

Pros & Cons

Pros

  • Tops the independent Artificial Analysis Intelligence Index v4.1 at 61, first of 190 models at max effort
  • Price held at 5 and 25 dollars per million tokens — identical to Opus 4.8, 4.7, 4.6, and 4.5, and half of Fable 5
  • May 2026 knowledge cutoff, the freshest in the Claude family by four months
  • Five-level effort ladder gives a real dial on cost and latency: 0.36 to 2.03 dollars per index task
  • 1M-token context as the default with no beta header and no long-context premium
  • Four times the input and output rate-limit allowance of Fable 5 on the same tier, in a dedicated bucket
  • No 30-day data retention requirement, which Fable 5 and Mythos 5 cannot offer
  • Prompt caching floor down to 512 tokens and lower tool-use overhead than Opus 4.8

Cons

  • 62.68 seconds to first token at max effort — unusable for interactive products at high effort
  • Anthropic documents a hallucination rate roughly 6 percent higher than Opus 4.8 alongside 11 percent better accuracy
  • Verbose: around 100 million output tokens on the index run against a 63 million peer median, and lowering effort does not reliably shorten responses
  • Cost per index task of 2.03 dollars is more than double Kimi K3 at 0.95 dollars for a four-point gap
  • Second rather than first on the FrontierCode coding evaluations and on ARC-AGI-2
  • Priority Tier is not supported
  • Migration from Opus 4.8 is an audit, not a string swap: thinking on by default, blocked sampling parameters, and a 400 error when disabling thinking above high effort

Best Use Cases

Hard reasoning problems where a wrong answer is expensive and a minute of latency is irrelevant
Asynchronous and batch pipelines, where the Batch API halves pricing and unlocks 300,000-token output
Long-context work on large codebases and document sets that previously required chunking
Agentic coding in Claude Code with the effort dial tuned deliberately rather than defaulted to max
Compliance-constrained deployments that cannot accept the 30-day retention attached to Fable 5
Defensive security research, given the step change in ExploitBench capability over Opus 4.8
Migrating workloads off Claude Opus 4.1 before its August 5, 2026 retirement

Platforms & Integrations

Available On

APIAWSAmazon BedrockGoogle CloudMicrosoft FoundryWebCLI

Integrations

Claude APIAmazon BedrockClaude Platform on AWSGoogle CloudMicrosoft FoundryClaude CodeClaude CoworkClaude Managed Agents
Anthony M. — Founder & Lead Reviewer
Anthony M.Verified Builder

We're developers and SaaS builders who use these tools daily in production. Every review comes from hands-on experience building real products — DealPropFirm, ThePlanetIndicator, PropFirmsCodes, and many more. We don't just review tools — we build and ship with them every day.

Written and tested by developers who build with these tools daily.

Was this review helpful?

Frequently Asked Questions

What is Claude Opus 5?

Anthropic's frontier reasoning model — top of the independent index at half the price of Fable 5.

How much does Claude Opus 5 cost?

Claude Opus 5 costs $5/month.

Is Claude Opus 5 free?

No, Claude Opus 5 starts at $5/month.

What are the best alternatives to Claude Opus 5?

Top-rated alternatives to Claude Opus 5 can be found in our WebApplication category, where we've reviewed and scored every tool on ThePlanetTools.ai.

Is Claude Opus 5 good for beginners?

Claude Opus 5 is rated 8.8/10 for ease of use.

What platforms does Claude Opus 5 support?

Claude Opus 5 is available on API, AWS, Amazon Bedrock, Google Cloud, Microsoft Foundry, Web, CLI.

Does Claude Opus 5 offer a free trial?

No, Claude Opus 5 does not offer a free trial.

Is Claude Opus 5 worth the price?

Claude Opus 5 scores 9.6/10 for value. We consider it excellent value.

Who should use Claude Opus 5?

Claude Opus 5 is ideal for: Hard reasoning problems where a wrong answer is expensive and a minute of latency is irrelevant, Asynchronous and batch pipelines, where the Batch API halves pricing and unlocks 300,000-token output, Long-context work on large codebases and document sets that previously required chunking, Agentic coding in Claude Code with the effort dial tuned deliberately rather than defaulted to max, Compliance-constrained deployments that cannot accept the 30-day retention attached to Fable 5, Defensive security research, given the step change in ExploitBench capability over Opus 4.8, Migrating workloads off Claude Opus 4.1 before its August 5, 2026 retirement.

What are the main limitations of Claude Opus 5?

Some limitations of Claude Opus 5 include: 62.68 seconds to first token at max effort — unusable for interactive products at high effort; Anthropic documents a hallucination rate roughly 6 percent higher than Opus 4.8 alongside 11 percent better accuracy; Verbose: around 100 million output tokens on the index run against a 63 million peer median, and lowering effort does not reliably shorten responses; Cost per index task of 2.03 dollars is more than double Kimi K3 at 0.95 dollars for a four-point gap; Second rather than first on the FrontierCode coding evaluations and on ARC-AGI-2; Priority Tier is not supported; Migration from Opus 4.8 is an audit, not a string swap: thinking on by default, blocked sampling parameters, and a 400 error when disabling thinking above high effort.

Ready to try Claude Opus 5?

Get started today

Try Claude Opus 5 Now