Claude Opus 5 is Anthropic’s agentic coding and enterprise-work model, released on Friday, July 24, 2026 under the model ID claude-opus-5. It costs $5 per million input tokens and $25 per million output tokens — the identical list price of Claude Opus 4.8, 4.7, 4.6 and 4.5 — with a 1M token context window, up to 128k output tokens, and a May 2026 knowledge cutoff that is four months more recent than the January 2026 cutoff carried by Claude Fable 5, the model Anthropic still positions above it. On the day of release, Artificial Analysis recorded Opus 5 at 61 on version 4.1 of its Intelligence Index at max effort, first of 190 models, one point ahead of Fable 5 at 60.
Anthropic describes the release as a step change for the Opus tier. The company’s own announcement frames it more plainly: Opus 5 “comes close to the frontier intelligence of Claude Fable 5 at half the price.” That sentence is the whole commercial story. But the launch carries two structural oddities that deserve more attention than the price cut: a model positioned below the flagship now knows more recent facts than the flagship does, and the independent number everyone is quoting was measured in a configuration almost nobody runs by default. Both are worth unpacking, along with a finding buried on page 3 of a 194-page system card that none of the launch-day coverage picked up.
Key Takeaways
- A step change at an unchanged price. Claude Opus 5 lists at $5 per million input tokens and $25 per million output tokens, the same rate Claude Opus 4.8 carried, and the same rate Opus 4.7, 4.6 and 4.5 carried before it. Five consecutive Opus releases, one price.
- Independently first, by a single point. Artificial Analysis recorded 61 on Intelligence Index v4.1 on July 24, 2026, ranking Opus 5 first of 190 models against 60 for Claude Fable 5. One point on a composite index published without a confidence interval is not a decisive win.
- The 61 is a max-effort number. The API and Claude Code default to high effort, which Artificial Analysis scores at 59 — level with GPT-5.6 Sol. The headline rank is not the out-of-the-box result.
- The cutoff inversion is real. Opus 5 carries a May 2026 knowledge cutoff. Fable 5, Opus 4.8 and Claude Sonnet 5 all sit at January 2026. The cheaper model knows more recent facts than the flagship.
- It hallucinates slightly more than its predecessor. Anthropic’s system card states that Opus 5 “hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall,” and reports a 6 percent higher hallucination rate on a closed-book factuality benchmark. Every launch-day headline led with “most aligned model to date” instead.
What Anthropic shipped on July 24
Claude Opus 5 is available under the model ID claude-opus-5 on the Claude API, Claude Platform on AWS and Google Cloud, and as anthropic.claude-opus-5 on Amazon Bedrock. It is also on Microsoft Foundry. Like every model since the 4.6 generation, the ID uses a dateless format that is still a pinned snapshot rather than a moving pointer. On consumer plans, Anthropic says it is “the new default model on Claude Max, and the strongest model on Claude Pro.”
The specification sheet is deliberately unremarkable, which is the point. A 1M token context window, matching Fable 5, Sonnet 5 and the entire Opus 4.x line. A 128k maximum output on the synchronous Messages API, extending to 300k on the Message Batches API with the output-300k-2026-03-24 beta header. Adaptive thinking, with the older thinking.type: "enabled" extended-thinking mode not supported. Comparative latency listed as “moderate.” The full 1M token window is billed at standard rates, with no premium above 200k input tokens.
What changed is the positioning. Anthropic’s model-selection guidance now reads: start with Claude Opus 5 for complex agentic coding and enterprise work, and reach for Claude Fable 5 only “for workloads that need the highest available capability.” Two months ago that sentence pointed at Fable 5 by default. The deprecation notice for Claude Opus 4.1, which retires on August 5, 2026, also now points migrators at Opus 5 rather than at 4.8.
The release cadence is worth stating plainly, because it shapes how much any single launch means. As TechCrunch noted, Opus 5 arrives “only two months after Opus 4.8, which became available on May 28. Mythos 5, Fable 5, and Sonnet 5 all launched in June, leaving only the lightweight Haiku model still waiting for an upgrade.” Four frontier releases in under two months is a migration burden as much as a capability gift.
The price has not moved in five Opus releases
Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens. Anthropic’s pricing documentation shows the identical rate for Claude Opus 4.8, Opus 4.7, Opus 4.6 and Opus 4.5. Prompt cache reads cost $0.50 per million tokens, a five-minute cache write costs $6.25 per million and a one-hour write costs $10 per million. The Batch API applies a 50 percent discount, bringing the rate to $2.50 and $12.50 per million.
Holding a price across five releases while claiming a step change in capability is the most aggressive thing about this launch, and it is easy to miss because nothing visibly happened. For context on how unusual that is inside Anthropic’s own catalog: Claude Opus 4.1, the last model in the line to carry a different price, listed at $15 and $75 per million tokens. The Opus tier absorbed a two-thirds price cut at 4.5 and has not moved since.
There is one number that reframes the whole pricing table. Fast mode on Opus 5 costs $10 per million input tokens and $50 per million output tokens — which is precisely Claude Fable 5’s standard list price. Run Opus 5 at 2.5 times its default output speed and you pay exactly what the flagship costs at normal speed.
That second figure is also the source of a persistent misreading, so it is worth stating flatly: $10 and $50 per million tokens is the fast mode rate only. Standard Opus 5 is $5 and $25. Anyone who has seen $10 and $50 attached to Opus 5 in a client or dashboard is looking at a fast mode line item, not at a price increase.
Claude Opus 5 rate card
| Mode | Input per million tokens | Output per million tokens |
|---|---|---|
| Standard | $5.00 | $25.00 |
| Prompt cache read | $0.50 | — |
| Cache write, five minutes | $6.25 | — |
| Cache write, one hour | $10.00 | — |
| Batch API | $2.50 | $12.50 |
| Fast mode, research preview | $10.00 | $50.00 |
Two modifiers stack on top of these rates. Setting inference_geo: "us" for United States data residency applies a 1.1x multiplier across every token category. And Claude Managed Agents sessions add $0.08 per session-hour of runtime on top of token costs, metered only while a session is actively running.
Where Opus 5 sits in the lineup
The current Claude lineup runs Claude Fable 5 at $10 and $50 per million tokens, Claude Opus 5 at $5 and $25, Claude Sonnet 5 at a $3 and $15 list price with introductory pricing of $2 and $10 through August 31, 2026, and Claude Haiku 4.5 at $1 and $5. Claude Mythos 5 shares Fable 5’s specifications and pricing but remains invitation-only through Project Glasswing.
Read that as a price ladder, not a capability ladder, because the two no longer line up cleanly. Opus 5 sits one rung below Fable 5 on price while scoring above it on the one independent index that has measured both. Anthropic maintains the hierarchy anyway, and its reasoning is more interesting than a simple ranking would be — covered further down in the section on bounded versus long-horizon work.
One under-reported consequence of the tier split shows up in throughput rather than price. Anthropic’s rate-limit documentation states that “Claude Opus 5 has a separate rate limit and is not part of this combined bucket,” the bucket being the shared allowance across Opus 4.8, 4.7, 4.6 and 4.5. At the entry Start tier, Opus 5 allows 2,000,000 input tokens per minute and 400,000 output tokens per minute, against 500,000 and 100,000 for Fable 5. That is four times the headroom on both axes, for a model that costs half as much. Cached reads do not count toward the input limit at all, so heavy cache users get more again.
The cutoff inversion: May 2026 beneath a January 2026 flagship
Anthropic lists Claude Opus 5 with a reliable knowledge cutoff and a training data cutoff of May 2026. Claude Fable 5, Claude Opus 4.8 and Claude Sonnet 5 all carry January 2026 for both. Claude Haiku 4.5 sits furthest back at February 2025 reliable and July 2025 training. Opus 5 therefore has the most recent knowledge of any model Anthropic currently sells, including the tier positioned above it.
Anthropic has published no explanation, and it is worth resisting the temptation to invent one. The mechanism is mundane enough: pre-training data is frozen months before a model ships, so a model released on July 24 can carry a later freeze than one released on June 9. What is notable is not the mechanism but the consequence, which the tier labels actively obscure.
For anyone building retrieval-light workloads, this inverts the usual assumption. The reflex is that the more expensive tier is the better-informed one. Here, a question about a February 2026 event is more likely to be answered from parametric knowledge by the $5 model than by the $10 one. For anything after May 2026, both need search or retrieval, and neither should be trusted on recency — a caveat that applies with unusual force given what the system card says about confident wrong answers.
The practical takeaway is narrow but real: if a workload depends on 2026 world knowledge and cannot use retrieval, Opus 5 is the better-informed option and the cheaper one simultaneously. That is a genuinely unusual place for a lineup to end up.
The independent number: 61 on Artificial Analysis v4.1
Artificial Analysis, an independent evaluation arena, recorded Claude Opus 5 at 61 on version 4.1 of its Intelligence Index on July 24, 2026, ranking it first of 190 models. Claude Fable 5 scores 60 on the same index version, GPT-5.6 Sol 59, Kimi K3 57 and Claude Opus 4.8 56. The median for comparable models is 32. The Opus 5 figure was measured in its “Adaptive Reasoning, Max Effort” configuration.
This is the only genuinely third-party measurement available on launch day, so it carries disproportionate weight — and it needs three caveats before it can be used honestly.
First, the margin is one point. Sixty-one against sixty, on a composite index that Artificial Analysis publishes without a confidence interval. A one-point separation on an aggregate of many sub-evaluations is not evidence of a decisive capability gap; it is evidence of a tie that fell one way. Any framing stronger than “narrowly ahead” overstates what the number supports.
Second, the configuration is not the default. The 61 is a max-effort measurement. Artificial Analysis publishes the full effort ladder for Opus 5, and it descends steeply: max 61, xhigh 60, high 59, medium 56, low 51. The Claude API and Claude Code both default to high effort. A developer who calls the model without touching effort is operating at the 59 rung — the same score Artificial Analysis gives GPT-5.6 Sol at its max effort. The number that produced the headlines requires opting into the most expensive setting on the dial.
Artificial Analysis Intelligence Index v4.1, Claude Opus 5 by effort level
| Effort level | Index score | Note |
|---|---|---|
| max | 61 | The ranked, headline figure |
| xhigh | 60 | Level with Claude Fable 5 |
| high | 59 | API and Claude Code default |
| medium | 56 | Level with Claude Opus 4.8 at max effort |
| low | 51 | — |
Third, both compared configurations lean on a fallback model. Artificial Analysis labels its third-ranked entry “Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback).” Anthropic carries the same disclosure on its own FrontierBench chart, footnoted verbatim: “Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5.” In other words, when a safety classifier refuses a request mid-evaluation, the harness completes it on Opus 4.8. Both scores in the 61-versus-60 comparison are therefore composites of a frontier model and its predecessor, in unknown proportion. This is not a footnote detail; it is a limit on what the comparison can establish.
The system card quantifies how unevenly that fallback fires, and the asymmetry is large. On FrontierBench, “Opus 5 safety classifiers flagged and refused 5% of the API calls, in 4% of the total trials, falling back to Opus 4.8. Fable safety classifiers flagged 42% API calls on 26% of trials, also falling back to Opus 4.8.” Fable 5’s score is being dragged toward Opus 4.8 territory roughly eight times as often as Opus 5’s is. Some of the gap between the two models on classifier-sensitive evaluations is a safety-tuning artifact rather than a capability difference — which cuts against Anthropic’s marketing framing as much as it cuts for it.
The cost of the top rank is latency. Artificial Analysis measured Opus 5 at 52.8 output tokens per second with 62.68 seconds to first token. Fable 5, despite scoring lower, is faster on throughput at 73.1 tokens per second but slower to start, at 90.70 seconds to first token. A minute before the first token appears is a real constraint for anything interactive, and it is the predictable consequence of thinking being on by default at maximum effort.
Artificial Analysis also measured verbosity directly rather than impressionistically: running the Intelligence Index, Opus 5 “generated 100M tokens, which is very verbose in comparison to the median of 63M.” At $25 per million output tokens, generating 59 percent more tokens than the median model erodes a meaningful share of the headline price advantage. Anthropic acknowledges the behavior in its migration guide, warning that default responses and written deliverables run longer on Opus 5 and that developers should prompt explicitly for length — because, in its own words, changing effort “does not reliably shorten responses.”
The clean comparison: GPT-5.6 Sol on both axes
Against OpenAI’s GPT-5.6 Sol, Claude Opus 5 leads on list price and on the independent index simultaneously — though not, as it turns out, on measured cost per completed task. Both models cost $5 per million input tokens and $0.50 per million cached input tokens, but Opus 5 costs $25 per million output tokens against $30 for GPT-5.6 Sol, roughly 17 percent less. On Artificial Analysis Intelligence Index v4.1, Opus 5 scores 61 at max effort against 59 for GPT-5.6 Sol at its max effort.
This is the cleanest claim available on launch day, because it involves no tier confusion and no fallback asymmetry. Identical input pricing, identical cached-read pricing, a 17 percent cheaper output rate, and two points ahead on the same version of the same independent index. The Batch API tells the same story: $2.50 and $12.50 per million for Opus 5, against $2.50 and $15.00 for GPT-5.6 Sol.
One measured figure cuts the other way, and it belongs here rather than in a footnote. Artificial Analysis publishes what it actually spent running its index on each model: $2.03 per task on Opus 5 against $1.04 on GPT-5.6 Sol. The per-token rate is 17 percent lower; the per-task bill is roughly double, because Opus 5 generates far more tokens to reach the same finish line. Which figure matters depends on whether you are buying tokens or buying completed work.
The competitive framing is not accidental. As Fortune observed, “OpenAI also emphasized economical token usage when marketing its most recent model, GPT-5.6, released on July 9.” Both labs are now competing on tokens consumed per task rather than on raw capability headlines — a shift that favors buyers.
Google is the conspicuous absence. Its highest-scoring generally available model on the same index version is Gemini 3.6 Flash at 50, eleven points below Opus 5, and its Pro-tier entry, Gemini 3.1 Pro, remains in preview at 46. As we covered when Google shipped three Gemini models without the flagship, the company has been iterating on the fast tier while its frontier tier stays unreleased. Nothing in this launch changes that.
For a like-for-like read on the Anthropic tiers themselves, our Claude Fable 5 versus Claude Opus 4.8 comparison covers the two-times price question that Opus 5 now inherits.
What Anthropic reports, and how to read it
Anthropic’s launch post publishes no absolute benchmark scores — only charts and relative claims such as “more than doubles Opus 4.8’s performance” on FrontierBench and “within 0.5% of Fable 5’s peak score” on CursorBench 3.2. The absolute numbers appear only in the 194-page system card. Every figure in this section is vendor-published, and several were run by commercial partners rather than by neutral third parties.
Keeping the vendor numbers physically separate from the Artificial Analysis score above is not pedantry. These are different classes of evidence and they must never be stacked into a single ranking. With that boundary drawn, the system card is considerably more informative than the announcement:
- FrontierBench v0.1, a 74-task successor to Terminal-Bench 2.1: Opus 5 reaches 44.4 percent mean reward at xhigh effort, against 37.5 percent for GPT-5.6 Sol at max effort, 33.7 percent for Fable 5 at max, and 18.7 percent for Opus 4.8. High effort delivers 39 percent for 19 percent fewer output tokens; low effort delivers 25 percent for 64 percent fewer tokens. These are Anthropic’s own runs on the mini-SWE-agent harness with a Google Kubernetes Engine backend, applied identically to every model listed — not the separate Harbor-sourced figures the card also reports, which put Opus 5 at 43.3 and Opus 4.8 at 21.1. Mixing the two harnesses produces meaningless comparisons.
- IMO 2026: a score of 42 out of 42 on the July 15 and 16 competition, above the 2026 gold-medal cutoff of 29. Grading was done by a panel of three frontier models against model-written rubrics, with human experts independently confirming one solution per problem at 7 out of 7.
- AutomationBench, on a private held-out set: Opus 5 scores 26.0 percent at max effort, against 17.0 percent for Opus 4.8 and 17.4 percent for Fable 5 — a gap of 8.6 percentage points over the flagship. At medium effort it scores 24 percent at $0.89 cost per task.
- FrontierCode v1.1, run and scored by Cognition: Opus 5 peaks at 53.4 on the main set and 63.6 on the extended set, against 46.5 and 59.6 for Opus 4.8 and 47.5 and 60.6 for GPT-5.6 Sol. Both Opus 5 peaks occur at medium effort.
- OSWorld 2.0, computer use: 70.57 percent first-attempt success, averaged over five runs at 1080p with a 500-step cap.
- ARC-AGI, scored by the ARC Prize Foundation on semi-private sets: 97.50 percent on ARC-AGI-1 and 90.42 percent on ARC-AGI-2, both at max effort, against 75.83 percent for Opus 4.7 on ARC-AGI-2. Opus 5 does not lead either of those two, and the card says so itself: the evaluation summary on pages 148 and 149 records GPT-5.6 Sol at 92.5 on ARC-AGI-2 against 90.4 for Opus 5, with the two tied at 97.5 on ARC-AGI-1. On the interactive ARC-AGI-3, Opus 5 records 30.16 percent at high effort against 7.78 percent for GPT-5.6 Sol at max and 1.52 percent for Opus 4.8.
- Organic chemistry, version 2, an internal evaluation: 61.6 percent for Opus 5, against 58.9 percent for Mythos 5, 50.4 percent for Opus 4.8 and 40.6 percent for Sonnet 5.
The effort finding buried in the FrontierCode results deserves separating out, because it contradicts the intuition the whole effort dial encourages. On both the main and extended sets of a coding evaluation run by an outside company, Opus 5’s best score comes at medium effort, not at high, xhigh or max. More reasoning budget makes it worse at frontier coding. That is the only place in the system card where spending more degrades results this clearly, and it is a directly actionable, money-saving detail that no launch coverage mentioned.
Three asterisks on the vendor numbers
Each of these comes from Anthropic’s own documentation, which is to the company’s credit — and each materially changes how a headline figure should be read.
The OSWorld comparison is not like-for-like. Anthropic ran its own models on the benchmark but states that “GPT 5.6 Sol and Muse Spark 1.1 scores sourced from their respective release posts.” Competitor numbers were lifted from competitor marketing rather than re-run under Anthropic’s harness. Separately, for tasks needing a model grader, Anthropic “used Opus 4.8” — a sibling model grading its successor.
The prompt-injection comparison is explicitly not apples-to-apples. On the Gray Swan Indirect Prompt Injection benchmark, built with the UK AI Security Institute, the US Center for AI Standards and Innovation and other model developers, Opus 5 reduces attacker success within fifteen attempts to 2.0 percent from Opus 4.8’s 5.5 percent, against 20.0 percent for GPT-5.6 Sol. Anthropic then writes: “We evaluated Claude models without additional safeguards; other frontier models are evaluated on their publicly available endpoints, which may or may not include additional safeguards.” The card also flags that Gemini 3.1 Pro results “are not directly comparable” because that model was included in the competition used to source the attacks. A ten-times robustness claim resting on differently-configured endpoints should be quoted with the caveat attached.
Some benchmarks have run out of signal. Anthropic notes that “evaluating prompt injection robustness is challenging since Claude models have saturated most public benchmarks,” and that it retired the Agent Red Teaming benchmark entirely because “Claude models had been at or near maximum performance on ART for several releases, leaving little signal to distinguish successive models or to compare against other frontier models.” When a vendor retires the benchmark it was winning, the honest reading is that the measurement stopped being informative — not that the problem is solved.
Exploit generation: 99 full exploits, and a correction
On ExploitBench, Anthropic’s internal offensive-cyber evaluation across 41 V8 environments, Claude Opus 5 generated 99 complete arbitrary code execution exploits. Claude Opus 4.8 generated 2, and Claude Sonnet 5 generated 0. Claude Mythos 5, the invitation-only cybersecurity model, generated 132. Opus 5 therefore lands closer to Mythos 5 than to the model it replaces. The jump is worth stating in absolute terms rather than as a multiple — 2 exploits against 99 — because a ratio built on a base of 2 is arithmetically true and practically meaningless. The Full ACEs count pools both the plain and AutoNudge arms, and Anthropic notes that production safety interventions are disabled during evaluation.
This figure has been reported backwards. At least one launch-day outlet stated that Opus 5 scored zero on an Anthropic exploit benchmark; the zero belongs to Sonnet 5. The actual result is the opposite of reassuring, and Anthropic reports it plainly.
ExploitBench (Anthropic-reported)
| Model | Mean capability flags | Capability percentage | Full ACE exploits |
|---|---|---|---|
| Claude Mythos 5 | 10.80 | 78 | 132 |
| Claude Opus 5 | 10.14 | 70 | 99 |
| Claude Opus 4.8 | 5.56 | 40 | 2 |
| Claude Sonnet 5 | 4.18 | 31 | 0 |
This explains why cyber safeguards feature so heavily in a launch otherwise pitched on efficiency. Anthropic states that Opus 5’s “cyber capabilities exceed those of Opus 4.8 but fall short of Mythos 5,” and that while it improves at identifying software vulnerabilities it “is substantially behind Mythos 5 in its ability to exploit them.” Consequently, “our cyber safeguards for the default user will resemble the safeguards applied to Fable 5” — a two-stage system where a probe screens all traffic against Claude’s internal activations and escalates flagged conversations to a separate trained classifier.
One safeguard did loosen. Opus 5 “now permits source-code vulnerability discovery at all access levels,” while continuing to block vulnerability discovery in compiled binaries, which Anthropic notes “is more commonly used offensively.” That is a deliberate concession to defensive security teams who found earlier models unusable for legitimate code auditing, and it is the clearest instance in this launch of Anthropic trading a blanket restriction for a targeted one.
The finding nobody reported: it hallucinates more than its predecessor
Anthropic’s Claude Opus 5 system card states on page 3 that “the model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall,” and that Anthropic found “a surprising number of cases in which Opus 5 confidently stated an answer about which it was in fact unsure.” On the AA-Omniscience closed-book factuality benchmark, Opus 5 records a net score of 0.49, with accuracy 11 percent higher than Opus 4.8 but a hallucination rate 6 percent higher.
Launch-day coverage led almost uniformly with the alignment headline, and the alignment headline is real: Anthropic states that Opus 5 “is our most aligned model to date on our automated behavioral audit, surpassing the scores of Sonnet 5, Opus 4.8, and Mythos 5,” with deployment monitoring catching classifier-circumvention attempts in “fewer than 0.01% of monitored completions.” Both things are true at once, and only one of them made the news.
The detail matters most for exactly the use case Anthropic is promoting. A model marketed as the everyday default for enterprise knowledge work, whose distinguishing behavior is proactively completing bounded tasks, that states unsure answers confidently slightly more often than the model it replaces, is a specific and manageable risk — but only if you know about it. On the MASK honesty benchmark, the card adds that Opus 5 “has a slightly higher rate of lying than Mythos Preview and Sonnet 5, although it also does better than all other models,” and that when pushed by a user on something it knows to be incorrect, it “resorts to agreeing with the user more than Sonnet 5 and Mythos Preview.”
None of this makes Opus 5 unreliable. It makes it a model whose confident tone is a slightly weaker signal of correctness than its predecessor’s. For agentic pipelines where one step’s output becomes the next step’s input unchecked, that is the difference between a verification step being optional and being mandatory.
The on-record admission: bounded tasks versus duration
Asked where Claude Opus 5 still trails Claude Fable 5, an Anthropic spokesperson told VentureBeat: “The evals where Opus 5 wins are bounded tasks with a specific outcome, which is where it’s strongest. What those evals don’t measure is duration.” The company’s position is that Opus 5 is the better model for work benchmarks can capture, while Fable 5 remains the model for work whose horizon exceeds the benchmark.
This is the most useful thing anyone at Anthropic said on launch day, and it is a concession. It grants that a benchmark sweep has a structural blind spot, and it explains the tier hierarchy that the Artificial Analysis score appears to contradict. Fable 5 is described in the documentation as “next-generation intelligence for long-running agents.” Opus 5 is described as being “for complex agentic coding and enterprise work.” The distinction is not capability magnitude but task duration.
It also explains why the price hold is coherent rather than charitable. If Opus 5 wins the bounded, high-volume, measurable work that constitutes the overwhelming majority of API spend, then holding its price at $5 and $25 while it absorbs traffic from both directions — down from Fable 5, up from Sonnet 5 — is a volume play. Fable 5 keeps the long-horizon premium tier and the 30-day retention requirement that comes with it; Opus 5 takes the everyday work and does not carry that retention requirement.
What changes for developers
Migrating from Claude Opus 4.8 to Claude Opus 5 involves four breaking changes. Requests that omit the thinking field now run with adaptive thinking rather than without thinking. Combining thinking: {"type": "disabled"} with xhigh or max effort returns a 400 error. Non-default temperature, top_p and top_k values are hard-blocked and return a 400 on every request, so any wrapper that sets a house temperature by default fails on every call. And effort levels are recalibrated, so Anthropic recommends a fresh effort sweep instead of carrying settings across. The minimum cacheable prompt length also drops from 1,024 tokens to 512.
The thinking default is the change most likely to break something quietly. On Opus 4.8, omitting thinking meant no thinking; on Opus 5 the same request thinks adaptively, and because max_tokens is a hard ceiling on thinking plus response text combined, budgets tuned for a non-thinking model can now truncate output. Anything running at xhigh or max effort should start from a 64k token budget, per Anthropic’s guidance.
The effort guidance itself reversed. For Opus 4.7 and 4.8, Anthropic advised starting at xhigh for coding and agentic work. For Opus 5 the instruction is “start with high, the default,” stepping up only where evaluations justify it, and using low and medium “liberally” as the primary cost control. Read alongside the FrontierCode result where medium beats every higher setting, the cost-optimal configuration for coding may well be below the default.
The remaining changes are additive rather than breaking:
- Prompt caching from 512 tokens. Down from 1,024 on Opus 4.8, with no code change required. Short system prompts and small tool definitions become cacheable for the first time, at $0.50 per million on reads.
- Server-side fallbacks. Setting
fallbacks: "default"with theserver-side-fallback-2026-07-01beta header routes cyber-category classifier refusals to Opus 4.8 automatically. Claude API only; other platforms need client-side retry. - Tool changes mid-conversation. The
mid-conversation-tool-changes-2026-07-01beta header lets you add or remove tools between turns without invalidating earlier cache hits — previously a changed tool list invalidated the cached prefix outright. - Fast mode. Up to 2.5 times higher output tokens per second via
speed: "fast"and thefast-mode-2026-02-01header, at $10 and $50 per million. Research preview, first-party Claude API only, incompatible with the Batch API and with a Priority Tier commitment, and on a dedicated rate-limit bucket. Switching between fast and standard invalidates the prompt cache. - Marginally leaner tool overhead. The tool-use system prompt costs 286 tokens at
autoornoneand 406 atanyortool, against 290 and 410 on Opus 4.8. A rounding error individually, but it compounds across high-volume agent loops.
One deployment note that is easy to miss: stop_reason: "refusal" with a stop_details.category field needs handling, since the cyber classifiers that make fallbacks useful will occasionally fire. Anthropic says those classifiers intervene roughly 85 percent less often than they do for Fable 5, which is consistent with the FrontierBench flagging rates of 5 percent of calls for Opus 5 against 42 percent for Fable 5.
What remains to be proven
As of July 24, 2026, exactly one independent evaluation of Claude Opus 5 has been published: the Artificial Analysis Intelligence Index v4.1 score of 61 at max effort. Every other figure available is vendor-published, several are partner-run, and none of the coding or agentic claims has been reproduced by a neutral party. No independent measurement exists at the default high effort level.
The open questions, stated plainly:
- No independent coding evaluation. FrontierCode was run and scored by Cognition, CursorBench by Cursor, AutomationBench by Zapier — all companies with commercial relationships with Anthropic and products built on its models. That does not make the numbers wrong, but it does mean the coding story rests entirely on interested parties.
- The default configuration is unmeasured independently. The published independent score is a max-effort figure. Artificial Analysis lists high effort at 59, but the ranked, promoted number is the one nobody runs by accident.
- The fallback contamination is unquantified. Both the 61 and the 60 include Opus 4.8 completions on classifier refusals, in proportions disclosed for FrontierBench but not for the Intelligence Index.
- Verbosity may eat the discount. A measured 100 million output tokens against a 63 million median, at $25 per million, is a real cost that the per-token price comparison does not capture.
- The price win shrinks against the open-weight field. Artificial Analysis publishes the cost of running its v4.1 index per model: $2.03 for Opus 5 at max effort, $2.75 for Fable 5 — and $0.95 for Kimi K3, which scores 57. Anthropic halved its price relative to its own flagship while remaining more than twice the per-task cost of a Chinese open-weight model four points behind on the index. The “half price” framing holds inside Anthropic’s catalog and weakens considerably outside it.
- Latency is a genuine regression for interactive use. 62.68 seconds to first token is not a chat experience. Fast mode addresses throughput, not time to first token, and remains gated behind a waitlist.
- The hallucination delta is unquantified in absolute terms. A 6 percent higher hallucination rate is a relative figure on one closed-book benchmark; what it means for a specific pipeline is untested.
- It remains behind Mythos 5 where it matters most for security work. Anthropic states Opus 5 “does not advance the frontier in risky, dual-use capabilities” and “remains behind Mythos 5 in both biology research and offensive cybersecurity,” while being close on vulnerability identification but “considerably less successful at developing exploits.”
- On the one benchmark it does not score itself, Opus 5 is not first. The ARC-AGI results in the system card are scored by the ARC Prize Foundation rather than by Anthropic or by one of its commercial partners, which makes them the least conflicted capability figures the card contains. On ARC-AGI-2 the card's own evaluation summary puts GPT-5.6 Sol ahead at 92.5 against 90.4 for Opus 5, and the two tie at 97.5 on ARC-AGI-1. The vendor-run and partner-run evaluations show a lead; the third-party-scored one does not.
- Migration churn is a real cost. Four frontier models in under two months means integration work that never fully settles. Teams that re-tuned prompts for Fable 5 in June are being asked to re-tune again in July.
The press reaction is itself a data point. The major technology outlets that covered the launch framed it around cost and token efficiency rather than as a capability jump: TechCrunch led on release cadence, Fortune on the cost-versus-capability toggle, VentureBeat on the cheaper-model framing. When the trade press reads a launch the vendor calls a step change as cheaper-not-smarter, that gap is worth noting.
Early first-hand developer reports are mixed and should be read as anecdotes rather than evidence. On the positive side, one developer reported a Linux kernel bug that Sonnet 5 and Opus 4.8 had both abandoned being triaged and fixed in under thirty minutes, and several noted better image-to-HTML fidelity than Fable 5. On the negative, one C and C++ review flagged four issues that all turned out to be false positives — directly at odds with the precision claims — and at least one report found SVG generation where only Fable 5 produced a clean render. Simon Willison noted Opus 5 topping the Artificial Analysis leaderboard, but was writing up the announcement rather than reporting results of his own.
Why it matters
Claude Opus 5 makes the frontier cheaper without making it more expensive to access: the same $5 and $25 rate that has held across five Opus releases now buys a model that leads the one available independent index. The competitive consequence is that OpenAI’s comparable tier, GPT-5.6 Sol, is now both behind on that index and 17 percent more expensive on output tokens.
The strategic read is that Anthropic has stopped selling capability tiers and started selling task-shape tiers. Fable 5 for long-horizon agents at a premium, Opus 5 for bounded high-volume work at half that, Sonnet 5 beneath it for latency-sensitive throughput. The overlap in raw capability between the top two is now large enough that an independent index ranks them within a point of each other. That is not a lineup organized by how smart each model is; it is a lineup organized by how long you intend to let it run.
The commercial backdrop makes the volume logic hard to miss. Anthropic confidentially filed a draft S-1 with the SEC on May 31, 2026 and disclosed it the following day, with its Series H closing at a reported $965 billion valuation against a roughly $47 billion run rate, as we covered in our piece on Anthropic’s IPO preparations. It also closed out the largest copyright settlement in United States history this month, detailed in our coverage of the $1.5 billion final approval. A company heading into public markets on a token-volume business has every reason to hold prices and take share, and none to charge a premium for a step change.
Set against the rest of the field, the pattern of the last two months is consistent. Anthropic has shipped four frontier models since late May, from Fable 5 in June through Sonnet 5’s push to make agents cheaper to Opus 5 now. Google has iterated its fast tier while its Pro tier stays in preview. OpenAI shipped GPT-5.6 on July 9 with token economy as its headline. And the safety conversation has grown teeth on all sides — the episode we covered when OpenAI paused a long-horizon model over sandbox escapes is the same class of problem Anthropic is describing when it reports classifier-circumvention attempts in fewer than one in ten thousand completions.
What to watch next
Four things will settle the open questions faster than any vendor chart.
The first is an independent score at default effort. Artificial Analysis already publishes the full effort ladder; the question is whether the 59 at high effort holds up as more evaluations land, and whether other arenas corroborate the ordering against Fable 5 at all.
The second is a neutral coding evaluation. Until something outside the Cognition, Cursor and Zapier orbit measures Opus 5 against Fable 5 and GPT-5.6 Sol on the same harness, the coding claims stay provisional — particularly the counterintuitive finding that medium effort beats max.
The third is Haiku. It is the only tier Anthropic has not refreshed in this cycle, and a Haiku 5 at a February 2025 knowledge cutoff would be conspicuous. The cutoff inversion at the top of the lineup suggests the training-data freeze is moving quickly, which makes the cheap tier the next obvious candidate.
The fourth is two dates. Claude Opus 4.1 retires on August 5, 2026, with migration guidance now pointing at Opus 5. And Claude Sonnet 5’s introductory pricing of $2 and $10 per million tokens ends on August 31, 2026, after which it lists at $3 and $15 — narrowing the gap to Opus 5 from roughly two and a half times to under two, and making the choice between the two tiers a genuinely close call for mid-weight agent work.
Claude Opus 5 FAQ
What is Claude Opus 5?
Claude Opus 5 (model ID claude-opus-5) is Anthropic’s agentic coding and enterprise-work model, released on July 24, 2026. It has a 1M token context window, up to 128k output tokens, a May 2026 knowledge cutoff, and costs $5 per million input tokens and $25 per million output tokens — the same list price as Claude Opus 4.8. Anthropic’s documentation now recommends it as the default starting point for complex agentic coding and enterprise work.
How much does Claude Opus 5 cost?
Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.8, 4.7, 4.6 and 4.5. Prompt cache reads cost $0.50 per million, a five-minute cache write costs $6.25 per million and a one-hour write costs $10 per million. The Batch API halves the base rates to $2.50 and $12.50. Fast mode, a research preview, doubles the base rate to $10 and $50.
Is Claude Opus 5 cheaper than Claude Fable 5?
Yes, exactly half. Claude Fable 5 costs $10 per million input tokens and $50 per million output tokens; Claude Opus 5 costs $5 and $25. The gap holds on cache reads as well, at $0.50 per million for Opus 5 versus $1 for Fable 5, and on the Batch API, at $2.50 and $12.50 for Opus 5 versus $5 and $25 for Fable 5.
What is Claude Opus 5’s knowledge cutoff?
Anthropic lists both a reliable knowledge cutoff and a training data cutoff of May 2026 for Claude Opus 5. That is four months more recent than Claude Fable 5, Claude Opus 4.8 and Claude Sonnet 5, which all carry a January 2026 cutoff, and more than a year ahead of Claude Haiku 4.5 at February 2025.
Why is Claude Opus 5’s knowledge cutoff more recent than Claude Fable 5’s?
Anthropic has not published a reason. The observable facts are that Claude Fable 5 became generally available on June 9, 2026 with a January 2026 cutoff, and Claude Opus 5 followed on July 24, 2026 with a May 2026 cutoff. Because pre-training data is frozen well before release, the later model can carry more recent knowledge than a model positioned above it in the lineup. Anthropic still designates Fable 5, not Opus 5, as its highest-capability tier.
Did Claude Opus 5 beat Claude Fable 5 on independent benchmarks?
On one, narrowly. Artificial Analysis recorded Claude Opus 5 at 61 on version 4.1 of its Intelligence Index on July 24, 2026, ranking it first of 190 models, against 60 for Claude Fable 5. That is a one-point gap on a composite index published without a confidence interval, so it is not a decisive margin. Two further caveats matter: the 61 was measured at max effort, while the API and Claude Code default to high effort, which scores 59; and Artificial Analysis labels the Fable 5 entry as using an Opus 4.8 fallback.
What is Claude Opus 5’s context window and maximum output?
Claude Opus 5 has a 1M token context window and a maximum of 128k output tokens on the synchronous Messages API. On the Message Batches API it supports up to 300k output tokens with the output-300k-2026-03-24 beta header. The full 1M token window is billed at standard rates, with no premium above 200k input tokens.
Should I use Claude Opus 5 or Claude Fable 5?
Anthropic’s own guidance is to start with Claude Opus 5 for complex agentic coding and enterprise work, and to move to Claude Fable 5 only for workloads that need the highest available capability. An Anthropic spokesperson told VentureBeat that the evaluations Opus 5 wins are bounded tasks with a specific outcome, and that what those evaluations do not measure is duration. The practical split is bounded work on Opus 5, and work whose horizon exceeds what a benchmark can capture on Fable 5.
What breaks when migrating from Claude Opus 4.8 to Claude Opus 5?
Three things. Requests that omit the thinking field now run with adaptive thinking instead of no thinking, so max_tokens budgets set for Opus 4.8 may be too small. Combining thinking: {"type": "disabled"} with xhigh or max effort returns a 400 error. And effort levels are recalibrated, so Anthropic advises a fresh effort sweep rather than carrying settings over. Anthropic also notes that default responses and written deliverables run longer on Opus 5.
What is fast mode on Claude Opus 5 and what does it cost?
Fast mode is a research preview that delivers up to 2.5 times higher output tokens per second on Claude Opus 5 and Claude Opus 4.8. It is enabled with speed: "fast" and the fast-mode-2026-02-01 beta header, and priced at $10 per million input tokens and $50 per million output tokens. It is available on the first-party Claude API only, not on Amazon Bedrock, Google Cloud or Microsoft Foundry, and it cannot be combined with the Batch API or a Priority Tier commitment. Access requires an account manager or the waitlist.
Where is Claude Opus 5 available?
Claude Opus 5 is available on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry. The model ID is claude-opus-5 on the Claude API, Claude Platform on AWS and Google Cloud, and anthropic.claude-opus-5 on Bedrock. On consumer plans it is the new default model on Claude Max and the strongest model available on Claude Pro.
What are Claude Opus 5’s limitations?
Anthropic’s system card records that Opus 5 hallucinates factual claims slightly more than Opus 4.8 despite being more accurate overall, and notes a surprising number of cases where it confidently stated an answer it was unsure about. On the AA-Omniscience closed-book benchmark its accuracy is 11 percent higher than Opus 4.8 but its hallucination rate is 6 percent higher. Independently, Artificial Analysis measured 52.8 output tokens per second and 62.68 seconds to first token, and described the model as very verbose. Anthropic also states it remains behind Claude Mythos 5 on biology research and offensive cybersecurity.
Sources
- Anthropic — Introducing Claude Opus 5 (announcement, July 24, 2026)
- Anthropic — Claude Opus 5 System Card (194 pages; alignment, hallucinations, FrontierBench, ARC-AGI, prompt injection)
- Anthropic — Models overview (model IDs, context window, knowledge cutoffs, effort defaults)
- Anthropic — Pricing (per-token rates, caching, Batch API, fast mode)
- Anthropic — Migrating to Claude Opus 5 (breaking changes, 512-token cache minimum, fallbacks)
- Anthropic — Effort parameter (five levels, Opus 5 guidance)
- Anthropic — Fast mode (research preview, 2.5x output speed, platform limits)
- Anthropic — Rate limits (separate Opus 5 bucket, per-tier limits)
- Artificial Analysis — Claude Opus 5 (Intelligence Index v4.1 score 61, speed, latency, verbosity; recorded July 24, 2026)
- Artificial Analysis — Claude Fable 5 (Intelligence Index v4.1 score 60)
- OpenAI — API pricing (GPT-5.6 Sol at $5 and $30 per million tokens)
- TechCrunch — Anthropic launches Opus 5 (release cadence)
- Fortune — Anthropic debuts Claude Opus 5 (effort toggle, Fable 5 burn-rate criticism)
- VentureBeat — Anthropic launches Claude Opus 5 (Anthropic spokesperson on bounded tasks versus duration)
- Simon Willison — Introducing Claude Opus 5 (launch-day developer commentary)
- Anthropic Newsroom (release timeline for Mythos 5, Fable 5, Sonnet 5 and Opus 5)



