Skip to content

Claude Opus 5 vs GPT-5.6 Luna: Where the Cheap Fast Model Is Actually Enough

Both hit 51 on AA v4.1. Luna gets there for $0.28 per task, Opus 5 for $0.36 — but 51 is Luna's ceiling and Opus 5's floor. Full effort ladders inside.

Claude Opus 5 at low effort and GPT-5.6 Luna at max effort both scoring 51 on the Artificial Analysis Intelligence Index v4.1, with cost per task measured July 28, 2026
Claude Opus 5 and GPT-5.6 Luna meet at exactly one score on the Artificial Analysis Intelligence Index v4.1 — and they arrive from opposite directions.

Feature Comparison

FeatureClaude Opus 5GPT-5.6 Luna
Peak Intelligence Index v4.161 at max effort51 at max effort
Cost per task at the shared score of 51$0.36 at low effort$0.28 at max effort
Headroom above 51Four levels, up to 61None — 51 is the ceiling
Input price per million tokens$5$0.20
Output price per million tokens$25$1.20
Cached input per million tokens$0.50$0.02
Long-prompt pricingFlat across the full 1M-token window2x input and 1.5x output above 272,000 tokens, applied to the full request
Context window1,000,000 tokens1,050,000 tokens
Time to first token at peak configuration67.72 seconds140.58 seconds
Output speed54.8 tokens per second200.9 tokens per second
Knowledge cutoffMay 2026February 16, 2026
Documented default efforthighNot published by OpenAI
AA-Omniscience Index31Not ranked on the leaderboard

Pricing Comparison

Claude Opus 5

$5 in / $25 out per M tokens
paid

GPT-5.6 Luna

$0.2 in / $1.2 out per M tokens
paid

Detailed Comparison

Claude Opus 5 scores 61 on the Artificial Analysis Intelligence Index v4.1 at max effort; GPT-5.6 Luna scores 51 at its own max. Ten points separate their ceilings, but the two models meet at exactly one number: Opus 5 also scores 51, at low effort, for $0.36 per task, while Luna's 51 costs $0.28 per task (both measured July 28, 2026, before OpenAI's July 30 price cut). Luna is therefore about 22 percent cheaper at the one score they share — but that 51 is Luna's ceiling and Opus 5's floor. Luna has nothing above it; Opus 5 has four levels and ten more points. Luna also charges double for input and 1.5 times for output on any prompt over 272,000 tokens, applied to the entire request, where Opus 5 bills its full 1M-token window at one flat rate.

The short answer

We researched both models against primary vendor documentation and the Artificial Analysis leaderboard rather than running our own evaluations. The finding that matters is narrower and stranger than a straight capability gap.

These are not competitors in the ordinary sense. Opus 5 is Anthropic's flagship, priced at $5 per million input tokens and $25 per million output tokens. Luna is the cheapest and fastest member of OpenAI's GPT-5.6 family, described by OpenAI as a "GPT-5.6 model optimized for cost-sensitive workloads," at $0.20 per million input tokens and $1.20 per million output tokens after OpenAI cut the GPT-5.6 rates by a factor of five on July 30, 2026. That cut widens the per-token gap to 25 times on input, about 20.8 times on output and 25 times on cached input. A ten-point index gap between them is the expected result, not a scandal.

What is worth writing down is the single point of contact. Both models register 51 on the Intelligence Index v4.1. For Luna that 51 is the best it can do — no configuration above max exists. For Opus 5 the same 51 is the worst it does, at low effort, with medium, high, xhigh and max stacked above it. And at that shared score, Luna is the cheaper of the two per task. The frontier model's cheapest setting does not undercut the budget model's most expensive one.

So the honest verdict is split, and we will not pretend otherwise: Opus 5 wins on capability, headroom and long-prompt cost predictability — its flat rate, not a lower bill. Luna wins the 51-point tier on price per task, decisively enough that anyone whose work genuinely tops out at 51 should not be paying for Opus 5.

What each model actually is

Claude Opus 5 launched July 24, 2026 at $5 per million input tokens and $25 per million output tokens, with a 1M-token context window and a May 2026 training cutoff. GPT-5.6 Luna is the entry tier of OpenAI's GPT-5.6 family at $0.20 per million input tokens and $1.20 per million output tokens, with a 1,050,000-token context window and a February 16, 2026 cutoff. Luna sits below Sol and Terra in its own family.

Claude Opus 5

Anthropic released Claude Opus 5 on July 24, 2026, describing it as "a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price." It kept the price of its predecessor exactly: $5 per million input tokens, $25 per million output tokens, cache hits at $0.50 per million, and Batch API rates of $2.50 and $12.50 per million. We covered the release in detail in our launch analysis.

Its training data cutoff is May 2026, which Anthropic also gives as its reliable knowledge cutoff — the most recent of any generally available Claude model, and three months later than Luna's. The context window is 1M tokens, and Anthropic's pricing page is unambiguous that no length premium applies: "Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.)" That sentence turns out to matter a great deal in this particular matchup.

Maximum output is 128k tokens on the synchronous Messages API, rising to 300k tokens on the Message Batches API behind the output-300k-2026-03-24 beta header.

GPT-5.6 Luna

GPT-5.6 Luna is the third and smallest member of the GPT-5.6 family, below Sol and Terra. OpenAI's model catalog describes Sol as a "Frontier model for complex professional work" and Luna as a "GPT-5.6 model optimized for cost-sensitive workloads." All three share a 1,050,000-token context window, a 128,000-token maximum output, and a February 16, 2026 knowledge cutoff. Our explainer on the family's naming covers how the three tiers relate.

Luna accepts text and image input and returns text only. It supports streaming, structured outputs, function calling, file search, image input, web search and prompt caching, and is exposed on the Chat Completions, Responses and Batch endpoints. Its published rates, cut by a factor of five across the board on July 30, 2026, are $0.20 per million input tokens, $0.02 per million cached input tokens, and $1.20 per million output tokens.

Those rates hold only up to a point. OpenAI's model page carries a line that reshapes the economics of the entire comparison: "Prompts with >272K input tokens are priced at 2x input and 1.5x output for the full request." We return to it below, because it is the single largest structural difference between these two models and it is almost never mentioned.

The effort ladders, side by side

Opus 5 exposes five effort levels — low, medium, high, xhigh, max — scoring 51, 56, 59, 60 and 61 on the Intelligence Index v4.1. Luna's Artificial Analysis ladder runs 33, 38, 46, 49 and 51 across the same five labels, plus a non-reasoning entry at 27. Luna's top score equals Opus 5's bottom score. The two ladders overlap at exactly one value.

This is the comparison in one table. Cost per task figures are measurements taken by Artificial Analysis and move over time; we recorded these on July 28, 2026, two days before OpenAI cut its GPT-5.6 list prices. Because those measurements predate the cut, Artificial Analysis is likely to republish a lower cost per task for Luna; Opus 5's figures are unaffected, since Anthropic's rates did not change. We report them as measured rather than recomputing an independent benchmark from list prices. Index scores are stable within a given index version, and both models here are measured on v4.1.

Effort ladders for Claude Opus 5 and GPT-5.6 Luna showing index scores at each of the five effort levels on Artificial Analysis Intelligence Index v4.1
The full effort ladders. Opus 5's floor and Luna's ceiling are the same number, 51 — and Luna's is the cheaper of the two.

Claude Opus 5 — five levels

EffortIntelligence Index v4.1Cost per task, USD (measured July 28, 2026)
max61$2.03
xhigh60$1.56
high (default)59$1.06
medium56$0.62
low51$0.36

GPT-5.6 Luna — five levels plus a non-reasoning entry

EffortIntelligence Index v4.1Cost per task, USD (measured July 28, 2026, before the July 30 price cut)
max51$0.28
xhigh49$0.18
high46$0.12
medium38$0.07
low33$0.06
Non-reasoning27$0.08

Two notes on that second table. First, "Non-reasoning" is Artificial Analysis's own label for the variant it measured, not a term OpenAI applies to Luna in its model documentation. Treat it as a third-party classification. Second, there is a genuine oddity in the measurements: the non-reasoning entry costs $0.08 per task while scoring 27, which is more per task than low at $0.06 scoring 33 and more than medium at $0.07 scoring 38. A configuration that costs more and scores less than two configurations beneath it is unusual. Cost per task is a measured quantity that depends on token consumption during the evaluation run, so this may reflect verbosity rather than list price, but we report it as measured rather than smoothing it.

What the overlap means

Read the two tables against each other and the structure is plain. Luna spends its entire ladder climbing to a number Opus 5 starts at. Every configuration Luna has below max — 49, 46, 38, 33, 27 — is below anything Opus 5 will produce at any setting. And at the meeting point, Luna is cheaper: $0.28 against $0.36, roughly 22 percent less per task. That 22 percent is a measurement, not a calculation from list prices: both figures were recorded July 28, 2026, two days before OpenAI cut Luna's rates by a factor of five. Since Opus 5's rates did not change, the gap can only have widened — treat 22 percent as the floor of Luna's price advantage at 51 rather than its current size.

That cost inversion is the part worth dwelling on, because the intuitive expectation is the opposite. One would expect a frontier model, throttled to its most economical setting, to undercut a budget model running flat out. It does not. Opus 5's dial does not reach far enough down. Anthropic's own guidance points the other way, telling developers to "use low and medium liberally as your primary control for token cost and response time" — but on this specific evaluation, low lands at a price Luna already beats.

The reverse framing is equally true and equally important: Luna buys that discount by having no reserve at all. If a workload occasionally needs more than 51, Luna has no answer. Opus 5 answers by changing one string in the request.

How the effort parameter differs in practice

Opus 5 takes effort as output_config.effort, nested rather than top-level, with five valid values and high as the API default. Anthropic states that setting high is identical to omitting the parameter. OpenAI documents a family-wide set of effort values but does not publish a default reasoning effort for GPT-5.6 Luna on either its model page or its reasoning guide.

On the Anthropic side the documentation is specific. Opus 5 "supports all five effort levels," the API default is high, and the parameter is nested inside output_config:

{
  "model": "claude-opus-5",
  "max_tokens": 4096,
  "messages": [{ "role": "user", "content": "..." }],
  "output_config": { "effort": "medium" }
}

Three details are easy to miss. Setting effort to "high" "produces exactly the same behavior as omitting the effort parameter entirely," so the default is a real level and not a separate adaptive mode. Effort is request-level, and because it reshapes the rendered prompt, changing it mid-conversation discards cached prefixes — Anthropic advises picking a level at the start of a cached session and holding it. And on Opus 5 specifically, thinking cannot be switched off at the top two levels: a request setting thinking: {"type": "disabled"} at xhigh or max returns a 400 error.

On the OpenAI side we have to be careful about what is and is not published. OpenAI's reasoning guide lists none, minimal, low, medium, high, xhigh and max as effort values, passed as "reasoning": { "effort": "..." }, and states that support is model-dependent. It gives a default for gpt-5.5 but does not state one for the GPT-5.6 family. Luna's own model page lists reasoning token support as a feature without enumerating levels or naming a default.

That gives us three distinct kinds of gap, and they should not be blurred together. Opus 5 has no minimal and no none — those levels do not exist for it, and Anthropic's five-level table says so. Luna's default effort is not published by OpenAI on the pages we checked — the setting exists, but the value is not documented. And Luna's AA-Omniscience result is a third case, covered below.

The 272,000-token cliff

OpenAI prices GPT-5.6 Luna prompts above 272,000 input tokens at $0.40 per million input tokens and $1.80 per million output tokens — double and 1.5 times its standard $0.20 and $1.20 — applied to the full request rather than the excess. Luna's window is 1,050,000 tokens, so roughly 74 percent of its advertised context sits beyond the threshold. Claude Opus 5 bills its entire 1M-token window at one flat rate with no length premium.

This is the largest structural difference between the two models, and it is the one we saw discussed least. OpenAI's wording is exact: "Prompts with >272K input tokens are priced at 2x input and 1.5x output for the full request." The phrase "for the full request" is doing the work. This is not a tiered rate where the first 272,000 tokens bill at the standard price and the excess bills higher. One token past the line and the whole prompt reprices.

The published long-context rates confirm the multipliers. Luna's standard rates are $0.20 input, $0.02 cached input and $1.20 output per million tokens. Its long-context rates are $0.40 input, $0.04 cached input and $1.80 output per million tokens — exactly double on input and cached input, exactly 1.5 times on output.

Work the arithmetic at the boundary. Every figure below is the total cost of one request with a 2,000-token response, on the rates in force after OpenAI's July 30, 2026 cut:

Prompt size (2,000-token response)GPT-5.6 LunaClaude Opus 5
272,000 input tokens$0.0568$1.41
272,001 input tokens$0.1124$1.41
1,000,000 input tokens$0.404$5.05

A single additional token takes a Luna request from $0.0568 to $0.1124 — a factor of 1.98, just under double, for one token, on the same 2,000-token response. Opus 5's figure does not move, because its rate does not change with length.

Now the honest counterweight, because the cliff is a real hazard but it is not a reversal. At 272,000 tokens Opus 5 costs about 24.8 times what Luna costs. One token later, Opus 5 costs about 12.5 times what Luna costs. The cliff halves Luna's advantage; it does not eliminate it. At a full million tokens of input, Luna is still roughly 12.5 times cheaper than Opus 5 — $0.404 against $5.05 on the same 2,000-token response. Anyone claiming the surcharge makes Opus 5 the cheaper long-context option has done the arithmetic wrong.

What the cliff does damage is predictability. A pipeline whose prompt size varies with the size of an input document — a codebase, a contract set, a log dump — will cross 272,000 tokens intermittently, and its unit economics will step by a factor of roughly two without warning at a boundary no one instrumented. Opus 5's flat window removes that class of surprise entirely. Whether that is worth paying roughly 12.5 times more for is a budgeting question, not a technical one, and the answer depends on how tightly a team needs to forecast.

Speed and latency, where the intuition breaks

At the configurations where each model performs best, Luna streams about 3.7 times faster than Opus 5 — 200.9 against 54.8 output tokens per second — but takes about 2.1 times longer to produce its first token, 140.58 seconds against 67.72 seconds. Total time to deliver 500 tokens is 143.07 seconds for Luna and 76.85 seconds for Opus 5. Luna at max effort is not a fast model. These are sliding measurements recorded July 28, 2026.

This is the finding that most contradicts the label on the tin. Luna is sold as the fast, cheap member of its family, and at low effort it certainly is. But the only configuration in which Luna reaches 51 is max, and the figures Artificial Analysis publishes for Luna correspond to that max-effort variant. At that setting, Luna spends 140.58 seconds thinking before it emits a single token — more than twice Opus 5's 67.72 seconds.

Measurement (recorded July 28, 2026)Claude Opus 5 (max)GPT-5.6 Luna (max)
Output speed, tokens per second54.8200.9
Time to first token, seconds67.72140.58
Total response time for 500 tokens, seconds76.85143.07
Blended price per 1M tokens, 7:2:1 ratio (list rates, July 30, 2026)$3.85$0.17
Output tokens consumed on the Index run100M130M

Two things follow. First, Luna's throughput advantage is real but arrives late — it streams roughly 3.7 times faster once it starts, yet still finishes a 500-token response in about 1.9 times the wall-clock time, because the reasoning phase dominates. For an interactive interface, Luna at max is the slower experience of the two despite the higher token rate.

Second, the verbosity numbers point the same way. Artificial Analysis records Luna consuming 130M output tokens over the Intelligence Index run against a median of 63M, while Opus 5 consumed 100M. The cheap model burned more tokens than the expensive one to reach a score ten points lower. Its low list price is what keeps its cost per task down, not restraint.

The practical reading: if latency is the reason to consider Luna, use it at high or below and accept a score of 46 or less. If 51 is the requirement, Luna surrenders the speed argument, and the comparison becomes a straight price question against Opus 5 at low.

Context, output and knowledge cutoff

SpecificationClaude Opus 5GPT-5.6 Luna
Context window1,000,000 tokens1,050,000 tokens
Long-prompt surchargeNone at any length2x input, 1.5x output above 272,000 tokens, applied to the full request
Max output, synchronous128k tokens128,000 tokens
Max output, batch300k tokens with the output-300k-2026-03-24 beta headerNot published as a separate limit
Training data cutoffMay 2026February 16, 2026
Input modalitiesText, imageText, image
Output modalitiesTextText

Luna's window is nominally the larger of the two by 50,000 tokens, which is a genuine if narrow advantage — provided the prompt stays under 272,000 tokens, where the extra headroom is free of the surcharge. Above that line the larger window is available but repriced.

The cutoff gap runs roughly three months in Opus 5's favor, May 2026 against February 16, 2026. Anthropic publishes its cutoff at month granularity, so the exact span is not determinable to the day. Anthropic distinguishes between a training data cutoff and a "reliable knowledge cutoff," giving May 2026 for both on Opus 5. OpenAI publishes a single cutoff date for Luna.

Factual recall and hallucination

Artificial Analysis scores Claude Opus 5 (Adaptive Reasoning, Max Effort) at 31 on the AA-Omniscience Index, a bounded metric from -100 to 100 that rewards correct answers, penalizes hallucinations and applies no penalty for declining to answer. GPT-5.6 Luna does not appear on that leaderboard at all, so no comparison between the two models on this measure is possible.

We checked the AA-Omniscience leaderboard directly for Luna and found no entry containing its name; the GPT-5.6 family is represented there by Sol variants. This is the third kind of gap described earlier, and it is worth naming precisely: Luna's AA-Omniscience result is not "zero" and not "poor." It is not measured. We cannot rank Luna against Opus 5 on factual recall, and we are not going to substitute a sibling model's score for it.

One further trap deserves flagging, because it is easy to fall into. Anthropic's own materials position Opus 5 relative to Opus 4.8 and Claude Fable 5 — the launch announcement claims Opus 5 "more than doubles Opus 4.8's performance at a lower cost per task" on Frontier-Bench v0.1, and that it comes "close to the frontier intelligence of Claude Fable 5 at half the price." None of those comparisons say anything whatsoever about Luna. A model that beats its predecessor by a wide margin has not thereby been compared to anything in OpenAI's lineup. The only common measuring stick we have for these two specific models is the Intelligence Index v4.1 and the cost and latency measurements attached to it.

Feature comparison

FeatureClaude Opus 5GPT-5.6 LunaEdge
Peak Intelligence Index v4.161 at max51 at maxOpus 5
Cost per task at the shared score of 51$0.36 at low$0.28 at maxLuna
Headroom above 51Four levels, up to 61NoneOpus 5
Input price per million tokens$5$0.20Luna
Output price per million tokens$25$1.20Luna
Cached input per million tokens$0.50$0.02Luna
Batch discount50 percent, $2.50 and $12.50Batch endpoint supported, rate not published on the model pageOpus 5
Long-prompt pricingFlat across the full 1M windowRepriced above 272,000 tokensOpus 5
Context window1,000,000 tokens1,050,000 tokensLuna
Time to first token at peak configuration67.72 seconds140.58 secondsOpus 5
Output speed54.8 tokens per second200.9 tokens per secondLuna
Knowledge cutoffMay 2026February 16, 2026Opus 5
Documented default efforthighNot publishedOpus 5
AA-Omniscience Index31Not rankedNot comparable

Claude Opus 5 — strengths and limits

Where Opus 5 is strong

  • Ten index points of headroom above Luna's absolute ceiling, reached by changing one parameter value rather than switching vendors.
  • Flat pricing across the entire 1M-token context window, with Anthropic stating explicitly that a 900k-token request bills at the same per-token rate as a 9k-token request.
  • A May 2026 knowledge cutoff, three months later than Luna's.
  • Roughly half the time to first token of Luna at each model's peak configuration, 67.72 seconds against 140.58 seconds.
  • A documented default effort level, five documented levels, and explicit guidance on when to use each.
  • 300k output tokens available on the Batch API behind a beta header, more than double Luna's published synchronous ceiling.
  • A published AA-Omniscience Index score of 31, where Luna has none.

Where Opus 5 falls short

  • 25 times Luna's input price, about 20.8 times its output price and 25 times its cached input price per token.
  • Its cheapest setting, low at $0.36 per task, is still more expensive than Luna's most expensive setting at $0.28 — the dial does not reach far enough down.
  • Output speed of 54.8 tokens per second, which Artificial Analysis characterizes as notably slow, roughly a quarter of Luna's throughput.
  • Blended price of $3.85 per million tokens against Luna's $0.17 on the same 7:2:1 ratio, computed on list rates as of July 30, 2026.
  • Thinking cannot be disabled at xhigh or max; such requests return a 400 error.
  • Changing effort mid-conversation invalidates cached prefixes, constraining how dynamically the parameter can be used in a cached session.

GPT-5.6 Luna — strengths and limits

Where Luna is strong

  • Reaches 51 on the Intelligence Index v4.1 for $0.28 per task, about 22 percent below Opus 5's cost at the same score.
  • $0.20 per million input tokens and $1.20 per million output tokens, 25 times and about 20.8 times below Opus 5's rates.
  • Cached input at $0.02 per million tokens, one twenty-fifth of Opus 5's cache rate.
  • Output speed of 200.9 tokens per second, roughly 3.7 times Opus 5's throughput.
  • A 1,050,000-token context window, 50,000 tokens larger than Opus 5's on paper.
  • A broad feature surface for its price tier: streaming, structured outputs, function calling, file search, image input, web search and prompt caching, across Chat Completions, Responses and Batch endpoints.
  • Even after the long-context surcharge applies it remains roughly 12.5 times cheaper than Opus 5, so it stays the cheaper option on million-token prompts.

Where Luna falls short

  • 51 is a hard ceiling. There is no configuration above max, so a workload that outgrows 51 requires a different model.
  • The 272,000-token threshold reprices the entire request at double input and 1.5 times output, so a single token can nearly double a bill.
  • Roughly 74 percent of its advertised context window sits beyond that threshold.
  • Time to first token of 140.58 seconds at max, more than twice Opus 5's, which negates its speed advantage in the only configuration where it ties.
  • Consumed 130M output tokens on the Intelligence Index run against Opus 5's 100M and a median of 63M — more verbose than the frontier model it ties.
  • A February 16, 2026 knowledge cutoff, roughly three months behind Opus 5's.
  • No published default reasoning effort on either the model page or the reasoning guide, leaving a production-relevant setting undocumented.
  • Absent from the AA-Omniscience leaderboard, so its factual-recall and hallucination behavior has no independent published measurement.

When to pick Claude Opus 5

Choose Claude Opus 5 when a workload sometimes needs more than 51 on the Intelligence Index, when prompts routinely exceed 272,000 tokens and cost predictability matters, when knowledge after February 2026 is relevant, or when time to first token drives the user experience. Opus 5 is the only one of the two with capability in reserve.

The work has a variable ceiling. This is the decisive case. If some fraction of requests are genuinely hard, Luna has no gear for them and Opus 5 has four. Routing by difficulty within one model — low for the bulk, xhigh for the hard tail — is operationally simpler than routing across two vendors with different parameter names, different defaults and different billing rules.

Prompts cross 272,000 tokens. Long-document analysis, whole-repository work and large log processing all live on the wrong side of Luna's threshold. Luna stays cheaper in absolute terms, but Opus 5's flat rate makes the bill a linear function of tokens instead of a step function with a cliff in the middle of the range.

Recency matters. Three months of additional training data is not decisive for most work, but for anything touching fast-moving tooling it is the difference between a model that knows a thing and one that does not.

First-token latency is the felt experience. In interactive use, 67.72 seconds against 140.58 seconds is the difference users actually perceive, and it runs opposite to what the models' price tiers suggest.

Factual recall needs an independent number. Opus 5 has an AA-Omniscience Index of 31. Luna has no published figure. If an evaluation process requires third-party evidence on hallucination behavior, only one of these models supplies it.

When to pick GPT-5.6 Luna

Choose GPT-5.6 Luna when the workload genuinely tops out at an Intelligence Index score of 51 or below, when prompts stay comfortably under 272,000 tokens, when volume makes per-token price the dominant cost, or when streaming throughput matters more than time to first token. At the score they share, Luna is about 22 percent cheaper per task on figures measured July 28, 2026, before OpenAI's July 30 price cut, and 25 times cheaper per input token on list rates.

The ceiling is known and low. Classification, extraction, summarization, routing, moderation and templated generation rarely need more than Luna delivers. If evaluations confirm quality holds at 51 — or at 46, where Luna costs $0.12 per task — there is no argument for paying Opus 5 rates.

Volume dominates. At $0.20 and $1.20 per million tokens against $5 and $25, the gap compounds. Add cached input at $0.02 per million against $0.50 and a high-volume pipeline with a stable system prompt tilts further toward Luna.

Prompts are short. Under 272,000 tokens the surcharge is irrelevant and Luna's economics are simply better, with a slightly larger window than Opus 5 as a bonus.

Throughput beats latency. For batch generation where nobody is watching a cursor blink, 200.9 tokens per second against 54.8 is a real advantage. Drop Luna to high or below and the latency penalty largely disappears too, at the cost of index points.

A sibling model is the real alternative. If Luna's ceiling is the only sticking point, the nearer comparison may not be Opus 5 at all. Terra reaches 55 at max for $0.82 per task, and we cover that trade-off in Terra versus Luna. Claude Sonnet 5 is the other obvious middle option, examined in Luna versus Sonnet 5.

The budget is the constraint. If the realistic choice is Luna or nothing, Luna at max reaching 51 is a serious model, and it reaches that number for less money than Opus 5 needs to reach the same number.

Frequently asked questions

Is GPT-5.6 Luna strictly worse than Claude Opus 5?

No. Luna's peak of 51 on the Intelligence Index v4.1 equals Opus 5's score at low effort, and Luna reaches it for $0.28 per task against Opus 5's $0.36 — about 22 percent cheaper on figures measured July 28, 2026, before OpenAI cut Luna's list prices on July 30. Opus 5 is more capable overall, with a ceiling of 61, but it does not dominate Luna at every point, because its cheapest configuration costs more than Luna's most expensive one.

Which model scores higher on the Artificial Analysis Intelligence Index?

Claude Opus 5, by ten points. Opus 5 scores 61 at max effort on Intelligence Index v4.1, the highest figure we recorded on that index version on July 28, 2026. GPT-5.6 Luna scores 51 at its own max effort. Comparing default configurations rather than peaks, Opus 5's documented default of high scores 59; OpenAI does not publish a default effort for Luna.

Does GPT-5.6 Luna charge more for long prompts?

Yes. OpenAI states that "Prompts with >272K input tokens are priced at 2x input and 1.5x output for the full request." The higher rate applies to the entire request, not only the tokens above the threshold, so a prompt of 272,001 tokens bills at double the input rate from its first token. Luna's long-context rates are $0.40 per million input tokens, $0.04 per million cached input tokens and $1.80 per million output tokens.

Does Claude Opus 5 have a long-context surcharge?

No. Anthropic's pricing documentation states that Claude 4.6 and later models "include the full 1M token context window at standard pricing" and that "a 900k-token request is billed at the same per-token rate as a 9k-token request." Prompt caching and batch discounts also apply at standard rates across the full window.

How much does each model cost per million tokens?

Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens, with cache hits at $0.50 per million and Batch API rates of $2.50 and $12.50 per million. GPT-5.6 Luna costs $0.20 per million input tokens and $1.20 per million output tokens, with cached input at $0.02 per million, after OpenAI cut its GPT-5.6 rates on July 30, 2026. Above 272,000 input tokens Luna's rates become $0.40 and $1.80 per million.

Which model responds faster?

It depends on which measurement matters. At each model's peak configuration, measured July 28, 2026, Luna streams at 200.9 output tokens per second against Opus 5's 54.8, roughly 3.7 times faster. But Luna takes 140.58 seconds to its first token against Opus 5's 67.72 seconds, and delivers 500 tokens in 143.07 seconds against Opus 5's 76.85. For interactive use Opus 5 is the faster of the two; for bulk streaming Luna is.

What effort levels does Claude Opus 5 support?

Five: low, medium, high, xhigh and max. The API default is high, and Anthropic states that setting high behaves identically to omitting the parameter. The value is passed nested as output_config.effort, not as a top-level field. There is no minimal and no none level for Opus 5.

What is GPT-5.6 Luna's default reasoning effort?

OpenAI does not publish it. Its reasoning guide lists none, minimal, low, medium, high, xhigh and max as effort values and notes that support is model-dependent, and it gives a default for gpt-5.5, but it does not state one for the GPT-5.6 family. Luna's model page lists reasoning token support without naming levels or a default. This is an undocumented value rather than an absent feature.

Which model hallucinates less?

There is no published basis for answering. Artificial Analysis scores Claude Opus 5 at 31 on the AA-Omniscience Index, a bounded metric from -100 to 100 that rewards correct answers and penalizes hallucinations while not penalizing abstention. GPT-5.6 Luna does not appear on that leaderboard, so it has no comparable published score. Anthropic's claims about Opus 5 improving on Opus 4.8 do not bear on Luna.

Which model has the larger context window?

GPT-5.6 Luna, at 1,050,000 tokens against Claude Opus 5's 1,000,000 — a difference of 50,000 tokens. The comparison flips on price rather than size: Luna's surcharge begins at 272,000 tokens, so roughly 74 percent of its window is priced above its headline rate, while all of Opus 5's window bills at one flat rate.

How recent is each model's knowledge?

Claude Opus 5 has a training data cutoff of May 2026, which Anthropic also gives as its reliable knowledge cutoff. GPT-5.6 Luna has a knowledge cutoff of February 16, 2026, shared with Sol and Terra across the GPT-5.6 family. The gap is roughly three months in Opus 5's favor, though Anthropic's month-level cutoff makes an exact span indeterminable.

Can I switch between Luna and Opus 5 to save money?

Only if the workload tolerates a hard ceiling of 51 on the portion routed to Luna, since Luna has no configuration above that score. Within Opus 5 alone, moving from max at $2.03 per task to low at $0.36 is a reduction of roughly 82 percent for a drop from 61 to 51, without changing vendors, parameter names or billing rules. Effort changes also invalidate prompt caching within a conversation, so Anthropic advises varying effort across workloads rather than inside a cached session.

Final verdict

Cost of one request at 272,000 then 272,001 input tokens with 2,000 output tokens: Claude Opus 5's total does not move while GPT-5.6 Luna's just under doubles
The 272,000-token cliff, priced on a request with 2,000 output tokens. One extra input token just under doubles Luna’s bill; Opus 5’s does not move.

Claude Opus 5 wins this comparison on capability, headroom, long-prompt cost predictability and time to first token. GPT-5.6 Luna wins decisively on price — 25 times cheaper per input token and about 20.8 times cheaper per output token on list rates since OpenAI's July 30, 2026 cut — and it also takes the one score they share, reaching 51 for $0.28 per task against Opus 5's $0.36 (both measured July 28, 2026). The deciding question is not which model is better but whether a workload ever needs more than 51 — because that is everything Luna has and the least Opus 5 gives.

We set out to test whether Opus 5 at low effort strictly dominates Luna, and it does not. The expected result was that a frontier model throttled down would undercut a budget model running flat out. The measurement says otherwise: $0.36 against $0.28, with the advantage to Luna. Opus 5's dial is excellent, but it does not reach the bottom of this market.

The asymmetry beneath that tie is what should drive the decision. Luna arrives at 51 having spent everything — its top effort setting, its latency advantage, 130M output tokens against Opus 5's 100M — and has nothing left. Opus 5 arrives at the same 51 from its floor, with medium at 56, high at 59, xhigh at 60 and max at 61 above it. Identical scores, opposite meanings.

Add the 272,000-token cliff and the picture sharpens for anyone working with large prompts. Luna remains the cheaper option on long inputs even after the surcharge, by roughly 12.5 times against 24.8 times below the threshold, so this is not a reversal. But a pricing model where one token can nearly double a bill is a forecasting problem that Opus 5's flat window simply does not have.

Our recommendation is unromantic. Evaluate the workload at 51. If quality holds, run Luna and keep the difference — the case for paying Opus 5 rates for a job that tops out at Luna's ceiling is weak, and Luna is the cheaper route to that ceiling. OpenAI's July 30, 2026 price cut widened that gap by a factor of five: on list rates Luna now costs 25 times less per input token and about 20.8 times less per output token than Opus 5. If quality does not hold, or if some fraction of requests is genuinely hard, or if prompts routinely pass 272,000 tokens, Opus 5 is not merely the better model but the only one of the two with an answer, and the ten points above 51 are available for the price of one parameter.

Sources and references

Index scores reflect Artificial Analysis Intelligence Index v4.1. Cost per task, output speed, time to first token and token consumption are measurements taken by Artificial Analysis and recorded by us on July 28, 2026; these values move over time, and Luna's cost per task predates OpenAI's July 30, 2026 price cut. Per-token pricing was verified directly against vendor documentation on July 30, 2026; other specifications were verified on July 28, 2026.

Our Verdict

Claude Opus 5 wins on capability, headroom, long-prompt cost predictability and time to first token; GPT-5.6 Luna wins decisively on price and ties the one score they share. Both reach 51 on Artificial Analysis Intelligence Index v4.1, but Luna gets there for $0.28 per task against Opus 5's $0.36 — roughly 22 percent cheaper, both measured July 28, 2026, before OpenAI's July 30 price cut — and that 51 is Luna's absolute ceiling while it is Opus 5's floor, with 56, 59, 60 and 61 stacked above it. Since OpenAI cut the GPT-5.6 rates by a factor of five on July 30, 2026, Luna's list price sits 25 times below Opus 5's on input and about 20.8 times below on output, so the price gap is far wider than the July 28 cost-per-task figures suggest. Luna also reprices an entire request at double input and 1.5 times output above 272,000 prompt tokens, where Opus 5 bills its full 1M-token window flat. Pick Luna if the workload genuinely tops out at 51 and prompts stay short; pick Opus 5 if the ceiling is variable, prompts run long, or first-token latency is the felt experience.

Winner:Claude Opus 5

Choose Claude Opus 5

Anthropic's frontier reasoning model — top of the independent index at half the price of Fable 5.

Try Claude Opus 5

Choose GPT-5.6 Luna

OpenAI's fastest, most economical GPT-5.6 tier — $0.20 per million input tokens, sub-second warm latency, and a 1.05M-token context for high-volume routine work.

Try GPT-5.6 Luna

Frequently Asked Questions

Is Claude Opus 5 better than GPT-5.6 Luna?

Claude Opus 5 wins on capability, headroom, long-prompt cost predictability and time to first token; GPT-5.6 Luna wins decisively on price and ties the one score they share. Both reach 51 on Artificial Analysis Intelligence Index v4.1, but Luna gets there for $0.28 per task against Opus 5's $0.36 — roughly 22 percent cheaper, both measured July 28, 2026, before OpenAI's July 30 price cut — and that 51 is Luna's absolute ceiling while it is Opus 5's floor, with 56, 59, 60 and 61 stacked above it. Since OpenAI cut the GPT-5.6 rates by a factor of five on July 30, 2026, Luna's list price sits 25 times below Opus 5's on input and about 20.8 times below on output, so the price gap is far wider than the July 28 cost-per-task figures suggest. Luna also reprices an entire request at double input and 1.5 times output above 272,000 prompt tokens, where Opus 5 bills its full 1M-token window flat. Pick Luna if the workload genuinely tops out at 51 and prompts stay short; pick Opus 5 if the ceiling is variable, prompts run long, or first-token latency is the felt experience.

Which is cheaper, Claude Opus 5 or GPT-5.6 Luna?

Claude Opus 5 is priced at $5 in / $25 out per M tokens. GPT-5.6 Luna is priced at $0.2 in / $1.2 out per M tokens. Check the pricing comparison section above for a full breakdown.

What are the main differences between Claude Opus 5 and GPT-5.6 Luna?

The key differences span across 13 features we compared. For Peak Intelligence Index v4.1, Claude Opus 5 offers 61 at max effort while GPT-5.6 Luna offers 51 at max effort. For Cost per task at the shared score of 51, Claude Opus 5 offers $0.36 at low effort while GPT-5.6 Luna offers $0.28 at max effort. For Headroom above 51, Claude Opus 5 offers Four levels, up to 61 while GPT-5.6 Luna offers None — 51 is the ceiling. See the full feature comparison table above for all details.

Related Comparisons