Skip to content

Grok 4.5 vs Kimi K2.7 Code: Measured Capability vs Open Weights (2026)

Grok 4.5 scores 54 to Kimi K2.7 Code's 42 on Artificial Analysis, for only 1.5x the output price. Grok takes it — but not in the EU.

Grok 4.5 versus Kimi K2.7 Code — both models are scored by Artificial Analysis on the Intelligence Index v4.1, read August 2, 2026: Grok 4.5 at 54 with a 500K-token context and closed weights, Kimi K2.7 Code at 42 with a 256K-token context and open weights under a Modified MIT license
Both models are measured on the same index. Read on August 2, 2026, Artificial Analysis puts Grok 4.5 at 54 and Kimi K2.7 Code at 42 on the Intelligence Index v4.1 — twelve points on one harness. Context and licensing are the other two structural differences. Per-token pricing and European availability are kept in the text, where they are corrected as they move. Illustration.

Feature Comparison

FeatureGrok 4.5Kimi K2.7 Code
Independent intelligence score (Artificial Analysis Intelligence Index)54 — measured by an independent evaluator, fourth overall42 on the v4.1 index, read August 2, 2026 — twelve points behind. The 54 that circulates belongs to the earlier Kimi K2.6 on an older index and does not transfer
Independent coding score (Artificial Analysis Coding Agent Index v1.3)64.4 — measured by an independent evaluator, via the Grok Build harness at high effort, read August 2, 2026None found published as of August 2, 2026 — its intelligence score exists, a third-party coding score does not
Independent SWE-bench Verified scoreNone. Not yet on the independent leaderboard (too new). No SWE-bench percentage should be attributed to this modelNone independently reproduced
Vendor self-reported coding claimNo self-reported coding benchmark published. The vendor framing ("Opus-class, much faster") is a qualitative claim, not a measurementSWE-bench Verified 60.4 percent and SWE-bench Pro 58.6 — self-reported by Moonshot AI on its own harness, not reproduced by any third party
Independent hallucination measurement (AA-Omniscience)52 percent accuracy with a 54 percent hallucination rate — a measured weakness, honestly a bad number, but it existsNo standalone figure published; AA-Omniscience feeds the index that scores it 42. Unreported rather than good
Maximum context window500,000 tokens256,000 tokens (262,144), with automatic context caching
Input price (per million tokens)USD 2.00USD 0.95 — roughly 2.1 times cheaper
Cached input price (per million tokens)USD 0.30USD 0.19 — roughly 1.6 times cheaper
Output price (per million tokens)USD 6.00USD 4.00 — only about 1.5 times cheaper, the narrowest output gap any independently scored flagship holds against this model
Model weights and licensingClosed and proprietary. API only, no weights releasedOpen weights under a Modified MIT license, published on day one
European Union availabilityBlocked at launch under the EU AI Act systemic-risk rules for general-purpose models. SpaceXAI signalled a staged EU opening expected around mid-July 2026 — verify for your region before building, this changes fastAvailable. Open weights cannot be geo-blocked, so it can be self-hosted inside the EU today
Self-hosting and data residencyNot possible — managed API onlyPossible — run the weights on your own hardware, in your own jurisdiction
Architectural transparencyNot disclosed by SpaceXAIFully published: mixture-of-experts, roughly 1 trillion total parameters with 32 billion active per token, 384 experts, multi-head latent attention, MoonViT vision encoder
Hosted metered APIAvailable from SpaceXAIAvailable from Moonshot AI

Pricing Comparison

Grok 4.5

$2 in / $6 out per M tokens
paid

Kimi K2.7 Code

$0.95 in / $4 out per M tokens
Free plan available
Free trial available
freemium

Detailed Comparison

Grok 4.5 vs Kimi K2.7 Code in 2026: Grok 4.5 is the flagship model from SpaceXAI (formerly xAI), priced at USD 2 per million input tokens, USD 0.30 cached, and USD 6 per million output tokens, with a 500,000-token context window. It scores 54 on the independent Artificial Analysis Intelligence Index and 64.4 on the independent Artificial Analysis Coding Agent Index v1.3, through the Grok Build harness at high effort. Kimi K2.7 Code is Moonshot AI's open-weight agentic coding model, priced at USD 0.95 per million input tokens, USD 0.19 cached, and USD 4 per million output tokens, with a 256,000-token context window and downloadable weights under a Modified MIT license. Kimi K2.7 Code is independently scored too: 42 on the Artificial Analysis Intelligence Index v4.1, read on August 2, 2026, twelve points behind Grok 4.5 on the same harness. That index is a generalist composite of nine evaluations, most of them reasoning and knowledge tests rather than coding, so part of the gap reflects scope rather than quality; Moonshot's own SWE-bench figures remain self-reported. Grok 4.5 costs roughly 2.1 times more on input but only about 1.5 times more on output — the narrowest premium any independently scored flagship charges over K2.7. Grok 4.5 takes the verdict wherever it is available. It was blocked in the European Union for its first nine days, and SpaceXAI opened EU access on July 17, 2026, so that no longer decides the matchup.

Quick Verdict

We ran both models side-by-side over a working week on the same coding briefs, the same agentic tool-use loops, and the same long-context retrieval tasks. Two things came out of it that make this the most interesting pairing we have written this year.

First: this is the tightest price duel we have seen between a closed flagship and an open-weight challenger. Input runs USD 2 against USD 0.95, roughly 2.1 times. But output — the line that actually dominates a coding bill — runs USD 6 against USD 4. That is about 1.5 times. Not ten times. Not six times. One and a half. A closed model with independently charted scores landing within touching distance of a Chinese open-weight model on the most expensive line item is, on its own, the story here.

Second: the two models are exact mirror images of each other, and the mirror runs in opposite directions. Grok 4.5 has the higher score but not the access. It is measured at 54 on the Artificial Analysis Intelligence Index v4.1. Kimi K2.7 Code has the access but not the lead. You can download its weights and run it in any jurisdiction on earth, and the same evaluator puts it at 42 on that index, read on August 2, 2026. Both models are independently measured; one is twelve points ahead and unavailable to part of the world, the other is behind and available everywhere. You are not choosing between a measured model and an unmeasured one. You are choosing what twelve points on a generalist index are worth to you.

  • Best independently measured capability: Grok 4.5 — 54 on the Artificial Analysis Intelligence Index and 64.4 on the Artificial Analysis Coding Agent Index v1.3 through the Grok Build harness at high effort, both produced by an outside evaluator. Kimi K2.7 Code is scored by the same evaluator at 42 on that index, read August 2, 2026 — twelve points back.
  • Best price on every single line: Kimi K2.7 Code — cheaper on input, cheaper on cached input, cheaper on output. The gap is real but unusually small.
  • Best context window: Grok 4.5 — 500,000 tokens against 256,000, close to double.
  • Best availability and control: Kimi K2.7 Code — open weights under a Modified MIT license, self-hostable in any jurisdiction, including every country where Grok 4.5 currently cannot be used.
  • Most honest about its own weaknesses: Grok 4.5, by accident of being measured — its 54 percent hallucination rate on the independent AA-Omniscience evaluation is a bad number, but at least it exists. No standalone hallucination figure is published for Kimi K2.7 Code, though AA-Omniscience is one of the nine evaluations feeding the index that scores it 42.
  • Overall: Grok 4.5, where you can use it. Twelve index points ahead on the same harness, for about 1.5 times the output price, is the narrowest capability premium we have priced this year. If you must self-host, or you are running hundreds of millions of output tokens a month, Kimi K2.7 Code is the answer instead.

Grok 4.5 vs Kimi K2.7 Code at a Glance

Every figure below comes from the vendors' own documentation for pricing and specifications, and from Artificial Analysis for the independent scores. Where a number is self-reported by the vendor, we label it as such and never present it as verified.

AttributeGrok 4.5Kimi K2.7 Code
VendorSpaceXAI (formerly xAI), United StatesMoonshot AI, Beijing, China
ReleasedPublic July 9, 2026June 12, 2026
Model typeClosed frontier, API onlyOpen-weight mixture-of-experts
ArchitectureNot disclosedRoughly 1 trillion total parameters, 32 billion active per token, 384 experts, multi-head latent attention, MoonViT vision encoder
LicenseProprietary, commercial API termsModified MIT, weights published on day one
Context window500,000 tokens256,000 tokens (262,144), automatic caching
Input priceUSD 2 per million tokensUSD 0.95 per million tokens
Cached input priceUSD 0.30 per million tokensUSD 0.19 per million tokens
Output priceUSD 6 per million tokensUSD 4 per million tokens
Independent intelligence score54 on the Artificial Analysis Intelligence Index v4.1 (Grok 4.5, high)42 on the Artificial Analysis Intelligence Index v4.1 (Kimi K2.7 Code), read August 2, 2026
Independent coding score64.4 on the Artificial Analysis Coding Agent Index v1.3, through the Grok Build harness at high effort, read August 2, 2026None found published by that evaluator as of August 2, 2026
Vendor self-reported coding claimNone published; the vendor framing is qualitativeSWE-bench Verified 60.4 percent, SWE-bench Pro 58.6 — self-reported by Moonshot AI
Independent hallucination measurementAA-Omniscience: 52 percent accuracy, 54 percent hallucination rateNo standalone figure published; AA-Omniscience is one of the nine evaluations inside its Intelligence Index score of 42
Self-hostingNoYes — download the weights
European Union availabilityBlocked at launch under the EU AI Act; a staged opening was signalled for around mid-July 2026 — verify before you buildAvailable, and self-hostable inside the EU

The Three Different 54s on This Page

Before anything else, a warning — because we have already seen aggregator pages get this wrong, and getting it wrong flips the entire conclusion.

Three separate numbers in this comparison are 54, and they mean three completely different things.

  • Grok 4.5 scores 54 on the independent Artificial Analysis Intelligence Index. That is a capability score, produced by an outside evaluator. Higher is better. This is a good number.
  • Grok 4.5 also posts a 54 percent hallucination rate on the independent AA-Omniscience evaluation (alongside 52 percent accuracy). That is a failure rate on an entirely different axis. Lower is better. This is a bad number, and it happens to collide with the good one.
  • Kimi K2.6 — the previous generation, a different model — is widely quoted at 54, but that figure comes from an earlier version of the Artificial Analysis index. On the current v4.1 index, the same one that scores Grok 4.5 at 54, Kimi K2.6 sits at 44. The two are not level, and the number you see repeated is stale. And it does not travel forward: Kimi K2.7 Code carries its own score on the current v4.1 index, and that score is 42, read on August 2, 2026 — not 54. If you see 54 attached to Kimi K2.7 Code anywhere on the web, it has been inherited by mistake from an older index and an older model.

So: the intelligence score of 54 on this page belongs to Grok 4.5. The hallucination rate of 54 percent also belongs to Grok 4.5, and is not a compliment. Kimi K2.7 Code's score on that same index is 42. We will not pretend otherwise in either direction.

Grok 4.5 Overview

Grok 4.5 is the current flagship from SpaceXAI, the company formerly known as xAI — the rebrand landed on July 6, 2026, and the model line-up did not change with it (we covered what the SpaceXAI rebrand actually changed separately). Grok 4.5 was announced on July 8 and went public on July 9, 2026, replacing Grok 4.3 as the flagship while the older models stay available.

The specifications are confirmed from the vendor's own documentation: a 500,000-token context window, text and image input to text output, function calling, structured outputs, and reasoning effort levels from low through high. Pricing is USD 2 per million input tokens, USD 0.30 per million cached input tokens, and USD 6 per million output tokens. Rate limits run to 150 requests per second.

What makes Grok 4.5 unusual is that its capability claims are not, in the main, its own. Artificial Analysis published independent results on July 8, 2026: an Intelligence Index of 54, placing it fourth overall, and — on the separate Coding Agent Index, version 1.3 when we read it on August 2, 2026 — 64.4 through the Grok Build harness at high effort. The evaluator also measured a cost of roughly USD 2.59 per task, which is a fraction of what the top-of-market models charge to complete the same work. Those are third-party numbers, not marketing.

Two caveats belong right here rather than buried at the bottom. First, Grok 4.5 does not have an independent SWE-bench Verified score — it is simply not on the leaderboard yet, being too new. Elon Musk's public framing of the model as "Opus-class, much faster" is a vendor claim, not a measurement, and we treat it as such. Second, the same Artificial Analysis run that produced the flattering Intelligence Index also produced an unflattering one: on AA-Omniscience, Grok 4.5 answers with 52 percent accuracy and hallucinates at a 54 percent rate. For a model you intend to point at a codebase unsupervised, that is a number to take seriously.

Kimi K2.7 Code Overview

Kimi K2.7 Code is Moonshot AI's open-weight agentic coding model, released on June 12, 2026. Architecturally it is fully public, which is refreshing after a year of frontier labs disclosing nothing: a mixture-of-experts design with roughly 1 trillion total parameters, only 32 billion of them active per token, 384 experts of which 8 are selected plus 1 shared, multi-head latent attention, and a MoonViT vision encoder so it can read screenshots and UI mockups inside a coding workflow.

The weights shipped on day one under a Modified MIT license. You can download them, run them on your own hardware, fine-tune them, and keep every token inside your own infrastructure. Metered API pricing from Moonshot is USD 0.95 per million input tokens, USD 0.19 per million cached input tokens, and USD 4 per million output tokens, with automatic context caching. The context window is 256,000 tokens (262,144 exactly).

And then the correction. Artificial Analysis does score this model: Kimi K2.7 Code sits at 42 on the Intelligence Index v4.1, read on August 2, 2026, alongside a published cost per task of USD 0.22. An earlier version of this page said no third-party result existed at all, and that was wrong — the score was already published when we wrote it. What is genuinely missing is narrower: we have found no independently reproduced SWE-bench Verified figure, no LMArena Elo rating, and no standalone Terminal-Bench, LiveCodeBench, GPQA or AIME result. The vendor's own coding figures remain exactly that, Moonshot AI's own: SWE-bench Verified at 60.4 percent and SWE-bench Pro at 58.6, self-reported on the vendor's own harness and described by Moonshot as a new high-water mark for open-source models.

We want to be precise about what that means, because it is easy to be unfair in either direction. It does not mean K2.7 is a weak model. Our hands-on runs suggest it is a strong one. It means Kimi K2.7 Code is measured, and behind, on a generalist index — which is a different claim, and a narrower one. That index weights reasoning and knowledge heavily, and this is a coding-specialized model, so a 42 understates the ground it was actually built for. Where it is genuinely untested is on independently reproduced coding benchmarks, and that gap can close at any time.

Why We Will Not Put These Benchmark Numbers Head-to-Head

This is the most important methodological point on the page, so we are stating it loudly rather than tucking it into a footnote.

Grok 4.5's charted coding figure is a Coding Agent Index v1.3 result of 64.4, produced by Artificial Analysis — an outside evaluator running Grok 4.5 inside the Grok Build harness at high effort. Kimi K2.7 Code's coding figures are SWE-bench Verified 60.4 percent and SWE-bench Pro 58.6, produced by Moonshot AI, running Moonshot AI's harness, on Moonshot AI's infrastructure.

Those two things differ on two axes at once, not one. They are different benchmarks — an agentic index and a SWE-bench suite are not measuring the same thing, and the scales are unrelated. And they come from different evidence regimes — one was produced by a party with nothing to gain, the other by the party selling the model. Putting those two figures side by side in a table would imply both a comparison and a verdict, and neither would be real. So we do not do it: not in a row, not in a sentence, and not in the infographic below, which deliberately carries no benchmark row at all.

What we can say is narrower, and true. On the one index that scores both, Grok 4.5 is at 54 and Kimi K2.7 Code at 42 — a real twelve-point gap, measured by the same evaluator on the same harness, read on August 2, 2026. That index is a generalist composite of nine evaluations, most of them reasoning and knowledge rather than coding, so on a coding-specialized model it measures breadth more than it measures the job. The gap is real, it is not the whole picture, and we score it accordingly.

Which model to pick between Grok 4.5 and Kimi K2.7 Code — choose Grok 4.5 if you want the higher index score, need the larger context window, or are happy on a closed API; choose Kimi K2.7 Code if you run high-volume coding loops, must host the model yourself, or want a coding specialist
The comparison read as a decision. Grok 4.5 is the answer when the higher index score and the larger context are what you are buying, and a closed API is no obstacle. Kimi K2.7 Code is the answer for high-volume coding loops, for anything that has to run on your own hardware, and when you would rather have a coding specialist than a generalist. Per-token rates and regional availability are kept in the text, where they are corrected as they move. Illustration.

Pricing: The Narrowest Gap of the Year

Both models bill per token, with input, cached input, and output metered separately. We pulled each rate directly from the vendor's own pricing documentation rather than from third-party aggregators, which drift.

Rate (per million tokens)Grok 4.5Kimi K2.7 CodeMultiple
InputUSD 2.00USD 0.95Grok is about 2.1 times more
Cached inputUSD 0.30USD 0.19Grok is about 1.6 times more
OutputUSD 6.00USD 4.00Grok is about 1.5 times more

Read the output row twice, because output tokens are what a coding agent actually burns. USD 6 against USD 4 is a premium of about 50 percent. To put that in context: in the comparisons we have run against the frontier tier, the same K2.7 comes out six to twelve times cheaper than the closed flagship it faces. Against Grok 4.5 it is one and a half times cheaper. The independently scored, closed, American flagship is charging a 50 percent surcharge over a Chinese open-weight model on the most expensive line on the invoice.

That is the fact that reframes the whole decision. In most closed-versus-open matchups, the question is whether verified capability is worth paying a multiple for. Here, the question is whether verified capability is worth paying a rounding error for — and once the premium gets that small, the answer for most teams flips.

Artificial Analysis publishes a cost per completed task for both models. Read on August 2, 2026, it lists USD 0.44 for Grok 4.5 (high) and USD 0.22 for Kimi K2.7 Code — so on the evaluator's own workload, Kimi finishes the job for about half the money, and that cuts against our verdict rather than for it. These figures move with token prices and with the index version behind them, so treat them as a dated reading rather than a constant. And a cheaper rate card does not automatically produce a cheaper bill, because a model that needs a second pass on a hard task burns its cheap tokens twice.

Real-World Cost Scenarios

Rate cards are abstract, so here is what the difference looks like on an actual monthly invoice. These are illustrative estimates using each vendor's published standard rates, with no cache hits assumed. Your real bill depends on caching, retries, and how often each model gets it right on the first pass.

Workload (monthly)Grok 4.5Kimi K2.7 CodeDifference
Solo developer: 5M input, 3M outputAbout USD 28About USD 17About USD 11
Small team agent: 30M input, 20M outputAbout USD 180About USD 109About USD 71
High-volume CI agent: 100M input, 80M outputAbout USD 680About USD 415About USD 265

Look at the solo-developer row. Eleven dollars a month is the entire cost of the twelve-point gap between the two on the Intelligence Index v4.1. For an individual developer or a small team, that difference is noise — it is less than a single hour of the time you would lose to a bad refactor. This is why the verdict tilts the way it does.

Now look at the bottom row. At high volume the gap becomes a real line item: about USD 265 a month, roughly USD 3,200 a year, and it scales linearly from there. And that is the metered comparison only. If you self-host K2.7 on your own GPUs, the per-token bill disappears entirely and is replaced by infrastructure cost, which at sufficient scale is the cheaper curve. There is a volume threshold above which Kimi wins on economics no matter how good Grok is, and teams running agents around the clock are above it.

Availability: The Mirror Runs Backwards

This is the section most comparisons would skip, and it is the one that will actually decide the question for a large share of readers.

Grok 4.5 was blocked in the European Union at launch, and that block was lifted on July 17, 2026. The reason is regulatory, not technical: under the EU AI Act, general-purpose models judged to pose systemic risk face obligations that SpaceXAI did not meet at ship time. So on July 9, 2026, a model with independently verified frontier-class coding scores became unavailable to every developer inside the bloc.

This is not a permanent ban, and we want to be exact about the temporality here, because it is the fastest-moving fact on this page. SpaceXAI signalled a staged European rollout expected around mid-July 2026, and it landed: the xAI release notes recorded on July 17, 2026 that "Grok 4.5 is now available in the API console for EU users." We are publishing in mid-July 2026. That means the situation described in this paragraph may already have changed by the time you read it, and it may change again. That rollout has now completed for the API console, so EU availability is no longer the open question it was at launch. If the rollout completes as signalled, the single hardest objection to Grok 4.5 in this comparison evaporates. If it stalls, that objection hardens into a wall.

Kimi K2.7 Code has the opposite property, and it is structural rather than granted. Because the weights are published under a Modified MIT license, the model cannot be geo-blocked in any meaningful sense. You can download it in Berlin, run it on GPUs in Frankfurt, and keep every token of customer data inside the European Economic Area. No vendor decision, no regulatory negotiation, and no rollout schedule sits between you and the model. For a team with data-residency obligations, or one that simply refuses to build on infrastructure that a policy dispute can switch off, that is not a feature — it is the whole argument.

So the chiasm completes: Grok 4.5 has the evidence and, for part of the world, not the access. K2.7 has the access everywhere and, for now, none of the evidence. Which gap you can tolerate is a question about your organization, not about the models.

Hands-On: We Ran Both Side-by-Side

We tested both models over a working week from outside the EU, on the same four task families: a multi-file refactor of a TypeScript service, an agentic loop that had to read a failing test, edit code, and re-run the suite until green, a long-context task that fed a large codebase and asked for a cross-file change, and a factual-recall task designed to probe how each model behaves when it does not know something.

Coding accuracy. Both models were genuinely strong, and the gap was narrower than the evidence asymmetry might lead you to expect. Grok 4.5 was quick and confident, and on the multi-file refactor it produced correct edits on the first pass more often than not. Kimi K2.7 Code was close behind and occasionally better on the tasks where its agentic tuning showed. If you handed us anonymized outputs, we would not reliably be able to tell you which model produced which.

Agentic tool use. Grok 4.5 was fast — noticeably so, and that speed compounds in a read-edit-rerun loop where every iteration costs wall-clock time. K2.7 was more deliberate and, on the longer loops, slightly more methodical about verifying its own work. Neither was clearly dominant. Kimi's self-reported strength is agentic tool use, and while we cannot confirm the vendor's numbers, our runs are consistent with the claim.

Long context. Grok 4.5's 500,000-token window let us drop substantially more of a codebase into a single prompt than Kimi's 256,000. On whole-repository tasks that is a concrete, structural advantage, not a matter of taste. If your prompts routinely exceed 256,000 tokens, this comparison is already decided.

Factual reliability. This is where the independent AA-Omniscience number stopped being abstract. Grok 4.5, asked about things at the edge of its knowledge, tended to answer rather than abstain — which is exactly the behavior a 54 percent hallucination rate describes. In a coding context this shows up as inventing an API surface that does not exist and stating it with total composure. K2.7 was, in our limited runs, more inclined to hedge. But we have to flag our own limits here: this is an impression from a week of use, not a measurement, and no standalone hallucination figure is published for Kimi K2.7 Code — AA-Omniscience feeds the index that scores it 42, but the evaluator does not break that component out. We are comparing a published number against a personal hunch, and we are not going to pretend that is a fair fight.

The pattern across the week: two models of very similar practical ability, separated by a modest price difference, a real context-window difference, and an enormous difference in how much anybody outside the vendor actually knows about them.

Ecosystem, Integration, and Deployment

Grok 4.5 lives inside SpaceXAI's managed platform: a first-party API with function calling, structured outputs, and reasoning-effort controls, served from US regions with generous rate limits. Nothing to provision, nothing to operate. The trade-off is total dependence on the vendor's platform decisions — including, as the nine-day EU gap demonstrated, whether the model is available to you at all.

Kimi K2.7 Code takes the open route. Its hosted API is OpenAI-compatible, which in practice means it drops into any agent framework, IDE plugin, or routing layer that already speaks that format, usually with a one-line base-URL change. Because the weights are public, it also runs inside self-hosted inference stacks and on GPU clouds, and a growing set of third-party providers serve it through their own gateways. The cost of that freedom is operational: if you self-host, you own the GPUs, the scaling, and the upgrades.

For a team that wants a managed endpoint and nothing else to think about, Grok 4.5 is the lower-friction option. For a team that has already standardized on the OpenAI API shape, or that needs to run inference inside its own perimeter, K2.7 is the drop-in. We line both up against the wider field in our roundup of the best AI coding tools of 2026.

Winner Per Category

CategoryWinnerWhy
Independently verified capabilityGrok 4.554 on the Artificial Analysis Intelligence Index v4.1 and 64.4 on the Artificial Analysis Coding Agent Index v1.3 through the Grok Build harness at high effort. Kimi K2.7 Code is measured at 42 on the Intelligence Index, read August 2, 2026 — twelve points back.
Context windowGrok 4.5500,000 tokens against 256,000 — close to double.
Speed in agentic loopsGrok 4.5Noticeably quicker per iteration in our read-edit-rerun runs.
Price on every lineKimi K2.7 CodeCheaper on input, cached input, and output — though output is only about 1.5 times apart.
Open weights and licensingKimi K2.7 CodeModified MIT, published day one. Grok 4.5 releases nothing.
Availability and jurisdictionKimi K2.7 CodeRuns anywhere, and can be self-hosted inside the EU, which Grok 4.5 cannot.
Self-hosting and data residencyKimi K2.7 CodeKeep every token inside your own perimeter. Grok 4.5 is API only.
Architectural transparencyKimi K2.7 CodeFull architecture published. SpaceXAI discloses nothing about Grok 4.5's internals.
Factual reliabilityNeither, honestlyGrok 4.5 has a measured 54 percent hallucination rate on AA-Omniscience — a bad number that at least exists. No standalone hallucination figure is published for Kimi K2.7 Code, though AA-Omniscience feeds the index that scores it 42. A quantified flaw against an unreported one.
Cost at very high volumeKimi K2.7 CodeSelf-hosting removes the per-token bill entirely above a certain scale.

Count the rows and K2.7 wins more of them. That is not an accident, and we are not going to hide it: on everything you can count, Kimi wins. Grok 4.5 wins the one thing you cannot buy, which is knowing what you are getting — and it does so at a premium small enough that most teams will pay it without noticing. Verdicts are weighed, not tallied.

Pros and Cons

Grok 4.5

Pros

  • Independently measured by an outside evaluator: 54 on the Artificial Analysis Intelligence Index, fourth overall, and 64.4 on the Artificial Analysis Coding Agent Index v1.3 through the Grok Build harness at high effort.
  • Independently measured cost of USD 0.44 per completed task, read on August 2, 2026 — a small fraction of what top-of-market models charge for the same work, though Kimi K2.7 Code is measured at half that again.
  • A 500,000-token context window, close to double K2.7's, which matters for whole-repository prompts.
  • Fast in practice, and that speed compounds across the iterations of an agentic loop.
  • Priced at USD 2 per million input and USD 6 per million output — remarkably low for a model with third-party frontier-class scores.
  • Fully managed: no GPUs to provision, no inference stack to run, generous rate limits.

Cons

  • Was blocked in the European Union at launch under the EU AI Act, until July 17, 2026. The staged opening signalled for around mid-July 2026 landed on July 17, 2026, and EU teams can now use the model.
  • A 54 percent hallucination rate on the independent AA-Omniscience evaluation, with 52 percent accuracy — a genuine reliability concern for unsupervised agents.
  • No independent SWE-bench Verified score. It is too new to be on that leaderboard, so the most widely cited coding benchmark simply has no entry for it.
  • Closed and API only — no weights, no self-hosting, no data residency control, and no recourse if platform availability changes again.
  • Architecture entirely undisclosed. The vendor's "Opus-class" framing is a claim, not a measurement.

Kimi K2.7 Code

Pros

  • Cheaper on every single line: USD 0.95 per million input, USD 0.19 cached, USD 4 per million output.
  • Open weights under a Modified MIT license, published on day one — download, self-host, and fine-tune today.
  • Available in every jurisdiction and self-hostable inside the European Union, which a closed API model cannot offer.
  • Fully published architecture: mixture-of-experts, roughly 1 trillion total parameters with 32 billion active, 384 experts, multi-head latent attention.
  • Includes a MoonViT vision encoder, so it reads screenshots and UI mockups inside coding workflows.
  • OpenAI-compatible API, so it drops into existing agent frameworks with a base-URL change.
  • Automatic context caching, which lowers the effective bill without any work on your side.

Cons

  • Behind on the one independent index that scores both models: 42 against Grok 4.5's 54 on the Artificial Analysis Intelligence Index v4.1, read August 2, 2026 — though that index is a generalist composite and this is a coding-specialized model.
  • No independently reproduced SWE-bench Verified result and no LMArena rating that we have been able to find, so its coding claim specifically still rests on the vendor's own harness.
  • Its coding figures, SWE-bench Verified 60.4 percent and SWE-bench Pro 58.6, are self-reported by Moonshot AI on its own harness and have not been reproduced by anyone.
  • No standalone hallucination figure published, so its factual reliability is unreported rather than demonstrated.
  • A 256,000-token context window, roughly half Grok 4.5's, which is a hard ceiling on whole-repository prompts.
  • Self-hosting shifts real operational cost onto you: GPUs, scaling, and upgrades all become your problem.
  • Modified MIT is not plain MIT — read the license if you deploy at hyperscale.

When to Pick Each Model

When to pick Grok 4.5

  • You want a closed frontier API rather than weights you host yourself.
  • You want capability that somebody other than the vendor has actually measured, and you want it at the smallest premium currently on offer.
  • Your prompts regularly exceed 256,000 tokens and you need the 500,000-token window.
  • Iteration speed matters to you: you are running tight agentic loops where wall-clock time per pass compounds.
  • You want a fully managed endpoint and have no interest in operating inference infrastructure.
  • Your workload is code, not open-domain factual recall — the hallucination rate hurts most where the model is asked to know things.

When to pick Kimi K2.7 Code

  • You must keep every token inside your own perimeter. This is not a preference, it is arithmetic: Grok 4.5 is not available to you, and K2.7 is.
  • You need open weights to self-host, fine-tune, or keep customer data inside your own perimeter.
  • Data residency, air-gapped deployment, or sovereignty requirements govern your architecture.
  • You run hundreds of millions of output tokens a month, where the price gap stops being noise and self-hosting removes the bill entirely.
  • You refuse to build on a platform that a regulatory dispute can switch off, and open weights are the only real insurance against that.
  • You are able to evaluate the model yourself on your own tasks, which matters more than a twelve-point gap on a generalist index.
The verdict for Grok 4.5 against Kimi K2.7 Code — both are scored by the same third-party evaluator, Grok 4.5 at 54 and Kimi K2.7 Code at 42 on the Artificial Analysis Intelligence Index v4.1 read August 2, 2026, a generalist composite rather than a coding suite; Grok 4.5 adds a larger context behind closed weights, Kimi K2.7 Code answers with open, self-hostable weights
Grok 4.5 takes the verdict on an exchange rate rather than a knockout. Both models are scored by the same evaluator on the same harness — 54 against 42 on the Artificial Analysis Intelligence Index v4.1, read on August 2, 2026 — but that index is a generalist composite, not a coding suite, and Kimi K2.7 Code is a coding-specialized model being measured off its own ground. Illustration.

What Would Change Our Verdict

Both models are weeks old. This page is a snapshot of a market that moves faster than we can publish, and three things would move it.

An independent coding result landing for Kimi K2.7 Code. The generalist index already scores it, at 42. What does not exist yet is a third-party coding measurement, and that is the number this matchup actually turns on. If an outside evaluator reproduces something near Moonshot AI's self-reported SWE-bench figures, then an open-weight model that is cheaper on every line, cheaper per completed task, and available in every jurisdiction becomes very hard to argue against — and our verdict flips. If it lands well below, the current verdict hardens instead.

The EU rollout, in either direction. Grok 4.5 opened in the European Union on July 17, 2026, as signalled, so the strongest objection to it has disappeared and the recommendation is simpler for a large slice of our readers. Were that access ever withdrawn, Grok 4.5 would become structurally unavailable to the EU market and the verdict there would not be close — it would be K2.7, without argument.

Any pricing move. A 50 percent output premium is what makes the proof cheap enough to buy. If SpaceXAI raises Grok 4.5's rates, or Moonshot cuts Kimi's further, that calculus changes quickly. This market discounts aggressively and without notice.

What would not change our verdict: another vendor-run benchmark from either side. We have enough self-reported numbers. What this comparison needs is someone independent to run K2.7, and someone independent to run Grok 4.5 on SWE-bench.

Final Verdict

Grok 4.5 wins this comparison, and it wins it on an exchange rate rather than a knockout. Both models are scored by the same independent evaluator, and Grok 4.5 is ahead: 54 against Kimi K2.7 Code's 42 on the Artificial Analysis Intelligence Index v4.1, read on August 2, 2026, plus 64.4 on the independent Artificial Analysis Coding Agent Index v1.3, through the Grok Build harness at high effort, where we have found no third-party equivalent published for Kimi. The twelve-point gap is real and it is measured on the same harness. It is also measured on a generalist composite of nine evaluations, most of them reasoning and knowledge rather than code — and Kimi K2.7 Code is a coding-specialized model, so the index is not scoring it on the ground it was built for. Grok 4.5 is ahead. It is not ahead by as much as an earlier version of this page implied.

What settles it is the price of that proof, and it has never been lower. Grok 4.5 costs about 2.1 times more on input and only about 1.5 times more on output — USD 6 against USD 4 on the line that dominates a real coding bill. For a solo developer that is roughly eleven dollars a month. Eleven dollars, for twelve points on the index that scores both models, plus close to double the context window. That is not a premium; that is a rounding error, and we would pay it.

But the win is scoped, and the scope is not a detail. Grok 4.5 has been available in the European Union since July 17, 2026. For an EU team this analysis now applies in full, where for nine days after launch it did not. That block was signalled to lift around mid-July 2026, which is now — so verify your region before you act on any of this, because it is the fastest-moving fact on the page. And Grok 4.5 carries a measured 54 percent hallucination rate on AA-Omniscience, which is a real reason to keep a human in the loop on anything that requires the model to know rather than to build.

K2.7 remains the right call, without hesitation, if any of the following is true: you must self-host, your data cannot leave your perimeter, you are burning hundreds of millions of output tokens a month, or you can properly evaluate the model yourself and therefore do not need anyone else to have done it for you. It is cheaper on every line, open on every axis, and available everywhere. Those are not consolation prizes.

Everyone else: take Grok 4.5, and revisit this page the moment an independent evaluator publishes a coding result for Kimi K2.7 Code. We will.

Frequently Asked Questions

Which is better, Grok 4.5 or Kimi K2.7 Code?

Grok 4.5, for most teams that can access it. Both models are scored by the same independent evaluator, and Grok 4.5 is ahead: 54 against 42 on the Artificial Analysis Intelligence Index v4.1, read on August 2, 2026, plus 64.4 on the Artificial Analysis Coding Agent Index v1.3 through the Grok Build harness at high effort. It also offers a 500,000-token context window against Kimi's 256,000, and charges only about 1.5 times more on output for it. Read the twelve-point gap with one caveat: that index is a generalist composite of mostly non-coding evaluations, and Kimi K2.7 Code is a coding-specialized model. But the win is scoped: if you must self-host or keep data in your own perimeter, K2.7 is the correct choice regardless of scores.

Is Grok 4.5 available in the EU?

Not at launch. Yes, since July 17, 2026. Grok 4.5 went public on July 9, 2026 and was blocked in the European Union at the time, a gap reported as following from the EU AI Act's obligations for general-purpose models judged to carry systemic risk. It was not a permanent ban: the staged European opening landed on July 17, 2026, when the xAI release notes recorded that "Grok 4.5 is now available in the API console for EU users." This is the fastest-moving fact on this page. Check whether Grok 4.5 is reachable from your region before you build anything on it, and re-check if your last look was more than a couple of weeks ago. Kimi K2.7 Code, by contrast, cannot be geo-blocked at all: its weights are downloadable, so it can be self-hosted inside the EU today.

Does Kimi K2.7 Code have an Artificial Analysis Intelligence Index score of 54?

No — it has an Intelligence Index score, but the number is 42, not 54. Read on August 2, 2026, Artificial Analysis scores Kimi K2.7 Code at 42 on the current v4.1 index. The 54 that circulates comes from an earlier version of the index, where it was attached to Kimi K2.6, the previous generation. Scores do not transfer between model versions, and they do not transfer between index versions either. Separately, Grok 4.5 scores 54 on the current index, which makes the confusion easy to fall into. An earlier version of this page said K2.7 had no Intelligence Index at all; that was wrong, and we have corrected it.

How much cheaper is Kimi K2.7 Code than Grok 4.5?

Kimi K2.7 Code costs USD 0.95 per million input tokens, USD 0.19 cached, and USD 4 per million output tokens. Grok 4.5 costs USD 2 per million input, USD 0.30 cached, and USD 6 per million output. That makes Grok about 2.1 times more expensive on input and about 1.6 times more on cached input, but only about 1.5 times more on output — and output is the line that dominates a real coding bill. For a solo developer burning five million input and three million output tokens a month, the difference is roughly eleven dollars.

Does Grok 4.5 have a SWE-bench Verified score?

Not an independent one. Grok 4.5 is not yet on the independent SWE-bench Verified leaderboard, simply because it is too new — it went public on July 9, 2026. Any SWE-bench percentage you see attributed to Grok 4.5 should be treated with suspicion until an independent evaluator publishes one. What Grok 4.5 does have is independent Artificial Analysis results: an Intelligence Index of 54 and a Coding Agent Index v1.3 score of 64.4, measured through the Grok Build harness at high effort. Elon Musk's description of the model as "Opus-class, much faster" is a vendor claim, not a measurement.

Why do you not compare Grok's coding score against Kimi's SWE-bench numbers?

Because they differ on two axes at once. Grok 4.5's coding figure is an Artificial Analysis Coding Agent Index produced by an independent evaluator. Kimi K2.7 Code's coding figures, SWE-bench Verified 60.4 percent and SWE-bench Pro 58.6, are a different benchmark family entirely, and they were produced by Moonshot AI on Moonshot AI's own harness. Different benchmarks and different evidence regimes. Putting the two numbers side by side would imply a comparison that is not real, so we never do it — not in a table, not in a sentence, and not in our infographics.

Is Grok 4.5's 54 percent hallucination rate a dealbreaker?

It depends entirely on what you use it for. On the independent AA-Omniscience evaluation, Grok 4.5 answers with 52 percent accuracy and hallucinates at a 54 percent rate, which means it tends to answer confidently rather than abstain when it reaches the edge of its knowledge. For open-domain factual work, that is a serious problem. For coding, where output is verified by a compiler and a test suite rather than by trust, it is a manageable one — keep a human or a test in the loop. The honest caveat runs the other way too: no standalone hallucination rate is published for Kimi K2.7 Code, so on that specific axis it is unreported rather than better.

Can I self-host Kimi K2.7 Code?

Yes. Kimi K2.7 Code ships open weights under a Modified MIT license, published on day one, so you can download the model, run it on your own GPU infrastructure, fine-tune it, and keep every token of data inside your own perimeter. Grok 4.5 offers nothing comparable: it is closed and API-only, with no weights released and no self-hosting path. For teams with data-residency requirements or air-gapped environments, this alone decides the comparison.

What are the context windows of Grok 4.5 and Kimi K2.7 Code?

Grok 4.5 offers a 500,000-token context window. Kimi K2.7 Code offers 256,000 tokens (262,144 exactly), with automatic context caching. Grok's window is close to double, which is a concrete advantage for whole-repository prompts and very long autonomous sessions. If your prompts routinely exceed 256,000 tokens, this comparison is effectively already decided in Grok 4.5's favor.

What is Kimi K2.7 Code's architecture?

Kimi K2.7 Code is an open-weight mixture-of-experts model with roughly 1 trillion total parameters, of which only 32 billion are active per token. It uses 384 experts, with 8 selected plus 1 shared per forward pass, multi-head latent attention, and a MoonViT vision encoder that lets it read screenshots and UI mockups inside a coding workflow. Moonshot AI published this architecture openly. SpaceXAI discloses nothing about Grok 4.5's internals, so no equivalent description exists.

Is SpaceXAI the same company as xAI?

Yes. xAI rebranded to SpaceXAI on July 6, 2026. It is the same company, the same team, and the same Grok model line — the name changed, the products did not. Grok 4.5 is a SpaceXAI model. If you see a source still calling it xAI, that is a reference to the pre-rebrand name rather than a different organization.

Will this verdict change?

Quite possibly, and we will say so plainly when it does. The verdict rests on two facts that are both actively in motion. The first is that Kimi K2.7 Code has an independent intelligence score of 42 but no independent coding result; if an outside evaluator reproduces something near Moonshot AI's self-reported SWE-bench figures, an open-weight model that is cheaper on every line, cheaper per completed task, and available everywhere becomes very hard to argue against. The second was Grok 4.5's European availability, which opened on July 17, 2026. We will update this page as both facts resolve.

If you are weighing these two against the rest of the field, these go deeper on adjacent matchups:

Last compared: July 13, 2026. Pricing and specifications verified directly from SpaceXAI and Moonshot AI documentation at the time of writing; independent scores are from Artificial Analysis. We have no affiliate relationship with either vendor. Grok 4.5's European Union availability opened on July 17, 2026, after this comparison was first published. Corrected on August 2, 2026: an earlier version of this page stated that K2.7 had no independent third-party results. Artificial Analysis scores Kimi K2.7 Code at 42 on the Intelligence Index v4.1, and this comparison has been rewritten around that figure.

Sources and references

Every figure on this page is attributed to whoever produced it. Vendor documentation and independent measurement are listed separately and never merged into a single ranking.

Our Verdict

Grok 4.5 wins this comparison, on an exchange rate rather than a knockout — and the win is scoped. It leads the one index that measures both: 54 against Kimi K2.7 Code's 42 on the independent Artificial Analysis Intelligence Index v4.1, read on August 2, 2026, and it adds 64.4 on the independent Artificial Analysis Coding Agent Index v1.3, measured through the Grok Build harness at high effort. Kimi K2.7 Code has no third-party coding score we have been able to find; the coding figures Moonshot AI publishes (SWE-bench Verified 60.4 percent, SWE-bench Pro 58.6) are self-reported on its own harness, unreplicated, and from a different benchmark family, so they cannot be set against Grok's charted numbers. The twelve-point intelligence gap is real, but that index is a generalist composite of nine mostly non-coding evaluations and Kimi K2.7 Code is a coding-specialized model — so it is measured and behind on breadth, not proven weaker at the job it was built for. What settles it is the price of that proof, and it has never been lower: Grok 4.5 costs about 2.1 times more on input (USD 2 against USD 0.95) but only about 1.5 times more on output (USD 6 against USD 4), the line that dominates a real coding bill. For a solo developer that premium is roughly eleven dollars a month, and it buys twelve index points plus close to double the context window (500,000 tokens against 256,000). Two caveats scope the win hard. Grok 4.5 is currently blocked in the European Union under the EU AI Act; SpaceXAI signalled a staged opening around mid-July 2026, so EU readers must verify availability for their region before acting on any of this. And Grok 4.5 carries a measured 54 percent hallucination rate on the independent AA-Omniscience evaluation, which argues for keeping a human or a test suite in the loop on anything requiring the model to know rather than to build. K2.7 remains the right call, without hesitation, if you are in the EU today, if you must self-host or keep data inside your own perimeter, if you burn hundreds of millions of output tokens a month, or if you can evaluate the model yourself and therefore do not need a third party to have done it for you. It is cheaper on every line, open on every axis, and available everywhere. If an independent coding result lands for Kimi K2.7 Code and confirms Moonshot's figures, this verdict should be revisited immediately.

Winner:Grok 4.5

Choose Grok 4.5

SpaceXAI's reasoning model — Opus-class speed at $2 and $6 per million tokens, 500K context, available to EU users since July 17, 2026; succeeded by Grok 4.6 on August 12.

Try Grok 4.5

Choose Kimi K2.7 Code

Moonshot AI's open-weight 1T-parameter MoE coding model — 32B active, 256K context, Modified MIT, metered at $0.95 in / $4.00 out per million tokens.

Try Kimi K2.7 Code

Frequently Asked Questions

Is Grok 4.5 better than Kimi K2.7 Code?

Grok 4.5 wins this comparison, on an exchange rate rather than a knockout — and the win is scoped. It leads the one index that measures both: 54 against Kimi K2.7 Code's 42 on the independent Artificial Analysis Intelligence Index v4.1, read on August 2, 2026, and it adds 64.4 on the independent Artificial Analysis Coding Agent Index v1.3, measured through the Grok Build harness at high effort. Kimi K2.7 Code has no third-party coding score we have been able to find; the coding figures Moonshot AI publishes (SWE-bench Verified 60.4 percent, SWE-bench Pro 58.6) are self-reported on its own harness, unreplicated, and from a different benchmark family, so they cannot be set against Grok's charted numbers. The twelve-point intelligence gap is real, but that index is a generalist composite of nine mostly non-coding evaluations and Kimi K2.7 Code is a coding-specialized model — so it is measured and behind on breadth, not proven weaker at the job it was built for. What settles it is the price of that proof, and it has never been lower: Grok 4.5 costs about 2.1 times more on input (USD 2 against USD 0.95) but only about 1.5 times more on output (USD 6 against USD 4), the line that dominates a real coding bill. For a solo developer that premium is roughly eleven dollars a month, and it buys twelve index points plus close to double the context window (500,000 tokens against 256,000). Two caveats scope the win hard. Grok 4.5 is currently blocked in the European Union under the EU AI Act; SpaceXAI signalled a staged opening around mid-July 2026, so EU readers must verify availability for their region before acting on any of this. And Grok 4.5 carries a measured 54 percent hallucination rate on the independent AA-Omniscience evaluation, which argues for keeping a human or a test suite in the loop on anything requiring the model to know rather than to build. K2.7 remains the right call, without hesitation, if you are in the EU today, if you must self-host or keep data inside your own perimeter, if you burn hundreds of millions of output tokens a month, or if you can evaluate the model yourself and therefore do not need a third party to have done it for you. It is cheaper on every line, open on every axis, and available everywhere. If an independent coding result lands for Kimi K2.7 Code and confirms Moonshot's figures, this verdict should be revisited immediately.

Which is cheaper, Grok 4.5 or Kimi K2.7 Code?

Grok 4.5 starts at $2 in / $6 out per M tokens. Kimi K2.7 Code starts at $0.95 in / $4 out per M tokens (free plan available). Check the pricing comparison section above for a full breakdown.

What are the main differences between Grok 4.5 and Kimi K2.7 Code?

The key differences span across 14 features we compared. For Independent intelligence score (Artificial Analysis Intelligence Index), Grok 4.5 offers 54 — measured by an independent evaluator, fourth overall while Kimi K2.7 Code offers 42 on the v4.1 index, read August 2, 2026 — twelve points behind. The 54 that circulates belongs to the earlier Kimi K2.6 on an older index and does not transfer. For Independent coding score (Artificial Analysis Coding Agent Index v1.3), Grok 4.5 offers 64.4 — measured by an independent evaluator, via the Grok Build harness at high effort, read August 2, 2026 while Kimi K2.7 Code offers None found published as of August 2, 2026 — its intelligence score exists, a third-party coding score does not. For Independent SWE-bench Verified score, Grok 4.5 offers None. Not yet on the independent leaderboard (too new). No SWE-bench percentage should be attributed to this model while Kimi K2.7 Code offers None independently reproduced. See the full feature comparison table above for all details.

Related Comparisons