Grok 4.5 vs Kimi K2.7: Measured Capability vs Open Weights (2026)
Grok 4.5 vs Kimi K2.7: independently scored 54 and 76 against zero third-party results, for only 1.5x the output price. Grok takes it — but not in the EU.
Feature Comparison
| Feature | Grok 4.5 | Kimi K2.7 |
|---|---|---|
| Independent intelligence score (Artificial Analysis Intelligence Index) | 54 — measured by an independent evaluator, fourth overall | None. Not yet on any independent leaderboard (too new). The Index figure of 54 that circulates for the earlier Kimi K2.6 belongs to a different model and does not transfer |
| Independent coding score (Artificial Analysis Coding Agent Index) | 76 — measured by an independent evaluator | None. No third-party coding result exists as of July 2026 |
| Independent SWE-bench Verified score | None. Not yet on the independent leaderboard (too new). No SWE-bench percentage should be attributed to this model | None independently reproduced |
| Vendor self-reported coding claim | No self-reported coding benchmark published. The vendor framing ("Opus-class, much faster") is a qualitative claim, not a measurement | SWE-bench Verified 60.4 percent and SWE-bench Pro 58.6 — self-reported by Moonshot AI on its own harness, not reproduced by any third party |
| Independent hallucination measurement (AA-Omniscience) | 52 percent accuracy with a 54 percent hallucination rate — a measured weakness, honestly a bad number, but it exists | Never measured by anyone outside the vendor. Unknown rather than good |
| Maximum context window | 500,000 tokens | 256,000 tokens (262,144), with automatic context caching |
| Input price (per million tokens) | USD 2.00 | USD 0.95 — roughly 2.1 times cheaper |
| Cached input price (per million tokens) | USD 0.50 | USD 0.19 — roughly 2.6 times cheaper |
| Output price (per million tokens) | USD 6.00 | USD 4.00 — only about 1.5 times cheaper, the narrowest output gap any independently scored flagship holds against this model |
| Model weights and licensing | Closed and proprietary. API only, no weights released | Open weights under a Modified MIT license, published on day one |
| European Union availability | Blocked at launch under the EU AI Act systemic-risk rules for general-purpose models. SpaceXAI signalled a staged EU opening expected around mid-July 2026 — verify for your region before building, this changes fast | Available. Open weights cannot be geo-blocked, so it can be self-hosted inside the EU today |
| Self-hosting and data residency | Not possible — managed API only | Possible — run the weights on your own hardware, in your own jurisdiction |
| Architectural transparency | Not disclosed by SpaceXAI | Fully published: mixture-of-experts, roughly 1 trillion total parameters with 32 billion active per token, 384 experts, multi-head latent attention, MoonViT vision encoder |
| Hosted metered API | Available from SpaceXAI | Available from Moonshot AI |
Pricing Comparison
Grok 4.5
Kimi K2.7
Detailed Comparison
Grok 4.5 vs Kimi K2.7 in 2026: Grok 4.5 is the flagship model from SpaceXAI (formerly xAI), priced at USD 2 per million input tokens, USD 0.50 cached, and USD 6 per million output tokens, with a 500,000-token context window. It scores 54 on the independent Artificial Analysis Intelligence Index and 76 on the independent Artificial Analysis Coding Index. Kimi K2.7 is Moonshot AI's open-weight agentic coding model, priced at USD 0.95 per million input tokens, USD 0.19 cached, and USD 4 per million output tokens, with a 256,000-token context window and downloadable weights under a Modified MIT license. Kimi K2.7 is not yet on any independent leaderboard: the coding figures Moonshot publishes are self-reported and have not been reproduced by a third party. Grok 4.5 costs roughly 2.1 times more on input but only about 1.5 times more on output — the narrowest premium any independently scored flagship charges over Kimi K2.7. Grok 4.5 takes the verdict wherever it is available. In the European Union it is currently blocked, and there Kimi K2.7 wins by default rather than by preference.
Quick Verdict
We ran both models side-by-side over a working week on the same coding briefs, the same agentic tool-use loops, and the same long-context retrieval tasks. Two things came out of it that make this the most interesting pairing we have written this year.
First: this is the tightest price duel we have seen between a closed flagship and an open-weight challenger. Input runs USD 2 against USD 0.95, roughly 2.1 times. But output — the line that actually dominates a coding bill — runs USD 6 against USD 4. That is about 1.5 times. Not ten times. Not six times. One and a half. A closed model with independently charted scores landing within touching distance of a Chinese open-weight model on the most expensive line item is, on its own, the story here.
Second: the two models are exact mirror images of each other, and the mirror runs in opposite directions. Grok 4.5 has proof but not access. It has been measured by an outside evaluator, and it is currently blocked in the European Union. Kimi K2.7 has access but not proof. You can download its weights and run it in any jurisdiction on earth, and nobody outside Moonshot AI has ever scored it. One model is verified and unavailable to part of the world; the other is available everywhere and unverified. You are not really choosing between two coding models. You are choosing which kind of uncertainty you can live with.
- Best independently measured capability: Grok 4.5 — 54 on the Artificial Analysis Intelligence Index and 76 on the Artificial Analysis Coding Index, both produced by an outside evaluator. Kimi K2.7 has no third-party score of any kind.
- Best price on every single line: Kimi K2.7 — cheaper on input, cheaper on cached input, cheaper on output. The gap is real but unusually small.
- Best context window: Grok 4.5 — 500,000 tokens against 256,000, close to double.
- Best availability and control: Kimi K2.7 — open weights under a Modified MIT license, self-hostable in any jurisdiction, including every country where Grok 4.5 currently cannot be used.
- Most honest about its own weaknesses: Grok 4.5, by accident of being measured — its 54 percent hallucination rate on the independent AA-Omniscience evaluation is a bad number, but at least it exists. Kimi K2.7's hallucination behavior has never been measured by anyone.
- Overall: Grok 4.5, where you can use it. The premium for independently verified capability has never been this small. If you are in the EU today, or you must self-host, or you are running hundreds of millions of output tokens a month, Kimi K2.7 is the answer instead.
Grok 4.5 vs Kimi K2.7 at a Glance
Every figure below comes from the vendors' own documentation for pricing and specifications, and from Artificial Analysis for the independent scores. Where a number is self-reported by the vendor, we label it as such and never present it as verified.
| Attribute | Grok 4.5 | Kimi K2.7 |
|---|---|---|
| Vendor | SpaceXAI (formerly xAI), United States | Moonshot AI, Beijing, China |
| Released | Public July 9, 2026 | June 12, 2026 |
| Model type | Closed frontier, API only | Open-weight mixture-of-experts |
| Architecture | Not disclosed | Roughly 1 trillion total parameters, 32 billion active per token, 384 experts, multi-head latent attention, MoonViT vision encoder |
| License | Proprietary, commercial API terms | Modified MIT, weights published on day one |
| Context window | 500,000 tokens | 256,000 tokens (262,144), automatic caching |
| Input price | USD 2 per million tokens | USD 0.95 per million tokens |
| Cached input price | USD 0.50 per million tokens | USD 0.19 per million tokens |
| Output price | USD 6 per million tokens | USD 4 per million tokens |
| Independent intelligence score | 54 on the Artificial Analysis Intelligence Index | None — not yet on any independent leaderboard |
| Independent coding score | 76 on the Artificial Analysis Coding Index | None — no third-party coding result exists |
| Vendor self-reported coding claim | None published; the vendor framing is qualitative | SWE-bench Verified 60.4 percent, SWE-bench Pro 58.6 — self-reported by Moonshot AI |
| Independent hallucination measurement | AA-Omniscience: 52 percent accuracy, 54 percent hallucination rate | Never measured by anyone outside Moonshot AI |
| Self-hosting | No | Yes — download the weights |
| European Union availability | Blocked at launch under the EU AI Act; a staged opening was signalled for around mid-July 2026 — verify before you build | Available, and self-hostable inside the EU |
The Three Different 54s on This Page
Before anything else, a warning — because we have already seen aggregator pages get this wrong, and getting it wrong flips the entire conclusion.
Three separate numbers in this comparison are 54, and they mean three completely different things.
- Grok 4.5 scores 54 on the independent Artificial Analysis Intelligence Index. That is a capability score, produced by an outside evaluator. Higher is better. This is a good number.
- Grok 4.5 also posts a 54 percent hallucination rate on the independent AA-Omniscience evaluation (alongside 52 percent accuracy). That is a failure rate on an entirely different axis. Lower is better. This is a bad number, and it happens to collide with the good one.
- Kimi K2.6 — the previous generation, a different model — is widely quoted at 54, but that figure comes from an earlier version of the Artificial Analysis index. On the current v4.1 index, the same one that scores Grok 4.5 at 54, Kimi K2.6 sits at 44. The two are not level, and the number you see repeated is stale. Either way it does not travel forward: Kimi K2.7 has never been assessed by Artificial Analysis, so it has no Intelligence Index at all. If you see an Index figure attached to Kimi K2.7 anywhere on the web, it has been inherited by mistake from its predecessor.
So: the intelligence score of 54 on this page belongs to Grok 4.5. The hallucination rate of 54 percent also belongs to Grok 4.5, and is not a compliment. Kimi K2.7 has no score. We will not pretend otherwise in either direction.
Grok 4.5 Overview
Grok 4.5 is the current flagship from SpaceXAI, the company formerly known as xAI — the rebrand landed on July 6, 2026, and the model line-up did not change with it (we covered what the SpaceXAI rebrand actually changed separately). Grok 4.5 was announced on July 8 and went public on July 9, 2026, replacing Grok 4.3 as the flagship while the older models stay available.
The specifications are confirmed from the vendor's own documentation: a 500,000-token context window, text and image input to text output, function calling, structured outputs, and reasoning effort levels from low through high. Pricing is USD 2 per million input tokens, USD 0.50 per million cached input tokens, and USD 6 per million output tokens. Rate limits run to 150 requests per second.
What makes Grok 4.5 unusual is that its capability claims are not, in the main, its own. Artificial Analysis published independent results on July 8, 2026: an Intelligence Index of 54, placing it fourth overall, and a Coding Agent Index of 76. The evaluator also measured a cost of roughly USD 2.49 per task, which is a fraction of what the top-of-market models charge to complete the same work. Those are third-party numbers, not marketing.
Two caveats belong right here rather than buried at the bottom. First, Grok 4.5 does not have an independent SWE-bench Verified score — it is simply not on the leaderboard yet, being too new. Elon Musk's public framing of the model as "Opus-class, much faster" is a vendor claim, not a measurement, and we treat it as such. Second, the same Artificial Analysis run that produced the flattering Intelligence Index also produced an unflattering one: on AA-Omniscience, Grok 4.5 answers with 52 percent accuracy and hallucinates at a 54 percent rate. For a model you intend to point at a codebase unsupervised, that is a number to take seriously.
Kimi K2.7 Overview
Kimi K2.7, and the coding-tuned K2.7-Code variant, is Moonshot AI's open-weight agentic coding model, released on June 12, 2026. Architecturally it is fully public, which is refreshing after a year of frontier labs disclosing nothing: a mixture-of-experts design with roughly 1 trillion total parameters, only 32 billion of them active per token, 384 experts of which 8 are selected plus 1 shared, multi-head latent attention, and a MoonViT vision encoder so it can read screenshots and UI mockups inside a coding workflow.
The weights shipped on day one under a Modified MIT license. You can download them, run them on your own hardware, fine-tune them, and keep every token inside your own infrastructure. Metered API pricing from Moonshot is USD 0.95 per million input tokens, USD 0.19 per million cached input tokens, and USD 4 per million output tokens, with automatic context caching. The context window is 256,000 tokens (262,144 exactly).
And then the hole. As of July 2026 — we last checked on July 13 — Kimi K2.7 has no independent third-party results of any kind. Not an Artificial Analysis Intelligence Index. Not an Artificial Analysis Coding Index. Not an LMArena Elo rating. Not an independently reproduced SWE-bench Verified score. Not Terminal-Bench, not LiveCodeBench, not GPQA, not AIME. Nothing. The only capability figures that exist for this model are Moonshot AI's own: SWE-bench Verified at 60.4 percent and SWE-bench Pro at 58.6, self-reported on the vendor's own harness and described by Moonshot as a new high-water mark for open-source models.
We want to be precise about what that means, because it is easy to be unfair in either direction. It does not mean Kimi K2.7 is a weak model. Our hands-on runs suggest it is a strong one. It means Kimi K2.7 is currently an unverifiable model — which is a different problem, and a temporary one. Evaluators usually reach a model of this profile within weeks. This page may need rewriting soon, and we would welcome that.
Why We Will Not Put These Benchmark Numbers Head-to-Head
This is the most important methodological point on the page, so we are stating it loudly rather than tucking it into a footnote.
Grok 4.5's charted coding figure is a Coding Agent Index of 76, produced by Artificial Analysis — an outside evaluator running its own harness. Kimi K2.7's coding figures are SWE-bench Verified 60.4 percent and SWE-bench Pro 58.6, produced by Moonshot AI, running Moonshot AI's harness, on Moonshot AI's infrastructure.
Those two things differ on two axes at once, not one. They are different benchmarks — an agentic index and a SWE-bench suite are not measuring the same thing, and the scales are unrelated. And they come from different evidence regimes — one was produced by a party with nothing to gain, the other by the party selling the model. Putting those two figures side by side in a table would imply both a comparison and a verdict, and neither would be real. So we do not do it: not in a row, not in a sentence, and not in the infographic below, which deliberately carries no benchmark row at all.
What we can say is narrower, and true. Grok 4.5's capability has been independently measured and is strong. Kimi K2.7's capability has not been independently measured. That is an asymmetry of evidence, not a proven gap in ability, and we score it accordingly.
Pricing: The Narrowest Gap of the Year
Both models bill per token, with input, cached input, and output metered separately. We pulled each rate directly from the vendor's own pricing documentation rather than from third-party aggregators, which drift.
| Rate (per million tokens) | Grok 4.5 | Kimi K2.7 | Multiple |
|---|---|---|---|
| Input | USD 2.00 | USD 0.95 | Grok is about 2.1 times more |
| Cached input | USD 0.50 | USD 0.19 | Grok is about 2.6 times more |
| Output | USD 6.00 | USD 4.00 | Grok is about 1.5 times more |
Read the output row twice, because output tokens are what a coding agent actually burns. USD 6 against USD 4 is a premium of about 50 percent. To put that in context: in the comparisons we have run against the frontier tier, the same Kimi K2.7 comes out six to twelve times cheaper than the closed flagship it faces. Against Grok 4.5 it is one and a half times cheaper. The independently scored, closed, American flagship is charging a 50 percent surcharge over a Chinese open-weight model on the most expensive line on the invoice.
That is the fact that reframes the whole decision. In most closed-versus-open matchups, the question is whether verified capability is worth paying a multiple for. Here, the question is whether verified capability is worth paying a rounding error for — and once the premium gets that small, the answer for most teams flips.
Grok 4.5's independently measured cost per completed task lands at roughly USD 2.49, according to Artificial Analysis. Kimi K2.7 has no equivalent independent cost-per-task figure, for the same reason it has no capability figure: nobody outside Moonshot has run it. What we can say is that a cheaper rate card does not automatically produce a cheaper bill, because a model that needs a second pass on a hard task burns its cheap tokens twice.
Real-World Cost Scenarios
Rate cards are abstract, so here is what the difference looks like on an actual monthly invoice. These are illustrative estimates using each vendor's published standard rates, with no cache hits assumed. Your real bill depends on caching, retries, and how often each model gets it right on the first pass.
| Workload (monthly) | Grok 4.5 | Kimi K2.7 | Difference |
|---|---|---|---|
| Solo developer: 5M input, 3M output | About USD 28 | About USD 17 | About USD 11 |
| Small team agent: 30M input, 20M output | About USD 180 | About USD 109 | About USD 71 |
| High-volume CI agent: 100M input, 80M output | About USD 680 | About USD 415 | About USD 265 |
Look at the solo-developer row. Eleven dollars a month is the entire cost of choosing an independently measured model over an unmeasured one. For an individual developer or a small team, that difference is noise — it is less than a single hour of the time you would lose to a bad refactor. This is why the verdict tilts the way it does.
Now look at the bottom row. At high volume the gap becomes a real line item: about USD 265 a month, roughly USD 3,200 a year, and it scales linearly from there. And that is the metered comparison only. If you self-host Kimi K2.7 on your own GPUs, the per-token bill disappears entirely and is replaced by infrastructure cost, which at sufficient scale is the cheaper curve. There is a volume threshold above which Kimi wins on economics no matter how good Grok is, and teams running agents around the clock are above it.
Availability: The Mirror Runs Backwards
This is the section most comparisons would skip, and it is the one that will actually decide the question for a large share of readers.
Grok 4.5 was blocked in the European Union at launch. The reason is regulatory, not technical: under the EU AI Act, general-purpose models judged to pose systemic risk face obligations that SpaceXAI did not meet at ship time. So on July 9, 2026, a model with independently verified frontier-class coding scores became unavailable to every developer inside the bloc.
This is not a permanent ban, and we want to be exact about the temporality here, because it is the fastest-moving fact on this page. SpaceXAI signalled a staged European rollout expected around mid-July 2026. We are publishing in mid-July 2026. That means the situation described in this paragraph may already have changed by the time you read it, and it may change again. Treat EU availability as a live variable, not a settled fact: check whether Grok 4.5 is reachable from your region before you design anything around it, and re-check if your last look was more than a couple of weeks ago. If the rollout completes as signalled, the single hardest objection to Grok 4.5 in this comparison evaporates. If it stalls, that objection hardens into a wall.
Kimi K2.7 has the opposite property, and it is structural rather than granted. Because the weights are published under a Modified MIT license, the model cannot be geo-blocked in any meaningful sense. You can download it in Berlin, run it on GPUs in Frankfurt, and keep every token of customer data inside the European Economic Area. No vendor decision, no regulatory negotiation, and no rollout schedule sits between you and the model. For a team with data-residency obligations, or one that simply refuses to build on infrastructure that a policy dispute can switch off, that is not a feature — it is the whole argument.
So the chiasm completes: Grok 4.5 has the evidence and, for part of the world, not the access. Kimi K2.7 has the access everywhere and, for now, none of the evidence. Which gap you can tolerate is a question about your organization, not about the models.
Hands-On: We Ran Both Side-by-Side
We tested both models over a working week from outside the EU, on the same four task families: a multi-file refactor of a TypeScript service, an agentic loop that had to read a failing test, edit code, and re-run the suite until green, a long-context task that fed a large codebase and asked for a cross-file change, and a factual-recall task designed to probe how each model behaves when it does not know something.
Coding accuracy. Both models were genuinely strong, and the gap was narrower than the evidence asymmetry might lead you to expect. Grok 4.5 was quick and confident, and on the multi-file refactor it produced correct edits on the first pass more often than not. Kimi K2.7 was close behind and occasionally better on the tasks where its agentic tuning showed. If you handed us anonymized outputs, we would not reliably be able to tell you which model produced which.
Agentic tool use. Grok 4.5 was fast — noticeably so, and that speed compounds in a read-edit-rerun loop where every iteration costs wall-clock time. Kimi K2.7 was more deliberate and, on the longer loops, slightly more methodical about verifying its own work. Neither was clearly dominant. Kimi's self-reported strength is agentic tool use, and while we cannot confirm the vendor's numbers, our runs are consistent with the claim.
Long context. Grok 4.5's 500,000-token window let us drop substantially more of a codebase into a single prompt than Kimi's 256,000. On whole-repository tasks that is a concrete, structural advantage, not a matter of taste. If your prompts routinely exceed 256,000 tokens, this comparison is already decided.
Factual reliability. This is where the independent AA-Omniscience number stopped being abstract. Grok 4.5, asked about things at the edge of its knowledge, tended to answer rather than abstain — which is exactly the behavior a 54 percent hallucination rate describes. In a coding context this shows up as inventing an API surface that does not exist and stating it with total composure. Kimi K2.7 was, in our limited runs, more inclined to hedge. But we have to flag our own limits here: this is an impression from a week of use, not a measurement, and Kimi K2.7's hallucination behavior has never been independently quantified by anyone. We are comparing a published number against a personal hunch, and we are not going to pretend that is a fair fight.
The pattern across the week: two models of very similar practical ability, separated by a modest price difference, a real context-window difference, and an enormous difference in how much anybody outside the vendor actually knows about them.
Ecosystem, Integration, and Deployment
Grok 4.5 lives inside SpaceXAI's managed platform: a first-party API with function calling, structured outputs, and reasoning-effort controls, served from US regions with generous rate limits. Nothing to provision, nothing to operate. The trade-off is total dependence on the vendor's platform decisions — including, as the EU situation demonstrates, whether the model is available to you at all.
Kimi K2.7 takes the open route. Its hosted API is OpenAI-compatible, which in practice means it drops into any agent framework, IDE plugin, or routing layer that already speaks that format, usually with a one-line base-URL change. Because the weights are public, it also runs inside self-hosted inference stacks and on GPU clouds, and a growing set of third-party providers serve it through their own gateways. The cost of that freedom is operational: if you self-host, you own the GPUs, the scaling, and the upgrades.
For a team that wants a managed endpoint and nothing else to think about, Grok 4.5 is the lower-friction option. For a team that has already standardized on the OpenAI API shape, or that needs to run inference inside its own perimeter, Kimi K2.7 is the drop-in. We line both up against the wider field in our roundup of the best AI coding tools of 2026.
Winner Per Category
| Category | Winner | Why |
|---|---|---|
| Independently verified capability | Grok 4.5 | 54 on the Artificial Analysis Intelligence Index and 76 on the Artificial Analysis Coding Index, both third-party. Kimi K2.7 has no independent score at all. |
| Context window | Grok 4.5 | 500,000 tokens against 256,000 — close to double. |
| Speed in agentic loops | Grok 4.5 | Noticeably quicker per iteration in our read-edit-rerun runs. |
| Price on every line | Kimi K2.7 | Cheaper on input, cached input, and output — though output is only about 1.5 times apart. |
| Open weights and licensing | Kimi K2.7 | Modified MIT, published day one. Grok 4.5 releases nothing. |
| Availability and jurisdiction | Kimi K2.7 | Runs anywhere, including the EU, where Grok 4.5 is currently blocked. |
| Self-hosting and data residency | Kimi K2.7 | Keep every token inside your own perimeter. Grok 4.5 is API only. |
| Architectural transparency | Kimi K2.7 | Full architecture published. SpaceXAI discloses nothing about Grok 4.5's internals. |
| Factual reliability | Neither, honestly | Grok 4.5 has a measured 54 percent hallucination rate on AA-Omniscience — a bad number that at least exists. Kimi K2.7 has never been measured. A known flaw against an unknown one. |
| Cost at very high volume | Kimi K2.7 | Self-hosting removes the per-token bill entirely above a certain scale. |
Count the rows and Kimi K2.7 wins more of them. That is not an accident, and we are not going to hide it: on everything you can count, Kimi wins. Grok 4.5 wins the one thing you cannot buy, which is knowing what you are getting — and it does so at a premium small enough that most teams will pay it without noticing. Verdicts are weighed, not tallied.
Pros and Cons
Grok 4.5
Pros
- Independently measured by an outside evaluator: 54 on the Artificial Analysis Intelligence Index, fourth overall, and 76 on the Artificial Analysis Coding Index.
- Independently measured cost of roughly USD 2.49 per completed task — a small fraction of what top-of-market models charge for the same work.
- A 500,000-token context window, close to double Kimi K2.7's, which matters for whole-repository prompts.
- Fast in practice, and that speed compounds across the iterations of an agentic loop.
- Priced at USD 2 per million input and USD 6 per million output — remarkably low for a model with third-party frontier-class scores.
- Fully managed: no GPUs to provision, no inference stack to run, generous rate limits.
Cons
- Currently blocked in the European Union under the EU AI Act. A staged opening was signalled for around mid-July 2026, but until it lands, EU teams simply cannot use this model.
- A 54 percent hallucination rate on the independent AA-Omniscience evaluation, with 52 percent accuracy — a genuine reliability concern for unsupervised agents.
- No independent SWE-bench Verified score. It is too new to be on that leaderboard, so the most widely cited coding benchmark simply has no entry for it.
- Closed and API only — no weights, no self-hosting, no data residency control, and no recourse if platform availability changes again.
- Architecture entirely undisclosed. The vendor's "Opus-class" framing is a claim, not a measurement.
Kimi K2.7
Pros
- Cheaper on every single line: USD 0.95 per million input, USD 0.19 cached, USD 4 per million output.
- Open weights under a Modified MIT license, published on day one — download, self-host, and fine-tune today.
- Available in every jurisdiction, including the entire European Union, where Grok 4.5 currently is not.
- Fully published architecture: mixture-of-experts, roughly 1 trillion total parameters with 32 billion active, 384 experts, multi-head latent attention.
- Includes a MoonViT vision encoder, so it reads screenshots and UI mockups inside coding workflows.
- OpenAI-compatible API, so it drops into existing agent frameworks with a base-URL change.
- Automatic context caching, which lowers the effective bill without any work on your side.
Cons
- No independent third-party score of any kind as of July 2026 — no Intelligence Index, no Coding Index, no LMArena rating, no independently reproduced SWE-bench result.
- Its coding figures, SWE-bench Verified 60.4 percent and SWE-bench Pro 58.6, are self-reported by Moonshot AI on its own harness and have not been reproduced by anyone.
- Hallucination behavior has never been independently quantified, so its reliability is unknown rather than good.
- A 256,000-token context window, roughly half Grok 4.5's, which is a hard ceiling on whole-repository prompts.
- Self-hosting shifts real operational cost onto you: GPUs, scaling, and upgrades all become your problem.
- Modified MIT is not plain MIT — read the license if you deploy at hyperscale.
When to Pick Each Model
When to pick Grok 4.5
- You are outside the European Union, or the EU rollout has completed by the time you read this — check before you commit.
- You want capability that somebody other than the vendor has actually measured, and you want it at the smallest premium currently on offer.
- Your prompts regularly exceed 256,000 tokens and you need the 500,000-token window.
- Iteration speed matters to you: you are running tight agentic loops where wall-clock time per pass compounds.
- You want a fully managed endpoint and have no interest in operating inference infrastructure.
- Your workload is code, not open-domain factual recall — the hallucination rate hurts most where the model is asked to know things.
When to pick Kimi K2.7
- You are in the European Union today. This is not a preference, it is arithmetic: Grok 4.5 is not available to you, and Kimi K2.7 is.
- You need open weights to self-host, fine-tune, or keep customer data inside your own perimeter.
- Data residency, air-gapped deployment, or sovereignty requirements govern your architecture.
- You run hundreds of millions of output tokens a month, where the price gap stops being noise and self-hosting removes the bill entirely.
- You refuse to build on a platform that a regulatory dispute can switch off, and open weights are the only real insurance against that.
- You are able to evaluate the model yourself on your own tasks, which makes the absence of third-party scores far less costly to you.
What Would Change Our Verdict
Both models are weeks old. This page is a snapshot of a market that moves faster than we can publish, and three things would move it.
Independent scores landing for Kimi K2.7. This is the big one, and it is likely within weeks. If Artificial Analysis charts Kimi K2.7 and the results confirm Moonshot AI's self-reported figures, then an open-weight model with credible verified numbers, cheaper on every line, and available in every jurisdiction becomes very hard to argue against — and our verdict flips. If the independent numbers come in well below the self-reported ones, the current verdict hardens instead. Either way, the single biggest fact on this page today is an absence, and absences get filled.
The EU rollout, in either direction. If Grok 4.5 opens in the European Union as signalled for around mid-July 2026, the strongest objection to it disappears and the recommendation gets simpler for a large slice of our readers. If the rollout stalls or reverses, Grok 4.5 becomes structurally unavailable to the EU market and the verdict there is not close — it is Kimi K2.7, without argument.
Any pricing move. A 50 percent output premium is what makes the proof cheap enough to buy. If SpaceXAI raises Grok 4.5's rates, or Moonshot cuts Kimi's further, that calculus changes quickly. This market discounts aggressively and without notice.
What would not change our verdict: another vendor-run benchmark from either side. We have enough self-reported numbers. What this comparison needs is someone independent to run Kimi K2.7, and someone independent to run Grok 4.5 on SWE-bench.
Final Verdict
Grok 4.5 wins this comparison, and it wins it on an exchange rate rather than a knockout. It is the only one of the two whose ability anybody outside the vendor has measured: 54 on the independent Artificial Analysis Intelligence Index and 76 on the independent Artificial Analysis Coding Index. Kimi K2.7 has no third-party score of any kind — the coding figures Moonshot AI publishes are self-reported, unreplicated, and drawn from a different benchmark family, so they cannot be set against Grok's charted numbers. That does not make Kimi K2.7 a worse model. It makes it an unverified one, which is a different and temporary condition.
What settles it is the price of that proof, and it has never been lower. Grok 4.5 costs about 2.1 times more on input and only about 1.5 times more on output — USD 6 against USD 4 on the line that dominates a real coding bill. For a solo developer that is roughly eleven dollars a month. Eleven dollars, for the difference between a model an outside evaluator has scored and a model nobody has, plus close to double the context window. That is not a premium; that is a rounding error, and we would pay it.
But the win is scoped, and the scope is not a detail. Grok 4.5 is blocked in the European Union today. For an EU team, this entire analysis is academic: the model is not available to you, and Kimi K2.7 is the answer by arithmetic rather than by preference. That block was signalled to lift around mid-July 2026, which is now — so verify your region before you act on any of this, because it is the fastest-moving fact on the page. And Grok 4.5 carries a measured 54 percent hallucination rate on AA-Omniscience, which is a real reason to keep a human in the loop on anything that requires the model to know rather than to build.
Kimi K2.7 remains the right call, without hesitation, if any of the following is true: you are in the EU, you must self-host, your data cannot leave your perimeter, you are burning hundreds of millions of output tokens a month, or you can properly evaluate the model yourself and therefore do not need anyone else to have done it for you. It is cheaper on every line, open on every axis, and available everywhere. Those are not consolation prizes.
Everyone else: take Grok 4.5, and revisit this page the moment somebody independent finally runs Kimi K2.7. We will.
Frequently Asked Questions
Which is better, Grok 4.5 or Kimi K2.7?
Grok 4.5, for most teams that can access it. It is the only model of the two that has been measured by an independent evaluator, scoring 54 on the Artificial Analysis Intelligence Index and 76 on the Artificial Analysis Coding Index, and it offers a 500,000-token context window against Kimi's 256,000. It charges only about 1.5 times more on output for that, which is the smallest premium any independently scored flagship has held over Kimi K2.7. But the win is scoped: Grok 4.5 is currently blocked in the European Union, and if you must self-host or keep data in your own perimeter, Kimi K2.7 is the correct choice regardless of scores.
Is Grok 4.5 available in the EU?
Not at launch. Grok 4.5 went public on July 9, 2026 and was blocked in the European Union, because under the EU AI Act general-purpose models judged to carry systemic risk face obligations SpaceXAI had not met at ship time. This is not a permanent ban: SpaceXAI signalled a staged European opening expected around mid-July 2026, so the situation is actively changing as we publish. This is the fastest-moving fact on this page. Check whether Grok 4.5 is reachable from your region before you build anything on it, and re-check if your last look was more than a couple of weeks ago. Kimi K2.7, by contrast, cannot be geo-blocked at all: its weights are downloadable, so it can be self-hosted inside the EU today.
Does Kimi K2.7 have an Artificial Analysis Intelligence Index score of 54?
No, and this is the most common error we see repeated about this model. The 54 that circulates comes from an earlier version of the Artificial Analysis index, where it was attached to Kimi K2.6, the previous generation — a different model that was actually evaluated. On the current v4.1 index, Kimi K2.6 sits at 44, not 54. Scores do not transfer between model versions, and they do not transfer between index versions either. Kimi K2.7 has never been assessed by Artificial Analysis and therefore has no Intelligence Index at all as of July 2026. Separately, Grok 4.5 scores 54 on the current Index, which makes the confusion easy to fall into. If you see an Index figure attached to Kimi K2.7 anywhere, it has been inherited by mistake.
How much cheaper is Kimi K2.7 than Grok 4.5?
Kimi K2.7 costs USD 0.95 per million input tokens, USD 0.19 cached, and USD 4 per million output tokens. Grok 4.5 costs USD 2 per million input, USD 0.50 cached, and USD 6 per million output. That makes Grok about 2.1 times more expensive on input and about 2.6 times more on cached input, but only about 1.5 times more on output — and output is the line that dominates a real coding bill. For a solo developer burning five million input and three million output tokens a month, the difference is roughly eleven dollars.
Does Grok 4.5 have a SWE-bench Verified score?
Not an independent one. Grok 4.5 is not yet on the independent SWE-bench Verified leaderboard, simply because it is too new — it went public on July 9, 2026. Any SWE-bench percentage you see attributed to Grok 4.5 should be treated with suspicion until an independent evaluator publishes one. What Grok 4.5 does have is independent Artificial Analysis results: an Intelligence Index of 54 and a Coding Agent Index of 76. Elon Musk's description of the model as "Opus-class, much faster" is a vendor claim, not a measurement.
Why do you not compare Grok's coding score against Kimi's SWE-bench numbers?
Because they differ on two axes at once. Grok 4.5's coding figure is an Artificial Analysis Coding Agent Index produced by an independent evaluator. Kimi K2.7's coding figures, SWE-bench Verified 60.4 percent and SWE-bench Pro 58.6, are a different benchmark family entirely, and they were produced by Moonshot AI on Moonshot AI's own harness. Different benchmarks and different evidence regimes. Putting the two numbers side by side would imply a comparison that is not real, so we never do it — not in a table, not in a sentence, and not in our infographics.
Is Grok 4.5's 54 percent hallucination rate a dealbreaker?
It depends entirely on what you use it for. On the independent AA-Omniscience evaluation, Grok 4.5 answers with 52 percent accuracy and hallucinates at a 54 percent rate, which means it tends to answer confidently rather than abstain when it reaches the edge of its knowledge. For open-domain factual work, that is a serious problem. For coding, where output is verified by a compiler and a test suite rather than by trust, it is a manageable one — keep a human or a test in the loop. The honest caveat runs the other way too: Kimi K2.7's hallucination rate has never been measured by anyone, so it is unknown rather than better.
Can I self-host Kimi K2.7?
Yes. Kimi K2.7 ships open weights under a Modified MIT license, published on day one, so you can download the model, run it on your own GPU infrastructure, fine-tune it, and keep every token of data inside your own perimeter. Grok 4.5 offers nothing comparable: it is closed and API-only, with no weights released and no self-hosting path. For teams with data-residency requirements or air-gapped environments, this alone decides the comparison.
What are the context windows of Grok 4.5 and Kimi K2.7?
Grok 4.5 offers a 500,000-token context window. Kimi K2.7 offers 256,000 tokens (262,144 exactly), with automatic context caching. Grok's window is close to double, which is a concrete advantage for whole-repository prompts and very long autonomous sessions. If your prompts routinely exceed 256,000 tokens, this comparison is effectively already decided in Grok 4.5's favor.
What is Kimi K2.7's architecture?
Kimi K2.7 is an open-weight mixture-of-experts model with roughly 1 trillion total parameters, of which only 32 billion are active per token. It uses 384 experts, with 8 selected plus 1 shared per forward pass, multi-head latent attention, and a MoonViT vision encoder that lets it read screenshots and UI mockups inside a coding workflow. Moonshot AI published this architecture openly. SpaceXAI discloses nothing about Grok 4.5's internals, so no equivalent description exists.
Is SpaceXAI the same company as xAI?
Yes. xAI rebranded to SpaceXAI on July 6, 2026. It is the same company, the same team, and the same Grok model line — the name changed, the products did not. Grok 4.5 is a SpaceXAI model. If you see a source still calling it xAI, that is a reference to the pre-rebrand name rather than a different organization.
Will this verdict change?
Quite possibly, and we will say so plainly when it does. The verdict rests on two facts that are both actively in motion. The first is that Kimi K2.7 has no independent verification as of July 2026; independent evaluators typically reach a model of this profile within weeks, so Kimi K2.7 is likely to be charted before long, and if those results confirm Moonshot AI's self-reported figures, an open-weight model that is cheaper on every line and available everywhere becomes very hard to argue against. The second is Grok 4.5's European availability, which was signalled to open around mid-July 2026. We will update this page as both facts resolve.
Related Comparisons
If you are weighing these two against the rest of the field, these go deeper on adjacent matchups:
- Claude Opus 4.8 vs Kimi K2.7 — the same open challenger against a frontier flagship that charges six times more.
- Kimi K2.7 vs GPT-5.5 — open weights against OpenAI's closed workhorse.
- Kimi K2.7 vs DeepSeek V4 — two Chinese open-weight models head-to-head.
- GLM-5.2 vs Kimi K2.7-Code — the open-weight coding race in detail.
- Claude Opus 4.8 — the frontier model both of these are priced against.
Last compared: July 13, 2026. Pricing and specifications verified directly from SpaceXAI and Moonshot AI documentation at the time of writing; independent scores are from Artificial Analysis. We have no affiliate relationship with either vendor. Grok 4.5's European Union availability is changing as we publish — verify it for your region before you build. Kimi K2.7 had no independent third-party benchmark results at the time of writing, and we will revise this page when that changes.
Our Verdict
Grok 4.5 wins this comparison, on an exchange rate rather than a knockout — and the win is scoped. It is the only one of the two whose capability anybody outside the vendor has measured: 54 on the independent Artificial Analysis Intelligence Index and 76 on the independent Artificial Analysis Coding Agent Index. Kimi K2.7 has no third-party score of any kind as of July 2026; the coding figures Moonshot AI publishes (SWE-bench Verified 60.4 percent, SWE-bench Pro 58.6) are self-reported on its own harness, unreplicated, and from a different benchmark family, so they cannot be set against Grok's charted numbers. That does not make Kimi K2.7 a worse model — it makes it an unverified one, which is a different and temporary condition. What settles it is the price of that proof, and it has never been lower: Grok 4.5 costs about 2.1 times more on input (USD 2 against USD 0.95) but only about 1.5 times more on output (USD 6 against USD 4), the line that dominates a real coding bill. For a solo developer that premium is roughly eleven dollars a month, and it buys an independently measured model plus close to double the context window (500,000 tokens against 256,000). Two caveats scope the win hard. Grok 4.5 is currently blocked in the European Union under the EU AI Act; SpaceXAI signalled a staged opening around mid-July 2026, so EU readers must verify availability for their region before acting on any of this. And Grok 4.5 carries a measured 54 percent hallucination rate on the independent AA-Omniscience evaluation, which argues for keeping a human or a test suite in the loop on anything requiring the model to know rather than to build. Kimi K2.7 remains the right call, without hesitation, if you are in the EU today, if you must self-host or keep data inside your own perimeter, if you burn hundreds of millions of output tokens a month, or if you can evaluate the model yourself and therefore do not need a third party to have done it for you. It is cheaper on every line, open on every axis, and available everywhere. If independent results land for Kimi K2.7 and confirm Moonshot's figures, this verdict should be revisited immediately.
Choose Grok 4.5
SpaceXAI's flagship reasoning model — Opus-class speed at $2 and $6 per million tokens, 500K context, blocked in the EU.
Try Grok 4.5 →Choose Kimi K2.7
Moonshot AI's open-weight 1T-parameter MoE coding model — 32B active, 256K context, Modified MIT, metered at $0.95 in / $4.00 out per million tokens.
Try Kimi K2.7 →Frequently Asked Questions
Is Grok 4.5 better than Kimi K2.7?
Grok 4.5 wins this comparison, on an exchange rate rather than a knockout — and the win is scoped. It is the only one of the two whose capability anybody outside the vendor has measured: 54 on the independent Artificial Analysis Intelligence Index and 76 on the independent Artificial Analysis Coding Agent Index. Kimi K2.7 has no third-party score of any kind as of July 2026; the coding figures Moonshot AI publishes (SWE-bench Verified 60.4 percent, SWE-bench Pro 58.6) are self-reported on its own harness, unreplicated, and from a different benchmark family, so they cannot be set against Grok's charted numbers. That does not make Kimi K2.7 a worse model — it makes it an unverified one, which is a different and temporary condition. What settles it is the price of that proof, and it has never been lower: Grok 4.5 costs about 2.1 times more on input (USD 2 against USD 0.95) but only about 1.5 times more on output (USD 6 against USD 4), the line that dominates a real coding bill. For a solo developer that premium is roughly eleven dollars a month, and it buys an independently measured model plus close to double the context window (500,000 tokens against 256,000). Two caveats scope the win hard. Grok 4.5 is currently blocked in the European Union under the EU AI Act; SpaceXAI signalled a staged opening around mid-July 2026, so EU readers must verify availability for their region before acting on any of this. And Grok 4.5 carries a measured 54 percent hallucination rate on the independent AA-Omniscience evaluation, which argues for keeping a human or a test suite in the loop on anything requiring the model to know rather than to build. Kimi K2.7 remains the right call, without hesitation, if you are in the EU today, if you must self-host or keep data inside your own perimeter, if you burn hundreds of millions of output tokens a month, or if you can evaluate the model yourself and therefore do not need a third party to have done it for you. It is cheaper on every line, open on every axis, and available everywhere. If independent results land for Kimi K2.7 and confirm Moonshot's figures, this verdict should be revisited immediately.
Which is cheaper, Grok 4.5 or Kimi K2.7?
Grok 4.5 is priced at $2 in / $6 out per M tokens. Kimi K2.7 is priced at $0.95 in / $4 out per M tokens (free plan available). Check the pricing comparison section above for a full breakdown.
What are the main differences between Grok 4.5 and Kimi K2.7?
The key differences span across 14 features we compared. For Independent intelligence score (Artificial Analysis Intelligence Index), Grok 4.5 offers 54 — measured by an independent evaluator, fourth overall while Kimi K2.7 offers None. Not yet on any independent leaderboard (too new). The Index figure of 54 that circulates for the earlier Kimi K2.6 belongs to a different model and does not transfer. For Independent coding score (Artificial Analysis Coding Agent Index), Grok 4.5 offers 76 — measured by an independent evaluator while Kimi K2.7 offers None. No third-party coding result exists as of July 2026. For Independent SWE-bench Verified score, Grok 4.5 offers None. Not yet on the independent leaderboard (too new). No SWE-bench percentage should be attributed to this model while Kimi K2.7 offers None independently reproduced. See the full feature comparison table above for all details.

