Skip to content

GPT-5.6 Terra vs Kimi K2.7 Code: Verified Scores vs Open-Weight Price (2026)

GPT-5.6 Terra scores 55 to Kimi K2.7 Code's 42 on Artificial Analysis. Kimi is cheaper, but only 2.1x. The measured gap takes it.

GPT-5.6 Terra and Kimi K2.7 Code are both scored on the Artificial Analysis Intelligence Index v4.1, where Terra leads 55 to 42 — thirteen points apart on the same harness, read August 2, 2026
Both models sit on the same independent index. On the Artificial Analysis Intelligence Index v4.1, GPT-5.6 Terra scores 55 and Kimi K2.7 Code scores 42 — thirteen points apart on the same harness, read August 2, 2026.

Feature Comparison

FeatureGPT-5.6 TerraKimi K2.7 Code
Independent intelligence score (Artificial Analysis Intelligence Index)55 (independent)42 on the v4.1 index, read August 2, 2026 — thirteen points behind. Figures circulating for Kimi K2.6 belong to a different model and do not transfer
Independent coding score (Artificial Analysis Coding Agent Index v1.3, read August 2, 2026)62.28 (Codex harness, max reasoning effort)Not charted
Evidence regime behind the coding claimIndependently charted: 62.28 on the AA Coding Agent Index v1.3 (Codex harness, max reasoning effort)SWE-bench Verified 60.4 percent and SWE-bench Pro 58.6, self-reported by Moonshot AI on its own harness and not reproduced by any third party
Maximum context window1,050,000 tokens256,000 tokens (262,144), with automatic context caching
Input price (per million tokens)USD 2.00USD 0.95 — roughly 2.1 times cheaper
Cached input price (per million tokens)USD 0.20USD 0.19 — roughly 1.3 times cheaper, close to level
Output price (per million tokens)USD 12.00USD 4.00 — roughly 3.0 times cheaper
Model weights and licensingClosed, proprietary; API only, no weights releasedOpen weights under a Modified MIT license, downloadable from day one
Self-hosting and data residencyNot possible — managed API onlyPossible — run the weights on your own hardware, in your own jurisdiction
Architectural transparencyNot disclosed by OpenAIFully published: mixture-of-experts, roughly 1 trillion total parameters with 32 billion active, 384 experts, multi-head latent attention, MoonViT vision encoder
Hosted metered APIAvailable from OpenAIAvailable from Moonshot AI

Pricing Comparison

GPT-5.6 Terra

$2 in / $12 out per M tokens
paid

Kimi K2.7 Code

$0.95 in / $4 out per M tokens
Free plan available
Free trial available
freemium

Detailed Comparison

GPT-5.6 Terra vs Kimi K2.7 Code in 2026: GPT-5.6 Terra is OpenAI's balanced tier, priced at USD 2.00 per million input tokens, USD 0.20 cached, and USD 12 per million output tokens, with a 1,050,000-token context window. It scores 55 on the independent Artificial Analysis Intelligence Index and 62.28 on the Artificial Analysis Coding Agent Index v1.3, through the Codex harness at max reasoning effort. Kimi K2.7 Code is Moonshot AI's open-weight agentic coding model, priced at USD 0.95 per million input tokens, USD 0.19 cached, and USD 4 per million output tokens, with a 256,000-token context and downloadable weights under a Modified MIT license. Kimi K2.7 Code is independently scored too: 42 on the Artificial Analysis Intelligence Index v4.1, read on August 2, 2026, thirteen points behind Terra on the same harness. That index is a generalist composite of nine evaluations, most of them reasoning and knowledge tests rather than coding, so part of the gap reflects scope rather than quality; Moonshot's own SWE-bench figures remain self-reported. Terra costs roughly 2.1 times more on input and roughly 3.0 times more on output — the smallest premium any independently scored flagship charges over Kimi. Terra takes the verdict, because here the price of proof is unusually cheap.

Quick Verdict

Most of the closed-versus-open matchups this year come down to the same trade: you pay ten or twenty times more for a model somebody outside the vendor has actually measured. That framing collapses in this particular pairing, and that is what makes it worth reading.

GPT-5.6 Terra is the closest thing in the closed field to Kimi K2.7 Code's price. Input runs USD 2.00 against USD 0.95 — roughly 2.1 times. Output runs USD 12 against USD 4 — roughly 3.0 times. Cached input is USD 0.20 against USD 0.19, a gap of about 1.1 times, which is close to noise. Compare that with the frontier tier, where the same open model is ten to twelve times cheaper than the flagship it faces, and the shape of the decision changes completely.

Meanwhile the capability gap does not shrink. Both models sit on the same independent Artificial Analysis intelligence index, and Terra is ahead on it: 55 against Kimi K2.7 Code's 42, read on August 2, 2026 — a thirteen-point gap on the same harness. On the coding side neither carries a third-party figure at all, so both models' coding cases rest on their makers' own harnesses or on nothing.

So GPT-5.6 Terra takes the overall verdict, and unlike the frontier matchups, it does not take it by a hair. It is measured ahead, not merely measured: thirteen index points, plus roughly four times the context, for about 2.1 times more on input. That is a small enough premium that most teams should simply pay it.

But Kimi K2.7 Code wins real things, and it wins them outright:

  • Kimi K2.7 Code wins price, on every line. USD 0.95 against USD 2.00 on input, USD 0.19 against USD 0.20 cached, USD 4 against USD 12 on output. Cheaper is cheaper, and at volume it compounds.
  • Kimi K2.7 Code wins control. Open weights under a Modified MIT license, downloadable on day one, runnable on your own hardware. Terra cannot do this at any price.
  • Kimi K2.7 Code wins architectural transparency. Moonshot publishes the full design. OpenAI publishes nothing comparable for Terra.
  • Terra wins measured intelligence. 55 on the Artificial Analysis Intelligence Index v4.1 against Kimi K2.7 Code's 42, read August 2, 2026 — thirteen points, same harness.
  • Independently verified coding covers one side only. The Artificial Analysis Coding Agent Index v1.3 charts GPT-5.6 Terra at 62.28, through the Codex harness at max reasoning effort, and carries no Kimi K2.7 Code entry — and we have found no third-party coding result published for Kimi K2.7 Code anywhere else, so that axis has no head-to-head.
  • Terra wins context. 1,050,000 tokens against 256,000 — roughly four times the room.

How We Compared Them, and the Rule We Hold To

We ran both models side by side on their respective APIs — the same refactoring passes, the same long-document work, the same agentic tool-calling loops — to get a feel for how each behaves in practice. That hands-on time informs the judgment calls here: how a model holds a long context, how it recovers when a tool call fails, how much supervision it needs before you can leave it running.

What it does not do is produce benchmark numbers. We do not publish our own scores, because a handful of prompts run by one team is not a benchmark. Every figure on this page comes either from a named independent evaluator or from the vendor, and we label which is which, every single time.

That labeling rule is usually a footnote. In this comparison it is the whole story, so here it is spelled out:

  • Independent means measured by a third party with no stake in the result — Artificial Analysis, LMArena, or vals.ai. Those numbers are comparable across models when the same harness ran them under the same conditions — which is why, on an index that scores a harness paired with a model, we name the harness every time.
  • Vendor self-reported means the company that built the model ran the benchmark itself, on its own harness, and published the outcome. That is useful signal. It is not verification. It is not comparable to an independent score, and it is not reliably comparable to another vendor's self-reported score either, because no two vendor harnesses are alike.

There is a second rule that applies specifically to this pairing. Terra's charted coding score and Kimi's published coding claim are not the same benchmark, and they were not produced under the same evidentiary regime. Two independent reasons to never place them side by side, so we never do — not in a table row, not in a sentence, not in the infographic. Any page that stacks them is telling you a lie by layout.

GPT-5.6 Terra and Kimi K2.7 Code at a Glance

GPT-5.6 Terra is the middle tier of OpenAI's GPT-5.6 family — the balanced model, positioned beneath GPT-5.6 Sol and above GPT-5.6 Luna. It is priced at USD 2.00 per million input tokens, USD 0.20 per million cached input tokens, and USD 12 per million output tokens, with a 1,050,000-token context window. It is a closed model: no weights, no self-hosting, no published architecture. What OpenAI offers in place of transparency is external measurement. Terra carries a score of 55 on the independent Artificial Analysis Intelligence Index, produced by an evaluator with no commercial stake in the outcome; the same evaluator charts it at 62.28 on the Artificial Analysis Coding Agent Index v1.3, through the Codex harness at max reasoning effort.

Kimi K2.7 Code is Moonshot AI's open-weight agentic coding model, released on June 12, 2026 from the company's Beijing headquarters. It is a mixture-of-experts design: roughly 1 trillion total parameters with about 32 billion active per token, spread across 384 experts of which 8 are selected per token plus 1 shared, using multi-head latent attention and a MoonViT vision encoder for native image input. The weights shipped under a Modified MIT license on day one — you can download them, run them on your own hardware, and put them inside your own product. It offers a 256,000-token context window (262,144 tokens exactly) with automatic context caching, at USD 0.95 per million input tokens, USD 0.19 cached, and USD 4 per million output tokens.

Read those two paragraphs again and notice what is missing from each. OpenAI will not tell you how Terra is built. Moonshot will not show you a number anyone else produced. Each company is opaque about precisely the thing the other is open about — and which opacity you can live with is, in the end, this entire comparison.

SpecificationGPT-5.6 TerraKimi K2.7 Code
VendorOpenAIMoonshot AI (Beijing)
PositioningBalanced tier of the GPT-5.6 familyOpen-weight agentic coding model, released June 12, 2026
Input, per million tokensUSD 2.00USD 0.95
Cached input, per million tokensUSD 0.20USD 0.19
Output, per million tokensUSD 12.00USD 4.00
Context window1,050,000 tokens256,000 tokens (262,144), with automatic caching
Independent intelligence score55 on the Artificial Analysis Intelligence Index v4.1 (GPT-5.6 Terra, max)42 on the Artificial Analysis Intelligence Index v4.1 (Kimi K2.7 Code), read August 2, 2026
Independent coding score62.28 on the Artificial Analysis Coding Agent Index v1.3, Codex harness at max reasoning effortNot charted; no third-party coding result found as of August 2, 2026
Weights and licensingClosed, API onlyOpen weights, Modified MIT, downloadable day one
Self-hostingNot possibleYes, on your own hardware
Published architectureNot disclosedMixture-of-experts, roughly 1T total and 32B active, 384 experts, multi-head latent attention, MoonViT vision encoder

The Part That Matters: Only One of These Models Has a Coding Score

Here is the situation in plain terms.

GPT-5.6 Terra, independently measured:

  • Artificial Analysis Intelligence Index: 55 (independent).
  • Artificial Analysis Coding Agent Index v1.3: 62.28, Codex harness at max reasoning effort (read August 2, 2026).

Kimi K2.7 Code, independently measured:

  • Artificial Analysis Intelligence Index: 42 on the v4.1 index, read August 2, 2026 — thirteen points behind Terra.
  • LMArena: not yet ranked.
  • SWE-bench Verified, SWE-bench Pro, Terminal-Bench, LiveCodeBench, GPQA, AIME: no independent third-party results found as of August 2, 2026.

An earlier version of this page said that second list was the complete state of independent evidence for Kimi K2.7 Code, and that there was none. That was wrong: the Intelligence Index score existed when we published. The accurate statement is narrower — Kimi K2.7 Code is independently scored on the generalist index at 42, and it is on the coding-specific evidence that nothing third-party has been published.

What Moonshot AI has published, on its own harness, is a SWE-bench Pro result of 58.6 and a SWE-bench Verified result of 60.4 percent, which the company describes as a new high-water mark for open-source models. Those figures may well be accurate. Moonshot has a reasonable track record, and the prior model in the line was independently evaluated in due course. But as of this writing, nobody outside Moonshot has reproduced them, and Terra's independent coding measurement belongs to a different benchmark family entirely — which is the second, separate reason we never put the two in one row.

One clarification that trips up a lot of coverage: Kimi K2.6, the predecessor, does carry an independent Artificial Analysis Intelligence score. K2.7 is a different model. That number does not transfer, and any article that quietly reuses it for K2.7 is inventing a data point. We are not going to do that. K2.7 has its own v4.1 figure, and it is 42, read on August 2, 2026 — so this page uses that number rather than borrowing its predecessor's.

Why a generalist index is not the same as a verdict on coding

It would be lazy — and wrong — to read a 42 as "Kimi K2.7 Code is worse." The evidence says it is thirteen points behind on a composite of nine evaluations, most of which are reasoning and knowledge tests rather than coding — and this is a coding-specialized model. Behind on breadth and worse at the job are very different claims.

The honest position on Kimi K2.7 Code today is that it might be excellent. Its architecture is serious, its price is aggressive, its predecessor scored respectably when independently tested, and its self-reported numbers are not outlandish for a model of this design. Nothing here suggests Moonshot is exaggerating.

What it means is that buying Kimi K2.7 Code today is buying a measured generalist deficit and an unmeasured coding claim, while buying GPT-5.6 Terra is buying a third-party number on both axes. For a hobbyist, that distinction is close to meaningless: run it, see if you like it, the price makes the experiment nearly free. For an engineering lead who has to stand in front of a room and justify why a model was chosen for a production system, it is the whole ballgame.

And the missing half is temporary. Kimi K2.7 Code is already on the Artificial Analysis Intelligence Index; what is still absent is an independent coding result and an LMArena rating. If a third-party coding number lands close to what Moonshot claims, the case for a model at roughly a third of Terra's input price gets strong very quickly, and this verdict should be revisited. We will update this page when that happens.

Artificial Analysis Coding Agent Index v1.3: GPT-5.6 Terra is charted at 62.28 at max reasoning effort through the Codex harness, while Kimi K2.7 Code is not charted and no third-party coding result was found for it — one side measured, the other not
The Artificial Analysis Coding Agent Index v1.3 charts GPT-5.6 Terra at 62.28 through the Codex harness at max reasoning effort and carries no Kimi K2.7 Code entry, and we found no third-party coding result published for Kimi K2.7 Code elsewhere as of August 2, 2026. One side measured, the other not — which settles nothing about coding ability.

Pricing: The Narrowest Gap in the Closed Field

These are the published metered rates, per million tokens:

Metered rate (per million tokens)GPT-5.6 TerraKimi K2.7 Code
InputUSD 2.00USD 0.95
Cached inputUSD 0.20USD 0.19
OutputUSD 12.00USD 4.00

Kimi K2.7 Code is roughly 2.1 times cheaper on input, roughly 1.3 times cheaper on cached input, and roughly 3.0 times cheaper on output. It wins every line. It is genuinely the cheaper model and nothing below takes that away.

But hold those multiples next to the rest of the market for a second, because the comparison only makes sense in context. Against the premium closed flagships, Kimi K2.7 Code is ten to twelve times cheaper. Against Terra it is under three times cheaper on input and under four on output. On cached input, the two are within a rounding error of each other — USD 0.20 against USD 0.19 is a difference of six cents per million tokens.

Put a number on it. A workload burning 50 million output tokens in a month — not extreme for a team running coding agents continuously — costs USD 750 on Terra and USD 200 on Kimi K2.7 Code. That is a gap of USD 550 a month. It is real money and you should not wave it away, but it is not the kind of gap that reorganizes a budget, and it is a fraction of one engineer's monthly cost. In the frontier tier the same workload produces a gap ten times larger.

This is the crux of the whole page. The question is never "is the closed model cheaper?" — it never is. The question is what you are paying for the difference, and what you get back. Here you are paying a couple of hundred dollars a month at moderate volume, and getting back thirteen index points, a coding score the other model does not have, and roughly four times the context. That is an unusually good exchange rate.

It stops being a good exchange rate at scale. At 500 million output tokens a month the gap becomes USD 5,500 a month, and at that point the arithmetic starts arguing for Kimi loudly. Multiply your own volume before you take our word for any of this.

Context Window: 1.05M Against 256K

Terra carries a 1,050,000-token context window. Kimi K2.7 Code carries 256,000 tokens (262,144 exactly), with automatic context caching that lowers the cost of repeated prefixes without you having to manage it.

Roughly four times the context is a real advantage, but be honest about when it actually binds. A 256,000-token window is already large enough for the overwhelming majority of coding work: a substantial repository slice, a long design document plus the code it describes, a full day of agent scrollback. If your workflow fits comfortably inside 256,000 tokens today, Terra's extra headroom is capacity you are paying for and not using.

Where it does bind: whole-monorepo reasoning, very long agentic sessions that accumulate hundreds of tool calls without compaction, and document work at the scale of full legal or research corpora. If that is your work, the gap is decisive — and Kimi's automatic caching does not close it. Caching makes a window cheaper to reuse; it does not make it bigger.

Open Weights, Self-Hosting, and What OpenAI Will Not Sell You

Kimi K2.7 Code's weights are published under a Modified MIT license and were downloadable on day one. This is not a marketing detail. It changes what you are legally and operationally able to do:

  • Run it on your own hardware. No inference data leaves your infrastructure. For teams with hard data-residency requirements, that is the difference between usable and not usable, and no vendor assurance substitutes for it.
  • Fix your cost ceiling. Self-hosted inference converts a metered bill into a hardware amortization. At high volume that changes the economics entirely — and it is the one route by which Kimi's price advantage grows rather than shrinks.
  • Never get deprecated out from under you. A downloaded weight file does not get sunset on a vendor's schedule. Anyone who has had a model retired mid-product knows what that is worth.
  • Inspect and modify. Full architecture disclosure — mixture-of-experts, roughly 1 trillion total parameters with 32 billion active, 384 experts, multi-head latent attention, the MoonViT vision encoder — means you can reason about the model's behavior instead of guessing at it.

GPT-5.6 Terra offers none of this, and OpenAI does not pretend otherwise. It is a closed, API-only model with an undisclosed architecture. If open weights are a requirement rather than a preference, this comparison is over before it starts and Kimi K2.7 Code wins by default — Terra is not a candidate at any price.

One honest caveat on the openness point, because it cuts both ways: open weights under a Modified MIT license is a genuine grant, but open-weight is not the same as open-source. Moonshot publishes the finished model, not the training data or the training code. You get the artifact, not the recipe. For deployment freedom that distinction rarely matters; for auditability and reproducibility, it does.

A second caveat that matters for the vision claim: Kimi K2.7 Code ships a native vision encoder, and Terra's published specification does not describe its multimodal handling in comparable terms. We are not going to score that as a clean win for either side, because we would be comparing a disclosed capability against an undisclosed one — the same mistake as stacking an independent score against a vendor one.

Winner by Category

CategoryWinnerWhy
Best measured intelligenceGPT-5.6 Terra55 against Kimi K2.7 Code's 42 on the Artificial Analysis Intelligence Index v4.1, read August 2, 2026 — thirteen points on the same harness.
Best independently verified codingNo head-to-headThe Artificial Analysis Coding Agent Index v1.3 charts GPT-5.6 Terra at 62.28 through the Codex harness at max reasoning effort and carries no Kimi K2.7 Code entry, and we have found no third-party coding result for Kimi K2.7 Code elsewhere — one side measured, the other not, which settles nothing about coding ability.
Best cost per tokenKimi K2.7 CodeCheaper on input, cached input, and output. It wins every line.
Best for self-hosting and data residencyKimi K2.7 CodeOpen weights, Modified MIT, downloadable. Terra cannot do this at all.
Best for very long contextGPT-5.6 Terra1,050,000 tokens against 256,000.
Best value for a verified modelGPT-5.6 TerraThe smallest premium any independently scored flagship charges over Kimi K2.7 Code — roughly 2.1 times on input.
Best for a very high-volume agentic fleetKimi K2.7 CodeAt hundreds of millions of output tokens a month the gap compounds, and self-hosting caps it.
Best when you must justify the choice to a review boardGPT-5.6 TerraThird-party numbers on both axes. For Kimi K2.7 Code you can cite a 42 on intelligence, but nothing external on coding.
Best architectural transparencyKimi K2.7 CodeFull architecture published. OpenAI discloses nothing comparable for Terra.

Pros and Cons

GPT-5.6 Terra

Pros

  • Independently verified at 55 on the Artificial Analysis Intelligence Index and 62.28 on the Coding Agent Index v1.3, through the Codex harness at max reasoning effort, by an evaluator with no stake in the result.
  • The smallest price premium of any independently scored flagship over Kimi K2.7 Code: roughly 2.1 times on input, roughly 3.0 times on output.
  • Cached input at USD 0.20 per million tokens is nearly level with Kimi K2.7 Code's USD 0.19, which makes prompt-heavy, cache-friendly workloads cost almost the same on both.
  • 1,050,000-token context window, roughly four times Kimi K2.7 Code's.
  • Managed API with no infrastructure to run, and a clear ladder to Sol above it or Luna below it if your needs shift.

Cons

  • Still the more expensive model on every single line, and output at USD 12 per million tokens hurts on agentic workloads that generate long traces.
  • Closed weights. No self-hosting, no data residency through hardware control, no protection against deprecation.
  • Undisclosed architecture — you cannot inspect or reason about how it works.
  • At very high volume the cost gap compounds into real money, and there is no self-hosted escape hatch.

Kimi K2.7 Code

Pros

  • Cheaper on every line: USD 0.95 per million input tokens, USD 0.19 cached, USD 4 per million output tokens.
  • Open weights under a Modified MIT license, downloadable from day one and self-hostable on your own hardware.
  • Full architecture disclosure — mixture-of-experts, roughly 1 trillion total parameters with 32 billion active, 384 experts, multi-head latent attention.
  • Native vision through the MoonViT encoder.
  • Purpose-built for agentic coding, with automatic context caching that needs no management.

Cons

  • Behind on the one index that scores both: 42 against Terra's 55 on the Artificial Analysis Intelligence Index v4.1, read August 2, 2026 — though that index is a generalist composite and this is a coding-specialized model.
  • No LMArena rating and no third-party coding result that we have been able to find, so its coding case specifically still rests on the vendor's own harness.
  • Its published coding figures are self-reported by Moonshot AI on its own harness and have not been reproduced by anyone else.
  • 256,000-token context, roughly a quarter of Terra's.
  • Its cost advantage over Terra specifically is narrower than its reputation suggests — under three times on input, and only about 1.3 times on cached input.
  • The hosted API is operated from Beijing, which carries its own compliance considerations. Self-hosting the weights avoids that; using the hosted API does not.

When to Pick GPT-5.6 Terra

  • You have to justify the choice to someone. A review board, a client, a regulator, a CTO who asks how you know the model is any good. Terra is the higher of the two on the one third-party index that scores both, 55 against 42, rather than on a vendor's announcement.
  • Your volume is moderate. At tens of millions of output tokens a month, the price gap is in the hundreds of dollars. Verified capability plus four times the context for that is a bargain, and you should take it.
  • Your workload is cache-heavy. If most of your tokens are repeated prefixes hitting the cache, you are comparing USD 0.20 against USD 0.19 — and at that point the price argument for Kimi has nearly evaporated.
  • You routinely exceed 256,000 tokens of context. Whole-monorepo work, very long agentic sessions, large document corpora.
  • You cannot run your own evaluations. If you have no internal harness, the independent scores are the only real information you have, and only one of these models has any.

When to Pick Kimi K2.7 Code

  • You need to self-host. Data residency, air-gapped environments, regulated industries. Terra simply cannot do this, so the comparison ends there.
  • Your volume is very high. At 500 million output tokens a month the gap is USD 5,500 a month, and self-hosting the weights caps it entirely. Scale is where Kimi's case gets loud.
  • You can run your own evaluations. This is the key one. A thirteen-point gap on a generalist index matters far less if you have an internal harness on your own tasks, because then you are measuring the model on the only benchmark that truly matters, which is yours.
  • You want protection from deprecation. A downloaded weight file is yours. An API endpoint is not.
  • You need to inspect or modify the model. The architecture is public and the license permits it. Terra offers neither.
Split verdict by category: GPT-5.6 Terra scores 55 on the Intelligence Index v4.1 and 62.28 on the Coding Agent Index v1.3 with the larger context window, while Kimi K2.7 Code scores 42 on the Intelligence Index v4.1 and ships open weights under a Modified MIT license that can be self-hosted
Both models are independently scored. GPT-5.6 Terra takes measured intelligence at 55 on the Artificial Analysis Intelligence Index v4.1 against 42 for Kimi K2.7 Code, and is charted at 62.28 on the Coding Agent Index v1.3; Kimi K2.7 Code takes open weights under a Modified MIT license and self-hosting. Read August 2, 2026.

Final Verdict

GPT-5.6 Terra wins this comparison, and it wins it on an exchange rate rather than on a knockout.

Kimi K2.7 Code takes more rows in our table than Terra does, and we still call it for Terra, so we owe you the reasoning rather than a shrug.

Three of Kimi's row wins — input, cached input, output — are the same axis counted three times. It is one advantage, cost, and it is a genuine one. Two more, open weights and self-hosting, are also closely related: they are the control axis. Kimi K2.7 Code's case, honestly stated, is that it is cheaper and you can own it. For a large number of teams that is decisive, and we are not going to pretend otherwise.

Terra's row wins are of a different kind. They are the rows that answer the question "does this model actually do the job?" — and both are answerable because somebody other than the vendor checked. Terra is not merely the measured one; it is the one measured ahead, by thirteen points on the index that scores both. That is a real result, and it is the hardest information on the table — with the caveat that the index rewards breadth, and Kimi K2.7 Code is built narrow.

What tips it decisively here, more than in any other closed-versus-open pairing this year, is the size of the bill. Terra is the closest a verified model gets to Kimi K2.7 Code's price. Roughly 2.1 times on input and 3.0 times on output is a premium most teams can absorb without a conversation, and cached input is close to level. You are buying proof and four times the context for a few hundred dollars a month at moderate volume. Elsewhere in the market that same proof costs ten to twelve times more. Take the deal.

So the rule we would give you. If you can measure the model yourself on your own tasks, or you need to self-host, or your volume is enormous, pick Kimi K2.7 Code — your own evaluation is worth more than thirteen points on a generalist index, and the weights are yours. Otherwise pick GPT-5.6 Terra, because the price of proof has rarely been this low.

Where our verdict is wrong: if independent results land for Kimi K2.7 Code in the coming weeks and confirm what Moonshot claims, then an open-weight model at roughly a third of Terra's input price with credible verified numbers reshapes this comparison and quite a few others. That outcome is plausible. It has simply not happened yet, and we do not publish verdicts on things that have not happened.

To see how each model fares against the rest of the field, we have Kimi K2.7 Code against Claude Opus 4.8, against GPT-5.5, and against DeepSeek V4 — the other open-weight contender, whose page is worth reading if price is your binding constraint. For the wider field, see our roundup of the best AI coding tools in 2026.

Frequently Asked Questions

Is GPT-5.6 Terra or Kimi K2.7 Code cheaper?

Kimi K2.7 Code, on every line, but by less than its reputation suggests. It charges USD 0.95 per million input tokens against Terra's USD 2.00, which is roughly 2.1 times cheaper. On output it charges USD 4 per million tokens against USD 12, roughly 3.0 times cheaper. On cached input the two are nearly level: USD 0.19 against USD 0.20. That makes GPT-5.6 Terra the closest an independently scored closed model comes to Kimi K2.7 Code on price — elsewhere in the market the same open model is ten to twelve times cheaper than the flagship it faces.

Does Kimi K2.7 Code have an Artificial Analysis Intelligence Index score?

Yes. Read on August 2, 2026, Artificial Analysis scores Kimi K2.7 Code at 42 on the v4.1 Intelligence Index, thirteen points behind GPT-5.6 Terra's 55 on the same harness. An earlier version of this page said no such score existed, and that was wrong. Be careful with articles that quote other figures for it — numbers around 44 belong to Kimi K2.6, the predecessor model, and they do not transfer. What Kimi K2.7 Code still lacks is an LMArena rating and any independently reproduced SWE-bench, Terminal-Bench, LiveCodeBench, GPQA or AIME result.

Why do you not compare the two models' coding scores directly?

Two separate reasons, and either one alone would be enough. First, they are not the same benchmark: GPT-5.6 Terra is charted on the Artificial Analysis Coding Agent Index, while Moonshot AI reports SWE-bench results for Kimi K2.7 Code. Second, and more importantly, they were not produced the same way: Terra's figure comes from an independent evaluator, and Kimi's comes from the company that built the model, running its own harness. Putting the two numbers in one row would imply a like-for-like measurement that does not exist, so you will not find them side by side anywhere on this page.

Does the absence of independent scores mean Kimi K2.7 Code is a bad model?

No, and it is important to be precise. A 42 places it thirteen points behind Terra on a composite of nine evaluations, most of which test reasoning and knowledge rather than coding — and Kimi K2.7 Code is a coding-specialized model, so the index measures it on ground it was not built for. Its architecture is serious and its self-reported coding figures are not implausible for a model of this design. The honest position is that it is measurably behind on breadth, and untested by anyone independent on the coding work it is actually sold for.

Which model won your overall verdict, and why?

GPT-5.6 Terra, and less narrowly than in most closed-versus-open matchups. Both models are independently scored, and Terra is ahead: 55 against Kimi K2.7 Code's 42 on the Artificial Analysis Intelligence Index v4.1, read on August 2, 2026, and the Artificial Analysis Coding Agent Index v1.3 charts Terra at 62.28 while carrying no Kimi K2.7 Code entry — a gap in what has been measured, not a coding verdict. What tips it is the price. Terra is roughly 2.1 times more expensive on input and 3.0 times on output — the smallest premium any independently scored flagship charges over Kimi K2.7 Code. Thirteen index points plus roughly four times the context for that surcharge is a good exchange rate. Kimi K2.7 Code still wins price, open weights, self-hosting, and architectural transparency outright.

Can I self-host either of these models?

Kimi K2.7 Code, yes. Its weights were published under a Modified MIT license on release day and can be downloaded and run on your own hardware, which makes it viable for data-residency requirements, air-gapped environments, and fixed-cost inference. GPT-5.6 Terra, no. It is a closed, API-only model from OpenAI with no downloadable weights and no self-hosting path. If self-hosting is a hard requirement, Terra is not a candidate at any price and this comparison resolves immediately in Kimi K2.7 Code's favor.

What is the context window on each model?

GPT-5.6 Terra offers a 1,050,000-token context window. Kimi K2.7 Code offers 256,000 tokens (262,144 exactly), with automatic context caching that lowers the cost of reusing repeated prefixes. Terra therefore has roughly four times the room. Whether that matters depends entirely on your work: 256,000 tokens already covers most coding tasks comfortably, and if your workflow fits inside it, Terra's extra headroom is capacity you are paying for and not using.

At what volume does Kimi K2.7 Code become the obvious choice on cost?

Run your own arithmetic, but the shape is this. A workload generating 50 million output tokens a month costs about USD 750 on GPT-5.6 Terra and about USD 200 on Kimi K2.7 Code — a gap of roughly USD 550, real but not budget-reorganizing. Multiply by ten and the gap becomes about USD 5,500 a month, at which point the cost argument gets loud, and self-hosting the open weights caps it entirely. Below moderate volume, the difference is small enough that the verified scores and the larger context are worth more than the savings.

Is Kimi K2.7 Code open source?

It is open-weight, which is not quite the same thing. Moonshot AI publishes the model weights under a Modified MIT license, and they were downloadable from day one — you can use them commercially, run them on your own hardware, and ship them inside your own product. What Moonshot does not publish is the training data or the training code. You get the finished model, not the recipe. For deployment freedom that distinction rarely matters in practice; for auditability and reproducibility, it does.

What is Kimi K2.7 Code's architecture?

It is a mixture-of-experts model with roughly 1 trillion total parameters and about 32 billion active per token. It uses 384 experts, of which 8 are selected per token plus 1 shared expert, and it employs multi-head latent attention. It also ships a native vision encoder called MoonViT. Notably, this level of architectural detail is public for Kimi K2.7 Code and has no equivalent for GPT-5.6 Terra, whose architecture OpenAI does not disclose — one of the few areas where Moonshot is clearly the more transparent of the two companies.

How does GPT-5.6 Terra compare to the other models in its family?

Terra is the balanced middle tier of the GPT-5.6 family. GPT-5.6 Sol sits above it, priced higher and scoring higher on the independent Artificial Analysis indexes; GPT-5.6 Luna sits below it as the cheaper, lighter option. Terra is the tier we would point most teams at for this particular matchup, because it is the one whose price lands close enough to Kimi K2.7 Code that the premium for independently verified numbers stops being an argument.

Will this verdict change?

Quite possibly, and we will say so plainly when it does. The verdict now rests on a measured thirteen-point gap on a generalist index, plus the absence of any independent coding result for either model. If a third party publishes a coding score that confirms Moonshot AI's self-reported figures, then an open-weight model at roughly a third of GPT-5.6 Terra's input price changes this comparison substantially. We will update this page when that lands.

Sources and references

Every figure on this page is attributed to whoever produced it. Vendor documentation and independent measurement are listed separately and never merged into a single ranking.

Our Verdict

GPT-5.6 Terra wins this comparison, on an exchange rate rather than a knockout. It is the higher-scoring of the two on the index that measures both: 55 against Kimi K2.7 Code's 42 on the independent Artificial Analysis Intelligence Index v4.1, read on August 2, 2026. On coding, the independent Artificial Analysis Coding Agent Index v1.3 charts GPT-5.6 Terra at 62.28, through the Codex harness at max reasoning effort, and carries no Kimi K2.7 Code entry; nor have we found an independent coding score for Kimi K2.7 Code anywhere else — the coding figures Moonshot AI publishes are self-reported on its own harness, unreplicated, and drawn from a different benchmark family, so they cannot be set against Terra's charted score. The thirteen-point intelligence gap is real, but it sits on a generalist composite of nine mostly non-coding evaluations, and Kimi K2.7 Code is a coding-specialized model. That does not make Kimi K2.7 Code a bad model; it makes it an unverifiable one, which is a different and temporary problem. What tips the verdict is price. Terra costs roughly 2.1 times more on input (USD 2.00 against USD 0.95) and roughly 3.0 times more on output (USD 12 against USD 4), with cached input nearly level at USD 0.20 against USD 0.19 — the smallest premium any independently scored flagship charges over Kimi K2.7 Code, where the frontier tier charges ten to twelve times more. Thirteen index points plus roughly four times the context (1,050,000 tokens against 256,000) for that surcharge is an unusually good deal. Kimi K2.7 Code still wins price on every line, open weights under a Modified MIT license, self-hosting, and architectural transparency. The rule: if you can evaluate the model yourself, or you must self-host, or your volume runs to hundreds of millions of output tokens a month, pick Kimi K2.7 Code. Otherwise pick GPT-5.6 Terra, because the price of proof has rarely been this low. If an independent coding result lands for Kimi K2.7 Code and confirms Moonshot's figures, this verdict should be revisited.

Winner:GPT-5.6 Terra

Choose GPT-5.6 Terra

OpenAI's balanced GPT-5.6 tier — GPT-5.5-competitive quality at 40 percent of the GPT-5.5 rate, with a 1.05M-token context and the full agentic toolbox.

Try GPT-5.6 Terra

Choose Kimi K2.7 Code

Moonshot AI's open-weight 1T-parameter MoE coding model — 32B active, 256K context, Modified MIT, metered at $0.95 in / $4.00 out per million tokens.

Try Kimi K2.7 Code

Frequently Asked Questions

Is GPT-5.6 Terra better than Kimi K2.7 Code?

GPT-5.6 Terra wins this comparison, on an exchange rate rather than a knockout. It is the higher-scoring of the two on the index that measures both: 55 against Kimi K2.7 Code's 42 on the independent Artificial Analysis Intelligence Index v4.1, read on August 2, 2026. On coding, the independent Artificial Analysis Coding Agent Index v1.3 charts GPT-5.6 Terra at 62.28, through the Codex harness at max reasoning effort, and carries no Kimi K2.7 Code entry; nor have we found an independent coding score for Kimi K2.7 Code anywhere else — the coding figures Moonshot AI publishes are self-reported on its own harness, unreplicated, and drawn from a different benchmark family, so they cannot be set against Terra's charted score. The thirteen-point intelligence gap is real, but it sits on a generalist composite of nine mostly non-coding evaluations, and Kimi K2.7 Code is a coding-specialized model. That does not make Kimi K2.7 Code a bad model; it makes it an unverifiable one, which is a different and temporary problem. What tips the verdict is price. Terra costs roughly 2.1 times more on input (USD 2.00 against USD 0.95) and roughly 3.0 times more on output (USD 12 against USD 4), with cached input nearly level at USD 0.20 against USD 0.19 — the smallest premium any independently scored flagship charges over Kimi K2.7 Code, where the frontier tier charges ten to twelve times more. Thirteen index points plus roughly four times the context (1,050,000 tokens against 256,000) for that surcharge is an unusually good deal. Kimi K2.7 Code still wins price on every line, open weights under a Modified MIT license, self-hosting, and architectural transparency. The rule: if you can evaluate the model yourself, or you must self-host, or your volume runs to hundreds of millions of output tokens a month, pick Kimi K2.7 Code. Otherwise pick GPT-5.6 Terra, because the price of proof has rarely been this low. If an independent coding result lands for Kimi K2.7 Code and confirms Moonshot's figures, this verdict should be revisited.

Which is cheaper, GPT-5.6 Terra or Kimi K2.7 Code?

GPT-5.6 Terra starts at $2 in / $12 out per M tokens. Kimi K2.7 Code starts at $0.95 in / $4 out per M tokens (free plan available). Check the pricing comparison section above for a full breakdown.

What are the main differences between GPT-5.6 Terra and Kimi K2.7 Code?

The key differences span across 11 features we compared. For Independent intelligence score (Artificial Analysis Intelligence Index), GPT-5.6 Terra offers 55 (independent) while Kimi K2.7 Code offers 42 on the v4.1 index, read August 2, 2026 — thirteen points behind. Figures circulating for Kimi K2.6 belong to a different model and do not transfer. For Independent coding score (Artificial Analysis Coding Agent Index v1.3, read August 2, 2026), GPT-5.6 Terra offers 62.28 (Codex harness, max reasoning effort) while Kimi K2.7 Code offers Not charted. For Evidence regime behind the coding claim, GPT-5.6 Terra offers Independently charted: 62.28 on the AA Coding Agent Index v1.3 (Codex harness, max reasoning effort) while Kimi K2.7 Code offers SWE-bench Verified 60.4 percent and SWE-bench Pro 58.6, self-reported by Moonshot AI on its own harness and not reproduced by any third party. See the full feature comparison table above for all details.

Related Comparisons