Claude Fable 5 vs Kimi K2.7: Verified Proof vs Unbeatable Price (2026)
Fable 5 vs Kimi K2.7: a 95% coding score checked by a third party against 60.4% Moonshot checked itself. Kimi is 10x cheaper. Why proof still takes it.
Feature Comparison
| Feature | Claude Fable 5 | Kimi K2.7 |
|---|---|---|
| Independent intelligence score (Artificial Analysis) | 60 — the highest score of any model on the index (independent) | None. Not yet on the independent leaderboard (too new). The score of 54 belongs to Kimi K2.6, a different model, and does not transfer |
| Independently verified coding result | SWE-bench Verified 95 percent, measured by vals.ai (independent) | None. Its SWE-bench Verified 60.4 percent and SWE-bench Pro 58.6 are self-reported by Moonshot AI and have not been reproduced by any third party |
| Human preference leaderboard (LMArena) | 1509 (independent) | Not yet ranked (too new) |
| Maximum context window | 1,000,000 tokens | 256,000 tokens (262,144), with automatic context caching |
| Input price (per million tokens) | USD 10.00 | USD 0.95 — roughly 10.5 times cheaper |
| Cached input price (per million tokens) | USD 1.00 | USD 0.19 — roughly 5.3 times cheaper |
| Output price (per million tokens) | USD 50.00 | USD 4.00 — roughly 12.5 times cheaper |
| Model weights and licensing | Closed, proprietary; API only, no weights released | Open weights under a Modified MIT license, downloadable from day one |
| Self-hosting and data residency | Not possible — managed API only | Possible — run the weights on your own hardware, in your own jurisdiction |
Pricing Comparison
Claude Fable 5
Kimi K2.7
Detailed Comparison
Claude Fable 5 vs Kimi K2.7 in 2026: Claude Fable 5 is Anthropic's premium flagship, priced at USD 10 per million input tokens and USD 50 per million output tokens, with a 1,000,000-token context. It scores 60 on the independent Artificial Analysis Intelligence Index (the highest of any model), 1509 on LMArena, and 95 percent on SWE-bench Verified as measured independently by vals.ai. Kimi K2.7 is Moonshot AI's open-weight agentic coding model, priced at USD 0.95 per million input tokens and USD 4 per million output tokens, with a 256,000-token context and downloadable weights under a Modified MIT license. Kimi K2.7 is not yet on any independent leaderboard: its SWE-bench Verified figure of 60.4 percent is self-reported by Moonshot AI and has not been reproduced by a third party. Fable 5 costs roughly 10.5 times more on input and 12.5 times more on output. It wins this comparison on verified evidence — not on value.
Quick Verdict
Both of these models publish a SWE-bench Verified score. That coincidence is the single most misleading thing about this matchup, and it is the reason this comparison exists.
Claude Fable 5's 95 percent on SWE-bench Verified was measured by vals.ai, an independent third party. Kimi K2.7's 60.4 percent on SWE-bench Verified was measured by Moonshot AI, the company that built it. Same benchmark name. Completely different evidentiary regimes. You will see those two numbers set against each other all over the internet this month, and every time you do, the comparison is broken — because one of them has been checked and the other has only been claimed.
So we are not going to put them in the same row. Not in the table below, not in the infographic, not in a sentence. What we are going to do instead is treat that asymmetry as the actual decision you are making, because it is.
Claude Fable 5 takes the overall verdict — narrowly, and on one axis only: it is the only model here whose capability claims anyone outside the vendor has verified. It also happens to lead every independent leaderboard that exists. That is a real thing to buy, and right now Kimi K2.7 cannot offer it at any price.
But read the next line carefully, because it matters as much as the verdict:
Kimi K2.7 wins more rows in our comparison table than Fable 5 does. It is roughly 10.5 times cheaper on input, roughly 12.5 times cheaper on output, its weights are downloadable and self-hostable, and its architecture is published in full. If your constraint is cost, data residency, or control, Kimi K2.7 is the correct pick and Fable 5 cannot come close. We explain exactly how we weighted this — and the conditions under which our call is wrong — further down.
- Fable 5 wins verified capability. 60 on the Artificial Analysis Intelligence Index and 1509 on LMArena, both independent. Kimi K2.7 appears on neither.
- Fable 5 wins verified coding. 95 percent on SWE-bench Verified, confirmed by vals.ai. Kimi's coding case rests entirely on its own harness.
- Fable 5 wins context. 1,000,000 tokens against 256,000 — roughly four times the room.
- Kimi K2.7 wins price, decisively. USD 0.95 against USD 10 on input; USD 4 against USD 50 on output. This is not a close call.
- Kimi K2.7 wins control. Open weights under a Modified MIT license, downloadable from day one, self-hostable on your own hardware.
- Kimi K2.7 wins transparency of design. Moonshot publishes the full architecture. Anthropic publishes nothing comparable.
How We Compared Them, and the Rule We Hold To
We ran both models side by side on their respective APIs — the same refactoring tasks, the same long-document work, the same agentic tool-calling loops — to get a feel for how each one behaves in practice. That hands-on time informs the judgment calls in this piece: how each model handles a long context, how it recovers when a tool call fails, how much hand-holding it needs.
What it does not do is produce benchmark numbers. We do not publish our own scores, because a handful of prompts run by one team is not a benchmark. Every number in this comparison comes either from a named independent third party or from the vendor, and we label which is which, every single time.
That labeling rule is usually a footnote. In this comparison it is the whole story, so here it is explicitly:
- Independent means measured by a third party with no stake in the outcome — Artificial Analysis, LMArena, or vals.ai. These numbers are comparable across models, because the same harness ran them all under the same conditions.
- Vendor self-reported means the company that built the model ran the benchmark itself, on its own harness, and published the result. This is useful signal. It is not verification. It is not comparable to an independent score, and it is not reliably comparable to another vendor's self-reported score either, because no two vendor harnesses are the same.
We never stack the two in a single row. A table that puts "95%" in one column and "60.4%" in the next, under a shared heading of "SWE-bench Verified," is telling you a lie by layout. It implies the two numbers were produced the same way. They were not.
Claude Fable 5 and Kimi K2.7 at a Glance
Claude Fable 5 is Anthropic's premium flagship — the most capable model the company sells, and priced accordingly at USD 10 per million input tokens and USD 50 per million output tokens, with cached input at USD 1 per million. It carries a 1,000,000-token context window. It is a closed model: there are no weights to download, no self-hosting option, and no published architecture. What Anthropic offers instead of transparency is verification. Fable 5 sits at the top of the independent Artificial Analysis Intelligence Index with a score of 60, holds an LMArena rating of 1509, and its 95 percent on SWE-bench Verified was measured by vals.ai rather than by Anthropic. On every independent scoreboard that exists, it is first.
Kimi K2.7 is Moonshot AI's open-weight agentic coding model, released on June 12, 2026 from the company's Beijing headquarters. It is a mixture-of-experts architecture: roughly 1 trillion total parameters with about 32 billion active per token, spread across 384 experts of which 8 are selected per token plus 1 shared, using multi-head latent attention and a MoonViT vision encoder for native image input. The weights shipped under a Modified MIT license on day one — you can download them, run them on your own hardware, and put them inside your own product. It offers a 256,000-token context window (262,144 tokens exactly) with automatic context caching, priced at USD 0.95 per million input tokens, USD 0.19 cached, and USD 4 per million output tokens. There is no subscription tier; it is metered only.
Read those two paragraphs again and notice what is missing from each. Anthropic will not tell you how Fable 5 is built. Moonshot will not show you an independent score. Each company is opaque about exactly the thing the other is open about — and which opacity you can live with is, in the end, this entire comparison.
The Part That Matters: One of These Numbers Has Been Checked
Here is the situation in plain terms.
Claude Fable 5, independently measured:
- Artificial Analysis Intelligence Index: 60 — the highest score of any model on the index (independent).
- LMArena: 1509 (independent).
- SWE-bench Verified: 95 percent — measured by vals.ai, an independent evaluator (independent).
Kimi K2.7, independently measured:
- Artificial Analysis Intelligence Index: not yet on the independent leaderboard (too new).
- LMArena: not yet ranked.
- SWE-bench Verified, SWE-bench Pro, Terminal-Bench, LiveCodeBench, GPQA, AIME: no independent third-party results exist as of June 15, 2026.
That second list is not a rhetorical device. It is the complete state of independent evidence for Kimi K2.7: there is none. Not a low score — no score.
What Moonshot AI has published, on its own harness, is a SWE-bench Pro result of 58.6 and a SWE-bench Verified result of 60.4 percent, which the company describes as a new high-water mark for open-source models. Those figures may well be accurate. Moonshot has a decent track record, and the prior model in the line was independently evaluated in due course. But as of this writing, nobody outside Moonshot has reproduced them.
One clarification that trips up a lot of coverage: Kimi K2.6, the predecessor, does carry an independent Artificial Analysis Intelligence score — 44 on the current v4.1 index. A figure of 54 also circulates for K2.6, but it comes from an earlier version of that index and no longer reflects the current methodology. Either way, K2.7 is a different model: no score transfers to it, and any article that quietly reuses one is inventing a data point. We are not going to do that, so where an independent number for K2.7 would go, this page says the honest thing: it does not exist yet.
Why an unverified benchmark is not the same as a bad one
It would be lazy — and wrong — to read all of this as "Kimi K2.7 is worse." That is not what the evidence says. The evidence says we cannot tell yet, and those are very different claims.
The honest position on Kimi K2.7 today is: it might be excellent. Its architecture is serious, its price is extraordinary, its predecessor scored respectably when independently tested, and its self-reported numbers are not outlandish for a model of this design. Nothing here suggests Moonshot is exaggerating.
What it means is that buying Kimi K2.7 today is buying on the vendor's word, while buying Claude Fable 5 is buying on someone else's proof. For a hobbyist, that distinction is close to meaningless — run it, see if you like it, the price makes the experiment nearly free. For an engineering lead who has to stand in front of a room and justify why a model was chosen for a production system, it is the whole ballgame.
And it is temporary. Independent evaluators typically get to a model of this profile within weeks of release, not months. Kimi K2.7 will almost certainly appear on the Artificial Analysis Intelligence Index and on LMArena in the near future, and when it does, this section becomes obsolete and this verdict should be revisited. If the independent numbers land close to Moonshot's self-reported figures, the case for Kimi at a tenth of the price gets very strong very quickly. We will update this page when that happens.
Pricing: Not Close, Not Even Slightly
Where the benchmark picture is murky, the pricing picture is brutally clear. These are the published metered rates, per million tokens:
| Metered rate (per million tokens) | Claude Fable 5 | Kimi K2.7 |
|---|---|---|
| Input | USD 10.00 | USD 0.95 |
| Cached input | USD 1.00 | USD 0.19 |
| Output | USD 50.00 | USD 4.00 |
Kimi K2.7 is roughly 10.5 times cheaper on input, roughly 5.3 times cheaper on cached input, and roughly 12.5 times cheaper on output. Output is where the gap bites hardest, and output is what agentic coding workloads generate in volume — long diffs, long tool-call chains, long reasoning traces.
Put a number on it. A workload burning 50 million output tokens in a month — not an extreme figure for a team running coding agents continuously — costs USD 2,500 on Fable 5 and USD 200 on Kimi K2.7. That is a difference of USD 2,300 every month, on output alone, before you count a single input token.
There is no framing of that gap that makes it small. Anyone telling you Fable 5's price is justified "because quality" is skipping the part where they have to show that the quality difference is worth twelve and a half times the money for your specific workload. For a lot of workloads, it plainly is not. Fable 5 is a premium product at a premium price, and premium pricing has to be earned per use case, not asserted.
Context Window: 1M Against 256K
Fable 5 carries a 1,000,000-token context window. Kimi K2.7 carries 256,000 tokens (262,144 exactly), with automatic context caching that reduces the cost of repeated prefixes without you managing it.
Four times the context is a real advantage, but be honest about when it actually binds. A 256,000-token window is already large enough for the overwhelming majority of coding work: a substantial repository slice, a long design document plus the code it describes, a full day of agent scrollback. If your workflow fits comfortably in 256,000 tokens today, Fable 5's extra headroom is a feature you are paying for and not using.
Where it does bind: whole-monorepo reasoning, very long agentic sessions that accumulate hundreds of tool calls without compaction, and document work at the scale of full legal or research corpora. If that is your work, the gap is decisive, and Kimi's automatic caching does not close it — caching makes a window cheaper to reuse, it does not make it bigger.
Open Weights, Self-Hosting, and What Anthropic Will Not Sell You
Kimi K2.7's weights are published under a Modified MIT license and were downloadable on day one. This is not a marketing detail. It changes what you are legally and operationally able to do:
- Run it on your own hardware. No inference data leaves your infrastructure. For teams with hard data-residency requirements, this is the difference between "usable" and "not usable," and no amount of vendor assurance substitutes for it.
- Fix your cost ceiling. Self-hosted inference converts a metered bill into a hardware amortization. At high volume, that changes the economics entirely.
- Never get deprecated out from under you. A downloaded weight file does not get sunset on a vendor's schedule. Anyone who has had a model retired mid-product knows exactly what this is worth.
- Inspect and modify. Full architecture disclosure — mixture-of-experts, 1 trillion total parameters with 32 billion active, 384 experts, multi-head latent attention, the MoonViT vision encoder — means you can reason about the model's behavior instead of guessing at it.
Claude Fable 5 offers exactly none of this, and Anthropic does not pretend otherwise. It is a closed, API-only model with an undisclosed architecture. If open weights are a requirement rather than a preference, this comparison is over before it starts and Kimi K2.7 wins by default — Fable 5 is not a candidate at any price.
One honest caveat on the openness point, because it cuts both ways: open weights under a Modified MIT license is a genuine grant, but open-weight is not the same as open-source. Moonshot publishes the finished model, not the training data or the training code. You get the artifact, not the recipe. For deployment freedom that distinction rarely matters; for auditability and reproducibility, it does.
And a second caveat that matters for the vision claim: Kimi K2.7 ships a native vision encoder, which Fable 5's published specification does not describe in comparable terms. We are not going to score that as a clean win for either side, because we would be comparing a disclosed capability against an undisclosed one — which is the same mistake as stacking an independent score against a vendor one.
Winner by Category
| Category | Winner | Why |
|---|---|---|
| Best independently verified intelligence | Claude Fable 5 | Artificial Analysis Intelligence Index of 60, the highest on the board. Kimi K2.7 is not on it. |
| Best independently verified coding | Claude Fable 5 | 95 percent on SWE-bench Verified, measured by vals.ai. Kimi's coding figures are vendor self-reported. |
| Best cost per token | Kimi K2.7 | Roughly 10.5 times cheaper on input, roughly 12.5 times cheaper on output. Not a close call. |
| Best for self-hosting and data residency | Kimi K2.7 | Open weights, Modified MIT, downloadable. Fable 5 cannot do this at all. |
| Best for very long context | Claude Fable 5 | 1,000,000 tokens against 256,000. |
| Best for a high-volume agentic coding fleet | Kimi K2.7 | Built for agentic coding, and the output pricing is the only thing that makes running agents continuously affordable. |
| Best when you must justify the choice to a review board | Claude Fable 5 | Third-party verified numbers exist. For Kimi K2.7, there is nothing external to cite yet. |
| Best architectural transparency | Kimi K2.7 | Full architecture published. Anthropic discloses nothing comparable. |
Pros and Cons
Claude Fable 5
Pros
- The only model in this comparison with independently verified capability — Artificial Analysis Intelligence Index 60, LMArena 1509, SWE-bench Verified 95 percent via vals.ai.
- Highest score on the independent intelligence index of any model currently measured.
- 1,000,000-token context window, roughly four times Kimi K2.7's.
- Managed API — no weights to serve, no hardware to maintain, no inference stack to operate.
- Cached input at USD 1 per million tokens materially softens the cost of repeated prefixes.
Cons
- Extremely expensive: USD 10 per million input tokens and USD 50 per million output tokens, roughly 10.5 and 12.5 times Kimi K2.7's rates.
- Closed weights. No self-hosting, no data residency guarantee through hardware control, no protection against deprecation.
- Undisclosed architecture — you cannot inspect or reason about how it works.
- Output pricing makes continuous agentic workloads genuinely painful to run at scale.
Kimi K2.7
Pros
- Extraordinary pricing: USD 0.95 per million input tokens, USD 0.19 cached, USD 4 per million output tokens.
- Open weights under a Modified MIT license, downloadable from day one and self-hostable.
- Full architecture disclosure — mixture-of-experts, 1 trillion total parameters with 32 billion active, 384 experts, multi-head latent attention.
- Native vision through the MoonViT encoder.
- Purpose-built for agentic coding, with automatic context caching.
Cons
- No independent verification of any kind. Not on the Artificial Analysis Intelligence Index, not on LMArena, no third-party coding result.
- Its SWE-bench figures — 60.4 percent Verified and 58.6 on Pro — are self-reported by Moonshot AI and have not been reproduced.
- 256,000-token context, roughly a quarter of Fable 5's.
- Metered billing only — there is no flat-rate subscription tier.
- The hosted API is operated from Beijing, which carries its own compliance considerations — self-hosting the weights avoids this, using the hosted API does not.
When to Pick Claude Fable 5
- You have to justify the choice to someone. A review board, a client, a regulator, a CTO who asks "how do you know it's good?" Fable 5 is the only one of the two where the answer is a third-party number rather than a vendor's press release.
- Your work needs the top of the market and you can measure the difference. If your evaluations show the hardest 5 percent of your tasks failing on cheaper models, the independently verified 95 percent on SWE-bench Verified is exactly what you are paying for.
- You routinely exceed 256,000 tokens of context. Whole-monorepo work, very long agentic sessions, large document corpora.
- Your token volume is low enough that the price gap is noise. If you spend USD 40 a month on inference, a 12.5 times multiplier on output is an argument about USD 500, not USD 5,000 — take the better model.
When to Pick Kimi K2.7
- Your volume is high and your budget is real. At 50 million output tokens a month, Kimi K2.7 costs USD 200 where Fable 5 costs USD 2,500. That difference funds an engineer.
- You need to self-host. Data residency, air-gapped environments, regulated industries. Fable 5 simply cannot do this, so the comparison ends here.
- You can run your own evaluations. This is the key one. The missing independent scores hurt far less if you have an internal evaluation harness on your own tasks — because then you are not taking Moonshot's word for anything, you are measuring it yourself on the only benchmark that actually matters, which is yours.
- You are running agentic coding fleets continuously. The model is designed for it and, more importantly, the output price is the only thing that makes it affordable to leave running.
- You want protection from deprecation. A downloaded weight file is yours. An API endpoint is not.
Final Verdict
Claude Fable 5 wins this comparison. It wins it on proof, and on nothing else.
You will have noticed that Kimi K2.7 takes more rows in our comparison table than Fable 5 does. We still call it for Fable 5, and we owe you the reasoning rather than a shrug.
Three of Kimi's row wins — input price, cached input price, output price — are the same axis counted three times. It is one advantage, cost, and it is an enormous one. Two more, open weights and self-hosting, are also closely related: they are the control axis. So Kimi K2.7's case, honestly stated, is: it is far cheaper and you can own it. That is a strong case and for many teams it is the deciding one.
Fable 5's row wins are of a different kind. They are the rows that answer the question "does this model actually do the job?" — and three of the four are only answerable because somebody other than the vendor checked. Against a competitor that no independent evaluator has looked at yet, that is not a tie-breaker. It is the only hard information on the table.
So the rule we would give you is this. If you can measure the model yourself on your own tasks, pick Kimi K2.7 — the price is not close, the weights are yours, and your own evaluation replaces the missing independent one. If you cannot measure it yourself, pick Claude Fable 5, because then you are choosing between a number a third party verified and a number a vendor asserted, and that is not really a choice.
Where our verdict is wrong: if independent results land for Kimi K2.7 in the coming weeks and they confirm Moonshot's self-reported figures, then a model at a tenth of the price with credible verified numbers reshapes this entire comparison, and quite a few others. That outcome is plausible. It has simply not happened yet, and we do not publish verdicts on things that have not happened.
If you want to see how each of these models fares against the rest of the field, we have Fable 5 against Claude Opus 4.8, against GPT-5.5, and against DeepSeek V4; and Kimi K2.7 against Claude Opus 4.8 and against DeepSeek V4. For the wider field, see our roundup of the best AI coding tools in 2026.
Frequently Asked Questions
Can I compare Claude Fable 5's 95 percent to Kimi K2.7's 60.4 percent on SWE-bench Verified?
No. They are the same benchmark name produced under completely different evidentiary regimes, and putting them side by side is misleading. Claude Fable 5's 95 percent was measured by vals.ai, an independent third party. Kimi K2.7's 60.4 percent was measured by Moonshot AI, the company that built the model, on its own harness, and no third party has reproduced it. One number has been verified; the other has been claimed. Setting them against each other implies a like-for-like measurement that does not exist, which is why you will not find those two figures in the same table row anywhere on this page.
Does Kimi K2.7 have an Artificial Analysis Intelligence Index score?
No. As of June 15, 2026, Kimi K2.7 is not yet on the independent leaderboard (too new). There are no third-party results for it on the Artificial Analysis Intelligence Index, LMArena, SWE-bench Verified, SWE-bench Pro, Terminal-Bench, LiveCodeBench, GPQA, or AIME. Be careful with articles that quote a score of 54 for it — that figure is an older-index score for Kimi K2.6, the predecessor model, which sits at 44 on the current v4.1 index, and no score transfers to K2.7 either way. Claude Fable 5, by contrast, scores 60 on the Artificial Analysis Intelligence Index, the highest of any model measured.
Does the absence of independent scores mean Kimi K2.7 is a bad model?
No, and it is important to be precise here. It means Kimi K2.7 is currently unverifiable, not that it is weak. Its architecture is serious, its self-reported figures are not implausible for a model of its design, and its predecessor performed respectably when independently tested. The honest position is that we cannot yet tell how good it is from outside evidence. Independent evaluators usually reach models of this profile within weeks, so this is likely a temporary state, and we will revisit this comparison when third-party results appear.
How much cheaper is Kimi K2.7 than Claude Fable 5?
Substantially. Kimi K2.7 charges USD 0.95 per million input tokens against Fable 5's USD 10, which is roughly 10.5 times cheaper. On output it charges USD 4 per million tokens against USD 50, which is roughly 12.5 times cheaper. Cached input is USD 0.19 against USD 1, roughly 5.3 times cheaper. To make that concrete: a workload generating 50 million output tokens per month costs about USD 200 on Kimi K2.7 and about USD 2,500 on Claude Fable 5.
Which model won your overall verdict, and why?
Claude Fable 5, narrowly, and specifically on verified evidence rather than on value. It is the only model of the two whose capability has been confirmed by third parties: 60 on the Artificial Analysis Intelligence Index, 1509 on LMArena, and 95 percent on SWE-bench Verified via vals.ai. Kimi K2.7 actually wins more rows in our comparison table than Fable 5 does, and it wins them on price, open weights, self-hosting, and architectural transparency. If your constraint is cost or control, Kimi K2.7 is the correct choice and Fable 5 cannot compete on those axes at all.
Can I self-host either of these models?
Kimi K2.7, yes. Its weights were published under a Modified MIT license on release day and can be downloaded and run on your own hardware, which makes it viable for data-residency requirements, air-gapped environments, and fixed-cost inference. Claude Fable 5, no. It is a closed, API-only model from Anthropic with no downloadable weights and no self-hosting path. If self-hosting is a hard requirement for you, Fable 5 is not a candidate at any price and this comparison resolves immediately in Kimi K2.7's favor.
What is the context window on each model?
Claude Fable 5 offers a 1,000,000-token context window. Kimi K2.7 offers 256,000 tokens (262,144 exactly), with automatic context caching that lowers the cost of reusing repeated prefixes. Fable 5 therefore has roughly four times the room. Whether that matters depends entirely on your work: 256,000 tokens already covers most coding tasks comfortably, and if your workflow fits inside it, Fable 5's extra headroom is capacity you are paying for and not using.
Is Kimi K2.7 open source?
It is open-weight, which is not quite the same thing. Moonshot AI publishes the model weights under a Modified MIT license, and they were downloadable from day one — you can use them commercially, run them on your own hardware, and ship them inside your own product. What Moonshot does not publish is the training data or the training code. You get the finished model, not the recipe. For deployment freedom that distinction rarely matters in practice; for auditability and reproducibility, it does.
What is Kimi K2.7's architecture?
It is a mixture-of-experts model with roughly 1 trillion total parameters and about 32 billion active per token. It uses 384 experts, of which 8 are selected per token plus 1 shared expert, and it employs multi-head latent attention. It also ships a native vision encoder called MoonViT. Notably, this level of architectural detail is public for Kimi K2.7 and has no equivalent for Claude Fable 5, whose architecture Anthropic does not disclose — which is one of the few areas where Moonshot is clearly the more transparent of the two companies.
Which model is better for high-volume agentic coding?
Kimi K2.7, for most teams, on economics. It is purpose-built for agentic coding, and agentic workloads generate output tokens in enormous volume — long diffs, long tool-call chains, long reasoning traces. At USD 4 per million output tokens against Fable 5's USD 50, Kimi K2.7 is often the only one of the two that is affordable to leave running continuously. The caveat is that its coding quality is not independently verified, so you should validate it on your own tasks before committing a fleet to it.
If I can run my own evaluations, does that change the recommendation?
Yes, and it is the single biggest factor. The case for Claude Fable 5 rests on third-party verification standing in for evidence you do not have. If you have an internal evaluation harness that measures both models on your actual tasks, you no longer need that substitute — you are measuring the thing directly, on the only benchmark that truly matters, which is yours. In that situation the missing independent scores hurt far less, and Kimi K2.7's roughly 10.5 times cheaper input and 12.5 times cheaper output become very hard to argue against.
Will this verdict change?
Quite possibly, and we will say so plainly when it does. The verdict rests on Kimi K2.7 having no independent verification as of June 15, 2026. Independent evaluators typically reach models of this profile within weeks of release, so Kimi K2.7 is likely to appear on the Artificial Analysis Intelligence Index and LMArena before long. If those results confirm Moonshot AI's self-reported figures, then a model priced at roughly a tenth of Claude Fable 5 with credible verified numbers changes this comparison substantially. We will update this page when independent results land.
Our Verdict
Claude Fable 5 wins this comparison, narrowly, and on verified evidence rather than on value. It is the only model of the two whose capability has been confirmed by anyone other than its maker: 60 on the independent Artificial Analysis Intelligence Index (the highest of any model), 1509 on LMArena, and 95 percent on SWE-bench Verified as measured independently by vals.ai. Kimi K2.7 has no independent score of any kind as of June 15, 2026 — its SWE-bench Verified figure of 60.4 percent is self-reported by Moonshot AI and unreplicated, so it cannot be set against Fable 5's verified 95 percent. That does not make Kimi K2.7 a bad model; it makes it an unverifiable one, which is a different and temporary problem. Kimi K2.7 wins more rows than Fable 5 here, and wins them decisively: roughly 10.5 times cheaper on input (USD 0.95 against USD 10), roughly 12.5 times cheaper on output (USD 4 against USD 50), open weights under a Modified MIT license, self-hostable, and fully documented architecturally. The rule: if you can evaluate the model yourself on your own tasks, pick Kimi K2.7 — your own measurement replaces the missing independent one and the price is not close. If you cannot, pick Claude Fable 5, because you are then choosing between a number a third party verified and a number a vendor asserted. If independent results land for Kimi K2.7 and confirm Moonshot's figures, this verdict should be revisited.
Choose Claude Fable 5
Anthropic's most capable widely released model — the public, safety-classified Mythos-class frontier tier.
Try Claude Fable 5 →Choose Kimi K2.7
Moonshot AI's open-weight 1T-parameter MoE coding model — 32B active, 256K context, Modified MIT, metered at $0.95 in / $4.00 out per million tokens.
Try Kimi K2.7 →Frequently Asked Questions
Is Claude Fable 5 better than Kimi K2.7?
Claude Fable 5 wins this comparison, narrowly, and on verified evidence rather than on value. It is the only model of the two whose capability has been confirmed by anyone other than its maker: 60 on the independent Artificial Analysis Intelligence Index (the highest of any model), 1509 on LMArena, and 95 percent on SWE-bench Verified as measured independently by vals.ai. Kimi K2.7 has no independent score of any kind as of June 15, 2026 — its SWE-bench Verified figure of 60.4 percent is self-reported by Moonshot AI and unreplicated, so it cannot be set against Fable 5's verified 95 percent. That does not make Kimi K2.7 a bad model; it makes it an unverifiable one, which is a different and temporary problem. Kimi K2.7 wins more rows than Fable 5 here, and wins them decisively: roughly 10.5 times cheaper on input (USD 0.95 against USD 10), roughly 12.5 times cheaper on output (USD 4 against USD 50), open weights under a Modified MIT license, self-hostable, and fully documented architecturally. The rule: if you can evaluate the model yourself on your own tasks, pick Kimi K2.7 — your own measurement replaces the missing independent one and the price is not close. If you cannot, pick Claude Fable 5, because you are then choosing between a number a third party verified and a number a vendor asserted. If independent results land for Kimi K2.7 and confirm Moonshot's figures, this verdict should be revisited.
Which is cheaper, Claude Fable 5 or Kimi K2.7?
Claude Fable 5 is priced at $10 in / $50 out per M tokens. Kimi K2.7 is priced at $0.95 in / $4 out per M tokens (free plan available). Check the pricing comparison section above for a full breakdown.
What are the main differences between Claude Fable 5 and Kimi K2.7?
The key differences span across 9 features we compared. For Independent intelligence score (Artificial Analysis), Claude Fable 5 offers 60 — the highest score of any model on the index (independent) while Kimi K2.7 offers None. Not yet on the independent leaderboard (too new). The score of 54 belongs to Kimi K2.6, a different model, and does not transfer. For Independently verified coding result, Claude Fable 5 offers SWE-bench Verified 95 percent, measured by vals.ai (independent) while Kimi K2.7 offers None. Its SWE-bench Verified 60.4 percent and SWE-bench Pro 58.6 are self-reported by Moonshot AI and have not been reproduced by any third party. For Human preference leaderboard (LMArena), Claude Fable 5 offers 1509 (independent) while Kimi K2.7 offers Not yet ranked (too new). See the full feature comparison table above for all details.

