Benchmarks say one thing, people using the models say another. We pull both — plus real market pricing and measured speed — and score 207 models on one scale. 113 of them carry a machine score and a human score, which is where the disagreements show up.
24 of these models have a hands-on review on this site.
What feeds the index
OpenRouter207 models
live market pricing
Artificial Analysis196 models
benchmark indices
LMArena124 models
human preference votes
113 models scored by machines and humans
Rank by
The current top 20
Benchmark indices cross-checked against human preference votes. The bar under each score splits it into the five things we weigh — same total, different shape means a different model for a different job.
The raw measurements come from third parties, credited below. What we add is the cross-source aggregation and the scoring: five axes on absolute 0-100 scales, weighted into one number. Absolute scales matter — a model scoring 78 today means the same as a model scoring 78 six months ago, so the history stays comparable as new models arrive.
Intelligence
40%
Average of every published capability index we can cross-check — general intelligence, coding and agentic benchmarks.
Price / performance
20%
Real market price per million tokens weighted against measured capability. A cheap weak model and an expensive strong one can score the same.
Capabilities
15%
Context window, maximum output, modalities, reasoning support, prompt caching and how recent the knowledge cutoff is.
Speed
10%
Measured output throughput (tokens per second) and time to first token. A model you wait ten seconds for is a different product from one that answers instantly.
Ecosystem
15%
How many providers serve the model, and whether the weights are open — availability and independence from a single vendor.
What changed recently
Every collection is archived, so price cuts and ranking moves stay traceable. The sources publish today's numbers; this is the trail behind them.
The default ranking is capability, so the most expensive model can sit above a free one. Switch the criterion at the top of the page and the order changes completely — that is the point. Four questions come up constantly, and here are the answers.
Why is a $10 model ahead of a free one?
Because the default ranking measures what a model can do, not what it costs. Price is a separate axis, and it has its own ranking one click away. Ranking on the composite score instead buried the frontier models behind cheap ones and turned the whole index into a price chart, which is not what a ranked list of models is for. Today that puts Claude Opus 5 first with 90.0 on capability at $25.00 per million output tokens.
What happens when benchmarks and humans disagree?
Both count, equally. Benchmarks test verifiable problems; the arena records which answer people actually preferred. When a model carries both, each contributes half its intelligence score. The widest gap in the index right now is Qwen: Qwen3 235B A22B Thinking 2507: 19 from benchmarks against 69 from human votes — a 50-point disagreement that a single-source ranking would hide entirely.
What is the coloured bar under each score?
The score itself, split into the five things we weigh. Each segment is that axis's contribution, so the full bar length is the composite score. Two models can share a total and have opposite shapes — one fast and cheap, the other slow and brilliant.
Intelligence
Price / performance
Capabilities
Speed
Ecosystem
Why are some cells empty?
Because that measurement does not exist yet, and we would rather show a dash than invent a number. Of the 207 models listed, 196 carry benchmark indices and 124 carry human ratings. A missing value never counts as zero — it simply leaves that axis out of the model's score, and the “Double-validated” filter isolates the 113 models measured both ways.
Frequently asked questions
What is the ThePlanetTools AI Index?+
A ranking of 207 large language models built by cross-referencing three independent sources: live market pricing from OpenRouter, benchmark indices from Artificial Analysis, and human preference votes from LMArena. We aggregate those measurements and compute our own score across five weighted axes. The raw measurements belong to their sources; the scoring and the ranking are ours.
How is the ranking ordered by default?+
By measured capability, not by price. A model that costs ten times more but answers harder questions ranks higher by default. Three alternative orderings are one click away — price/performance, speed, and the weighted composite score — because "best" means something different depending on whether you are paying the bill or waiting for the response.
What does the intelligence score actually measure?+
Two things, weighted equally when both exist. Machine benchmarks contribute the general intelligence index (50%), the coding index (25%) and the agentic index (25%), each normalised against its own ceiling. Human preference contributes the LMArena Elo rating, mapped onto the same 0-100 scale. A model measured by only one of the two is scored on that one alone.
What does "double-validated" mean?+
A model carrying both a machine benchmark score and a human preference rating. 113 of the 207 models in the index qualify. It matters because the two methods disagree more often than you would expect: benchmarks measure what is testable, the arena measures what people actually prefer after using the model.
Why do benchmark scores and human votes disagree?+
They measure different things. Benchmarks test verifiable problems with known answers — maths, code that compiles, factual recall. Human votes reward tone, instruction-following, formatting and refusal behaviour, which no benchmark captures. A model can top the benchmarks and lose blind head-to-head votes, and the reverse happens just as often.
Why is the maths index excluded from the score?+
Because nearly every current model scores in the 90s on it, so it separates nothing. Worse, averaging it in rewarded models that had been measured on maths and almost nothing else — one model with no coding or agentic score reached a higher intelligence score than models measured on everything. The maths index is still displayed as a raw measurement on each model.
How is price/performance calculated?+
Measured capability divided by the blended price per million tokens, weighted one part input to three parts output, which reflects typical real-world usage. The result is placed on a logarithmic scale because prices span four orders of magnitude. Free or unpriced models receive a neutral 50 rather than a perfect score, since they are not comparable.
Are the prices real?+
Yes. They come from OpenRouter, which publishes live per-token rates charged by the providers actually serving each model, in US dollars per million tokens. Input and output are shown separately because the gap between them is often large and matters more than the headline figure.
How often is the index updated?+
Every 48 hours. Each collection re-fetches all three sources, recalculates every score, and writes a historical snapshot so price changes and ranking movements stay traceable. Models that disappear from the sources are marked inactive rather than deleted, so their history survives.
What does "open weights" mean in this index?+
A model whose weights can be downloaded and self-hosted, as opposed to one reachable only through a vendor API. 111 of the 207 models qualify. Open weights add points on the ecosystem axis because they remove vendor lock-in, but a closed model served by many providers is not penalised to zero.
Why do the scores use absolute scales instead of ranking against each other?+
So the numbers stay comparable over time. If scores were normalised against the current field, every new model release would silently move every existing score, and a snapshot from six months ago would mean nothing. On an absolute scale, a model scoring 78 today means the same as a model scoring 78 last year.
Which models here have been reviewed hands-on?+
24 of the 207 models link to a full review on this site, written after actually using the model. Those links appear directly on each model card and in the comparison table. The rest are scored from published measurements only, which is stated rather than hidden.
Data sources
OpenRouter — live market pricing, context windows, modalities and availability across providers.
Artificial Analysis — independent intelligence, coding, agentic and math indices, plus measured throughput and time to first token.
LMArena — human-preference Elo ratings from their public leaderboard dataset.
All raw metrics belong to their respective sources and are reproduced with attribution. The composite scores, weighting and rankings on this page are ThePlanetTools.ai's own computation and should be attributed to us, not to the data providers.
We're developers and SaaS builders who use these tools daily in production. Every review comes from hands-on experience building real products — DealPropFirm, ThePlanetIndicator, PropFirmsCodes, and many more. We don't just review tools — we build and ship with them every day.
Written and tested by developers who build with these tools daily.