Skip to content
Refreshed every 48 hours· August 2, 2026

The AI Model Index

Benchmarks say one thing, people using the models say another. We pull both — plus real market pricing and measured speed — and score 207 models on one scale. 113 of them carry a machine score and a human score, which is where the disagreements show up.

24 of these models have a hands-on review on this site.

What feeds the index

  • OpenRouter207 models

    live market pricing

  • Artificial Analysis196 models

    benchmark indices

  • LMArena124 models

    human preference votes

113 models scored by machines and humans

Rank by

The current top 20

Benchmark indices cross-checked against human preference votes. The bar under each score splits it into the five things we weigh — same total, different shape means a different model for a different job.

  • 1

    Claude Opus 5

    anthropic

    90.0intelligence
    Intelligence
    60.7
    Human Elo
    1512
    Blended /M
    $10
    Throughput
    54 t/s
    Review
  • 2

    Claude Fable 5

    anthropic

    86.7intelligence
    Intelligence
    59.9
    Human Elo
    1494
    Blended /M
    $20
    Throughput
    61 t/s
    Review
  • 3

    GPT-5.6 Sol

    openai

    82.9intelligence
    Intelligence
    58.9
    Human Elo
    Blended /M
    $11
    Throughput
    73 t/s
    Review
  • 4OPEN

    Kimi K3

    moonshotai

    82.5intelligence
    Intelligence
    57.1
    Human Elo
    1473
    Blended /M
    $6.0
    Throughput
    33 t/s
    Review
  • 5

    Claude Opus 4.7

    anthropic

    82.0intelligence
    Intelligence
    53.5
    Human Elo
    1490
    Blended /M
    $10
    Throughput
    Review
  • 6

    GPT-5.5

    openai

    80.0intelligence
    Intelligence
    54.8
    Human Elo
    1469
    Blended /M
    $11
    Throughput
    Review
  • 7

    Claude Opus 4.8

    anthropic

    79.8intelligence
    Intelligence
    55.7
    Human Elo
    1463
    Blended /M
    $10
    Throughput
    Review
  • 8

    Muse Spark 1.1

    meta

    78.0intelligence
    Intelligence
    50.6
    Human Elo
    1479
    Blended /M
    $2.0
    Throughput
    133 t/s
    Review
  • 9

    Qwen3.6 Max Preview

    qwen

    78.0intelligence
    Intelligence
    Human Elo
    1446
    Blended /M
    Throughput
  • 10

    Gemini 3.5 Flash

    google

    77.8intelligence
    Intelligence
    50.2
    Human Elo
    1480
    Blended /M
    $3.4
    Throughput
    221 t/s
    Review
  • 11

    GPT-5.4

    openai

    77.7intelligence
    Intelligence
    51.4
    Human Elo
    1470
    Blended /M
    $5.6
    Throughput
    Review
  • 12

    GPT-5.6 Terra

    openai

    77.5intelligence
    Intelligence
    55.0
    Human Elo
    Blended /M
    $4.5
    Throughput
    123 t/s
    Review
  • 13

    Gemini 3.6 Flash

    google

    77.5intelligence
    Intelligence
    50.1
    Human Elo
    1478
    Blended /M
    $3.0
    Throughput
    195 t/s
  • 14

    Grok 4.5

    x-ai

    77.1intelligence
    Intelligence
    53.8
    Human Elo
    1452
    Blended /M
    $3.0
    Throughput
    53 t/s
    Review
  • 15OPEN

    GLM 5.2

    z-ai

    76.9intelligence
    Intelligence
    51.1
    Human Elo
    1464
    Blended /M
    $2.1
    Throughput
    142 t/s
    Review
  • 16

    Claude Sonnet 5

    anthropic

    75.6intelligence
    Intelligence
    53.4
    Human Elo
    1440
    Blended /M
    $4.0
    Throughput
    75 t/s
    Review
  • 17

    Qwen3.7 Max

    qwen

    73.8intelligence
    Intelligence
    46.0
    Human Elo
    1475
    Blended /M
    $3.8
    Throughput
    198 t/s
  • 18OPEN

    DeepSeek V4 Flash

    deepseek

    73.6intelligence
    Intelligence
    49.9
    Human Elo
    1432
    Blended /M
    $0.17
    Throughput
  • 19

    Claude Opus 4.6

    anthropic

    73.4intelligence
    Intelligence
    37.8
    Human Elo
    1503
    Blended /M
    $10
    Throughput
    Review
  • 20

    Gemini 3.1 Pro Preview

    google

    73.3intelligence
    Intelligence
    46.5
    Human Elo
    1480
    Blended /M
    $4.5
    Throughput
    120 t/s
    Review

Put two models head to head

Benchmarks, human votes, speed and real pricing on the same screen. The highlighted side wins that line.

Claude Opus 5

anthropic

90.0

intelligence · wins 11 lines

VS

Claude Fable 5

anthropic

86.7

intelligence · wins 2 lines

Our composite score

Our composite score
63.3Overall59.4
90Intelligence40% of the total87
15Price / perf20% of the total1
85Capabilities15% of the total85
34Speed10% of the total37
55Ecosystem15% of the total55

Measured intelligence

Measured intelligence
60.7Intelligence indexArtificial Analysis59.9
78.0Coding indexArtificial Analysis76.5
55.3Agentic indexArtificial Analysis52.8
Math indexArtificial Analysis
1512 · 2k votesHuman preferenceLMArena Elo1494 · 16k votes

Speed

Speed
54 t/sOutput throughputtokens/second61 t/s
34.11 sTime to first token47.51 s

Price per million tokens

Price per million tokens
$5.0Input$10
$25Output$50
$10Blended 3:1typical real-world mix$20

Capabilities

Capabilities
1MContext window1M
128kMax output128k
YesReasoningYes
NoOpen weightsNo
text, image, fileModalitiestext, image, file
1Providers serving it1
Knowledge cutoff

Every model we track

Filter it down to the models that fit your constraints, then open a row for the full measurement set.

207 of 207 models

#Modelintelligence & profileIn / OutActions
1anthropic
90.0
$5.0 / $25
2anthropic
86.7
$10 / $50
3openai
82.9
$5.0 / $30
4moonshotai
82.5
$3.0 / $15
5anthropic
82.0
$5.0 / $25
6openai
80.0
$5.0 / $30
7anthropic
79.8
$5.0 / $25
8meta
78.0
$1.3 / $4.3
9
Qwen3.6 Max Preview
qwen
78.0
$1.0 / $6.2
10google
77.8
$1.5 / $9.0
11openai
77.7
$2.5 / $15
12openai
77.5
$1.0 / $6.0
13
Gemini 3.6 Flash
google
77.5
$1.5 / $7.5
14x-ai
77.1
$2.0 / $6.0
15z-ai
76.9
$0.31 / $0.98
16anthropic
75.6
$2.0 / $10
17
Qwen3.7 Max
qwen
73.8
$1.5 / $4.4
18
DeepSeek V4 FlashOPEN
deepseek
73.6
$0.14 / $0.28
19anthropic
73.4
$5.0 / $25
20google
73.3
$2.0 / $12
21openai
72.7
$0.10 / $0.60
22
DeepSeek V4 Flash 0731OPEN
deepseek
72.6
$0.09 / $0.18
23
DeepSeek V4 ProOPEN
deepseek
70.1
$0.43 / $0.87
24moonshotai
69.9
$0.60 / $3.4
25
MiMo-V2.5-ProOPEN
xiaomi
69.7
$0.43 / $0.87
26anthropic
69.3
$3.0 / $15
27
Hy3OPEN
tencent
69.0
$0.13 / $0.53
28
GLM 5.1OPEN
z-ai
68.9
$0.97 / $3.0
29
Gemini 3 Flash Preview
google
68.6
$0.50 / $3.0
30minimax
67.9
$0.30 / $1.2
31
GLM 5OPEN
z-ai
67.1
$0.95 / $2.5
32
Hy3 previewOPEN
tencent
66.8
$0.06 / $0.21
33
InklingOPEN
thinkingmachines
66.0
$1.0 / $4.0
34
Qwen3.7 Plus
qwen
65.7
$0.32 / $1.3
35
Qwen3.5-Flash
qwen
65.2
$0.07 / $0.26
36
GPT-5.2
openai
65.2
$1.8 / $14
37
GPT-5.2 Chat
openai
65.2
$1.8 / $14
38
Qwen3.6 Plus
qwen
64.4
$0.33 / $1.9
39
GPT-5.3-Codex
openai
63.3
$1.8 / $14
40
Gemini 3.5 Flash Lite
google
62.6
$0.30 / $2.5
41
GPT-5.3 Chat
openai
62.5
$1.8 / $14
42
GLM 5V Turbo
z-ai
62.4
$1.2 / $4.0
43
GPT-5.1
openai
62.1
$1.3 / $10
44
GPT-5.4 Mini
openai
62.0
$0.75 / $4.5
45
MiMo-V2.5OPEN
xiaomi
61.9
$0.14 / $0.28
46
Kimi K2.5OPEN
moonshotai
61.9
$0.57 / $2.9
47
Qwen3 Max Thinking
qwen
60.7
$0.78 / $3.9
48
GLM 4.7OPEN
z-ai
60.5
$0.40 / $1.8
49
Qwen3.5 397B A17BOPEN
qwen
60.3
$0.39 / $2.3
50
DeepSeek V3.2 ExpOPEN
deepseek
59.5
$0.27 / $0.41
51
MiniMax M2.7OPEN
minimax
59.1
$0.25 / $1.0
52
Gemini 2.5 Pro Preview 06-05
google
58.9
$1.3 / $10
53
Claude Opus 4.5
anthropic
58.3
$5.0 / $25
54
Qwen3.5-27BOPEN
qwen
58.1
$0.20 / $1.6
55
GLM 4.6OPEN
z-ai
58.0
$0.50 / $2.0
56
Gemma 4 31BOPEN
google
57.5
$0.10 / $0.34
57
Kimi K2.7 CodeOPEN
moonshotai
57.4
$0.73 / $3.5
58
GPT-5.2-Codex
openai
57.3
$1.8 / $14
59
DeepSeek V3.2OPEN
deepseek
56.9
$0.27 / $0.40
60
Nex-N2-ProOPEN
nex-agi
56.8
$0.25 / $1.0
61
Qwen3.5-122B-A10BOPEN
qwen
56.8
$0.26 / $2.1
62x-ai
56.5
$1.3 / $2.5
63
Mistral Medium 3.5
mistralai
56.3
$1.5 / $7.5
64openai
55.9
$1.3 / $10
65
GPT-5.4 Nano
openai
55.7
$0.20 / $1.3
66
DeepSeek V3.1 TerminusOPEN
deepseek
55.6
$0.27 / $1.0
67
Gemini 2.5 Pro
google
55.6
$1.3 / $10
68
Qwen3 Max
qwen
55.2
$0.78 / $3.9
69
Inkling SmallOPEN
thinkingmachines
54.4
$0.50 / $1.2
70
GLM 5 Turbo
z-ai
54.4
$1.2 / $4.0
71
MiniMax M2.1OPEN
minimax
54.2
$0.30 / $1.2
72
Gemma 4 26B A4B OPEN
google
54.1
$0.07 / $0.34
73
Gemini 3.1 Flash Lite
google
53.2
$0.25 / $1.5
74x-ai
52.9
$1.3 / $2.5
75
Grok Build 0.1
x-ai
52.7
$1.0 / $2.0
76
Kimi K2 0711OPEN
moonshotai
52.4
$0.57 / $2.3
77
Step 3.5 FlashOPEN
stepfun
52.0
$0.10 / $0.30
78
MiniMax M2.5OPEN
minimax
51.5
$0.15 / $0.90
79
Qwen3.6 27BOPEN
qwen
51.1
$0.30 / $2.0
80
R1 0528OPEN
deepseek
50.9
$0.50 / $2.1
81
MiniMax M1
minimax
50.7
$0.55 / $2.2
82
GLM 4.5OPEN
z-ai
50.6
$0.60 / $2.2
83
Nemotron 3 Ultra (free)OPEN
nvidia
50.5
Free / Free
84
Nemotron 3 UltraOPEN
nvidia
50.5
$0.60 / $3.6
85
DeepSeek V3.1OPEN
deepseek
50.4
$0.25 / $0.95
86
Qwen3 VL 235B A22B InstructOPEN
qwen
50.4
$0.21 / $1.9
87
Qwen3 VL 235B A22B ThinkingOPEN
qwen
50.4
$0.40 / $4.0
88
Qwen3.5-35B-A3BOPEN
qwen
50.0
$0.14 / $1.0
89
Gemini 3.1 Flash Lite Preview
google
49.7
$0.25 / $1.5
90
Gemini 2.5 Flash
google
49.6
$0.30 / $2.5
91
GPT-5.1-Codex
openai
49.6
$1.3 / $10
92
GPT-5.1-Codex-Max
openai
49.6
$1.3 / $10
93
Claude Sonnet 4.5
anthropic
49.3
$3.0 / $15
94
Qwen3 235B A22B Instruct 2507OPEN
qwen
48.9
$0.09 / $0.55
95
LongCat 2.0OPEN
meituan
48.7
$0.30 / $1.2
96
Kimi K2 0905OPEN
moonshotai
47.0
$0.60 / $2.5
97
Qwen3 Next 80B A3B InstructOPEN
qwen
46.5
$0.10 / $1.1
98
o3 Pro
openai
46.4
$20 / $80
99
KAT-Coder-Pro V2
kwaipilot
46.1
$0.30 / $1.2
100
GPT-5 Mini
openai
46.1
$0.25 / $2.0
101
MiniMax M2OPEN
minimax
45.5
$0.26 / $1.0
102
o1
openai
45.2
$15 / $60
103
Qwen3 Next 80B A3B ThinkingOPEN
qwen
44.1
$0.15 / $1.2
104
GPT-5.1-Codex-Mini
openai
43.7
$0.25 / $2.0
105
gpt-oss-120bOPEN
openai
43.5
$0.04 / $0.17
106
o3
openai
43.4
$2.0 / $8.0
107
Qwen3 235B A22B Thinking 2507OPEN
qwen
43.2
$0.23 / $2.3
108
GLM 4.7 FlashOPEN
z-ai
43.1
$0.06 / $0.40
109
GLM 4.5 AirOPEN
z-ai
42.4
$0.13 / $0.85
110
Qwen3.6 35B A3BOPEN
qwen
41.9
$0.14 / $1.0
111
GLM 4.6VOPEN
z-ai
41.6
$0.30 / $0.90
112
Gemma 3n 4BOPEN
google
41.1
$0.06 / $0.12
113
Mercury 2
inception
41.0
$0.25 / $0.75
114
Ring-2.6-1T
inclusionai
40.5
$0.07 / $0.63
115
R1OPEN
deepseek
40.5
$0.70 / $2.5
116
Step 3.7 FlashOPEN
stepfun
40.3
$0.20 / $1.1
117
Qwen3 Coder 480B A35BOPEN
qwen
40.0
$0.30 / $1.0
118anthropic
39.2
$1.0 / $5.0
119
DeepSeek V3 0324OPEN
deepseek
38.3
$0.27 / $1.1
120
Gemma 4 31B (free)OPEN
google
38.2
Free / Free
121
Nova 2 Lite
amazon
38.2
$0.30 / $2.5
122
o3 Mini
openai
38.1
$1.1 / $4.4
123
Qwen3 235B A22BOPEN
qwen
38.0
$0.46 / $1.8
124
Qwen3 30B A3B Instruct 2507OPEN
qwen
37.4
$0.05 / $0.19
125
Ling-2.6-1T
inclusionai
37.3
$0.07 / $0.63
126
Claude Sonnet 4
anthropic
37.0
$3.0 / $15
127
GPT-5 Nano
openai
36.6
$0.05 / $0.40
128
o4 Mini High
openai
36.6
$1.1 / $4.4
129
o4 Mini
openai
36.6
$1.1 / $4.4
130
Trinity Large ThinkingOPEN
arcee-ai
36.0
$0.22 / $0.85
131
DeepSeek V3OPEN
deepseek
35.0
$0.26 / $1.0
132
o3 Mini High
openai
33.9
$1.1 / $4.4
133
Gemma 4 26B A4B (free)OPEN
google
33.2
Free / Free
134
Nemotron 3 SuperOPEN
nvidia
31.7
$0.09 / $0.40
135
Nemotron 3 Super (free)OPEN
nvidia
31.7
Free / Free
136
Qwen3 32BOPEN
qwen
31.6
$0.08 / $0.28
137
Gemma 3 27BOPEN
google
31.5
$0.08 / $0.45
138
GLM 4.5VOPEN
z-ai
30.7
$0.60 / $1.8
139
Kimi K2 ThinkingOPEN
moonshotai
29.8
$0.60 / $2.5
140
Llama 3.1 70B InstructOPEN
meta-llama
29.2
$0.40 / $0.40
141
Qwen3 30B A3BOPEN
qwen
28.6
$0.12 / $0.50
142
GPT-4o (2024-05-13)
openai
28.4
$5.0 / $15
143
Qwen3 Coder NextOPEN
qwen
28.3
$0.12 / $0.80
144
GPT-4.1
openai
27.7
$2.0 / $8.0
145
Qwen3 30B A3B Thinking 2507OPEN
qwen
27.3
$0.20 / $2.4
146
Gemma 3 12BOPEN
google
27.1
$0.05 / $0.15
147
o1-pro
openai
27.0
$150 / $600
148
gpt-oss-20bOPEN
openai
26.9
$0.03 / $0.13
149
Qwen3.5-9BOPEN
qwen
25.9
$0.10 / $0.15
150
Qwen3 VL 32B InstructOPEN
qwen
25.6
$0.10 / $0.42
151
North Mini Code (free)OPEN
cohere
25.4
Free / Free
152
Sonar Reasoning Pro
perplexity
25.4
$2.0 / $8.0
153
GPT-4o (2024-08-06)
openai
24.3
$2.5 / $10
154
Olmo 3 32B ThinkOPEN
allenai
24.2
$0.15 / $0.50
155
Granite 4.1 8BOPEN
ibm-granite
23.7
$0.05 / $0.10
156
GPT-4o-mini (2024-07-18)
openai
23.4
$0.15 / $0.60
157
Mistral Small 4OPEN
mistralai
23.1
$0.15 / $0.60
158
Llama 3.3 70B InstructOPEN
meta-llama
21.5
$0.13 / $0.40
159
Gemma 2 27BOPEN
google
21.4
$0.65 / $0.65
160
Command R+ (08-2024)
cohere
20.8
$2.5 / $10
161
Mistral Large 2407
mistralai
20.5
$2.0 / $6.0
162
Gemma 3 4BOPEN
google
19.6
$0.05 / $0.10
163
Qwen3 Coder 30B A3B InstructOPEN
qwen
19.4
$0.07 / $0.28
164
Qwen3 VL 30B A3B InstructOPEN
qwen
19.0
$0.13 / $0.52
165
Qwen3 VL 30B A3B ThinkingOPEN
qwen
19.0
$0.20 / $2.4
166
Mistral Medium 3.1
mistralai
18.4
$0.40 / $2.0
167
Ling-2.6-flash
inclusionai
17.9
$0.01 / $0.03
168
Mistral Medium 3
mistralai
17.9
$0.40 / $2.0
169
gpt-oss-20b (free)OPEN
openai
17.5
Free / Free
170
GPT-4.1 Mini
openai
16.8
$0.40 / $1.6
171
Sonar
perplexity
16.7
$1.0 / $1.0
172
Command AOPEN
cohere
16.5
$2.5 / $10
173
Gemini 2.5 Flash Lite
google
16.3
$0.10 / $0.40
174
GPT-4o
openai
16.0
$2.5 / $10
175
GPT-4o (2024-11-20)
openai
16.0
$2.5 / $10
176
Mistral Small 3OPEN
mistralai
15.9
$0.05 / $0.08
177
Solar Pro 3
upstage
15.5
$0.15 / $0.60
178
GPT-4 Turbo
openai
15.5
$10 / $30
179
GPT-4 Turbo Preview
openai
15.5
$10 / $30
180
Llama 4 MaverickOPEN
meta-llama
15.2
$0.20 / $0.80
181
Qwen3 VL 8B InstructOPEN
qwen
15.1
$0.12 / $0.46
182
Qwen3 VL 8B ThinkingOPEN
qwen
15.1
$0.18 / $2.1
183
Nemotron 3 Nano 30B A3BOPEN
nvidia
14.9
$0.05 / $0.20
184
Nemotron 3 Nano 30B A3B (free)OPEN
nvidia
14.9
Free / Free
185
R1 Distill Llama 70BOPEN
deepseek
14.1
$0.80 / $0.80
186
Sonar Pro
perplexity
13.3
$3.0 / $15
187
Ministral 3 14B 2512OPEN
mistralai
12.7
$0.20 / $0.20
188
Phi 4OPEN
microsoft
12.3
$0.07 / $0.14
189
Qwen3 14BOPEN
qwen
11.9
$0.23 / $0.91
190
GPT-4
openai
11.5
$30 / $60
191
Mistral Large 3 2512
mistralai
10.7
$0.50 / $1.5
192
GPT-4.1 Nano
openai
10.4
$0.10 / $0.40
193
Command R (08-2024)
cohere
9.9
$0.15 / $0.60
194
Llama 4 ScoutOPEN
meta-llama
9.8
$0.10 / $0.30
195
Ministral 3 8B 2512OPEN
mistralai
9.6
$0.15 / $0.15
196
Saba
mistralai
9.1
$0.20 / $0.60
197
Qwen3 8BOPEN
qwen
9.0
$0.12 / $0.46
198
GPT-4o-mini
openai
8.5
$0.15 / $0.60
199
Llama 3.1 8B InstructOPEN
meta-llama
8.4
$0.05 / $0.08
200
Hermes 3 70B InstructOPEN
nousresearch
7.3
$0.70 / $0.70
201
Ministral 3 3B 2512OPEN
mistralai
6.8
$0.10 / $0.10
202
Mistral Large
mistralai
6.3
$2.0 / $6.0
203
Reka Flash 3OPEN
rekaai
5.9
$0.10 / $0.20
204
Claude 3 Haiku
anthropic
5.6
$0.25 / $1.3
205
Granite 4.0 MicroOPEN
ibm-granite
3.4
$0.02 / $0.11
206
Llama 3.2 3B InstructOPEN
meta-llama
0.0
$0.05 / $0.33
207
Llama 3.2 1B InstructOPEN
meta-llama
0.0
$0.03 / $0.20

How we score

The raw measurements come from third parties, credited below. What we add is the cross-source aggregation and the scoring: five axes on absolute 0-100 scales, weighted into one number. Absolute scales matter — a model scoring 78 today means the same as a model scoring 78 six months ago, so the history stays comparable as new models arrive.

Intelligence

40%

Average of every published capability index we can cross-check — general intelligence, coding and agentic benchmarks.

Price / performance

20%

Real market price per million tokens weighted against measured capability. A cheap weak model and an expensive strong one can score the same.

Capabilities

15%

Context window, maximum output, modalities, reasoning support, prompt caching and how recent the knowledge cutoff is.

Speed

10%

Measured output throughput (tokens per second) and time to first token. A model you wait ten seconds for is a different product from one that answers instantly.

Ecosystem

15%

How many providers serve the model, and whether the weights are open — availability and independence from a single vendor.

What changed recently

Every collection is archived, so price cuts and ranking moves stay traceable. The sources publish today's numbers; this is the trail behind them.

Price moves · output per million tokens

  • NVIDIA: Nemotron 3 Ultra$2.20 → $3.60 64%
  • Z.ai: GLM 5.2$2.39 → $0.98 59%
  • DeepSeek: DeepSeek V4 Flash 0731$0.28 → $0.18 36%

How to read this index

The default ranking is capability, so the most expensive model can sit above a free one. Switch the criterion at the top of the page and the order changes completely — that is the point. Four questions come up constantly, and here are the answers.

Why is a $10 model ahead of a free one?

Because the default ranking measures what a model can do, not what it costs. Price is a separate axis, and it has its own ranking one click away. Ranking on the composite score instead buried the frontier models behind cheap ones and turned the whole index into a price chart, which is not what a ranked list of models is for. Today that puts Claude Opus 5 first with 90.0 on capability at $25.00 per million output tokens.

What happens when benchmarks and humans disagree?

Both count, equally. Benchmarks test verifiable problems; the arena records which answer people actually preferred. When a model carries both, each contributes half its intelligence score. The widest gap in the index right now is Qwen: Qwen3 235B A22B Thinking 2507: 19 from benchmarks against 69 from human votes — a 50-point disagreement that a single-source ranking would hide entirely.

What is the coloured bar under each score?

The score itself, split into the five things we weigh. Each segment is that axis's contribution, so the full bar length is the composite score. Two models can share a total and have opposite shapes — one fast and cheap, the other slow and brilliant.

  • Intelligence
  • Price / performance
  • Capabilities
  • Speed
  • Ecosystem

Why are some cells empty?

Because that measurement does not exist yet, and we would rather show a dash than invent a number. Of the 207 models listed, 196 carry benchmark indices and 124 carry human ratings. A missing value never counts as zero — it simply leaves that axis out of the model's score, and the “Double-validated” filter isolates the 113 models measured both ways.

Frequently asked questions

What is the ThePlanetTools AI Index?

A ranking of 207 large language models built by cross-referencing three independent sources: live market pricing from OpenRouter, benchmark indices from Artificial Analysis, and human preference votes from LMArena. We aggregate those measurements and compute our own score across five weighted axes. The raw measurements belong to their sources; the scoring and the ranking are ours.

How is the ranking ordered by default?

By measured capability, not by price. A model that costs ten times more but answers harder questions ranks higher by default. Three alternative orderings are one click away — price/performance, speed, and the weighted composite score — because "best" means something different depending on whether you are paying the bill or waiting for the response.

What does the intelligence score actually measure?

Two things, weighted equally when both exist. Machine benchmarks contribute the general intelligence index (50%), the coding index (25%) and the agentic index (25%), each normalised against its own ceiling. Human preference contributes the LMArena Elo rating, mapped onto the same 0-100 scale. A model measured by only one of the two is scored on that one alone.

What does "double-validated" mean?

A model carrying both a machine benchmark score and a human preference rating. 113 of the 207 models in the index qualify. It matters because the two methods disagree more often than you would expect: benchmarks measure what is testable, the arena measures what people actually prefer after using the model.

Why do benchmark scores and human votes disagree?

They measure different things. Benchmarks test verifiable problems with known answers — maths, code that compiles, factual recall. Human votes reward tone, instruction-following, formatting and refusal behaviour, which no benchmark captures. A model can top the benchmarks and lose blind head-to-head votes, and the reverse happens just as often.

Why is the maths index excluded from the score?

Because nearly every current model scores in the 90s on it, so it separates nothing. Worse, averaging it in rewarded models that had been measured on maths and almost nothing else — one model with no coding or agentic score reached a higher intelligence score than models measured on everything. The maths index is still displayed as a raw measurement on each model.

How is price/performance calculated?

Measured capability divided by the blended price per million tokens, weighted one part input to three parts output, which reflects typical real-world usage. The result is placed on a logarithmic scale because prices span four orders of magnitude. Free or unpriced models receive a neutral 50 rather than a perfect score, since they are not comparable.

Are the prices real?

Yes. They come from OpenRouter, which publishes live per-token rates charged by the providers actually serving each model, in US dollars per million tokens. Input and output are shown separately because the gap between them is often large and matters more than the headline figure.

How often is the index updated?

Every 48 hours. Each collection re-fetches all three sources, recalculates every score, and writes a historical snapshot so price changes and ranking movements stay traceable. Models that disappear from the sources are marked inactive rather than deleted, so their history survives.

What does "open weights" mean in this index?

A model whose weights can be downloaded and self-hosted, as opposed to one reachable only through a vendor API. 111 of the 207 models qualify. Open weights add points on the ecosystem axis because they remove vendor lock-in, but a closed model served by many providers is not penalised to zero.

Why do the scores use absolute scales instead of ranking against each other?

So the numbers stay comparable over time. If scores were normalised against the current field, every new model release would silently move every existing score, and a snapshot from six months ago would mean nothing. On an absolute scale, a model scoring 78 today means the same as a model scoring 78 last year.

Which models here have been reviewed hands-on?

24 of the 207 models link to a full review on this site, written after actually using the model. Those links appear directly on each model card and in the comparison table. The rest are scored from published measurements only, which is stated rather than hidden.

Data sources

  • OpenRouter — live market pricing, context windows, modalities and availability across providers.
  • Artificial Analysis — independent intelligence, coding, agentic and math indices, plus measured throughput and time to first token.
  • LMArena — human-preference Elo ratings from their public leaderboard dataset.

All raw metrics belong to their respective sources and are reproduced with attribution. The composite scores, weighting and rankings on this page are ThePlanetTools.ai's own computation and should be attributed to us, not to the data providers.

Anthony M. — Founder & Lead Reviewer
Anthony M.Verified Builder

We're developers and SaaS builders who use these tools daily in production. Every review comes from hands-on experience building real products — DealPropFirm, ThePlanetIndicator, PropFirmsCodes, and many more. We don't just review tools — we build and ship with them every day.

Written and tested by developers who build with these tools daily.