Skip to content
Refreshed every 48 hours· September 7, 2026

The AI Model Index

Benchmarks say one thing, people using the models say another. We pull both — plus real market pricing and measured speed — and score 227 models on one scale. 118 of them carry a machine score and a human score, which is where the disagreements show up.

28 of these models have a hands-on review on this site.

What feeds the index

  • OpenRouter227 models

    live market pricing

  • Artificial Analysis214 models

    benchmark indices

  • LMArena131 models

    human preference votes

118 models scored by machines and humans

Rank by

The current top 20

Benchmark indices cross-checked against human preference votes. The bar under each score splits it into the five things we weigh — same total, different shape means a different model for a different job.

  • 1

    Claude Fable 5.1

    anthropic

    89.9intelligence
    Intelligence
    56.8
    Human Elo
    1514
    Blended /M
    $40
    Throughput
    67 t/s
  • 2

    Claude Opus 5

    anthropic

    86.9intelligence
    Intelligence
    54.1
    Human Elo
    1505
    Blended /M
    $20
    Throughput
    58 t/s
    Review
  • 3

    Claude Fable 5.1 (batch)

    anthropic

    84.0intelligence
    Intelligence
    56.8
    Human Elo
    Blended /M
    Throughput
  • 4

    Claude Fable 5

    anthropic

    84.0intelligence
    Intelligence
    53.2
    Human Elo
    1494
    Blended /M
    $40
    Throughput
    70 t/s
    Review
  • 5OPEN

    Kimi K3

    moonshotai

    80.5intelligence
    Intelligence
    50.2
    Human Elo
    1476
    Blended /M
    $12
    Throughput
    40 t/s
    Review
  • 6

    Claude Opus 5 (batch)

    anthropic

    80.5intelligence
    Intelligence
    54.1
    Human Elo
    Blended /M
    Throughput
  • 7

    Gemini 3.8 Flash

    google

    80.2intelligence
    Intelligence
    47.1
    Human Elo
    1495
    Blended /M
    $3.0
    Throughput
  • 8OPEN

    GLM 5.3

    z-ai

    79.9intelligence
    Intelligence
    48.6
    Human Elo
    1474
    Blended /M
    $3.7
    Throughput
    84 t/s
    Review
  • 9

    Qwen3.8 Max (0902)

    qwen

    79.0intelligence
    Intelligence
    46.9
    Human Elo
    1480
    Blended /M
    $5.0
    Throughput
    40 t/s
  • 10

    Muse Spark 1.3

    meta

    78.9intelligence
    Intelligence
    53.0
    Human Elo
    Blended /M
    $3.5
    Throughput
    190 t/s
  • 11

    GPT-6 Astra (batch)

    openai

    78.9intelligence
    Intelligence
    54.7
    Human Elo
    Blended /M
    Throughput
  • 12

    GPT-6 Astra

    openai

    78.9intelligence
    Intelligence
    54.7
    Human Elo
    Blended /M
    $40
    Throughput
    62 t/s
  • 13

    Gemini 3.7 Flash

    google

    78.2intelligence
    Intelligence
    45.2
    Human Elo
    1491
    Blended /M
    $3.0
    Throughput
    309 t/s
  • 14

    Qwen3.6 Max Preview

    qwen

    78.0intelligence
    Intelligence
    Human Elo
    1446
    Blended /M
    Throughput
  • 15

    Claude Opus 4.7

    anthropic

    77.9intelligence
    Intelligence
    44.3
    Human Elo
    1490
    Blended /M
    $20
    Throughput
    52 t/s
    Review
  • 16OPEN

    GLM 5.3 Flash

    z-ai

    77.8intelligence
    Intelligence
    46.2
    Human Elo
    1471
    Blended /M
    $0.41
    Throughput
    47 t/s
  • 17

    Claude Fable 5 (batch)

    anthropic

    77.5intelligence
    Intelligence
    53.2
    Human Elo
    Blended /M
    Throughput
  • 18

    Grok 4.6

    x-ai

    76.6intelligence
    Intelligence
    50.6
    Human Elo
    Blended /M
    $5.0
    Throughput
    64 t/s
    Review
  • 19

    GPT-5.6 Sol (batch)

    openai

    76.3intelligence
    Intelligence
    51.3
    Human Elo
    Blended /M
    Throughput
  • 20

    GPT-5.6 Sol

    openai

    76.3intelligence
    Intelligence
    51.3
    Human Elo
    Blended /M
    $16
    Throughput
    85 t/s
    Review

Put two models head to head

Benchmarks, human votes, speed and real pricing on the same screen. The highlighted side wins that line.

Claude Fable 5.1

anthropic

89.9

intelligence · wins 6 lines

VS

Claude Opus 5

anthropic

86.9

intelligence · wins 7 lines

Our composite score

Our composite score
60.9Overall62.5
90Intelligence40% of the total87
1Price / perf20% of the total14
85Capabilities15% of the total85
39Speed10% of the total41
55Ecosystem15% of the total55

Measured intelligence

Measured intelligence
56.8Intelligence indexArtificial Analysis54.1
81.6Coding indexArtificial Analysis78.0
58.2Agentic indexArtificial Analysis56.4
1514 · 3k votesHuman preferenceLMArena Elo1505 · 35k votes

Speed

Speed
67 t/sOutput throughputtokens/second58 t/s
266.51 sTime to first token73.62 s

Price per million tokens

Price per million tokens
$10Input$5.0
$50Output$25
$40Blended 3:1typical real-world mix$20

Capabilities

Capabilities
1MContext window1M
128kMax output128k
YesReasoningYes
NoOpen weightsNo
text, image, fileModalitiestext, image, file
1Providers serving it1
Knowledge cutoff
No review published yetOur Claude Opus 5 review

Every model we track

Filter it down to the models that fit your constraints, then open a row for the full measurement set.

227 of 227 models

#Modelintelligence & profileIn / OutActions
1
Claude Fable 5.1
anthropic
89.9
$10 / $50
2anthropic
86.9
$5.0 / $25
3
Claude Fable 5.1 (batch)
anthropic
84.0
$5.0 / $25
4anthropic
84.0
$10 / $50
5moonshotai
80.5
$3.0 / $15
6
Claude Opus 5 (batch)
anthropic
80.5
$2.5 / $13
7
Gemini 3.8 Flash
google
80.2
$0.75 / $3.8
8z-ai
79.9
$1.4 / $4.4
9
Qwen3.8 Max (0902)
qwen
79.0
$2.0 / $6.0
10
Muse Spark 1.3
meta
78.9
$1.3 / $4.3
11
GPT-6 Astra (batch)
openai
78.9
$5.0 / $25
12
GPT-6 Astra
openai
78.9
$10 / $50
13
Gemini 3.7 Flash
google
78.2
$0.75 / $3.8
14
Qwen3.6 Max Preview
qwen
78.0
$1.0 / $6.2
15anthropic
77.9
$5.0 / $25
16
GLM 5.3 FlashOPEN
z-ai
77.8
$0.07 / $0.25
17
Claude Fable 5 (batch)
anthropic
77.5
$5.0 / $25
18x-ai
76.6
$2.0 / $6.0
19
GPT-5.6 Sol (batch)
openai
76.3
$1.0 / $5.0
20openai
76.3
$2.0 / $10
21openai
75.6
$2.5 / $15
22openai
75.6
$5.0 / $30
23anthropic
75.5
$5.0 / $25
24
Kimi K3 (batch)OPEN
moonshotai
75.2
$3.0 / $15
25meta
73.7
$1.3 / $4.3
26x-ai
73.7
$2.0 / $6.0
27z-ai
73.5
$0.97 / $3.0
28google
73.5
$1.5 / $9.0
29
MiMo-V2.5OPEN
xiaomi
73.0
$0.14 / $0.28
30anthropic
72.6
$2.0 / $10
31
Gemini 3.6 Flash
google
72.3
$0.75 / $3.8
32
DeepSeek V4 Pro 0423OPEN
deepseek
71.8
$1.0 / $2.1
33
DeepSeek V4 Pro 0813OPEN
deepseek
71.8
$1.1 / $3.3
34
Qwen3.8 2.4T A95B (batch)OPEN
qwen
71.4
$2.0 / $6.0
35
Qwen3.8 2.4T A95BOPEN
qwen
71.4
$2.0 / $6.0
36
GLM 5.3 Flash (batch)OPEN
z-ai
71.3
$0.15 / $0.50
37
Qwen3.8 27BOPEN
qwen
70.5
$0.42 / $3.0
38
GPT-5.6 Terra (batch)
openai
70.4
$1.0 / $6.0
39openai
70.4
$2.0 / $12
40
Gemini 3.8 Flash (batch)
google
69.6
$0.38 / $1.9
41
Muse Spark 1.2
meta
69.3
$1.3 / $4.3
42
Qwen3.7 Max
qwen
69.3
$1.5 / $4.4
43deepseek
68.7
$0.09 / $0.18
44
DeepSeek V4 Flash 0731OPEN
deepseek
68.7
$0.14 / $0.28
45anthropic
68.5
$5.0 / $25
46
Claude Sonnet 5 (batch)
anthropic
68.0
$1.0 / $5.0
47google
68.0
$2.0 / $12
48
Gemini 3.7 Flash (batch)
google
66.5
$0.38 / $1.9
49
GPT-5.6 Luna (batch)
openai
66.2
$0.10 / $0.60
50openai
66.2
$0.20 / $1.2
51anthropic
65.6
$3.0 / $15
52
MiMo-V2.5-ProOPEN
xiaomi
65.5
$0.43 / $0.87
53moonshotai
65.5
$0.95 / $4.0
54
Qwen3.5-Flash
qwen
65.2
$0.07 / $0.26
55
Qwen3.7 Plus
qwen
65.0
$0.32 / $1.3
56
Hy3OPEN
tencent
64.8
$0.13 / $0.53
57
Hy3 previewOPEN
tencent
64.8
$0.18 / $0.60
58
GLM 5.1OPEN
z-ai
64.8
$0.97 / $3.0
59
DeepSeek V4 Pro 0813 (batch)OPEN
deepseek
64.4
$1.3 / $4.0
60minimax
64.0
$0.30 / $1.2
61
DeepSeek V4 Flash Vision ExpOPEN
deepseek
63.6
$0.22 / $0.66
62
Gemini 3 Flash Preview
google
63.6
$0.50 / $3.0
63
DeepSeek V4 Flash 0731 (batch)OPEN
deepseek
63.3
$0.14 / $0.28
64
Qwen3.6 Plus
qwen
62.9
$0.33 / $1.9
65
GLM 5OPEN
z-ai
62.1
$0.60 / $1.9
66
GPT-5.1
openai
61.4
$1.3 / $10
67
InklingOPEN
thinkingmachines
61.1
$1.0 / $4.0
68
Kimi K2.5OPEN
moonshotai
60.7
$0.45 / $2.3
69openai
60.1
$1.8 / $14
70
GPT-5.2 Chat
openai
60.1
$1.8 / $14
71
Gemini 3.6 Flash (batch)
google
58.8
$0.38 / $1.9
72
GLM 4.7OPEN
z-ai
58.5
$0.40 / $1.8
73
Inkling SmallOPEN
thinkingmachines
58.1
$0.45 / $1.2
74
MiniMax M2.7OPEN
minimax
57.6
$0.30 / $1.2
75
GLM 5V Turbo
z-ai
57.5
$1.2 / $4.0
76
Gemini 3.5 Flash Lite
google
57.3
$0.30 / $2.5
77
GPT-5.4 Mini
openai
57.2
$0.75 / $4.5
78
GLM 4.6OPEN
z-ai
57.0
$0.43 / $1.8
79
DeepSeek V3.2OPEN
deepseek
56.4
$0.27 / $0.40
80
DeepSeek V3.2 ExpOPEN
deepseek
56.4
$0.27 / $0.41
81
Nova 2 Lite
amazon
55.9
$0.30 / $2.5
82
Qwen3 Max Thinking
qwen
55.9
$0.78 / $3.9
83
Qwen3.5 397B A17BOPEN
qwen
55.8
$0.39 / $2.3
84
Gemini 2.5 Pro
google
55.7
$1.3 / $10
85
Gemini 2.5 Pro Preview 06-05
google
55.7
$1.3 / $10
86
Qwen3.5-122B-A10BOPEN
qwen
55.5
$0.29 / $2.4
87
Solar Pro 4
upstage
55.3
$0.03 / $0.12
88
DeepSeek V3.1 TerminusOPEN
deepseek
54.8
$0.27 / $1.0
89
Kimi K2.7 CodeOPEN
moonshotai
54.8
$0.66 / $3.4
90
GPT-5.4 Nano
openai
54.4
$0.20 / $1.3
91
Gemma 4 26B A4B OPEN
google
53.9
$0.07 / $0.34
92
Gemma 4 31BOPEN
google
53.6
$0.09 / $0.34
93openai
53.6
$1.3 / $10
94
Qwen3.5-27BOPEN
qwen
53.2
$0.20 / $1.6
95
Nex-N2-ProOPEN
nex-agi
53.0
$0.25 / $1.0
96
MiniMax M3 (batch)OPEN
minimax
52.8
$0.30 / $1.2
97
MiniMax M3 (free)OPEN
minimax
52.8
Free / Free
98
GPT-5.3-Codex
openai
52.7
$1.8 / $14
99x-ai
52.1
$1.3 / $2.5
100
Mistral Medium 3.5
mistralai
52.0
$1.5 / $7.5
101
MiniMax M1
minimax
50.7
$0.55 / $2.2
102
Qwen3 Max
qwen
50.6
$0.78 / $3.9
103
Qwen3.5-35B-A3BOPEN
qwen
49.9
$0.31 / $1.3
104
MiniMax M2.1OPEN
minimax
49.3
$0.30 / $1.2
105
Gemini 3.1 Pro Preview (batch)
google
49.0
$1.0 / $6.0
106
Claude Opus 4.5
anthropic
48.0
$5.0 / $25
107
Kimi K2 0711OPEN
moonshotai
47.5
$0.57 / $2.3
108
Step 3.5 FlashOPEN
stepfun
47.3
$0.10 / $0.30
109
GPT-5.2-Codex
openai
47.1
$1.8 / $14
110
Gemini 3.1 Flash Lite
google
47.0
$0.25 / $1.5
111
Gemini 3.1 Flash Lite Preview
google
47.0
$0.25 / $1.5
112
MiniMax M2.5OPEN
minimax
46.7
$0.27 / $1.1
113
KAT-Coder-Pro V2
kwaipilot
46.4
$0.30 / $1.2
114
R1 0528OPEN
deepseek
46.4
$0.50 / $2.1
115
GLM 4.5OPEN
z-ai
46.3
$0.60 / $2.2
116
Inkling (free)OPEN
thinkingmachines
46.2
Free / Free
117
Inkling (batch)OPEN
thinkingmachines
46.2
$1.0 / $4.0
118
DeepSeek V3.1OPEN
deepseek
46.1
$0.55 / $1.6
119
Qwen3 VL 235B A22B InstructOPEN
qwen
45.9
$0.21 / $1.9
120
Qwen3 VL 235B A22B ThinkingOPEN
qwen
45.9
$0.40 / $4.0
121
Qwen3 235B A22B Instruct 2507OPEN
qwen
45.8
$0.09 / $0.55
122
Gemini 2.5 Flash
google
45.1
$0.30 / $2.5
123
GLM 5 Turbo
z-ai
44.3
$1.2 / $4.0
124
GPT-5 Mini
openai
43.7
$0.25 / $2.0
125
Qwen3 Next 80B A3B ThinkingOPEN
qwen
43.5
$0.15 / $1.2
126
Qwen3 Next 80B A3B InstructOPEN
qwen
43.5
$0.10 / $1.1
127x-ai
42.9
$1.3 / $2.5
128
Qwen3.6 27BOPEN
qwen
42.8
$0.30 / $2.0
129
Qwen3 235B A22B Thinking 2507OPEN
qwen
42.6
$0.23 / $2.3
130
Kimi K2 0905OPEN
moonshotai
42.5
$0.60 / $2.5
131
o1
openai
42.1
$15 / $60
132
LongCat 2.0OPEN
meituan
41.3
$0.30 / $1.2
133
Ling 3.0 FlashOPEN
inclusionai
41.2
$0.02 / $0.06
134
MiniMax M2OPEN
minimax
40.8
$0.26 / $1.0
135
Qwen3.6 35B A3BOPEN
qwen
40.5
$0.10 / $0.90
136
R1OPEN
deepseek
40.5
$0.70 / $2.5
137
GPT-5.1-Codex-Max
openai
39.9
$1.3 / $10
138
GPT-5.1-Codex
openai
39.9
$1.3 / $10
139
gpt-oss-120bOPEN
openai
39.3
$0.04 / $0.17
140
Gemini 3.5 Flash Lite (batch)
google
39.2
$0.15 / $1.3
141
GLM 4.7 FlashOPEN
z-ai
38.5
$0.06 / $0.40
142
GLM 4.5 AirOPEN
z-ai
38.2
$0.13 / $0.85
143
DeepSeek V3 0324OPEN
deepseek
37.9
$0.25 / $1.0
144
Mercury 2
inception
37.7
$0.25 / $0.75
145
GLM 4.6VOPEN
z-ai
37.4
$0.30 / $0.90
146
o3 Pro
openai
36.7
$20 / $80
147
Step 3.7 FlashOPEN
stepfun
36.5
$0.20 / $1.1
148
Trinity Large ThinkingOPEN
arcee-ai
35.8
$0.25 / $0.80
149
Qwen3 Coder 480B A35BOPEN
qwen
35.6
$0.30 / $1.0
150
GPT-5.1-Codex-Mini
openai
34.3
$0.25 / $2.0
151
o3
openai
33.9
$2.0 / $8.0
152
Qwen3 235B A22BOPEN
qwen
33.9
$0.46 / $1.8
153
o3 Mini High
openai
33.7
$1.1 / $4.4
154
o3 Mini
openai
33.7
$1.1 / $4.4
155
Qwen3 30B A3B Instruct 2507OPEN
qwen
33.3
$0.05 / $0.19
156
Llama 3.3 70B InstructOPEN
meta-llama
32.6
$0.10 / $0.32
157
Kimi K2 ThinkingOPEN
moonshotai
32.4
$0.60 / $2.5
158
DeepSeek V3OPEN
deepseek
32.2
$0.32 / $0.89
159
GPT-5 Nano
openai
32.1
$0.05 / $0.40
160
Claude Haiku 4.5 (batch)
anthropic
32.0
$0.50 / $2.5
161anthropic
32.0
$1.0 / $5.0
162
Granite 4.2 8BOPEN
ibm-granite
30.9
$0.10 / $0.15
163
Qwen3 32BOPEN
qwen
30.5
$0.08 / $0.28
164
Gemma 3 27BOPEN
google
30.1
$0.08 / $0.45
165
Llama 3.1 70B InstructOPEN
meta-llama
29.1
$0.40 / $0.40
166
o4 Mini High
openai
27.3
$1.1 / $4.4
167
o4 Mini
openai
27.3
$1.1 / $4.4
168
Qwen3 Coder NextOPEN
qwen
27.0
$0.12 / $0.80
169
GLM 4.5VOPEN
z-ai
26.6
$0.60 / $1.8
170
Qwen3 30B A3B Thinking 2507OPEN
qwen
25.9
$0.20 / $2.4
171
Gemma 3 12BOPEN
google
25.8
$0.05 / $0.15
172
GPT-4o (2024-05-13)
openai
25.7
$5.0 / $15
173
Nemotron 3.5 LightningOPEN
nvidia
25.5
$0.08 / $0.20
174
Qwen3.5-9BOPEN
qwen
24.7
$0.10 / $0.15
175
Qwen3 30B A3BOPEN
qwen
24.5
$0.12 / $0.50
176
gpt-oss-20bOPEN
openai
24.4
$0.03 / $0.13
177
Mistral Small 4OPEN
mistralai
22.0
$0.15 / $0.60
178
gpt-oss-120b (batch)OPEN
openai
21.8
$0.15 / $0.60
179
Gemma 2 27BOPEN
google
21.4
$0.65 / $0.65
180
GPT-4o-mini (2024-07-18)
openai
20.7
$0.15 / $0.60
181
Command R+ (08-2024)
cohere
20.7
$2.5 / $10
182
GPT-4o (2024-08-06)
openai
20.2
$2.5 / $10
183
Gemma 3 4BOPEN
google
19.5
$0.05 / $0.10
184
GPT-4.1
openai
18.9
$2.0 / $8.0
185
o1-pro
openai
18.3
$150 / $600
186
Devstral 2 2512OPEN
mistralai
17.9
$0.40 / $2.0
187
Qwen3 VL 32B InstructOPEN
qwen
17.0
$0.10 / $0.42
188
Sonar Reasoning Pro
perplexity
16.9
$2.0 / $8.0
189
Mistral Large 2407
mistralai
16.5
$2.0 / $6.0
190
GPT-4.1 Mini
openai
15.9
$0.40 / $1.6
191
Mistral Medium 3.1
mistralai
15.6
$0.40 / $2.0
192
Mistral Small 3OPEN
mistralai
12.0
$0.05 / $0.08
193
Solar Pro 3
upstage
11.1
$0.15 / $0.60
194
Qwen3 Coder 30B A3B InstructOPEN
qwen
11.0
$0.07 / $0.28
195
Llama 4 MaverickOPEN
meta-llama
10.8
$0.20 / $0.70
196
Qwen3 VL 30B A3B InstructOPEN
qwen
10.7
$0.15 / $0.60
197
Qwen3 VL 30B A3B ThinkingOPEN
qwen
10.7
$0.20 / $2.4
198
Command AOPEN
cohere
10.5
$2.5 / $10
199
GPT-4 Turbo
openai
10.2
$10 / $30
200
GPT-4 Turbo Preview
openai
10.2
$10 / $30
201
Command R (08-2024)
cohere
9.8
$0.15 / $0.60
202
Qwen3 14BOPEN
qwen
9.7
$0.12 / $0.24
203
Llama 3.1 8B InstructOPEN
meta-llama
9.6
$0.05 / $0.08
204
Mistral Medium 3
mistralai
9.6
$0.40 / $2.0
205
Phi 4OPEN
microsoft
9.5
$0.07 / $0.14
206
Sonar
perplexity
8.4
$1.0 / $1.0
207
Gemini 2.5 Flash Lite
google
8.0
$0.10 / $0.40
208
GPT-4.1 Nano
openai
8.0
$0.10 / $0.40
209
GPT-4o (2024-11-20)
openai
7.7
$2.5 / $10
210
GPT-4o
openai
7.7
$2.5 / $10
211
Llama 4 ScoutOPEN
meta-llama
7.4
$0.10 / $0.30
212
Mistral Large 3 2512
mistralai
7.2
$0.50 / $1.5
213
Qwen3 VL 8B InstructOPEN
qwen
7.0
$0.12 / $0.46
214
Qwen3 VL 8B ThinkingOPEN
qwen
7.0
$0.18 / $2.1
215
GPT-4
openai
6.3
$30 / $60
216
Qwen3 8BOPEN
qwen
6.0
$0.12 / $0.46
217
R1 Distill Llama 70BOPEN
deepseek
6.0
$0.80 / $0.80
218
GPT-4o-mini
openai
5.6
$0.15 / $0.60
219
Sonar Pro
perplexity
5.1
$3.0 / $15
220
Granite 4.0 MicroOPEN
ibm-granite
1.4
$0.02 / $0.11
221
Reka Flash 3OPEN
rekaai
1.4
$0.10 / $0.20
222
Hermes 3 70B InstructOPEN
nousresearch
1.4
$0.70 / $0.70
223
Saba
mistralai
1.4
$0.20 / $0.60
224
Mistral Large
mistralai
1.4
$2.0 / $6.0
225
Claude 3 Haiku
anthropic
1.4
$0.25 / $1.3
226
Llama 3.2 3B InstructOPEN
meta-llama
0.0
$0.05 / $0.33
227
Llama 3.2 1B InstructOPEN
meta-llama
0.0
$0.03 / $0.20

How we score

The raw measurements come from third parties, credited below. What we add is the cross-source aggregation and the scoring: five axes on absolute 0-100 scales, weighted into one number. Absolute scales matter — a model scoring 78 today means the same as a model scoring 78 six months ago, so the history stays comparable as new models arrive.

Intelligence

40%

Average of every published capability index we can cross-check — general intelligence, coding and agentic benchmarks.

Price / performance

20%

Real market price per million tokens weighted against measured capability. A cheap weak model and an expensive strong one can score the same.

Capabilities

15%

Context window, maximum output, modalities, reasoning support, prompt caching and how recent the knowledge cutoff is.

Speed

10%

Measured output throughput (tokens per second) and time to first token. A reasoning model thinks before it answers, so its first token can take minutes — that is the trade it makes for capability, and it scores low here by design.

Ecosystem

15%

How many providers serve the model, and whether the weights are open — availability and independence from a single vendor.

What changed recently

Every collection is archived, so price cuts and ranking moves stay traceable. The sources publish today's numbers; this is the trail behind them.

Price moves · output per million tokens

  • Z.ai: GLM 5.2$0.31 → $3.74 1114%
  • DeepSeek: DeepSeek V4 Pro 0423$0.87 → $2.34 169%
  • Qwen: Qwen3.5-35B-A3B$1.00 → $1.80 80%
  • Qwen: Qwen3 14B$0.91 → $0.24 74%
  • DeepSeek: DeepSeek V4 Flash 0731$0.18 → $0.28 56%
  • Google: Gemini 3.6 Flash$7.50 → $3.75 50%
  • Google: Gemini 3.6 Flash (batch)$3.75 → $1.88 50%
  • Qwen: Qwen3 VL 235B A22B Instruct$1.90 → $1.04 45%

Ranking moves · by capability

  • Amazon: Nova 2 Lite#150 → #94 56
  • Kwaipilot: KAT-Coder-Pro V2#127 → #102 25
  • Meta: Llama 3.3 70B Instruct#189 → #169 20
  • MoonshotAI: Kimi K2 Thinking#168 → #152 16
  • DeepSeek: DeepSeek V4 Pro 0423#37 → #26 11
  • Nex AGI: Nex-N2-Pro#78 → #67 11

How to read this index

The default ranking is capability, so the most expensive model can sit above a free one. Switch the criterion at the top of the page and the order changes completely — that is the point. Four questions come up constantly, and here are the answers.

Why is a $10 model ahead of a free one?

Because the default ranking measures what a model can do, not what it costs. Price is a separate axis, and it has its own ranking one click away. Ranking on the composite score instead buried the frontier models behind cheap ones and turned the whole index into a price chart, which is not what a ranked list of models is for. Today that puts Anthropic: Claude Fable 5.1 first with 89.9 on capability at $50.00 per million output tokens.

What happens when benchmarks and humans disagree?

Both count, equally. Benchmarks test verifiable problems; the arena records which answer people actually preferred. When a model carries both, each contributes half its intelligence score. The widest gap in the index right now is Qwen: Qwen3 235B A22B Thinking 2507: 11 from benchmarks against 70 from human votes — a 59-point disagreement that a single-source ranking would hide entirely.

What is the coloured bar under each score?

The score itself, split into the five things we weigh. Each segment is that axis's contribution, so the full bar length is the composite score. Two models can share a total and have opposite shapes — one fast and cheap, the other slow and brilliant.

  • Intelligence
  • Price / performance
  • Capabilities
  • Speed
  • Ecosystem

Why are some cells empty?

Because that measurement does not exist yet, and we would rather show a dash than invent a number. Of the 227 models listed, 214 carry benchmark indices and 131 carry human ratings. A missing value never counts as zero — it simply leaves that axis out of the model's score, and the “Double-validated” filter isolates the 118 models measured both ways.

Frequently asked questions

What is the ThePlanetTools AI Index?

A ranking of 227 large language models built by cross-referencing three independent sources: live market pricing from OpenRouter, benchmark indices from Artificial Analysis, and human preference votes from LMArena. We aggregate those measurements and compute our own score across five weighted axes. The raw measurements belong to their sources; the scoring and the ranking are ours.

How is the ranking ordered by default?

By measured capability, not by price. A model that costs ten times more but answers harder questions ranks higher by default. Three alternative orderings are one click away — price/performance, speed, and the weighted composite score — because "best" means something different depending on whether you are paying the bill or waiting for the response.

What does the intelligence score actually measure?

Two things, weighted equally when both exist. Machine benchmarks contribute the general intelligence index (50%), the coding index (25%) and the agentic index (25%), each normalised against its own ceiling. Human preference contributes the LMArena Elo rating, mapped onto the same 0-100 scale. A model measured by only one of the two is scored on that one alone.

What does "double-validated" mean?

A model carrying both a machine benchmark score and a human preference rating. 118 of the 227 models in the index qualify. It matters because the two methods disagree more often than you would expect: benchmarks measure what is testable, the arena measures what people actually prefer after using the model.

Why do benchmark scores and human votes disagree?

They measure different things. Benchmarks test verifiable problems with known answers — maths, code that compiles, factual recall. Human votes reward tone, instruction-following, formatting and refusal behaviour, which no benchmark captures. A model can top the benchmarks and lose blind head-to-head votes, and the reverse happens just as often.

Why is the maths index excluded from the score?

Because nearly every current model scored in the 90s on it, so it separated nothing. Worse, averaging it in rewarded models measured on maths and almost nothing else — one model with no coding or agentic score reached a higher intelligence score than models measured on everything. Artificial Analysis has since stopped publishing that index on the endpoint we use, which changes nothing for us: it was already out of the score.

How is price/performance calculated?

Measured capability divided by the blended price per million tokens, weighted one part input to three parts output, which reflects typical real-world usage. The result is placed on a logarithmic scale because prices span four orders of magnitude. Free or unpriced models receive a neutral 50 rather than a perfect score, since they are not comparable.

Are the prices real?

Yes. They come from OpenRouter, which publishes live per-token rates charged by the providers actually serving each model, in US dollars per million tokens. Input and output are shown separately because the gap between them is often large and matters more than the headline figure.

How often is the index updated?

Every 48 hours. Each collection re-fetches all three sources, recalculates every score, and writes a historical snapshot so price changes and ranking movements stay traceable. Models that disappear from the sources are marked inactive rather than deleted, so their history survives.

What does "open weights" mean in this index?

A model whose weights can be downloaded and self-hosted, as opposed to one reachable only through a vendor API. 115 of the 227 models qualify. Open weights add points on the ecosystem axis because they remove vendor lock-in, but a closed model served by many providers is not penalised to zero.

Why do the scores use absolute scales instead of ranking against each other?

So the numbers stay comparable over time. If scores were normalised against the current field, every new model release would silently move every existing score, and a snapshot from six months ago would mean nothing. On an absolute scale, a model scoring 78 today means the same as a model scoring 78 last year.

Which models here have been reviewed hands-on?

28 of the 227 models link to a full review on this site, written after actually using the model. Those links appear directly on each model card and in the comparison table. The rest are scored from published measurements only, which is stated rather than hidden.

Data sources

  • OpenRouter — live market pricing, context windows, modalities and availability across providers.
  • Artificial Analysis — independent intelligence, coding, agentic and math indices, plus measured throughput and time to first token.
  • LMArena — human-preference Elo ratings from their public leaderboard dataset.

All raw metrics belong to their respective sources and are reproduced with attribution. The composite scores, weighting and rankings on this page are ThePlanetTools.ai's own computation and should be attributed to us, not to the data providers.

Anthony M. — Founder & Lead Reviewer
Anthony M.Verified Builder

We're developers and SaaS builders who use these tools daily in production. Every review comes from hands-on experience building real products — DealPropFirm, ThePlanetIndicator, PropFirmsCodes, and many more. We don't just review tools — we build and ship with them every day.

Written and tested by developers who build with these tools daily.