Skip to content
Refreshed every 48 hours· September 17, 2026

The AI Model Index

Benchmarks say one thing, people using the models say another. We pull both — plus real market pricing and measured speed — and score 250 models on one scale. 122 of them carry a machine score and a human score, which is where the disagreements show up.

28 of these models have a hands-on review on this site.

What feeds the index

  • OpenRouter250 models

    live market pricing

  • Artificial Analysis240 models

    benchmark indices

  • LMArena132 models

    human preference votes

122 models scored by machines and humans

Rank by

The current top 20

Benchmark indices cross-checked against human preference votes. The bar under each score splits it into the five things we weigh — same total, different shape means a different model for a different job.

  • 1

    Claude Fable 5.1

    anthropic

    87.8intelligence
    Intelligence
    53.4
    Human Elo
    1508
    Blended /M
    $40
    Throughput
    67 t/s
  • 2

    Claude Opus 5

    anthropic

    85.7intelligence
    Intelligence
    50.7
    Human Elo
    1505
    Blended /M
    $20
    Throughput
    50 t/s
    Review
  • 3

    Claude Fable 5

    anthropic

    82.6intelligence
    Intelligence
    49.7
    Human Elo
    1493
    Blended /M
    $40
    Throughput
    67 t/s
    Review
  • 4

    Muse Spark 1.3

    meta

    82.4intelligence
    Intelligence
    48.2
    Human Elo
    1490
    Blended /M
    $3.5
    Throughput
    242 t/s
  • 5

    Claude Fable 5.1 (batch)

    anthropic

    81.5intelligence
    Intelligence
    53.4
    Human Elo
    Blended /M
    Throughput
  • 6

    Qwen3.8 Max (0902)

    qwen

    80.3intelligence
    Intelligence
    45.4
    Human Elo
    1481
    Blended /M
    $5.0
    Throughput
    40 t/s
  • 7OPEN

    GLM 5.3

    z-ai

    78.7intelligence
    Intelligence
    44.9
    Human Elo
    1475
    Blended /M
    $3.7
    Throughput
    67 t/s
    Review
  • 8

    Gemini 3.8 Flash

    google

    78.0intelligence
    Intelligence
    41.2
    Human Elo
    1495
    Blended /M
    $3.0
    Throughput
    341 t/s
  • 9

    Claude Opus 5 (batch)

    anthropic

    78.0intelligence
    Intelligence
    50.7
    Human Elo
    Blended /M
    Throughput
  • 10

    Qwen3.6 Max Preview

    qwen

    78.0intelligence
    Intelligence
    Human Elo
    1446
    Blended /M
    Throughput
  • 11OPEN

    Kimi K3

    moonshotai

    77.7intelligence
    Intelligence
    43.8
    Human Elo
    1472
    Blended /M
    $12
    Throughput
    36 t/s
    Review
  • 12

    GPT-6 Astra (batch)

    openai

    77.5intelligence
    Intelligence
    52.8
    Human Elo
    Blended /M
    Throughput
  • 13

    GPT-6 Astra

    openai

    77.4intelligence
    Intelligence
    52.8
    Human Elo
    1444
    Blended /M
    $40
    Throughput
    53 t/s
  • 14

    Claude Opus 4.7

    anthropic

    76.5intelligence
    Intelligence
    40.7
    Human Elo
    1490
    Blended /M
    $20
    Throughput
    48 t/s
    Review
  • 15OPEN

    GLM 5.3 Flash

    z-ai

    76.4intelligence
    Intelligence
    41.9
    Human Elo
    1472
    Blended /M
    $0.41
    Throughput
    116 t/s
  • 16

    Gemini 3.7 Flash

    google

    75.4intelligence
    Intelligence
    39.6
    Human Elo
    1491
    Blended /M
    $3.0
    Throughput
    300 t/s
  • 17

    Claude Fable 5 (batch)

    anthropic

    75.0intelligence
    Intelligence
    49.7
    Human Elo
    Blended /M
    Throughput
  • 18

    GPT-5.4

    openai

    73.8intelligence
    Intelligence
    39.0
    Human Elo
    1470
    Blended /M
    $12
    Throughput
    136 t/s
    Review
  • 19

    Claude Opus 4.8

    anthropic

    73.8intelligence
    Intelligence
    42.0
    Human Elo
    1461
    Blended /M
    $20
    Throughput
    57 t/s
    Review
  • 20

    GPT-5.6 Sol (batch)

    openai

    73.2intelligence
    Intelligence
    47.1
    Human Elo
    Blended /M
    Throughput

Put two models head to head

Benchmarks, human votes, speed and real pricing on the same screen. The highlighted side wins that line.

Claude Fable 5.1

anthropic

87.8

intelligence · wins 6 lines

VS

Claude Opus 5

anthropic

85.7

intelligence · wins 7 lines

Our composite score

Our composite score
60.0Overall61.8
88Intelligence40% of the total86
0Price / perf20% of the total13
85Capabilities15% of the total85
39Speed10% of the total40
55Ecosystem15% of the total55

Measured intelligence

Measured intelligence
53.4Intelligence indexArtificial Analysis50.7
81.6Coding indexArtificial Analysis78.0
58.0Agentic indexArtificial Analysis56.2
1508 · 6k votesHuman preferenceLMArena Elo1505 · 21k votes

Speed

Speed
67 t/sOutput throughputtokens/second50 t/s
281.94 sTime to first token52.58 s

Price per million tokens

Price per million tokens
$10Input$5.0
$50Output$25
$40Blended 3:1typical real-world mix$20

Capabilities

Capabilities
1MContext window1M
128kMax output128k
YesReasoningYes
NoOpen weightsNo
text, image, fileModalitiestext, image, file
1Providers serving it1
Knowledge cutoff
No review published yetOur Claude Opus 5 review

Every model we track

Filter it down to the models that fit your constraints, then open a row for the full measurement set.

250 of 250 models

#Modelintelligence & profileIn / OutActions
1
Claude Fable 5.1
anthropic
87.8
$10 / $50
2anthropic
85.7
$5.0 / $25
3anthropic
82.6
$10 / $50
4
Muse Spark 1.3
meta
82.4
$1.3 / $4.3
5
Claude Fable 5.1 (batch)
anthropic
81.5
$5.0 / $25
6
Qwen3.8 Max (0902)
qwen
80.3
$2.0 / $6.0
7z-ai
78.7
$1.4 / $4.4
8
Gemini 3.8 Flash
google
78.0
$0.75 / $3.8
9
Claude Opus 5 (batch)
anthropic
78.0
$2.5 / $13
10
Qwen3.6 Max Preview
qwen
78.0
$1.0 / $6.2
11moonshotai
77.7
$3.0 / $15
12
GPT-6 Astra (batch)
openai
77.5
$5.0 / $25
13
GPT-6 Astra
openai
77.4
$10 / $50
14anthropic
76.5
$5.0 / $25
15
GLM 5.3 FlashOPEN
z-ai
76.4
$0.09 / $0.30
16
Gemini 3.7 Flash
google
75.4
$0.75 / $3.8
17
Claude Fable 5 (batch)
anthropic
75.0
$5.0 / $25
18openai
73.8
$2.5 / $15
19anthropic
73.8
$5.0 / $25
20
GPT-5.6 Sol (batch)
openai
73.2
$1.0 / $5.0
21openai
73.2
$2.0 / $10
22openai
73.1
$5.0 / $30
23x-ai
72.9
$2.0 / $6.0
24
GLM 5.3 (batch)OPEN
z-ai
71.9
$0.70 / $2.2
25x-ai
71.0
$2.0 / $6.0
26meta
70.5
$1.3 / $4.3
27
Kimi K3 (batch)OPEN
moonshotai
70.5
$3.0 / $15
28z-ai
70.4
$1.4 / $4.4
29google
70.3
$1.5 / $9.0
30
Gemini 3.6 Flash
google
70.2
$0.75 / $3.8
31anthropic
70.0
$2.0 / $10
32
DeepSeek V4 Pro 0813OPEN
deepseek
69.6
$0.66 / $2.0
33
DeepSeek V4 Pro 0423OPEN
deepseek
69.6
$1.6 / $3.2
34
GLM 5.3 Flash (batch)OPEN
z-ai
68.1
$0.07 / $0.25
35
Qwen3.8 27BOPEN
qwen
67.9
$0.21 / $2.5
36
GPT-5.6 Terra (batch)
openai
67.1
$1.0 / $6.0
37openai
67.1
$2.0 / $12
38
Qwen3.7 Max
qwen
66.7
$1.5 / $4.4
39
Qwen3.8 2.4T A95B (batch)OPEN
qwen
66.5
$2.0 / $6.0
40
Qwen3.8 2.4T A95BOPEN
qwen
66.5
$2.0 / $6.0
41
DeepSeek V4 Flash 0731OPEN
deepseek
66.4
$0.06 / $0.12
42deepseek
66.4
$0.09 / $0.18
43
Claude Opus 4.8 (batch)
anthropic
65.9
$2.5 / $13
44google
65.7
$2.0 / $12
45
Gemini 3.8 Flash (batch)
google
65.3
$0.38 / $1.9
46anthropic
65.3
$5.0 / $25
47
Qwen3.5-Flash
qwen
65.2
$0.07 / $0.26
48
Muse Spark 1.2
meta
64.2
$1.3 / $4.3
49anthropic
64.0
$3.0 / $15
50
MiMo-V2.5-ProOPEN
xiaomi
63.3
$0.43 / $0.87
51
Claude Sonnet 5 (batch)
anthropic
63.1
$1.0 / $5.0
52
GLM 5.1OPEN
z-ai
62.8
$0.97 / $3.0
53moonshotai
62.5
$0.95 / $4.0
54
Gemini 3.7 Flash (batch)
google
62.3
$0.38 / $1.9
55
GPT-5.6 Luna (batch)
openai
61.9
$0.10 / $0.60
56openai
61.9
$0.20 / $1.2
57
GPT-5.5 (batch)
openai
61.7
$2.5 / $15
58minimax
61.5
$0.30 / $1.2
59
Qwen3.6 Plus
qwen
60.7
$0.33 / $1.9
60
Qwen3.7 Plus
qwen
60.5
$0.32 / $1.3
61
Gemini 3 Flash Preview
google
60.5
$0.50 / $3.0
62
DeepSeek V4 Flash Vision ExpOPEN
deepseek
60.2
$0.22 / $0.66
63
Hy3OPEN
tencent
60.2
$0.13 / $0.53
64
Hy3 previewOPEN
tencent
60.2
$0.18 / $0.60
65
DeepSeek V4 Pro 0813 (batch)OPEN
deepseek
60.1
$0.66 / $2.0
66
GPT-5.1
openai
59.3
$1.3 / $10
67
GLM 5OPEN
z-ai
58.9
$0.60 / $1.9
68
InklingOPEN
thinkingmachines
58.8
$1.0 / $4.0
69
Kimi K2.5OPEN
moonshotai
58.8
$0.45 / $2.3
70
DeepSeek V4 Flash 0731 (batch)OPEN
deepseek
58.7
$0.11 / $0.33
71
GLM 5.2 (batch)OPEN
z-ai
57.5
$0.70 / $2.2
72
GLM 5.2 (free)OPEN
z-ai
57.5
Free / Free
73openai
56.8
$1.8 / $14
74
GPT-5.2 Chat
openai
56.8
$1.8 / $14
75
GLM 4.7OPEN
z-ai
56.6
$0.40 / $1.8
76
DeepSeek V4.1 FlashOPEN
deepseek
56.4
$0.15 / $0.60
77
Nova 2 Lite
amazon
55.9
$0.30 / $2.5
78
Inkling SmallOPEN
thinkingmachines
55.6
$0.45 / $1.2
79
MiMo-V2.5OPEN
xiaomi
55.5
$0.14 / $0.28
80
GLM 4.6OPEN
z-ai
55.5
$0.43 / $1.8
81
Gemini 3.5 Flash Lite
google
55.4
$0.30 / $2.5
82
DeepSeek V3.2OPEN
deepseek
54.6
$0.27 / $0.40
83
DeepSeek V3.2 ExpOPEN
deepseek
54.6
$0.27 / $0.41
84
GPT-5.4 Mini
openai
54.6
$0.75 / $4.5
85
Gemini 3.6 Flash (batch)
google
54.5
$0.38 / $1.9
86
GLM 5V Turbo
z-ai
54.5
$1.2 / $4.0
87
Qwen3.5 397B A17BOPEN
qwen
53.3
$0.55 / $3.5
88
Qwen3 Max Thinking
qwen
53.3
$0.78 / $3.9
89
Solar Pro 4
upstage
53.1
$0.09 / $0.36
90
Gemini 3.5 Flash (batch)
google
52.8
$0.75 / $4.5
91
Gemma 4 26B A4B OPEN
google
52.7
$0.09 / $0.30
92
MiniMax M2.7OPEN
minimax
52.1
$0.30 / $1.2
93
Gemini 2.5 Pro
google
51.7
$1.3 / $10
94
Gemini 2.5 Pro Preview 06-05
google
51.7
$1.3 / $10
95openai
51.6
$1.3 / $10
96
Gemma 4 31BOPEN
google
51.1
$0.09 / $0.34
97
Claude Sonnet 4.6 (batch)
anthropic
51.1
$1.5 / $7.5
98
MiniMax M1
minimax
50.7
$0.40 / $2.2
99x-ai
50.6
$1.3 / $2.5
100
Qwen3.5-27BOPEN
qwen
50.3
$0.20 / $1.6
101
Qwen3 Max
qwen
49.2
$0.78 / $3.9
102
Qwen3.5-122B-A10BOPEN
qwen
49.1
$0.26 / $2.1
103
Mistral Medium 3.5
mistralai
49.1
$1.5 / $7.5
104
DeepSeek V3.1 TerminusOPEN
deepseek
48.7
$0.27 / $1.0
105
MiniMax M3 (batch)OPEN
minimax
48.4
$0.30 / $1.2
106
Qwen3.5-35B-A3BOPEN
qwen
48.3
$0.16 / $1.3
107
GPT-5.4 Nano
openai
47.9
$0.20 / $1.3
108
MiniMax M2.1OPEN
minimax
46.7
$0.30 / $1.2
109
GPT-5.3-Codex
openai
46.4
$1.8 / $14
110
Gemini 3.1 Flash Lite
google
46.0
$0.25 / $1.5
111
Gemini 3.1 Flash Lite Preview
google
46.0
$0.25 / $1.5
112
GLM 4.5OPEN
z-ai
46.0
$0.60 / $2.2
113
R1 0528OPEN
deepseek
45.9
$0.50 / $2.1
114
Step 3.5 FlashOPEN
stepfun
45.5
$0.10 / $0.30
115
DeepSeek V3.1OPEN
deepseek
45.3
$0.25 / $0.95
116
Qwen3 VL 235B A22B InstructOPEN
qwen
45.2
$0.21 / $1.9
117
Qwen3 VL 235B A22B ThinkingOPEN
qwen
45.2
$0.40 / $4.0
118
Kimi K2 0711OPEN
moonshotai
44.8
$0.57 / $2.3
119
Gemini 2.5 Flash
google
44.5
$0.30 / $2.5
120
Gemini 3.1 Pro Preview (batch)
google
44.5
$1.0 / $6.0
121
Ling 3.0 Flash VLOPEN
inclusionai
44.4
$0.06 / $0.18
122
Ling 3.0 Flash VL (free)OPEN
inclusionai
44.4
Free / Free
123
MiniMax M2.5OPEN
minimax
43.8
$0.27 / $1.1
124
Qwen3 Next 80B A3B ThinkingOPEN
qwen
43.8
$0.15 / $1.2
125
Qwen3 Next 80B A3B InstructOPEN
qwen
43.8
$0.09 / $1.1
126
Kimi K2.7 CodeOPEN
moonshotai
43.7
$0.71 / $3.2
127
Qwen3 235B A22B Instruct 2507OPEN
qwen
43.3
$0.09 / $0.35
128
KAT-Coder-Pro V2
kwaipilot
42.7
$0.30 / $1.2
129
Ling 3.0 Flash Fin
inclusionai
42.3
$0.06 / $0.18
130
Inkling Small (free)OPEN
thinkingmachines
42.3
Free / Free
131
Ling 3.0 Flash Fin (free)
inclusionai
42.3
Free / Free
132
Qwen3 235B A22B Thinking 2507OPEN
qwen
41.6
$0.23 / $2.3
133
Claude Opus 4.5
anthropic
41.6
$5.0 / $25
134
Inkling (free)OPEN
thinkingmachines
41.4
Free / Free
135
Inkling (batch)OPEN
thinkingmachines
41.4
$1.0 / $4.0
136
o1
openai
41.3
$15 / $60
137
Kimi K2 0905OPEN
moonshotai
41.1
$0.60 / $2.5
138
GPT-5.2-Codex
openai
40.7
$1.8 / $14
139
GPT-5 Mini
openai
40.5
$0.25 / $2.0
140
GPT-5.4 Mini (batch)
openai
40.2
$0.38 / $2.3
141
GLM 4.5 AirOPEN
z-ai
38.7
$0.13 / $0.85
142
MiniMax M2OPEN
minimax
38.6
$0.26 / $1.0
143
Nemotron 3 Ultra (free)OPEN
nvidia
38.2
Free / Free
144
Nemotron 3 UltraOPEN
nvidia
38.2
$0.63 / $3.1
145
gpt-oss-120bOPEN
openai
38.1
$0.04 / $0.17
146
GLM 5 Turbo
z-ai
38.0
$1.2 / $4.0
147
GLM 4.6VOPEN
z-ai
37.8
$0.30 / $0.90
148
Qwen3.6 27BOPEN
qwen
37.7
$0.30 / $2.0
149
R1OPEN
deepseek
37.6
$0.70 / $2.5
150
GLM 4.7 FlashOPEN
z-ai
37.2
$0.06 / $0.40
151
GPT-5.4 Nano (batch)
openai
37.0
$0.10 / $0.63
152x-ai
36.7
$1.3 / $2.5
153
Mercury 2
inception
36.5
$0.25 / $0.75
154
Ling 3.0 FlashOPEN
inclusionai
36.3
$0.02 / $0.06
155
Qwen3 30B A3B Instruct 2507OPEN
qwen
36.2
$0.05 / $0.19
156
DeepSeek V3 0324OPEN
deepseek
36.2
$0.25 / $1.0
157
Grok 4.3 (batch)
x-ai
36.0
$1.0 / $2.0
158
Claude Sonnet 4.5 (batch)
anthropic
35.9
$1.5 / $7.5
159
Claude Sonnet 4.5
anthropic
35.9
$3.0 / $15
160
Qwen3 Coder 480B A35BOPEN
qwen
35.7
$0.30 / $1.0
161
Gemini 3.5 Flash Lite (batch)
google
35.6
$0.15 / $1.3
162
Qwen3 235B A22BOPEN
qwen
35.2
$0.46 / $1.8
163
GPT-5.1-Codex-Max
openai
33.9
$1.3 / $10
164
GPT-5.1-Codex
openai
33.9
$1.3 / $10
165
o3 Mini
openai
33.5
$1.1 / $4.4
166
Step 3.7 FlashOPEN
stepfun
33.2
$0.20 / $1.1
167
Trinity Large ThinkingOPEN
arcee-ai
32.9
$0.25 / $0.80
168
Llama 3.3 70B InstructOPEN
meta-llama
32.6
$0.10 / $0.32
169
LongCat 2.0OPEN
meituan
32.3
$0.30 / $1.2
170
GPT-5 Nano
openai
31.6
$0.05 / $0.40
171
o3 Mini High
openai
31.4
$1.1 / $4.4
172
o3 Pro
openai
31.3
$20 / $80
173
Gemma 3 27BOPEN
google
30.5
$0.08 / $0.45
174
Qwen3.6 35B A3BOPEN
qwen
30.4
$0.10 / $0.90
175
DeepSeek V3OPEN
deepseek
30.4
$0.26 / $1.0
176
Granite 4.2 8BOPEN
ibm-granite
29.9
$0.06 / $0.25
177
Qwen3 32BOPEN
qwen
29.9
$0.08 / $0.28
178
GLM 4.5VOPEN
z-ai
29.5
$0.60 / $1.8
179
Llama 3.1 70B InstructOPEN
meta-llama
29.2
$0.40 / $0.40
180
GPT-5.1-Codex-Mini
openai
29.1
$0.25 / $2.0
181
o3
openai
28.9
$2.0 / $8.0
182
Kimi K2 ThinkingOPEN
moonshotai
28.7
$0.60 / $2.5
183
Claude Haiku 4.5 (batch)
anthropic
28.4
$0.50 / $2.5
184anthropic
28.4
$1.0 / $5.0
185
GPT-4o (2024-05-13)
openai
27.7
$5.0 / $15
186
Qwen3 30B A3BOPEN
qwen
27.4
$0.12 / $0.50
187
Mistral Medium 3.5 (batch)
mistralai
27.0
$0.75 / $3.8
188
Qwen3 30B A3B Thinking 2507OPEN
qwen
26.5
$0.20 / $2.4
189
Gemma 3 12BOPEN
google
26.4
$0.05 / $0.15
190
Gemma 4 31B (free)OPEN
google
25.4
Free / Free
191
gpt-oss-20bOPEN
openai
24.8
$0.03 / $0.13
192
o4 Mini High
openai
23.9
$1.1 / $4.4
193
o4 Mini
openai
23.9
$1.1 / $4.4
194
Qwen3.5-9BOPEN
qwen
23.7
$0.10 / $0.15
195
GPT-4o-mini (2024-07-18)
openai
23.2
$0.15 / $0.60
196
GPT-4o (2024-08-06)
openai
22.9
$2.5 / $10
197
Gemini 2.5 Pro (batch)
google
22.4
$0.63 / $5.0
198
Nemotron 3 SuperOPEN
nvidia
21.7
$0.08 / $0.45
199
Nemotron 3 Super (free)OPEN
nvidia
21.7
Free / Free
200
Gemma 2 27BOPEN
google
21.4
$0.65 / $0.65
201
Gemma 3 4BOPEN
google
21.3
$0.05 / $0.10
202
Command R+ (08-2024)
cohere
20.7
$2.5 / $10
203
Mistral Large 2407
mistralai
20.1
$2.0 / $6.0
204
GPT-5 Mini (batch)
openai
19.9
$0.13 / $1.0
205
gpt-oss-120b (batch)OPEN
openai
19.4
$0.15 / $0.60
206
Nemotron 3.5 LightningOPEN
nvidia
19.3
$0.08 / $0.20
207
Nemotron 3.5 Lightning (free)OPEN
nvidia
19.3
Free / Free
208
Qwen3 Coder NextOPEN
qwen
18.6
$0.12 / $0.80
209
GPT-4.1
openai
18.1
$2.0 / $8.0
210
o1-pro
openai
17.7
$150 / $600
211
North Mini Code (free)OPEN
cohere
17.6
Free / Free
212
GPT-4.1 Mini
openai
17.2
$0.40 / $1.6
213
Qwen3 VL 32B InstructOPEN
qwen
17.0
$0.10 / $0.42
214
Devstral 2 2512OPEN
mistralai
16.9
$0.40 / $2.0
215
Sonar Reasoning Pro
perplexity
16.9
$2.0 / $8.0
216
Mistral Small 4OPEN
mistralai
16.1
$0.15 / $0.60
217
Mistral Small 4 (batch)OPEN
mistralai
16.1
$0.07 / $0.30
218
Mistral Small 3OPEN
mistralai
15.8
$0.05 / $0.08
219
GPT-4 Turbo
openai
14.6
$10 / $30
220
Command AOPEN
cohere
14.0
$2.5 / $10
221
Qwen3 Coder 30B A3B InstructOPEN
qwen
13.7
$0.07 / $0.28
222
Qwen3 VL 30B A3B InstructOPEN
qwen
13.6
$0.13 / $0.52
223
Qwen3 VL 30B A3B ThinkingOPEN
qwen
13.6
$0.20 / $2.4
224
Mistral Medium 3.1 (batch)
mistralai
13.6
$0.20 / $1.0
225
Mistral Medium 3.1
mistralai
13.6
$0.40 / $2.0
226
Mistral Large 3 2512 (batch)
mistralai
13.4
$0.25 / $0.75
227
Phi 4OPEN
microsoft
13.0
$0.07 / $0.14
228
Mistral Medium 3
mistralai
12.9
$0.40 / $2.0
229
Sonar
perplexity
12.4
$1.0 / $1.0
230
Gemini 2.5 Flash Lite
google
12.1
$0.10 / $0.40
231
GPT-4o (2024-11-20)
openai
12.0
$2.5 / $10
232
GPT-4o
openai
12.0
$2.5 / $10
233
Qwen3 VL 8B InstructOPEN
qwen
11.7
$0.12 / $0.46
234
Qwen3 VL 8B ThinkingOPEN
qwen
11.7
$0.18 / $2.1
235
GPT-4.1 Nano
openai
11.5
$0.10 / $0.40
236
Llama 4 MaverickOPEN
meta-llama
11.4
$0.19 / $0.65
237
R1 Distill Llama 70BOPEN
deepseek
11.3
$0.80 / $0.80
238
GPT-4
openai
11.2
$30 / $60
239
Sonar Pro
perplexity
10.9
$3.0 / $15
240
Nemotron 3 Nano 30B A3BOPEN
nvidia
10.7
$0.06 / $0.24
241
Solar Pro 3
upstage
10.6
$0.15 / $0.60
242
GPT-4o-mini
openai
10.6
$0.15 / $0.60
243
Mistral Large 3 2512
mistralai
10.6
$0.50 / $1.5
244
Command R (08-2024)
cohere
9.8
$0.15 / $0.60
245
Llama 3.1 8B InstructOPEN
meta-llama
9.6
$0.05 / $0.08
246
Saba
mistralai
9.3
$0.20 / $0.60
247
Qwen3 14BOPEN
qwen
8.9
$0.12 / $0.24
248
Ministral 3 14B 2512OPEN
mistralai
8.7
$0.20 / $0.20
249
Hermes 3 70B InstructOPEN
nousresearch
8.6
$0.70 / $0.70
250
Mistral Large
mistralai
8.3
$2.0 / $6.0

How we score

The raw measurements come from third parties, credited below. What we add is the cross-source aggregation and the scoring: five axes on absolute 0-100 scales, weighted into one number. Absolute scales matter — a model scoring 78 today means the same as a model scoring 78 six months ago, so the history stays comparable as new models arrive.

Intelligence

40%

Average of every published capability index we can cross-check — general intelligence, coding and agentic benchmarks.

Price / performance

20%

Real market price per million tokens weighted against measured capability. A cheap weak model and an expensive strong one can score the same.

Capabilities

15%

Context window, maximum output, modalities, reasoning support, prompt caching and how recent the knowledge cutoff is.

Speed

10%

Measured output throughput (tokens per second) and time to first token. A reasoning model thinks before it answers, so its first token can take minutes — that is the trade it makes for capability, and it scores low here by design.

Ecosystem

15%

How many providers serve the model, and whether the weights are open — availability and independence from a single vendor.

What changed recently

Every collection is archived, so price cuts and ranking moves stay traceable. The sources publish today's numbers; this is the trail behind them.

Price moves · output per million tokens

  • Meta: Llama 3.3 70B Instruct$0.32 → $0.71 122%
  • DeepSeek: DeepSeek V3.1$0.95 → $1.65 74%
  • DeepSeek: DeepSeek V4 Flash 0731$0.28 → $0.12 57%
  • DeepSeek: DeepSeek V4 Pro 0423$3.20 → $1.74 46%
  • Qwen: Qwen3.5-35B-A3B$1.25 → $1.80 44%
  • Qwen: Qwen3 235B A22B Instruct 2507$0.55 → $0.35 36%
  • OpenAI: GPT-5.6 Sol (batch)$7.50 → $5.00 33%
  • OpenAI: GPT-5.6 Sol$15.00 → $10.00 33%

Ranking moves · by capability

  • NVIDIA: Nemotron 3 Ultra (free)#113 → #120 7
  • inclusionAI: Ling 3.0 Flash#111 → #116 5
  • DeepSeek: R1 0528#112 → #117 5
  • OpenAI: GPT-5.1-Codex-Max#116 → #121 5
  • OpenAI: GPT-5.1-Codex#117 → #122 5
  • Z.ai: GLM 4.5#118 → #123 5

How to read this index

The default ranking is capability, so the most expensive model can sit above a free one. Switch the criterion at the top of the page and the order changes completely — that is the point. Four questions come up constantly, and here are the answers.

Why is a $10 model ahead of a free one?

Because the default ranking measures what a model can do, not what it costs. Price is a separate axis, and it has its own ranking one click away. Ranking on the composite score instead buried the frontier models behind cheap ones and turned the whole index into a price chart, which is not what a ranked list of models is for. Today that puts Anthropic: Claude Fable 5.1 first with 87.8 on capability at $50.00 per million output tokens.

What happens when benchmarks and humans disagree?

Both count, equally. Benchmarks test verifiable problems; the arena records which answer people actually preferred. When a model carries both, each contributes half its intelligence score. The widest gap in the index right now is Google: Gemini 2.5 Pro: 24 from benchmarks against 81 from human votes — a 57-point disagreement that a single-source ranking would hide entirely.

What is the coloured bar under each score?

The score itself, split into the five things we weigh. Each segment is that axis's contribution, so the full bar length is the composite score. Two models can share a total and have opposite shapes — one fast and cheap, the other slow and brilliant.

  • Intelligence
  • Price / performance
  • Capabilities
  • Speed
  • Ecosystem

Why are some cells empty?

Because that measurement does not exist yet, and we would rather show a dash than invent a number. Of the 250 models listed, 240 carry benchmark indices and 132 carry human ratings. A missing value never counts as zero — it simply leaves that axis out of the model's score, and the “Double-validated” filter isolates the 122 models measured both ways.

Frequently asked questions

What is the ThePlanetTools AI Index?

A ranking of 250 large language models built by cross-referencing three independent sources: live market pricing from OpenRouter, benchmark indices from Artificial Analysis, and human preference votes from LMArena. We aggregate those measurements and compute our own score across five weighted axes. The raw measurements belong to their sources; the scoring and the ranking are ours.

How is the ranking ordered by default?

By measured capability, not by price. A model that costs ten times more but answers harder questions ranks higher by default. Three alternative orderings are one click away — price/performance, speed, and the weighted composite score — because "best" means something different depending on whether you are paying the bill or waiting for the response.

What does the intelligence score actually measure?

Two things, weighted equally when both exist. Machine benchmarks contribute the general intelligence index (50%), the coding index (25%) and the agentic index (25%), each normalised against its own ceiling. Human preference contributes the LMArena Elo rating, mapped onto the same 0-100 scale. A model measured by only one of the two is scored on that one alone.

What does "double-validated" mean?

A model carrying both a machine benchmark score and a human preference rating. 122 of the 250 models in the index qualify. It matters because the two methods disagree more often than you would expect: benchmarks measure what is testable, the arena measures what people actually prefer after using the model.

Why do benchmark scores and human votes disagree?

They measure different things. Benchmarks test verifiable problems with known answers — maths, code that compiles, factual recall. Human votes reward tone, instruction-following, formatting and refusal behaviour, which no benchmark captures. A model can top the benchmarks and lose blind head-to-head votes, and the reverse happens just as often.

Why is the maths index excluded from the score?

Because nearly every current model scored in the 90s on it, so it separated nothing. Worse, averaging it in rewarded models measured on maths and almost nothing else — one model with no coding or agentic score reached a higher intelligence score than models measured on everything. Artificial Analysis has since stopped publishing that index on the endpoint we use, which changes nothing for us: it was already out of the score.

How is price/performance calculated?

Measured capability divided by the blended price per million tokens, weighted one part input to three parts output, which reflects typical real-world usage. The result is placed on a logarithmic scale because prices span four orders of magnitude. Free or unpriced models receive a neutral 50 rather than a perfect score, since they are not comparable.

Are the prices real?

Yes. They come from OpenRouter, which publishes live per-token rates charged by the providers actually serving each model, in US dollars per million tokens. Input and output are shown separately because the gap between them is often large and matters more than the headline figure.

How often is the index updated?

Every 48 hours. Each collection re-fetches all three sources, recalculates every score, and writes a historical snapshot so price changes and ranking movements stay traceable. Models that disappear from the sources are marked inactive rather than deleted, so their history survives.

What does "open weights" mean in this index?

A model whose weights can be downloaded and self-hosted, as opposed to one reachable only through a vendor API. 124 of the 250 models qualify. Open weights add points on the ecosystem axis because they remove vendor lock-in, but a closed model served by many providers is not penalised to zero.

Why do the scores use absolute scales instead of ranking against each other?

So the numbers stay comparable over time. If scores were normalised against the current field, every new model release would silently move every existing score, and a snapshot from six months ago would mean nothing. On an absolute scale, a model scoring 78 today means the same as a model scoring 78 last year.

Which models here have been reviewed hands-on?

28 of the 250 models link to a full review on this site, written after actually using the model. Those links appear directly on each model card and in the comparison table. The rest are scored from published measurements only, which is stated rather than hidden.

Data sources

  • OpenRouter — live market pricing, context windows, modalities and availability across providers.
  • Artificial Analysis — independent intelligence, coding, agentic and math indices, plus measured throughput and time to first token.
  • LMArena — human-preference Elo ratings from their public leaderboard dataset.

All raw metrics belong to their respective sources and are reproduced with attribution. The composite scores, weighting and rankings on this page are ThePlanetTools.ai's own computation and should be attributed to us, not to the data providers.

Anthony M. — Founder & Lead Reviewer
Anthony M.Verified Builder

We're developers and SaaS builders who use these tools daily in production. Every review comes from hands-on experience building real products — DealPropFirm, ThePlanetIndicator, PropFirmsCodes, and many more. We don't just review tools — we build and ship with them every day.

Written and tested by developers who build with these tools daily.