Back to AI Coding

AI Model Ranking

AI Coding Model Rankings

Snapshot: 2026-09-16 · Last reviewed: 2026-09-16

Methodology

Metrics follow a snapshot of the Artificial Analysis public LLM leaderboard: Intelligence Index, cost per task, and speed reflect benchmark performance at the snapshot date. Provider pricing, availability, and rankings change frequently, and a high benchmark score does not guarantee the best result for a specific codebase or workflow.

Source
Artificial Analysis
Last reviewed
2026-09-16

Top 50 models

01

Current reasoning model

Anthropic - Claude Fable - 5.1 (max)

Claude Fable 5.1 max still leads this snapshot at Intelligence Index 53 with 1M context. First-chunk latency remains very high; AA scores are not comparable 1:1 with the previous snapshot.

Best for: Highest-effort Claude Fable work when Intelligence Index matters more than wait time.

Anthropic

Context
1M
AA Index
53
Cost per task
$7.63
Speed
66 tok/s
First chunk
234.98s
Total response
242.56s
02

Current reasoning model

Anthropic - Claude Fable - 5.1 (xhigh)

Claude Fable 5.1 xhigh rises to #2 at Intelligence Index 53 with 1M context. Cost per task is now public; first-chunk wait is much shorter than Fable 5.1 max.

Best for: High-effort Fable 5.1 coding when max first-chunk wait is too high and you still want near-top Intelligence Index.

Anthropic

Context
1M
AA Index
53
Cost per task
$5.98
Speed
58 tok/s
First chunk
85.53s
Total response
94.12s
03

Current reasoning model

OpenAI - GPT Astra - 6 (max)

GPT-6 Astra max is #3 with Intelligence Index 53 and 1M context. First-chunk wait is the longest in the top 10.

Best for: OpenAI-centered frontier agents that can absorb very long first-chunk delay.

OpenAI

Context
1M
AA Index
53
Cost per task
$3.26
Speed
54 tok/s
First chunk
282.50s
Total response
291.76s
04

Current reasoning model

OpenAI - GPT Astra - 6 (xhigh)

GPT-6 Astra xhigh sits just below max with Intelligence Index 53, lower cost per task, and still-long first-chunk latency.

Best for: OpenAI coding agents that need near-max Astra quality at a lower task cost than max.

OpenAI

Context
1M
AA Index
53
Cost per task
$2.31
Speed
55 tok/s
First chunk
117.71s
Total response
126.83s
05

Current reasoning model

Anthropic - Claude Fable - 5.1 (high)

Claude Fable 5.1 high jumps into the top 5 at Intelligence Index 51 with 1M context. First-chunk latency is much lower than Fable max/xhigh.

Best for: Claude Fable 5.1 coding that still needs a high Intelligence Index without the max-tier first-chunk wait.

Anthropic

Context
1M
AA Index
51
Cost per task
$3.91
Speed
51 tok/s
First chunk
10.95s
Total response
20.82s
06

Current reasoning model

OpenAI - GPT Astra - 6 (high)

GPT-6 Astra high holds Intelligence Index 51 with 1M context; first chunk is much shorter than Astra max/xhigh.

Best for: OpenAI Astra agents that need strong reasoning without the max-tier first-chunk wait.

OpenAI

Context
1M
AA Index
51
Cost per task
$1.72
Speed
53 tok/s
First chunk
33.30s
Total response
42.67s
07

Current reasoning model

Anthropic - Claude Opus - 5 (max)

Claude Opus 5 max drops to #7. Intelligence Index 51 with 1M context; first chunk is shorter than Fable 5.1 max and Astra max.

Best for: Frontier Claude repository analysis and high-value agents when Fable 5.1 max wait is too high.

Anthropic

Context
1M
AA Index
51
Cost per task
$5.86
Speed
50 tok/s
First chunk
43.51s
Total response
53.51s
08

Current reasoning model

Anthropic - Claude Fable - 5

Claude Fable 5 stays in the top 10 at Intelligence Index 50 with 1M context; first-chunk wait remains long and cost per task is the highest in the top 10.

Best for: Deep repository analysis and high-value Claude Fable agent tasks where latency is acceptable.

Anthropic

Context
1M
AA Index
50
Cost per task
$8.75
Speed
65 tok/s
First chunk
85.45s
Total response
93.19s
09

Current reasoning model

OpenAI - GPT Astra - 6 (medium)

GPT-6 Astra medium enters the top 10 at Intelligence Index 50 with 1M context and a 4.02s first chunk, much faster to first token than Astra max/xhigh.

Best for: Daily OpenAI Astra coding with 1M context when max/xhigh first-chunk wait is too high.

OpenAI

Context
1M
AA Index
50
Cost per task
$1.54
Speed
53 tok/s
First chunk
4.02s
Total response
13.50s
10

Current reasoning model

Anthropic - Claude Opus - 5 (xhigh)

Opus 5 xhigh keeps Intelligence Index 50 with lower first-chunk latency than Opus 5 max and a lower cost per task.

Best for: Deep multi-file edits and coding agents that still need frontier quality with better cost than max.

Anthropic

Context
1M
AA Index
50
Cost per task
$4.88
Speed
49 tok/s
First chunk
18.47s
Total response
28.58s
11

Current reasoning model

Anthropic - Claude Fable - 5.1 (medium)

Anthropic Claude Fable 5.1 (medium) ranks in this snapshot with Intelligence Index 49, 1M context, cost per task $2.98, median 51 tok/s, and first chunk 7.47s.

Best for: Anthropic coding agents, long-context work, and snapshot-based model comparison.

Anthropic

Context
1M
AA Index
49
Cost per task
$2.98
Speed
51 tok/s
First chunk
7.47s
Total response
17.31s
12

Current reasoning model

Anthropic - Claude Opus - 5 (high)

Anthropic Claude Opus 5 (high) ranks in this snapshot with Intelligence Index 48, 1M context, cost per task $3.61, median 49 tok/s, and first chunk 8.58s.

Best for: Daily Claude agent coding, review, and tool-heavy workflows that still need strong reasoning.

Anthropic

Context
1M
AA Index
48
Cost per task
$3.61
Speed
49 tok/s
First chunk
8.58s
Total response
18.72s
13

Current reasoning model

Meta - Muse Spark - 1.3 (max)

Meta Muse Spark 1.3 (max) ranks in this snapshot with Intelligence Index 48, 1M context, cost per task $1.60, median 221 tok/s, and first chunk 20.57s.

Best for: Fast Meta-stack coding with 1M context when you want high Intelligence Index and much lower wait than Fable/Astra max.

Meta

Context
1M
AA Index
48
Cost per task
$1.60
Speed
221 tok/s
First chunk
20.57s
Total response
31.89s
14

Current reasoning model

OpenAI - GPT Sol - 5.6 (max)

OpenAI GPT Sol 5.6 (max) ranks in this snapshot with Intelligence Index 47, 1M context, cost per task $1.99, median 65 tok/s, and first chunk 99.40s.

Best for: OpenAI coding agents, long-context work, and snapshot-based model comparison.

OpenAI

Context
1M
AA Index
47
Cost per task
$1.99
Speed
65 tok/s
First chunk
99.40s
Total response
107.14s
15

Current reasoning model

Anthropic - Claude Fable - 5.1 (low)

Anthropic Claude Fable 5.1 (low) ranks in this snapshot with Intelligence Index 47, 1M context, cost per task $2.37, median 50 tok/s, and first chunk 5.13s.

Best for: Anthropic coding agents, long-context work, and snapshot-based model comparison.

Anthropic

Context
1M
AA Index
47
Cost per task
$2.37
Speed
50 tok/s
First chunk
5.13s
Total response
15.04s
16

Current reasoning model

OpenAI - GPT Astra - 6 (low)

OpenAI GPT Astra 6 (low) ranks in this snapshot with Intelligence Index 46, 1M context, cost per task $0.82, median 51 tok/s, and first chunk 2.47s.

Best for: OpenAI coding agents, long-context work, and snapshot-based model comparison.

OpenAI

Context
1M
AA Index
46
Cost per task
$0.82
Speed
51 tok/s
First chunk
2.47s
Total response
12.29s
17

Current reasoning model

Alibaba - Qwen - 3.8 Max (0902)

Alibaba Qwen 3.8 Max (0902) ranks in this snapshot with Intelligence Index 45, 984k context, cost per task $5.41, median 41 tok/s, and first chunk 2.72s.

Best for: Qwen-centered coding agents that want the 0902 Max snapshot rather than the undated Max SKU.

Alibaba

Context
984k
AA Index
45
Cost per task
$5.41
Speed
41 tok/s
First chunk
2.72s
Total response
63.21s
18

Current reasoning model

Meta - Muse Spark - 1.3 (xhigh)

Meta Muse Spark 1.3 (xhigh) ranks in this snapshot with Intelligence Index 45, 1M context, cost per task $1.37, median 232 tok/s, and first chunk 22.12s.

Best for: Meta coding agents, long-context work, and snapshot-based model comparison.

Meta

Context
1M
AA Index
45
Cost per task
$1.37
Speed
232 tok/s
First chunk
22.12s
Total response
32.91s
19

Current reasoning model

Anthropic - Claude Opus - 5 (medium)

Anthropic Claude Opus 5 (medium) ranks in this snapshot with Intelligence Index 45, 1M context, cost per task $2.19, median 48 tok/s, and first chunk 5.24s.

Best for: Anthropic coding agents, long-context work, and snapshot-based model comparison.

Anthropic

Context
1M
AA Index
45
Cost per task
$2.19
Speed
48 tok/s
First chunk
5.24s
Total response
15.72s
20

Current reasoning model

Z AI - GLM - 5.3 (max)

Z AI GLM 5.3 (max) ranks in this snapshot with Intelligence Index 45, 1M context, cost per task $2.01, median 67 tok/s, and first chunk 3.09s.

Best for: Z AI coding agents, long-context work, and snapshot-based model comparison.

Z AI

Context
1M
AA Index
45
Cost per task
$2.01
Speed
67 tok/s
First chunk
3.09s
Total response
40.65s
21

Current reasoning model

SpaceXAI - Grok - 4.6 (high)

SpaceXAI Grok 4.6 (high) ranks in this snapshot with Intelligence Index 44, 500k context, cost per task $1.86, median 59 tok/s, and first chunk 36.19s.

Best for: SpaceXAI coding agents, long-context work, and snapshot-based model comparison.

SpaceXAI

Context
500k
AA Index
44
Cost per task
$1.86
Speed
59 tok/s
First chunk
36.19s
Total response
44.65s
22

Current reasoning model

SpaceXAI - Grok - 4.6 (xhigh)

SpaceXAI Grok 4.6 (xhigh) ranks in this snapshot with Intelligence Index 44, 500k context, cost per task $2.32, median 55 tok/s, and first chunk 33.14s.

Best for: SpaceXAI coding agents, long-context work, and snapshot-based model comparison.

SpaceXAI

Context
500k
AA Index
44
Cost per task
$2.32
Speed
55 tok/s
First chunk
33.14s
Total response
42.30s
23

Current reasoning model

OpenAI - GPT Sol - 5.6 (xhigh)

OpenAI GPT Sol 5.6 (xhigh) ranks in this snapshot with Intelligence Index 44, 1M context, cost per task $1.18, median 57 tok/s, and first chunk 33.29s.

Best for: OpenAI coding agents, long-context work, and snapshot-based model comparison.

OpenAI

Context
1M
AA Index
44
Cost per task
$1.18
Speed
57 tok/s
First chunk
33.29s
Total response
42.12s
24

Current reasoning model

Kimi - Kimi - K3 (max)

Kimi K3 (max) ranks in this snapshot with Intelligence Index 44, 1.05M context, cost per task $2.00, median 35 tok/s, and first chunk 4.52s.

Best for: Long-horizon coding agents, frontend generation, regional API stacks, and cost-aware 1M context work.

Kimi

Context
1.05M
AA Index
44
Cost per task
$2.00
Speed
35 tok/s
First chunk
4.52s
Total response
76.42s
25

Current reasoning model

SpaceXAI - Grok - 4.6 (medium)

SpaceXAI Grok 4.6 (medium) ranks in this snapshot with Intelligence Index 43, 500k context, cost per task $1.50, median 54 tok/s, and first chunk 32.75s.

Best for: SpaceXAI coding agents, long-context work, and snapshot-based model comparison.

SpaceXAI

Context
500k
AA Index
43
Cost per task
$1.50
Speed
54 tok/s
First chunk
32.75s
Total response
41.96s
26

Current reasoning model

OpenAI - GPT Sol - 5.6 (high)

OpenAI GPT Sol 5.6 (high) ranks in this snapshot with Intelligence Index 42, 1M context, cost per task $0.81, median 55 tok/s, and first chunk 9.34s.

Best for: OpenAI coding agents, long-context work, and snapshot-based model comparison.

OpenAI

Context
1M
AA Index
42
Cost per task
$0.81
Speed
55 tok/s
First chunk
9.34s
Total response
18.38s
27

Current reasoning model

OpenAI - GPT Terra - 5.6 (max)

OpenAI GPT Terra 5.6 (max) ranks in this snapshot with Intelligence Index 42, 1M context, cost per task $1.40, median 101 tok/s, and first chunk 139.76s.

Best for: OpenAI coding agents, long-context work, and snapshot-based model comparison.

OpenAI

Context
1M
AA Index
42
Cost per task
$1.40
Speed
101 tok/s
First chunk
139.76s
Total response
144.70s
28

Current reasoning model

Z AI - GLM - 5.3 Flash

Z AI GLM 5.3 Flash ranks in this snapshot with Intelligence Index 42, 1M context, cost per task $0.25, median 114 tok/s, and first chunk 2.45s.

Best for: Cost-sensitive GLM coding with 1M context and low first-chunk latency.

Z AI

Context
1M
AA Index
42
Cost per task
$0.25
Speed
114 tok/s
First chunk
2.45s
Total response
24.31s
29

Current reasoning model

Google - Gemini - 3.8 Flash (high)

Google Gemini 3.8 Flash (high) ranks in this snapshot with Intelligence Index 41, 1M context, cost per task $1.24, median 336 tok/s, and first chunk 16.28s.

Best for: Fast Google-stack coding loops, long-file skim, and high-throughput assistant work.

Google

Context
1M
AA Index
41
Cost per task
$1.24
Speed
336 tok/s
First chunk
16.28s
Total response
17.77s
30

Current reasoning model

Alibaba - Qwen - 3.8 Max

Alibaba Qwen 3.8 Max ranks in this snapshot with Intelligence Index 40, 1M context, cost per task $2.67, median 41 tok/s, and first chunk 2.59s.

Best for: Multilingual coding, regional API stacks, and Qwen-centered agent experiments.

Alibaba

Context
1M
AA Index
40
Cost per task
$2.67
Speed
41 tok/s
First chunk
2.59s
Total response
64.27s
31

Current reasoning model

Alibaba - Qwen - 3.8 2.4T A95B

Alibaba Qwen 3.8 2.4T A95B ranks in this snapshot with Intelligence Index 40, 984k context, cost per task $2.16, median 41 tok/s, and first chunk 2.77s.

Best for: Alibaba coding agents, long-context work, and snapshot-based model comparison.

Alibaba

Context
984k
AA Index
40
Cost per task
$2.16
Speed
41 tok/s
First chunk
2.77s
Total response
64.19s
32

Current model; partial public metrics

Google - Gemini - 3.8 Flash (medium)

Google Gemini 3.8 Flash (medium) ranks in this snapshot with Intelligence Index 40, 1M context, cost per task $0.93, speed not fully public, and latency not fully public.

Best for: Google coding agents, long-context work, and snapshot-based model comparison.

Google

Context
1M
AA Index
40
Cost per task
$0.93
Speed
tok/s
First chunk
Total response
33

Current reasoning model

Alibaba - Qwen - 3.8 Flash Next

Alibaba Qwen 3.8 Flash Next ranks in this snapshot with Intelligence Index 40, 256k context, cost per task $0.37, median 52 tok/s, and first chunk 2.67s.

Best for: Alibaba coding agents, long-context work, and snapshot-based model comparison.

Alibaba

Context
256k
AA Index
40
Cost per task
$0.37
Speed
52 tok/s
First chunk
2.67s
Total response
50.33s
34

Current reasoning model

Meta - Muse Spark - 1.2 (xhigh)

Meta Muse Spark 1.2 (xhigh) ranks in this snapshot with Intelligence Index 40, 1.05M context, cost per task $0.97, median 181 tok/s, and first chunk 16.46s.

Best for: Meta coding agents, long-context work, and snapshot-based model comparison.

Meta

Context
1.05M
AA Index
40
Cost per task
$0.97
Speed
181 tok/s
First chunk
16.46s
Total response
30.26s
35

Current reasoning model

Anthropic - Claude Opus - 5 (low)

Anthropic Claude Opus 5 (low) ranks in this snapshot with Intelligence Index 40, 1M context, cost per task $1.10, median 46 tok/s, and first chunk 2.52s.

Best for: Anthropic coding agents, long-context work, and snapshot-based model comparison.

Anthropic

Context
1M
AA Index
40
Cost per task
$1.10
Speed
46 tok/s
First chunk
2.52s
Total response
13.37s
36

Current model; partial public metrics

Google - Gemini - 3.7 Flash (medium)

Google Gemini 3.7 Flash (medium) ranks in this snapshot with Intelligence Index 40*, 1M context, cost per task is not fully public, median 293 tok/s, and first chunk 4.39s. Estimated Intelligence Index.

Best for: Google coding agents, long-context work, and snapshot-based model comparison.

Google

Context
1M
AA Index
40*
Cost per task
Speed
293 tok/s
First chunk
4.39s
Total response
6.09s
37

Current reasoning model

DeepSeek - DeepSeek - V4.1 Flash (max)

DeepSeek V4.1 Flash (max) ranks in this snapshot with Intelligence Index 40, 1M context, cost per task $0.27, median 212 tok/s, and first chunk 1.31s.

Best for: Low-cost DeepSeek coding with 1M context, fast first chunk, and high output speed.

DeepSeek

Context
1M
AA Index
40
Cost per task
$0.27
Speed
212 tok/s
First chunk
1.31s
Total response
13.10s
38

Current reasoning model

OpenAI - GPT Sol - 5.6 (medium)

OpenAI GPT Sol 5.6 (medium) ranks in this snapshot with Intelligence Index 39, 1M context, cost per task $0.50, median 58 tok/s, and first chunk 3.21s.

Best for: OpenAI coding agents, long-context work, and snapshot-based model comparison.

OpenAI

Context
1M
AA Index
39
Cost per task
$0.50
Speed
58 tok/s
First chunk
3.21s
Total response
11.90s
39

Current reasoning model

Google - Gemini - 3.7 Flash (high)

Google Gemini 3.7 Flash (high) ranks in this snapshot with Intelligence Index 39, 1M context, cost per task $0.93, median 292 tok/s, and first chunk 10.22s.

Best for: Google coding agents, long-context work, and snapshot-based model comparison.

Google

Context
1M
AA Index
39
Cost per task
$0.93
Speed
292 tok/s
First chunk
10.22s
Total response
11.93s
40

Current reasoning model

SpaceXAI - Grok - 4.5 (high)

SpaceXAI Grok 4.5 (high) ranks in this snapshot with Intelligence Index 39, 500k context, cost per task $1.04, median 60 tok/s, and first chunk 6.45s.

Best for: SpaceXAI coding agents, long-context work, and snapshot-based model comparison.

SpaceXAI

Context
500k
AA Index
39
Cost per task
$1.04
Speed
60 tok/s
First chunk
6.45s
Total response
14.71s
41

Current reasoning model

Anthropic - Claude Sonnet - 5 (max)

Anthropic Claude Sonnet 5 (max) ranks in this snapshot with Intelligence Index 38, 1M context, cost per task $5.09, median 84 tok/s, and first chunk 195.83s.

Best for: Anthropic coding agents, long-context work, and snapshot-based model comparison.

Anthropic

Context
1M
AA Index
38
Cost per task
$5.09
Speed
84 tok/s
First chunk
195.83s
Total response
201.79s
42

Current reasoning model

OpenAI - GPT Terra - 5.6 (xhigh)

OpenAI GPT Terra 5.6 (xhigh) ranks in this snapshot with Intelligence Index 38, 1M context, cost per task $0.63, median 96 tok/s, and first chunk 14.46s.

Best for: OpenAI coding agents, long-context work, and snapshot-based model comparison.

OpenAI

Context
1M
AA Index
38
Cost per task
$0.63
Speed
96 tok/s
First chunk
14.46s
Total response
19.66s
43

Current reasoning model

OpenAI - GPT Luna - 5.6 (max)

OpenAI GPT Luna 5.6 (max) ranks in this snapshot with Intelligence Index 38, 1M context, cost per task $0.18, median 117 tok/s, and first chunk 124.21s.

Best for: OpenAI coding agents, long-context work, and snapshot-based model comparison.

OpenAI

Context
1M
AA Index
38
Cost per task
$0.18
Speed
117 tok/s
First chunk
124.21s
Total response
128.50s
44

Current model; partial public metrics

Google - Gemini - 3.7 Flash (low)

Google Gemini 3.7 Flash (low) ranks in this snapshot with Intelligence Index 37*, 1M context, cost per task is not fully public, median 308 tok/s, and first chunk 0.84s. Estimated Intelligence Index.

Best for: Very fast Google Flash loops when estimated Intelligence Index is enough and cost is still being measured.

Google

Context
1M
AA Index
37*
Cost per task
Speed
308 tok/s
First chunk
0.84s
Total response
2.47s
45

Current reasoning model

DeepSeek - DeepSeek - V4 Pro 0813 (max)

DeepSeek V4 Pro 0813 (max) ranks in this snapshot with Intelligence Index 36, 1M context, cost per task $0.67, median 94 tok/s, and first chunk 1.72s.

Best for: DeepSeek coding agents, long-context work, and snapshot-based model comparison.

DeepSeek

Context
1M
AA Index
36
Cost per task
$0.67
Speed
94 tok/s
First chunk
1.72s
Total response
28.19s
46

Current model; partial public metrics

Sapiens AI - Agnes - 3.0 Flash

Sapiens AI Agnes 3.0 Flash ranks in this snapshot with Intelligence Index 36*, 1M context, cost per task is not fully public, median 238 tok/s, and first chunk 1.84s. Estimated Intelligence Index.

Best for: Fast Agnes coding when public cost is still incomplete and you need a 1M-context Flash SKU.

Sapiens AI

Context
1M
AA Index
36*
Cost per task
Speed
238 tok/s
First chunk
1.84s
Total response
12.32s
47

Current reasoning model

SpaceXAI - Grok - 4.6 (low)

SpaceXAI Grok 4.6 (low) ranks in this snapshot with Intelligence Index 35, 500k context, cost per task $0.48, median 52 tok/s, and first chunk 5.63s.

Best for: SpaceXAI coding agents, long-context work, and snapshot-based model comparison.

SpaceXAI

Context
500k
AA Index
35
Cost per task
$0.48
Speed
52 tok/s
First chunk
5.63s
Total response
15.22s
48

Current model; partial public metrics

Sapiens AI - Agnes - 2.5 Pro Beta

Sapiens AI Agnes 2.5 Pro Beta ranks in this snapshot with Intelligence Index 35*, 1M context, cost per task is not fully public, speed not fully public, and latency not fully public. Estimated Intelligence Index.

Best for: Experimental Agnes Pro Beta coding when only a partial public metric snapshot is available.

Sapiens AI

Context
1M
AA Index
35*
Cost per task
Speed
tok/s
First chunk
Total response
49

Current reasoning model

DeepSeek - DeepSeek - V4 Flash Vision (max)

DeepSeek V4 Flash Vision (max) ranks in this snapshot with Intelligence Index 35, 1M context, cost per task $0.31, median 216 tok/s, and first chunk 1.16s.

Best for: DeepSeek coding agents, long-context work, and snapshot-based model comparison.

DeepSeek

Context
1M
AA Index
35
Cost per task
$0.31
Speed
216 tok/s
First chunk
1.16s
Total response
12.74s
50

Current reasoning model

OpenAI - GPT Luna - 5.6 (xhigh)

OpenAI GPT Luna 5.6 (xhigh) ranks in this snapshot with Intelligence Index 35, 1M context, cost per task $0.09, median 117 tok/s, and first chunk 45.03s.

Best for: OpenAI coding agents, long-context work, and snapshot-based model comparison.

OpenAI

Context
1M
AA Index
35
Cost per task
$0.09
Speed
117 tok/s
First chunk
45.03s
Total response
49.32s