XBench

Home = our average · lab pages = their ranks

As of 2026-08-14 · XBench average

Ours · average of labs

XBench average.

Our board. Not theirs.

XBench scores are the mean of independent and aggregator labs after each lab is scaled 0–100. Rank here will not match Artificial Analysis or Arena.

XBench #1 is Fable 5. AA #1 is Opus 5. Arena #1 is Fable 5.

2026-08-14 · Z.ai

GLM-5.3

Same 743B base as GLM-5.2. Post-training only. Z.ai reports 28.3 on Terminal-Bench 3.0, 66.9 DeepSWE, 84.5% CyberGym. Weights staged.

Source

2026-08-13 · Google

Gemini 3.7 Flash

Flash workhorse, not a 4.7. Intelligence Index 56. Arena Text #9 preliminary. Intro $0.75 / $3.75 through 31 Dec 2026.

Source

2026-08-12 · SpaceXAI

Grok 4.6

Ties GPT-5.6 Sol at Intelligence 61, $2 / $6, $0.84 per AA task. Strong GDPval; trails GPT-5.6 Sol on DeepSWE and Terminal-Bench 3.0.

Source

Click a release row to highlight that model on the charts and score table. Click again to drop it. Purple is selected, not “this week.”

100.980.059.138.217.3Gemini 3.1 ProGPT-5.5Claude Opus 4.7Grok 4.3 / 4.20DeepSeek-V4Claude Opus 4.8Qwen3.7-MaxClaude Fable 5Claude Mythos 5Claude Sonnet 5GLM-5.2Doubao Seed 2.1 ProLongCat-2.0Grok 4.5GPT-5.6 SolGPT-5.6 TerraGPT-5.6 LunaGemini 3.6 FlashKimi K3Claude Opus 5Grok 4.6Gemini 3.7 FlashGLM-5.3Score

Left to right is the 2026 release timeline. Purple dots are models you picked in the table.

ReleaseCompanyModelRough positioning
Feb 2026Google DeepMindGemini 3.1 ProFrontier multimodal/reasoning
Apr 2026OpenAIGPT-5.5Frontier reasoning + agents
Apr 2026AnthropicClaude Opus 4.7Frontier coding/reasoning
Apr–May 2026SpaceXAIGrok 4.3 / 4.20Frontier reasoning/agents
Apr 2026DeepSeekDeepSeek-V4Frontier open/low-cost challenger
May 2026AnthropicClaude Opus 4.8Frontier reasoning/coding
May 2026AlibabaQwen3.7-MaxFrontier Chinese model
Jun 2026AnthropicClaude Fable 5New Claude 5 frontier tier
Jun 2026AnthropicClaude Mythos 5New Claude 5 frontier tier
Jun 2026AnthropicClaude Sonnet 5Frontier general-purpose/agentic
Jun 2026Zhipu AIGLM-5.2Frontier/open-weight coding
Jun 2026ByteDanceDoubao Seed 2.1 ProFrontier reasoning
Jun 2026MeituanLongCat-2.0Frontier open-weight agentic coding
Jul 2026SpaceXAIGrok 4.5Frontier coding/agents
Jul 2026OpenAIGPT-5.6 SolTop-tier flagship reasoning
Jul 2026OpenAIGPT-5.6 TerraBalanced GPT-5.6
Jul 2026OpenAIGPT-5.6 LunaCost-efficient GPT-5.6
Jul 2026Google DeepMindGemini 3.6 FlashFrontier-class fast reasoning
Jul 2026Moonshot AIKimi K3Frontier open-weight
Jul 2026AnthropicClaude Opus 5Anthropic’s flagship frontier model
Aug 2026SpaceXAIGrok 4.6Latest Grok frontier model
13 Aug 2026Google DeepMindGemini 3.7 FlashLatest Flash workhorse
14 Aug 2026Z.aiGLM-5.3Latest GLM coding / cyber-defense

XBench average

Each lab is min-max scaled to 0–100 on its own board, then averaged. Vendor first-party numbers are not in the average. XBench place needs at least two pillars (Intelligence, Agentic, Coding) so one SWE-Pro row cannot take #1. Placement is XBench’s, not a copy of AA or Arena.

Line · higher is better · same order as the table

108795021-8Fable 5Opus 5Qwen3.8 openQwen3.8 MaxGrok 4.6GPT-5.6 SolGPT-5.6 TerraKimi K3Opus 4.7Spark 1.23.7 FlashOpus 4.8V4 Pro 0813GPT-5.6 LunaV4 FlashSonnet 5GLM-5.2GPT-5.53.6 FlashGrok 4.53.1 ProMiniMax M3Qwen3.7 MaxMiMo V2.5 ProHy3InklingGrok 4.3Mythos 5Grok 4.20V4 ProDoubao SeedLongCat 2.0Score

Hover a column. Click a legend chip to hide a series. Purple dots are models picked in the release table.

Line · lower is better

$-0.22$0.68$1.58$2.49$3.39Fable 5Opus 5Qwen3.8 openQwen3.8 MaxGrok 4.6GPT-5.6 SolGPT-5.6 TerraKimi K3Opus 4.7Spark 1.23.7 FlashOpus 4.8V4 Pro 0813GPT-5.6 LunaV4 FlashSonnet 5GLM-5.2GPT-5.53.6 FlashGrok 4.53.1 ProMiniMax M3Qwen3.7 MaxMiMo V2.5 ProHy3InklingGrok 4.3Mythos 5Grok 4.20V4 ProDoubao SeedLongCat 2.0Lower better

Rank and cost. Top of the chart is better.

Line · API price USD / 1M

$-3.86$10.60$25.07$39.53$53.99Fable 5Opus 5Qwen3.8 openQwen3.8 MaxGrok 4.6GPT-5.6 SolGPT-5.6 TerraKimi K3Opus 4.7Spark 1.23.7 FlashOpus 4.8V4 Pro 0813GPT-5.6 LunaV4 FlashSonnet 5GLM-5.2GPT-5.53.6 FlashGrok 4.53.1 ProMiniMax M3Qwen3.7 MaxMiMo V2.5 ProHy3InklingGrok 4.3Mythos 5Grok 4.20V4 ProDoubao SeedLongCat 2.0$ / 1M

Input and output token price. Lower is cheaper.

1
Fable 5
Anthropic · closed
New Claude 5 frontier tier
95.1XBench · 3 labs98.6Intel · 2 labs92.8Agentic · 1 lab93.8Coding · 3 labs$3.14LabAA #2Arena #1BenchLM #2$10 / $50
2
Opus 5
Anthropic · closed
Anthropic’s flagship frontier model
94.6XBench · 3 labs86.7Intel · 2 labs100Agentic · 1 lab97Coding · 3 labs$2.34LabAA #1Arena #6BenchLM #3$5 / $25
3
Qwen3.8 open
Alibaba · open
Open-weight Qwen3.8 (AA #2 open)
89XBench · 3 labs86.5Intel · 1 lab97.8Agentic · 1 lab82.6Coding · 1 lab$1.09LabAA #6Arena BenchLM #13
4
Qwen3.8 Max
Alibaba · closed
Hosted 2.4T flagship; weights promised
87.3XBench · 3 labs81.5Intel · 2 labs97.8Agentic · 1 lab82.6Coding · 1 lab$1.13LabAA #6Arena #4BenchLM #13$2 / $6
5
Grok 4.6
SpaceXAI · closed
Latest Grok frontier model
87.1XBench · 3 labs64.5Intel · 2 labs98.6Agentic · 1 lab98.3Coding · 2 labs$0.84LabAA #3Arena #16BenchLM #4$2 / $6
6
GPT-5.6 Sol
OpenAI · closed
Top-tier flagship reasoning
85.6XBench · 3 labs94.6Intel · 1 lab96.1Agentic · 1 lab66.1Coding · 3 labs$1.23LabAA #3Arena BenchLM #23$5 / $30
7
GPT-5.6 Terra
OpenAI · closed
Balanced GPT-5.6
85.1XBench · 3 labs83.8Intel · 1 lab75.1Agentic · 1 lab96.4Coding · 1 lab$0.51LabAA #8Arena BenchLM #5$2 / $12
8
Kimi K3
Moonshot AI · open
Frontier open-weight
84XBench · 3 labs82.7Intel · 2 labs86.4Agentic · 1 lab82.9Coding · 2 labs$0.84LabAA #5Arena #6BenchLM #6$3 / $15
9
Opus 4.7
Anthropic · closed
Frontier coding/reasoning
81.9XBench · 3 labs93.8Intel · 1 lab64.3Agentic · 1 lab87.7Coding · 1 labAA Arena #2BenchLM #10$5 / $25
10
Spark 1.2
Meta · closed
Agent-native Meta model
80.7XBench · 3 labs85.6Intel · 2 labs72.6Agentic · 1 lab83.8Coding · 1 lab$0.40LabAA #8Arena #3BenchLM #12$1.25 / $4.25
11
3.7 Flash
Google DeepMind · closed
Latest Flash workhorse
77.9XBench · 3 labs78Intel · 2 labs60.9Agentic · 1 lab94.7Coding · 1 lab$0.40LabAA #10Arena #5BenchLM #7$0.75 / $3.75
12
Opus 4.8
Anthropic · closed
Frontier reasoning/coding
77.8XBench · 3 labs71Intel · 2 labs72.9Agentic · 1 lab89.6Coding · 1 lab$1.80LabAA #10Arena #11BenchLM #9$5 / $25
13
V4 Pro 0813
DeepSeek · open
Current DeepSeek-V4 open challenger
73.5XBench · 3 labs73Intel · 1 lab73.4Agentic · 1 lab74.2Coding · 1 lab$0.25LabAA #14Arena BenchLM #19
14
GPT-5.6 Luna
OpenAI · closed
Cost-efficient GPT-5.6
72.7XBench · 3 labs70.3Intel · 1 lab65.9Agentic · 1 lab81.8Coding · 1 lab$0.05LabAA #16Arena BenchLM #15$0.20 / $1.20
15
V4 Flash
DeepSeek · open
Open price-performance pick
71.8XBench · 3 labs70.3Intel · 1 lab70.1Agentic · 1 lab75.1Coding · 1 lab$0.03LabAA #16Arena BenchLM #18$0.14 / $0.28
16
Sonnet 5
Anthropic · closed
Frontier general-purpose/agentic
70.1XBench · 3 labs54.8Intel · 2 labs73.7Agentic · 1 lab81.8Coding · 1 lab$1.72LabAA #13Arena #17BenchLM #15$2 / $10
17
GLM-5.2
Zhipu AI · open
Frontier/open-weight coding
65.3XBench · 3 labs59.1Intel · 2 labs62.6Agentic · 1 lab74.2Coding · 1 lab$0.32LabAA #14Arena #14BenchLM #19$1.40 / $4.40
18
GPT-5.5
OpenAI · closed
Frontier reasoning + agents
64.6XBench · 3 labs35.3Intel · 2 labs67.3Agentic · 1 lab91.3Coding · 1 labAA #28Arena #10BenchLM #8$2.50 / $15
19
3.6 Flash
Google DeepMind · closed
Frontier-class fast reasoning
63.8XBench · 3 labs67.9Intel · 2 labs48.2Agentic · 1 lab75.4Coding · 1 lab$0.56LabAA #16Arena #9BenchLM #17$1.50 / $7.50
20
Grok 4.5
SpaceXAI · closed
Frontier coding/agents
58.5XBench · 3 labs61.6Intel · 2 labs71.5Agentic · 1 lab42.3Coding · 2 labs$0.36LabAA #10Arena #15BenchLM #11$2 / $6
21
3.1 Pro
Google DeepMind · closed
Frontier multimodal/reasoning
46.1XBench · 3 labs64.1Intel · 2 labs0Agentic · 1 lab74.2Coding · 1 lab$0.33LabAA #19Arena #8BenchLM #19$1 / $6
22
MiniMax M3
MiniMax · open
Open-weight frontier-adjacent
44.4XBench · 3 labs51.4Intel · 1 lab36Agentic · 1 lab45.7Coding · 1 lab$0.14LabAA #20Arena BenchLM #26$0.60 / $2.40
23
Qwen3.7 Max
Alibaba · closed
Frontier Chinese model
43.5XBench · 3 labs42.6Intel · 2 labs21.6Agentic · 1 lab66.4Coding · 1 lab$0.24LabAA #25Arena #13BenchLM #22$1.48 / $4.42
24
MiMo V2.5 Pro
Xiaomi · open
Open MIT, cheap per-task
37.9XBench · 3 labs45.9Intel · 1 lab17.7Agentic · 1 lab50.1Coding · 1 lab$0.03LabAA #22Arena BenchLM #24$0.43 / $0.87
25
Hy3
Tencent · open
Open Apache-2.0 coding/reasoning
37.5XBench · 3 labs43.2Intel · 1 lab23Agentic · 1 lab46.2Coding · 1 lab$0.04LabAA #23Arena BenchLM #25$0.13 / $0.53
26
Inkling
Thinking Machines · open
New open lab, Apache-2.0
33.7XBench · 3 labs43.2Intel · 1 lab30.5Agentic · 1 lab27.5Coding · 1 lab$0.34LabAA #23Arena BenchLM #27$0.50 / $2.02
27
Grok 4.3
SpaceXAI · closed
Frontier reasoning/agents
6XBench · 3 labs14.9Intel · 2 labs3Agentic · 1 lab0Coding · 1 labAA #26Arena #18BenchLM #28$1.25 / $2.50
Mythos 5
Anthropic · closed
New Claude 5 frontier tier
96.4Coding · 2 labsAA Arena BenchLM #1$10 / $50
Grok 4.20
SpaceXAI · closed
Frontier reasoning/agents
51.6Intel · 1 labAA Arena #12BenchLM $1.25 / $2.50
V4 Pro
DeepSeek · open
Frontier open/low-cost challenger
51.4Intel · 1 lab$0.05LabAA #20Arena BenchLM $0.44 / $0.87
Doubao Seed
ByteDance · closed
Frontier reasoning
0Intel · 1 labAA #29Arena BenchLM
LongCat 2.0
Meituan · open
Frontier open-weight agentic coding
21.6Intel · 1 labAA #27Arena BenchLM

32 models on this board · # is this page’s rank · purple = picked in the release-dates table

Line graph of rank on this page versus XBench, AA, Arena, and BenchLM. Top of the chart is #1.

#-1#7#15#22#30Fable 5Opus 5Qwen3.8 openQwen3.8 MaxGrok 4.6GPT-5.6 SolGPT-5.6 TerraKimi K3Opus 4.7Spark 1.23.7 FlashOpus 4.8V4 Pro 0813GPT-5.6 LunaV4 FlashSonnet 5GLM-5.2GPT-5.53.6 FlashGrok 4.53.1 ProMiniMax M3Qwen3.7 MaxMiMo V2.5 ProHy3InklingGrok 4.3Place

Place across boards. Top of the chart is #1. Purple dots are models picked in the release table.

XBench composite bars

Mean of Intelligence, Agentic, and Coding pillars when present. Click a release-date row to highlight that model on this page.

102.276.450.524.7-1.1Fable 5Opus 5Qwen3.8 openQwen3.8 MaxGrok 4.6GPT-5.6 SolGPT-5.6 TerraKimi K3Opus 4.7Spark 1.23.7 FlashOpus 4.8V4 Pro 0813GPT-5.6 LunaV4 FlashSonnet 5GLM-5.2GPT-5.53.6 FlashGrok 4.53.1 ProMiniMax M3Qwen3.7 MaxMiMo V2.5 ProHy3InklingGrok 4.3

Models left to right by this metric. Purple dots are models picked in the release table.