XBench

Home = our average · lab pages = their ranks

As of 2026-08-14

BenchLM.

Their board. Their ranks.

BenchLM hosts AA Agentic and Coding mirrors plus SWE-bench Verified and SWE-bench Pro. Rank follows SWE-Pro when present, else AA Coding.

BenchLM #1 is Mythos 5

Open the lab’s original

BenchLM

Aggregator snapshot, 14 Aug 2026. AA mirrors are the same numbers as the AA page — SWE boards are not.

1
Mythos 5
Anthropic · closed
New Claude 5 frontier tier
80.3%Roundup95.5%Roundup88%TB 2.080.3%SWE-Pro
2
Fable 5
Anthropic · closed
New Claude 5 frontier tier
80%Roundup95%Roundup56.6AA Agentic76.5AA Coding
3
Opus 5
Anthropic · closed
Anthropic’s flagship frontier model
79.2%Roundup96%Roundup59.2AA Agentic78AA Coding
23
GPT-5.6 Sol
OpenAI · closed
Top-tier flagship reasoning
64.6%Roundup57.8AA Agentic77.4AA Coding
4
Grok 4.6
SpaceXAI · closed
Latest Grok frontier model
58.7AA Agentic76.8AA Coding
6
Kimi K3
Moonshot AI · open
Frontier open-weight
54.3AA Agentic76.2AA Coding
13
Qwen3.8 Max
Alibaba · closed
Hosted 2.4T flagship; weights promised
58.4AA Agentic71.8AA Coding
13
Qwen3.8 open
Alibaba · open
Open-weight Qwen3.8 (AA #2 open)
58.4AA Agentic71.8AA Coding
12
Spark 1.2
Meta · closed
Agent-native Meta model
49.3AA Agentic72.2AA Coding
5
GPT-5.6 Terra
OpenAI · closed
Balanced GPT-5.6
50.2AA Agentic76.7AA Coding
7
3.7 Flash
Google DeepMind · closed
Latest Flash workhorse
45.1AA Agentic76.1AA Coding
11
Grok 4.5
SpaceXAI · closed
Frontier coding/agents
48.9AA Agentic72.5AA Coding
15
Sonnet 5
Anthropic · closed
Frontier general-purpose/agentic
49.7AA Agentic71.5AA Coding
19
GLM-5.2
Zhipu AI · open
Frontier/open-weight coding
45.7AA Agentic68.8AA Coding
19
V4 Pro 0813
DeepSeek · open
Current DeepSeek-V4 open challenger
49.6AA Agentic68.8AA Coding
18
V4 Flash
DeepSeek · open
Open price-performance pick
48.4AA Agentic69.1AA Coding
15
GPT-5.6 Luna
OpenAI · closed
Cost-efficient GPT-5.6
46.9AA Agentic71.5AA Coding
17
3.6 Flash
Google DeepMind · closed
Frontier-class fast reasoning
40.5AA Agentic69.2AA Coding
19
3.1 Pro
Google DeepMind · closed
Frontier multimodal/reasoning
23.1AA Agentic68.8AA Coding
9
Opus 4.8
Anthropic · closed
Frontier reasoning/coding
49.4AA Agentic74.3AA Coding
10
Opus 4.7
Anthropic · closed
Frontier coding/reasoning
46.3AA Agentic73.6AA Coding
8
GPT-5.5
OpenAI · closed
Frontier reasoning + agents
47.4AA Agentic74.9AA Coding
Grok 4.20
SpaceXAI · closed
Frontier reasoning/agents
47.1%TB 2.076.7%SWE Verified
28
Grok 4.3
SpaceXAI · closed
Frontier reasoning/agents
24.2AA Agentic42.3AA Coding
22
Qwen3.7 Max
Alibaba · closed
Frontier Chinese model
30.9AA Agentic66AA Coding
26
MiniMax M3
MiniMax · open
Open-weight frontier-adjacent
36.1AA Agentic58.6AA Coding
27
Inkling
Thinking Machines · open
New open lab, Apache-2.0
34.1AA Agentic52.1AA Coding
25
Hy3
Tencent · open
Open Apache-2.0 coding/reasoning
31.4AA Agentic58.8AA Coding
24
MiMo V2.5 Pro
Xiaomi · open
Open MIT, cheap per-task
29.5AA Agentic60.2AA Coding

29 models on this board · # is this page’s rank