XBench

Home = our average · lab pages = their ranks

As of 2026-08-14

Vendor launch tables.

First-party. Not a lab.

Provider blogs and model cards. Rank on this page is DeepSWE when the vendor published it.

Vendor DeepSWE #1 is GPT-5.6 Sol

Vendor launch tables

First-party launch posts. Harnesses differ. Independent boards beat these when they disagree.

1
GPT-5.6 Sol
OpenAI · closed
Top-tier flagship reasoning
72.7%Vendor88.8%Vendor34.6%Vendor1728Vendor
2
GPT-5.6 Terra
OpenAI · closed
Balanced GPT-5.6
69.6%Vendor87.4%Vendor20.8%Vendor63.4%Vendor
3
Opus 5
Anthropic · closed
Anthropic’s flagship frontier model
68.8%Vendor
4
Kimi K3
Moonshot AI · open
Frontier open-weight
67.3%Vendor88.3%Vendor1668Vendor
5
GPT-5.6 Luna
OpenAI · closed
Cost-efficient GPT-5.6
67.2%Vendor84.7%Vendor62.7%Vendor
6
GLM-5.3
Z.ai · staged
Latest GLM coding / cyber-defense
66.9%Vendor28.3%Vendor1769Vendor
7
Grok 4.6
SpaceXAI · closed
Latest Grok frontier model
65.9%Vendor26%Vendor
8
3.7 Flash
Google DeepMind · closed
Latest Flash workhorse
65.3%Vendor85.8%Vendor14.9%Vendor1525Vendor
9
Qwen3.8 Max
Alibaba · closed
Hosted 2.4T flagship; weights promised
56.6%Vendor86.6%Vendor67.7%Vendor
10
Grok 4.5
SpaceXAI · closed
Frontier coding/agents
53%Vendor83.3%Vendor64.7%Vendor
11
3.6 Flash
Google DeepMind · closed
Frontier-class fast reasoning
49%Vendor78%Vendor1421Vendor
12
GLM-5.2
Zhipu AI · open
Frontier/open-weight coding
46.2%Vendor81%Vendor4.6%Vendor62.1%Vendor
Mythos 5
Anthropic · closed
New Claude 5 frontier tier
Doubao Seed
ByteDance · closed
Frontier reasoning
LongCat 2.0
Meituan · open
Frontier open-weight agentic coding
MiniMax M3
MiniMax · open
Open-weight frontier-adjacent
66%Vendor59%Vendor

16 models on this board · # is this page’s rank

SpaceXAI — Grok 4.6 launch table

2026-08-12 · x.ai/news/grok-4-6

Vendor table. Third-party scores are the best of self-reported or public results.

EvalGrok 4.6 HighGrok 4.5 HighGPT-5.6 Sol MaxFable 5 Max
AA Intelligence Index61566162
GDPVal-AA v2 (Elo)1753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.1 Extended61.3%56.6%60.6%63.6%
APEX-Agents57.5%47.1%56.7%59.2%
Terminal-Bench v3.026%15.7%34.6%34.1%
AA-Briefcase (Elo)1577131315021574
Open original

Google DeepMind — Gemini 3.7 Flash comparison

2026-08-13 · deepmind.google/models/gemini/flash

Mid-tier comparison (Flash / Sonnet / GPT-5.6 Terra / Spark), not the Opus / GPT-5.6 Sol / Grok table.

Eval3.7 Flash3.6 FlashSonnet 5GPT-5.6 TerraSpark 1.2
AA Intelligence Index5652555757
DeepSWE v1.165.3%48.6%53.8%69.6%54.9%
Code Arena Elo15881538154115231535
Terminal-bench 2.185.8%78.0%80.4%87.4%82.9%
Terminal-bench 3.014.9%5.4%14.6%20.8%
GDPVal-AA v215251422159815781628
Open original

Z.ai — GLM-5.3 launch chart

2026-08-14 · z.ai/blog/glm-5.3

First-party. Mythos/Fable is a combined competitor column. No independent AA score yet.

EvalGLM-5.3GLM-5.2Kimi K3Mythos/Fable 5GPT-5.6 Sol
Terminal-Bench 3.028.3%4.6%17.4%33.7%34.6%
DeepSWE v1.166.9%46.2%67.5%69.7%72.7%
GDPVal-AA v217691508168217431730
CyberGym84.5%77.2%80.0%83.8%83.6%
ExploitBench54.4%24.4%32.2%78.0%76.5%
HLE w/ tools62.5%54.7%59.8%63.9%64.5%
Open original

OpenAI — GPT-5.6 Sol / Terra / Luna (GA table)

2026-07-09 · openai.com/index/gpt-5-6

First-party GA table. OpenAI’s HTML is JS-walled here; numbers match the published table as quoted in contemporaneous coverage and AA’s GPT-5.6 article. Current AA Intelligence Index has moved (Sol 61 / Terra 57 / Luna 52).

EvalGPT-5.6 SolGPT-5.6 TerraGPT-5.6 LunaGPT-5.5Fable 5
AA Coding Agent Index v1.18077.474.676.477.2
DeepSWE v1.172.7%69.6%67.2%67%69.7%
Terminal-Bench 2.188.8%87.4%84.7%85.6%83.1%
SWE-bench Pro64.6%63.4%62.7%59.4%80%
BrowseComp90.4%87.5%83.3%84.4%84.3%
OSWorld 2.062.6%50.2%45.6%47.5%
AA Intelligence (launch v4.1)58.95551.254.859.9
Open original

Z.ai — GLM-5.2 launch table

2026-06-16 · z.ai/blog/glm-5.2

First-party full table. Terminus-2 harness unless noted. Stars on HLE are full-set scores. Qwen3.7-Max and DeepSeek-V4-Pro columns omitted here for width; they are on the original post.

EvalGLM-5.2GLM-5.1Opus 4.8GPT-5.5Gemini 3.1 ProMiniMax M3
SWE-bench Pro62.1%58.4%69.2%58.6%54.2%59.0%
DeepSWE46.2%18.0%58.0%70.0%10.0%20.0%
Terminal-Bench 2.1 (Terminus-2)81.0%63.5%85.0%84.0%74.0%65.0%
FrontierSWE dominance (16 Jun)74.430.575.172.639.6
PostTrainBench34.320.137.228.421.6
MCP-Atlas public76.8%71.8%77.8%75.3%69.2%74.2%
HLE w/ tools54.7%52.3%57.9%*52.2%*51.4%*
Open original

Qwen — Qwen3.8-Max launch table

2026-08-03 · qwen.ai/blog?id=qwen3.8

First-party. Table vs Opus 4.8 / Fable 5 / GPT-5.6 Sol / Qwen3.7-Max as published on the Qwen3.8-Max post (chart on qwen.ai; numbers from that table).

EvalQwen3.8-MaxOpus 4.8Fable 5GPT-5.6 SolQwen3.7-Max
Terminal-Bench 2.186.6%84.6%84.6%88.8%74.5%
SWE-bench Pro67.7%69.2%80.0%64.6%60.6%
DeepSWE 1.156.6%59.0%70.0%73.0%21.6%
PaperBench93.0%80.3%88.8%90.5%64.8%
OSWorld-Verified86.1%
JobBench53.4%31.3%
FrontierSWE73.57088.840.7
Open original

Google — Gemini 3.6 Flash launch numbers

2026-07-21 · blog.google Gemini 3.6 Flash

First-party vs Gemini 3.5 Flash. 3.6 Flash priced $1.50 / $7.50.

Eval3.6 Flash3.5 Flash
DeepSWE49%37%
MLE-Bench63.9%49.7%
OSWorld-Verified83.0%78.4%
GDPval-AA v214211349
AA output tokens vs 3.5−17%baseline
Open original

MiniMax — M3 launch numbers

2026-06-16 · minimax.io/blog/minimax-m3

First-party. SWE-bench Pro and Terminal-Bench 2.1 on MiniMax infra (Claude Code / Terminus 2).

EvalMiniMax M3
SWE-bench Pro59.0%
Terminal-Bench 2.166.0%
SWE-fficiency34.8%
KernelBench Hard28.8%
MCP Atlas74.2%
Open original

SpaceXAI — Grok 4.5 launch chart

2026-07-16 · x.ai/news/grok-4-5

First-party charts on x.ai/news/grok-4-5. DeepSWE 1.0 extracted from the page (AA harnesses). Other tabs (DeepSWE 1.1, SWE Marathon, Terminal-Bench 2.1, SWE-bench Pro) are the same post’s chart series as transcribed in contemporaneous coverage of those tabs. Grok 4.6’s later table restates 4.5 High DeepSWE v1.1 at 54%.

EvalGrok 4.5Fable maxGPT-5.5 xhighOpus 4.8 max
DeepSWE 1.0 (AA harnesses)62.0%66.1%64.3%55.8%
DeepSWE 1.153%70%67%59%
SWE Marathon pass@129.0%24.0%26.0%
Terminal-Bench 2.183.3%84.3%83.4%78.9%
SWE-bench Pro64.7%80.4%58.6%69.2%
Price in/out$2 / $6$10 / $50$5 / $25
Open original

Moonshot — Kimi K3 tech blog

2026-07-16 · kimi.com/blog/kimi-k3

First-party. Full comparison chart is an image; numbers below are from the blog footnotes (Kimi Code harness unless noted).

EvalKimi K3Note
DeepSWE v1.1 (mini-SWE-agent)67.3%Official DeepSWE board; Kimi Code harness is the blog headline run
BrowseComp (1M, no compaction)90.4%Blog footnote 5
Price$3 / $15Cache hit $0.30 / 1M
Open original

Anthropic — Claude Opus 5 launch

2026-07-24 · anthropic.com/news/claude-opus-5

First-party. Most charts did not extract as numbers. Claims below are the ones stated in prose on the launch post.

ClaimOpus 5
Price$5 / $25 — same as Opus 4.8, half of Fable 5
Frontier-Bench v0.1SOTA; more than 2× Opus 4.8 at lower cost/task
CursorBench 3.2 (max)Within 0.5% of Fable 5 peak, half the cost/task
ARC-AGI 3~3× the next-best model
Zapier AutomationBench~1.5× next-best pass rate at the same cost/task
Cyber vs Mythos 5Close on finding vulns; far behind on exploit development (OSS-Fuzz)
Open original

ByteDance Seed — Seed 2.1 Pro launch

2026-06-23 · research.doubao.com/en/seed2_1

First-party. Independent AA Intelligence for this row is Doubao Seed Code 26*.

EvalSeed 2.1 ProSeed 2.1 TurboOpus 4.7GPT-5.5Gemini 3.1 Pro
Terminal-Bench 2.171.0%67.6%71.7%73.8%70.7%
SWE-Atlas35.2%30.6%38.7%44.7%23.6%
NL2Repo-Bench47.0%43.7%58.2%45.1%33.4%
Workspace Bench53.0%54.7%55.1%58.7%32.8%
Open original

Meituan — LongCat-2.0 model card

2026-06-01 · huggingface.co/meituan-longcat/LongCat-2.0

In-house harness unless marked *. Compared on the card to Gemini 3.1 Pro, GPT-5.5, Opus 4.7/4.8.

EvalLongCat-2.0Gemini 3.1 ProGPT-5.5Opus 4.7Opus 4.8
Terminal-Bench 2.170.870.7*73.8*71.7*78.9*
SWE-bench Pro59.554.2*58.6*64.3*69.2*
SWE-bench Multilingual77.376.9*80.5*84.8*
Open original