2026-08-14 · Z.ai
GLM-5.3
Same 743B base as GLM-5.2. Post-training only. Z.ai reports 28.3 on Terminal-Bench 3.0, 66.9 DeepSWE, 84.5% CyberGym. Weights staged.
SourceHome = our average · lab pages = their ranks
As of 2026-08-14 · XBench average
Ours · average of labs
XBench scores are the mean of independent and aggregator labs after each lab is scaled 0–100. Rank here will not match Artificial Analysis or Arena.
XBench #1 is Fable 5. AA #1 is Opus 5. Arena #1 is Fable 5.
2026-08-14 · Z.ai
Same 743B base as GLM-5.2. Post-training only. Z.ai reports 28.3 on Terminal-Bench 3.0, 66.9 DeepSWE, 84.5% CyberGym. Weights staged.
Source2026-08-13 · Google
Flash workhorse, not a 4.7. Intelligence Index 56. Arena Text #9 preliminary. Intro $0.75 / $3.75 through 31 Dec 2026.
Source2026-08-12 · SpaceXAI
Ties GPT-5.6 Sol at Intelligence 61, $2 / $6, $0.84 per AA task. Strong GDPval; trails GPT-5.6 Sol on DeepSWE and Terminal-Bench 3.0.
SourceClick a release row to highlight that model on the charts and score table. Click again to drop it. Purple is selected, not “this week.”
Left to right is the 2026 release timeline. Purple dots are models you picked in the table.
| Release | Company | Model | Rough positioning |
|---|---|---|---|
| Feb 2026 | Google DeepMind | Gemini 3.1 Pro | Frontier multimodal/reasoning |
| Apr 2026 | OpenAI | GPT-5.5 | Frontier reasoning + agents |
| Apr 2026 | Anthropic | Claude Opus 4.7 | Frontier coding/reasoning |
| Apr–May 2026 | SpaceXAI | Grok 4.3 / 4.20 | Frontier reasoning/agents |
| Apr 2026 | DeepSeek | DeepSeek-V4 | Frontier open/low-cost challenger |
| May 2026 | Anthropic | Claude Opus 4.8 | Frontier reasoning/coding |
| May 2026 | Alibaba | Qwen3.7-Max | Frontier Chinese model |
| Jun 2026 | Anthropic | Claude Fable 5 | New Claude 5 frontier tier |
| Jun 2026 | Anthropic | Claude Mythos 5 | New Claude 5 frontier tier |
| Jun 2026 | Anthropic | Claude Sonnet 5 | Frontier general-purpose/agentic |
| Jun 2026 | Zhipu AI | GLM-5.2 | Frontier/open-weight coding |
| Jun 2026 | ByteDance | Doubao Seed 2.1 Pro | Frontier reasoning |
| Jun 2026 | Meituan | LongCat-2.0 | Frontier open-weight agentic coding |
| Jul 2026 | SpaceXAI | Grok 4.5 | Frontier coding/agents |
| Jul 2026 | OpenAI | GPT-5.6 Sol | Top-tier flagship reasoning |
| Jul 2026 | OpenAI | GPT-5.6 Terra | Balanced GPT-5.6 |
| Jul 2026 | OpenAI | GPT-5.6 Luna | Cost-efficient GPT-5.6 |
| Jul 2026 | Google DeepMind | Gemini 3.6 Flash | Frontier-class fast reasoning |
| Jul 2026 | Moonshot AI | Kimi K3 | Frontier open-weight |
| Jul 2026 | Anthropic | Claude Opus 5 | Anthropic’s flagship frontier model |
| Aug 2026 | SpaceXAI | Grok 4.6 | Latest Grok frontier model |
| 13 Aug 2026 | Google DeepMind | Gemini 3.7 Flash | Latest Flash workhorse |
| 14 Aug 2026 | Z.ai | GLM-5.3 | Latest GLM coding / cyber-defense |
Each lab is min-max scaled to 0–100 on its own board, then averaged. Vendor first-party numbers are not in the average. XBench place needs at least two pillars (Intelligence, Agentic, Coding) so one SWE-Pro row cannot take #1. Placement is XBench’s, not a copy of AA or Arena.
Line · higher is better · same order as the table
Hover a column. Click a legend chip to hide a series. Purple dots are models picked in the release table.
Line · lower is better
Rank and cost. Top of the chart is better.
Line · API price USD / 1M
Input and output token price. Lower is cheaper.
| 1 | Fable 5 Anthropic · closed New Claude 5 frontier tier | 95.1XBench · 3 labs | 98.6Intel · 2 labs | 92.8Agentic · 1 lab | 93.8Coding · 3 labs | $3.14Lab | AA #2Arena #1BenchLM #2 | $10 / $50 |
| 2 | Opus 5 Anthropic · closed Anthropic’s flagship frontier model | 94.6XBench · 3 labs | 86.7Intel · 2 labs | 100Agentic · 1 lab | 97Coding · 3 labs | $2.34Lab | AA #1Arena #6BenchLM #3 | $5 / $25 |
| 3 | Qwen3.8 open Alibaba · open Open-weight Qwen3.8 (AA #2 open) | 89XBench · 3 labs | 86.5Intel · 1 lab | 97.8Agentic · 1 lab | 82.6Coding · 1 lab | $1.09Lab | AA #6Arena —BenchLM #13 | — |
| 4 | Qwen3.8 Max Alibaba · closed Hosted 2.4T flagship; weights promised | 87.3XBench · 3 labs | 81.5Intel · 2 labs | 97.8Agentic · 1 lab | 82.6Coding · 1 lab | $1.13Lab | AA #6Arena #4BenchLM #13 | $2 / $6 |
| 5 | Grok 4.6 SpaceXAI · closed Latest Grok frontier model | 87.1XBench · 3 labs | 64.5Intel · 2 labs | 98.6Agentic · 1 lab | 98.3Coding · 2 labs | $0.84Lab | AA #3Arena #16BenchLM #4 | $2 / $6 |
| 6 | GPT-5.6 Sol OpenAI · closed Top-tier flagship reasoning | 85.6XBench · 3 labs | 94.6Intel · 1 lab | 96.1Agentic · 1 lab | 66.1Coding · 3 labs | $1.23Lab | AA #3Arena —BenchLM #23 | $5 / $30 |
| 7 | GPT-5.6 Terra OpenAI · closed Balanced GPT-5.6 | 85.1XBench · 3 labs | 83.8Intel · 1 lab | 75.1Agentic · 1 lab | 96.4Coding · 1 lab | $0.51Lab | AA #8Arena —BenchLM #5 | $2 / $12 |
| 8 | Kimi K3 Moonshot AI · open Frontier open-weight | 84XBench · 3 labs | 82.7Intel · 2 labs | 86.4Agentic · 1 lab | 82.9Coding · 2 labs | $0.84Lab | AA #5Arena #6BenchLM #6 | $3 / $15 |
| 9 | Opus 4.7 Anthropic · closed Frontier coding/reasoning | 81.9XBench · 3 labs | 93.8Intel · 1 lab | 64.3Agentic · 1 lab | 87.7Coding · 1 lab | — | AA —Arena #2BenchLM #10 | $5 / $25 |
| 10 | Spark 1.2 Meta · closed Agent-native Meta model | 80.7XBench · 3 labs | 85.6Intel · 2 labs | 72.6Agentic · 1 lab | 83.8Coding · 1 lab | $0.40Lab | AA #8Arena #3BenchLM #12 | $1.25 / $4.25 |
| 11 | 3.7 Flash Google DeepMind · closed Latest Flash workhorse | 77.9XBench · 3 labs | 78Intel · 2 labs | 60.9Agentic · 1 lab | 94.7Coding · 1 lab | $0.40Lab | AA #10Arena #5BenchLM #7 | $0.75 / $3.75 |
| 12 | Opus 4.8 Anthropic · closed Frontier reasoning/coding | 77.8XBench · 3 labs | 71Intel · 2 labs | 72.9Agentic · 1 lab | 89.6Coding · 1 lab | $1.80Lab | AA #10Arena #11BenchLM #9 | $5 / $25 |
| 13 | V4 Pro 0813 DeepSeek · open Current DeepSeek-V4 open challenger | 73.5XBench · 3 labs | 73Intel · 1 lab | 73.4Agentic · 1 lab | 74.2Coding · 1 lab | $0.25Lab | AA #14Arena —BenchLM #19 | — |
| 14 | GPT-5.6 Luna OpenAI · closed Cost-efficient GPT-5.6 | 72.7XBench · 3 labs | 70.3Intel · 1 lab | 65.9Agentic · 1 lab | 81.8Coding · 1 lab | $0.05Lab | AA #16Arena —BenchLM #15 | $0.20 / $1.20 |
| 15 | V4 Flash DeepSeek · open Open price-performance pick | 71.8XBench · 3 labs | 70.3Intel · 1 lab | 70.1Agentic · 1 lab | 75.1Coding · 1 lab | $0.03Lab | AA #16Arena —BenchLM #18 | $0.14 / $0.28 |
| 16 | Sonnet 5 Anthropic · closed Frontier general-purpose/agentic | 70.1XBench · 3 labs | 54.8Intel · 2 labs | 73.7Agentic · 1 lab | 81.8Coding · 1 lab | $1.72Lab | AA #13Arena #17BenchLM #15 | $2 / $10 |
| 17 | GLM-5.2 Zhipu AI · open Frontier/open-weight coding | 65.3XBench · 3 labs | 59.1Intel · 2 labs | 62.6Agentic · 1 lab | 74.2Coding · 1 lab | $0.32Lab | AA #14Arena #14BenchLM #19 | $1.40 / $4.40 |
| 18 | GPT-5.5 OpenAI · closed Frontier reasoning + agents | 64.6XBench · 3 labs | 35.3Intel · 2 labs | 67.3Agentic · 1 lab | 91.3Coding · 1 lab | — | AA #28Arena #10BenchLM #8 | $2.50 / $15 |
| 19 | 3.6 Flash Google DeepMind · closed Frontier-class fast reasoning | 63.8XBench · 3 labs | 67.9Intel · 2 labs | 48.2Agentic · 1 lab | 75.4Coding · 1 lab | $0.56Lab | AA #16Arena #9BenchLM #17 | $1.50 / $7.50 |
| 20 | Grok 4.5 SpaceXAI · closed Frontier coding/agents | 58.5XBench · 3 labs | 61.6Intel · 2 labs | 71.5Agentic · 1 lab | 42.3Coding · 2 labs | $0.36Lab | AA #10Arena #15BenchLM #11 | $2 / $6 |
| 21 | 3.1 Pro Google DeepMind · closed Frontier multimodal/reasoning | 46.1XBench · 3 labs | 64.1Intel · 2 labs | 0Agentic · 1 lab | 74.2Coding · 1 lab | $0.33Lab | AA #19Arena #8BenchLM #19 | $1 / $6 |
| 22 | MiniMax M3 MiniMax · open Open-weight frontier-adjacent | 44.4XBench · 3 labs | 51.4Intel · 1 lab | 36Agentic · 1 lab | 45.7Coding · 1 lab | $0.14Lab | AA #20Arena —BenchLM #26 | $0.60 / $2.40 |
| 23 | Qwen3.7 Max Alibaba · closed Frontier Chinese model | 43.5XBench · 3 labs | 42.6Intel · 2 labs | 21.6Agentic · 1 lab | 66.4Coding · 1 lab | $0.24Lab | AA #25Arena #13BenchLM #22 | $1.48 / $4.42 |
| 24 | MiMo V2.5 Pro Xiaomi · open Open MIT, cheap per-task | 37.9XBench · 3 labs | 45.9Intel · 1 lab | 17.7Agentic · 1 lab | 50.1Coding · 1 lab | $0.03Lab | AA #22Arena —BenchLM #24 | $0.43 / $0.87 |
| 25 | Hy3 Tencent · open Open Apache-2.0 coding/reasoning | 37.5XBench · 3 labs | 43.2Intel · 1 lab | 23Agentic · 1 lab | 46.2Coding · 1 lab | $0.04Lab | AA #23Arena —BenchLM #25 | $0.13 / $0.53 |
| 26 | Inkling Thinking Machines · open New open lab, Apache-2.0 | 33.7XBench · 3 labs | 43.2Intel · 1 lab | 30.5Agentic · 1 lab | 27.5Coding · 1 lab | $0.34Lab | AA #23Arena —BenchLM #27 | $0.50 / $2.02 |
| 27 | Grok 4.3 SpaceXAI · closed Frontier reasoning/agents | 6XBench · 3 labs | 14.9Intel · 2 labs | 3Agentic · 1 lab | 0Coding · 1 lab | — | AA #26Arena #18BenchLM #28 | $1.25 / $2.50 |
| — | Mythos 5 Anthropic · closed New Claude 5 frontier tier | — | — | — | 96.4Coding · 2 labs | — | AA —Arena —BenchLM #1 | $10 / $50 |
| — | Grok 4.20 SpaceXAI · closed Frontier reasoning/agents | — | 51.6Intel · 1 lab | — | — | — | AA —Arena #12BenchLM — | $1.25 / $2.50 |
| — | V4 Pro DeepSeek · open Frontier open/low-cost challenger | — | 51.4Intel · 1 lab | — | — | $0.05Lab | AA #20Arena —BenchLM — | $0.44 / $0.87 |
| — | Doubao Seed ByteDance · closed Frontier reasoning | — | 0Intel · 1 lab | — | — | — | AA #29Arena —BenchLM — | — |
| — | LongCat 2.0 Meituan · open Frontier open-weight agentic coding | — | 21.6Intel · 1 lab | — | — | — | AA #27Arena —BenchLM — | — |
32 models on this board · # is this page’s rank · purple = picked in the release-dates table
Line graph of rank on this page versus XBench, AA, Arena, and BenchLM. Top of the chart is #1.
Place across boards. Top of the chart is #1. Purple dots are models picked in the release table.
Mean of Intelligence, Agentic, and Coding pillars when present. Click a release-date row to highlight that model on this page.
Models left to right by this metric. Purple dots are models picked in the release table.