As of 2026-08-14
Vendor launch tables.
First-party. Not a lab.
Provider blogs and model cards. Rank on this page is DeepSWE when the vendor published it.
Vendor DeepSWE #1 is GPT-5.6 Sol
Vendor launch tables
First-party launch posts. Harnesses differ. Independent boards beat these when they disagree.
| 1 | GPT-5.6 Sol OpenAI · closed Top-tier flagship reasoning | 72.7%Vendor | 88.8%Vendor | 34.6%Vendor | — | 1728Vendor |
| 2 | GPT-5.6 Terra OpenAI · closed Balanced GPT-5.6 | 69.6%Vendor | 87.4%Vendor | 20.8%Vendor | 63.4%Vendor | — |
| 3 | Opus 5 Anthropic · closed Anthropic’s flagship frontier model | 68.8%Vendor | — | — | — | — |
| 4 | Kimi K3 Moonshot AI · open Frontier open-weight | 67.3%Vendor | 88.3%Vendor | — | — | 1668Vendor |
| 5 | GPT-5.6 Luna OpenAI · closed Cost-efficient GPT-5.6 | 67.2%Vendor | 84.7%Vendor | — | 62.7%Vendor | — |
| 6 | GLM-5.3 Z.ai · staged Latest GLM coding / cyber-defense | 66.9%Vendor | — | 28.3%Vendor | — | 1769Vendor |
| 7 | Grok 4.6 SpaceXAI · closed Latest Grok frontier model | 65.9%Vendor | — | 26%Vendor | — | — |
| 8 | 3.7 Flash Google DeepMind · closed Latest Flash workhorse | 65.3%Vendor | 85.8%Vendor | 14.9%Vendor | — | 1525Vendor |
| 9 | Qwen3.8 Max Alibaba · closed Hosted 2.4T flagship; weights promised | 56.6%Vendor | 86.6%Vendor | — | 67.7%Vendor | — |
| 10 | Grok 4.5 SpaceXAI · closed Frontier coding/agents | 53%Vendor | 83.3%Vendor | — | 64.7%Vendor | — |
| 11 | 3.6 Flash Google DeepMind · closed Frontier-class fast reasoning | 49%Vendor | 78%Vendor | — | — | 1421Vendor |
| 12 | GLM-5.2 Zhipu AI · open Frontier/open-weight coding | 46.2%Vendor | 81%Vendor | 4.6%Vendor | 62.1%Vendor | — |
| — | Mythos 5 Anthropic · closed New Claude 5 frontier tier | — | — | — | — | — |
| — | Doubao Seed ByteDance · closed Frontier reasoning | — | — | — | — | — |
| — | LongCat 2.0 Meituan · open Frontier open-weight agentic coding | — | — | — | — | — |
| — | MiniMax M3 MiniMax · open Open-weight frontier-adjacent | — | 66%Vendor | — | 59%Vendor | — |
16 models on this board · # is this page’s rank
SpaceXAI — Grok 4.6 launch table
2026-08-12 · x.ai/news/grok-4-6
Vendor table. Third-party scores are the best of self-reported or public results.
| Eval | Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Fable 5 Max |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 (Elo) | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1 Extended | 61.3% | 56.6% | 60.6% | 63.6% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| AA-Briefcase (Elo) | 1577 | 1313 | 1502 | 1574 |
Google DeepMind — Gemini 3.7 Flash comparison
2026-08-13 · deepmind.google/models/gemini/flash
Mid-tier comparison (Flash / Sonnet / GPT-5.6 Terra / Spark), not the Opus / GPT-5.6 Sol / Grok table.
| Eval | 3.7 Flash | 3.6 Flash | Sonnet 5 | GPT-5.6 Terra | Spark 1.2 |
|---|---|---|---|---|---|
| AA Intelligence Index | 56 | 52 | 55 | 57 | 57 |
| DeepSWE v1.1 | 65.3% | 48.6% | 53.8% | 69.6% | 54.9% |
| Code Arena Elo | 1588 | 1538 | 1541 | 1523 | 1535 |
| Terminal-bench 2.1 | 85.8% | 78.0% | 80.4% | 87.4% | 82.9% |
| Terminal-bench 3.0 | 14.9% | 5.4% | 14.6% | 20.8% | — |
| GDPVal-AA v2 | 1525 | 1422 | 1598 | 1578 | 1628 |
Z.ai — GLM-5.3 launch chart
2026-08-14 · z.ai/blog/glm-5.3
First-party. Mythos/Fable is a combined competitor column. No independent AA score yet.
| Eval | GLM-5.3 | GLM-5.2 | Kimi K3 | Mythos/Fable 5 | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 3.0 | 28.3% | 4.6% | 17.4% | 33.7% | 34.6% |
| DeepSWE v1.1 | 66.9% | 46.2% | 67.5% | 69.7% | 72.7% |
| GDPVal-AA v2 | 1769 | 1508 | 1682 | 1743 | 1730 |
| CyberGym | 84.5% | 77.2% | 80.0% | 83.8% | 83.6% |
| ExploitBench | 54.4% | 24.4% | 32.2% | 78.0% | 76.5% |
| HLE w/ tools | 62.5% | 54.7% | 59.8% | 63.9% | 64.5% |
OpenAI — GPT-5.6 Sol / Terra / Luna (GA table)
2026-07-09 · openai.com/index/gpt-5-6
First-party GA table. OpenAI’s HTML is JS-walled here; numbers match the published table as quoted in contemporaneous coverage and AA’s GPT-5.6 article. Current AA Intelligence Index has moved (Sol 61 / Terra 57 / Luna 52).
| Eval | GPT-5.6 Sol | GPT-5.6 Terra | GPT-5.6 Luna | GPT-5.5 | Fable 5 |
|---|---|---|---|---|---|
| AA Coding Agent Index v1.1 | 80 | 77.4 | 74.6 | 76.4 | 77.2 |
| DeepSWE v1.1 | 72.7% | 69.6% | 67.2% | 67% | 69.7% |
| Terminal-Bench 2.1 | 88.8% | 87.4% | 84.7% | 85.6% | 83.1% |
| SWE-bench Pro | 64.6% | 63.4% | 62.7% | 59.4% | 80% |
| BrowseComp | 90.4% | 87.5% | 83.3% | 84.4% | 84.3% |
| OSWorld 2.0 | 62.6% | 50.2% | 45.6% | 47.5% | — |
| AA Intelligence (launch v4.1) | 58.9 | 55 | 51.2 | 54.8 | 59.9 |
Z.ai — GLM-5.2 launch table
2026-06-16 · z.ai/blog/glm-5.2
First-party full table. Terminus-2 harness unless noted. Stars on HLE are full-set scores. Qwen3.7-Max and DeepSeek-V4-Pro columns omitted here for width; they are on the original post.
| Eval | GLM-5.2 | GLM-5.1 | Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro | MiniMax M3 |
|---|---|---|---|---|---|---|
| SWE-bench Pro | 62.1% | 58.4% | 69.2% | 58.6% | 54.2% | 59.0% |
| DeepSWE | 46.2% | 18.0% | 58.0% | 70.0% | 10.0% | 20.0% |
| Terminal-Bench 2.1 (Terminus-2) | 81.0% | 63.5% | 85.0% | 84.0% | 74.0% | 65.0% |
| FrontierSWE dominance (16 Jun) | 74.4 | 30.5 | 75.1 | 72.6 | 39.6 | — |
| PostTrainBench | 34.3 | 20.1 | 37.2 | 28.4 | 21.6 | — |
| MCP-Atlas public | 76.8% | 71.8% | 77.8% | 75.3% | 69.2% | 74.2% |
| HLE w/ tools | 54.7% | 52.3% | 57.9%* | 52.2%* | 51.4%* | — |
Qwen — Qwen3.8-Max launch table
2026-08-03 · qwen.ai/blog?id=qwen3.8
First-party. Table vs Opus 4.8 / Fable 5 / GPT-5.6 Sol / Qwen3.7-Max as published on the Qwen3.8-Max post (chart on qwen.ai; numbers from that table).
| Eval | Qwen3.8-Max | Opus 4.8 | Fable 5 | GPT-5.6 Sol | Qwen3.7-Max |
|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 86.6% | 84.6% | 84.6% | 88.8% | 74.5% |
| SWE-bench Pro | 67.7% | 69.2% | 80.0% | 64.6% | 60.6% |
| DeepSWE 1.1 | 56.6% | 59.0% | 70.0% | 73.0% | 21.6% |
| PaperBench | 93.0% | 80.3% | 88.8% | 90.5% | 64.8% |
| OSWorld-Verified | 86.1% | — | — | — | — |
| JobBench | 53.4% | — | — | — | 31.3% |
| FrontierSWE | 73.5 | 70 | 88.8 | — | 40.7 |
Google — Gemini 3.6 Flash launch numbers
2026-07-21 · blog.google Gemini 3.6 Flash
First-party vs Gemini 3.5 Flash. 3.6 Flash priced $1.50 / $7.50.
| Eval | 3.6 Flash | 3.5 Flash |
|---|---|---|
| DeepSWE | 49% | 37% |
| MLE-Bench | 63.9% | 49.7% |
| OSWorld-Verified | 83.0% | 78.4% |
| GDPval-AA v2 | 1421 | 1349 |
| AA output tokens vs 3.5 | −17% | baseline |
MiniMax — M3 launch numbers
2026-06-16 · minimax.io/blog/minimax-m3
First-party. SWE-bench Pro and Terminal-Bench 2.1 on MiniMax infra (Claude Code / Terminus 2).
| Eval | MiniMax M3 |
|---|---|
| SWE-bench Pro | 59.0% |
| Terminal-Bench 2.1 | 66.0% |
| SWE-fficiency | 34.8% |
| KernelBench Hard | 28.8% |
| MCP Atlas | 74.2% |
SpaceXAI — Grok 4.5 launch chart
2026-07-16 · x.ai/news/grok-4-5
First-party charts on x.ai/news/grok-4-5. DeepSWE 1.0 extracted from the page (AA harnesses). Other tabs (DeepSWE 1.1, SWE Marathon, Terminal-Bench 2.1, SWE-bench Pro) are the same post’s chart series as transcribed in contemporaneous coverage of those tabs. Grok 4.6’s later table restates 4.5 High DeepSWE v1.1 at 54%.
| Eval | Grok 4.5 | Fable max | GPT-5.5 xhigh | Opus 4.8 max |
|---|---|---|---|---|
| DeepSWE 1.0 (AA harnesses) | 62.0% | 66.1% | 64.3% | 55.8% |
| DeepSWE 1.1 | 53% | 70% | 67% | 59% |
| SWE Marathon pass@1 | 29.0% | 24.0% | — | 26.0% |
| Terminal-Bench 2.1 | 83.3% | 84.3% | 83.4% | 78.9% |
| SWE-bench Pro | 64.7% | 80.4% | 58.6% | 69.2% |
| Price in/out | $2 / $6 | $10 / $50 | — | $5 / $25 |
Moonshot — Kimi K3 tech blog
2026-07-16 · kimi.com/blog/kimi-k3
First-party. Full comparison chart is an image; numbers below are from the blog footnotes (Kimi Code harness unless noted).
| Eval | Kimi K3 | Note |
|---|---|---|
| DeepSWE v1.1 (mini-SWE-agent) | 67.3% | Official DeepSWE board; Kimi Code harness is the blog headline run |
| BrowseComp (1M, no compaction) | 90.4% | Blog footnote 5 |
| Price | $3 / $15 | Cache hit $0.30 / 1M |
Anthropic — Claude Opus 5 launch
2026-07-24 · anthropic.com/news/claude-opus-5
First-party. Most charts did not extract as numbers. Claims below are the ones stated in prose on the launch post.
| Claim | Opus 5 |
|---|---|
| Price | $5 / $25 — same as Opus 4.8, half of Fable 5 |
| Frontier-Bench v0.1 | SOTA; more than 2× Opus 4.8 at lower cost/task |
| CursorBench 3.2 (max) | Within 0.5% of Fable 5 peak, half the cost/task |
| ARC-AGI 3 | ~3× the next-best model |
| Zapier AutomationBench | ~1.5× next-best pass rate at the same cost/task |
| Cyber vs Mythos 5 | Close on finding vulns; far behind on exploit development (OSS-Fuzz) |
ByteDance Seed — Seed 2.1 Pro launch
2026-06-23 · research.doubao.com/en/seed2_1
First-party. Independent AA Intelligence for this row is Doubao Seed Code 26*.
| Eval | Seed 2.1 Pro | Seed 2.1 Turbo | Opus 4.7 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 71.0% | 67.6% | 71.7% | 73.8% | 70.7% |
| SWE-Atlas | 35.2% | 30.6% | 38.7% | 44.7% | 23.6% |
| NL2Repo-Bench | 47.0% | 43.7% | 58.2% | 45.1% | 33.4% |
| Workspace Bench | 53.0% | 54.7% | 55.1% | 58.7% | 32.8% |
Meituan — LongCat-2.0 model card
2026-06-01 · huggingface.co/meituan-longcat/LongCat-2.0
In-house harness unless marked *. Compared on the card to Gemini 3.1 Pro, GPT-5.5, Opus 4.7/4.8.
| Eval | LongCat-2.0 | Gemini 3.1 Pro | GPT-5.5 | Opus 4.7 | Opus 4.8 |
|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 70.8 | 70.7* | 73.8* | 71.7* | 78.9* |
| SWE-bench Pro | 59.5 | 54.2* | 58.6* | 64.3* | 69.2* |
| SWE-bench Multilingual | 77.3 | 76.9* | — | 80.5* | 84.8* |