AI Benchmark Index
Global Model Rankings
Which AI model is #1 right now?
New #1 · half the price of Fable 5
Claude Opus 5
61
AA Index
Release Radar
What shipped this month
The busiest month on record — a new frontier #1, a 2.8T open model, three Gemini tiers and a triple Qwen drop, all inside four weeks.
JUL 2026
Anthropic
Claude Opus 5
AA 611M ctx$5 / $25
Takes the overall crown — AA Intelligence 61, Agentic 55.3 — at half the price of Fable 5. SWE-bench Pro 79.2%, adds a low/medium/high effort toggle. Anthropic's 4th model in under two months.
Alibaba
Qwen3.8-Max + Image 3.0 + Audio 3.0
LLMImageTTS
Triple drop in 72 hours: Qwen3.8-Max-Preview flagship LLM, Qwen-Image 3.0 for text-to-image, and Qwen-Audio 3.0 TTS — pushing Alibaba across all three modalities in one week.
Google
Gemini 3.6 Flash + Flash-Lite + Cyber
1M ctx$1.50 / $7.5017% fewer
Workhorse refresh burns 17% fewer output tokens, output price cut from $9. Cutoff jumps to Mar 2026. Shipped with high-throughput Flash-Lite and gated security model Flash Cyber.
Moonshot AI
Kimi K3
2.8T params1M ctxMoE 16/896
First open 3T-class model — #1 on Arena's Frontend Code board, above Fable 5 and GPT-5.6 Sol. AA Index 57, level with Opus 4.8. Full weights released Jul 27.
ByteDance
Seedream 5.0 Pro + Seedance 2.5
Image4K videoNative audio
Multilingual text-and-layout image model with region-precise editing; sibling Seedance 2.5 does native 30s 4K video with audio and up to 50 reference images.
Meta
Muse Spark 1.1
1M ctxMultimodalAA 51
Agentic + video understanding and computer-use, competitive with GPT-5.5 / Opus 4.8 on agentic evals. Meta's first paid model at $1.25 / $4.25.
OpenAI
GPT-5.6 · Sol / Terra / Luna
1M ctx128K out3 tiers
Flagship Sol ties the Coding Index (78) and leads Agents' Last Exam; Terra balanced, Luna cheap. Ships the ChatGPT Work agent + parallel Ultra mode. $1–$5 per 1M input.
xAI
Grok 4.5
V9 arch$2 / $6
Cursor-trained coding model tuned to be roughly 5× cheaper per task than peers. Strong on price-per-intelligence rather than raw ceiling.