claudecodex

Board

Our measured standings first, then coding results from public leaderboards — reproduced with source and as-of date, never merged with ours.

Measured on this site: 33 trials, 624 locked tests, six models, one attempt each. A trial win is the fastest verified pass. One run per cell — treat tempo as signal, not gospel.

#ModelScore
1ChatGPT 5.6 Sol10
2Claude Fable 59
3ChatGPT 5.6 Terra8
4Claude Sonnet 55
5ChatGPT 5.51
6Claude Opus 4.80
Trial wins · as of 2026-07-13measured on this site

Model catalog

compiled from public announcements · refreshed 2026-07-30
ModelVendorReleasedClassStatus
Claude Opus 5NewAnthropic's newest Opus model for difficult agentic coding and enterprise work.Anthropic2026-07-24frontierScouting
Gemini 3.6 FlashNewA production Flash model optimized for rapid code-generation and agentic loops.Google2026-07-21fastScouting
Gemini 3.5 Flash-LiteNewGoogle's low-latency, low-cost model for high-volume automation and subagents.Google2026-07-21fastScouting
Grok 4.5NewSpaceXAI's frontier model for coding, agentic tasks, and knowledge work.SpaceXAI (xAI)2026-07-16frontierScouting
Kimi K3NewMoonshot's native multimodal open-weight frontier model for long-horizon coding and knowledge work.Moonshot AI2026-07-16open-weightScouting
GPT-5.6 SolOpenAI's flagship GPT-5.6 tier for complex reasoning, coding, and professional work.OpenAI2026-07-09frontierFielded here
GPT-5.6 TerraThe balanced GPT-5.6 tier, trading some peak capability for lower cost.OpenAI2026-07-09frontierFielded here
GPT-5.6 LunaNewThe fastest and most affordable GPT-5.6 tier for high-volume workloads.OpenAI2026-07-09fastScouting
Muse Spark 1.1NewMeta's proprietary multimodal reasoning model for coding, tool use, and multi-agent work.Meta2026-07-09frontierScouting
Claude Sonnet 5Anthropic's speed-and-intelligence tier with stronger autonomous coding and tool use.Anthropic2026-06-30fastFielded here
GLM-5.2NewAn MIT-licensed flagship model with a long context and emphasis on sustained coding tasks.Z.ai (Zhipu AI)2026-06-17open-weightScouting
Claude Fable 5Anthropic's highest-capability Claude tier for long-running agents.Anthropic2026-06-09frontierFielded here
Claude Opus 4.8A high-capability Opus release for complex agentic coding and enterprise work.Anthropic2026-05-28frontierFielded here
Qwen3.7-MaxNewAlibaba's current flagship Qwen model for long-context reasoning and complex engineering agents.Alibaba2026-05-20frontierScouting
Gemini 3.5 FlashNewGoogle's agentic Flash model combining frontier-level capability with high speed.Google2026-05-19fastScouting
Grok Build 0.1NewA specialized early-access model trained for agentic coding workflows.SpaceXAI (xAI)2026-05-19fastScouting
DeepSeek V4 ProNewDeepSeek's large open-weight V4 model for reasoning and agentic coding.DeepSeek2026-04-24open-weightScouting
DeepSeek V4 FlashNewThe smaller, faster, and more economical open-weight member of DeepSeek V4.DeepSeek2026-04-24open-weightScouting
GPT-5.5A frontier model for agentic coding, computer use, and long-running professional work.OpenAI2026-04-23frontierFielded here
MiMo-V2.5-ProNewXiaomi's open-weight model optimized for complex agent and coding workloads.Xiaomi2026-04-23open-weightScouting
MiMo-V2.5NewXiaomi's open-weight multimodal agent model for general-purpose tool-driven work.Xiaomi2026-04-23open-weightScouting
Qwen3.6-35B-A3BNewA sparse open-weight Qwen3.6 model aimed at stable real-world utility and agentic coding.Alibaba2026-04-16open-weightScouting
Gemini 3.1 Pro PreviewNewGoogle's current Pro preview for demanding reasoning, coding, and multimodal tasks.Google2026-02-19frontierScouting
Devstral 2NewMistral's large open-weight code-agent model for repository exploration and multi-file editing.Mistral AI2025-12-09open-weightScouting
Devstral Small 2NewA compact Apache-licensed coding-agent model designed for local or single-GPU use.Mistral AI2025-12-09open-weightScouting
Llama 4 MaverickNewMeta's current downloadable Llama flagship, retained despite predating the main catalog window.Meta2025-04-05open-weightScouting