Board
Our measured standings first, then coding results from public leaderboards — reproduced with source and as-of date, never merged with ours.
Measured on this site: 33 trials, 624 locked tests, six models, one attempt each. A trial win is the fastest verified pass. One run per cell — treat tempo as signal, not gospel.
| # | Model | Score | |
|---|---|---|---|
| 1 | ChatGPT 5.6 Sol33/33 verified · 49m 21s | 10 | |
| 2 | Claude Fable 533/33 verified · 45m 40s | 9 | |
| 3 | ChatGPT 5.6 Terra33/33 verified · 51m 36s | 8 | |
| 4 | Claude Sonnet 533/33 verified · 64m 41s | 5 | |
| 5 | ChatGPT 5.533/33 verified · 70m 59s | 1 | |
| 6 | Claude Opus 4.833/33 verified · 58m 45s | 0 |
Trial wins · as of 2026-07-13measured on this site
Model catalog
compiled from public announcements · refreshed 2026-07-30| Model | Vendor | Released | Class | Status |
|---|---|---|---|---|
| Claude Opus 5NewAnthropic's newest Opus model for difficult agentic coding and enterprise work. | Anthropic | 2026-07-24 | frontier | Scouting |
| Gemini 3.6 FlashNewA production Flash model optimized for rapid code-generation and agentic loops. | 2026-07-21 | fast | Scouting | |
| Gemini 3.5 Flash-LiteNewGoogle's low-latency, low-cost model for high-volume automation and subagents. | 2026-07-21 | fast | Scouting | |
| Grok 4.5NewSpaceXAI's frontier model for coding, agentic tasks, and knowledge work. | SpaceXAI (xAI) | 2026-07-16 | frontier | Scouting |
| Kimi K3NewMoonshot's native multimodal open-weight frontier model for long-horizon coding and knowledge work. | Moonshot AI | 2026-07-16 | open-weight | Scouting |
| GPT-5.6 SolOpenAI's flagship GPT-5.6 tier for complex reasoning, coding, and professional work. | OpenAI | 2026-07-09 | frontier | Fielded here |
| GPT-5.6 TerraThe balanced GPT-5.6 tier, trading some peak capability for lower cost. | OpenAI | 2026-07-09 | frontier | Fielded here |
| GPT-5.6 LunaNewThe fastest and most affordable GPT-5.6 tier for high-volume workloads. | OpenAI | 2026-07-09 | fast | Scouting |
| Muse Spark 1.1NewMeta's proprietary multimodal reasoning model for coding, tool use, and multi-agent work. | Meta | 2026-07-09 | frontier | Scouting |
| Claude Sonnet 5Anthropic's speed-and-intelligence tier with stronger autonomous coding and tool use. | Anthropic | 2026-06-30 | fast | Fielded here |
| GLM-5.2NewAn MIT-licensed flagship model with a long context and emphasis on sustained coding tasks. | Z.ai (Zhipu AI) | 2026-06-17 | open-weight | Scouting |
| Claude Fable 5Anthropic's highest-capability Claude tier for long-running agents. | Anthropic | 2026-06-09 | frontier | Fielded here |
| Claude Opus 4.8A high-capability Opus release for complex agentic coding and enterprise work. | Anthropic | 2026-05-28 | frontier | Fielded here |
| Qwen3.7-MaxNewAlibaba's current flagship Qwen model for long-context reasoning and complex engineering agents. | Alibaba | 2026-05-20 | frontier | Scouting |
| Gemini 3.5 FlashNewGoogle's agentic Flash model combining frontier-level capability with high speed. | 2026-05-19 | fast | Scouting | |
| Grok Build 0.1NewA specialized early-access model trained for agentic coding workflows. | SpaceXAI (xAI) | 2026-05-19 | fast | Scouting |
| DeepSeek V4 ProNewDeepSeek's large open-weight V4 model for reasoning and agentic coding. | DeepSeek | 2026-04-24 | open-weight | Scouting |
| DeepSeek V4 FlashNewThe smaller, faster, and more economical open-weight member of DeepSeek V4. | DeepSeek | 2026-04-24 | open-weight | Scouting |
| GPT-5.5A frontier model for agentic coding, computer use, and long-running professional work. | OpenAI | 2026-04-23 | frontier | Fielded here |
| MiMo-V2.5-ProNewXiaomi's open-weight model optimized for complex agent and coding workloads. | Xiaomi | 2026-04-23 | open-weight | Scouting |
| MiMo-V2.5NewXiaomi's open-weight multimodal agent model for general-purpose tool-driven work. | Xiaomi | 2026-04-23 | open-weight | Scouting |
| Qwen3.6-35B-A3BNewA sparse open-weight Qwen3.6 model aimed at stable real-world utility and agentic coding. | Alibaba | 2026-04-16 | open-weight | Scouting |
| Gemini 3.1 Pro PreviewNewGoogle's current Pro preview for demanding reasoning, coding, and multimodal tasks. | 2026-02-19 | frontier | Scouting | |
| Devstral 2NewMistral's large open-weight code-agent model for repository exploration and multi-file editing. | Mistral AI | 2025-12-09 | open-weight | Scouting |
| Devstral Small 2NewA compact Apache-licensed coding-agent model designed for local or single-GPU use. | Mistral AI | 2025-12-09 | open-weight | Scouting |
| Llama 4 MaverickNewMeta's current downloadable Llama flagship, retained despite predating the main catalog window. | Meta | 2025-04-05 | open-weight | Scouting |