claudecodex

News

Releases, benchmark movement, and tooling from named public sources. Green items are this site's own run log. External items refreshed 2026-07-30.

July 2026

Jul 30siteSite rebuilt as a comparison sheet; leaderboard and news sections addedNew layout: prompt and side-by-side results first. External benchmark data from six public sources, dated and linked.
Jul 24releaseAnthropic released Claude Opus 5Claude Opus 5 replaced Opus 4.8 as Anthropic's everyday high-end model, leading several coding and knowledge-work evaluations while keeping the prior Opus price.
Jul 21releaseGoogle released Gemini 3.6 Flash and 3.5 Flash-LiteGoogle made both models generally available, positioning 3.6 Flash for efficient coding and agentic planning and Flash-Lite for high-volume subagent execution.
Jul 16releaseSpaceXAI launched Grok 4.5Grok 4.5 became SpaceXAI's frontier coding and agentic model and the default model in Grok Build.
Jul 16releaseMoonshot AI launched Kimi K3Kimi K3 debuted as a native multimodal, long-context model for long-horizon coding and knowledge work, with API access at launch and weights promised later in July.
Jul 13siteBench series 004–006 added: 33 trials, 624 locked testsProtocol, interface, and engine trials joined the corpus across Python, JavaScript, Delphi, and C#. 198/198 runs verified.
Jul 12siteBench went polyglotThe measured bench expanded from Python-only to four languages.
Jul 09releaseOpenAI made the GPT-5.6 family generally availableOpenAI launched GPT-5.6 Sol, Terra, and Luna across ChatGPT, Codex, and the API, introducing durable capability tiers at three price and performance levels.
Jul 09releaseMeta opened Muse Spark 1.1 and its Model API previewMeta released Muse Spark 1.1 with stronger coding, tool use, and multi-agent orchestration, alongside public-preview developer access through the Meta Model API.

June 2026

May 2026

April 2026