News
Releases, benchmark movement, and tooling from named public sources. Green items are this site's own run log. External items refreshed 2026-07-30.
July 2026
Jul 30siteSite rebuilt as a comparison sheet; leaderboard and news sections addedNew layout: prompt and side-by-side results first. External benchmark data from six public sources, dated and linked.
Jul 24releaseAnthropic released Claude Opus 5Claude Opus 5 replaced Opus 4.8 as Anthropic's everyday high-end model, leading several coding and knowledge-work evaluations while keeping the prior Opus price.
Jul 21releaseGoogle released Gemini 3.6 Flash and 3.5 Flash-LiteGoogle made both models generally available, positioning 3.6 Flash for efficient coding and agentic planning and Flash-Lite for high-volume subagent execution.
Jul 16releaseSpaceXAI launched Grok 4.5Grok 4.5 became SpaceXAI's frontier coding and agentic model and the default model in Grok Build.
Jul 16releaseMoonshot AI launched Kimi K3Kimi K3 debuted as a native multimodal, long-context model for long-horizon coding and knowledge work, with API access at launch and weights promised later in July.
Jul 13siteBench series 004–006 added: 33 trials, 624 locked testsProtocol, interface, and engine trials joined the corpus across Python, JavaScript, Delphi, and C#. 198/198 runs verified.
Jul 12siteBench went polyglotThe measured bench expanded from Python-only to four languages.
Jul 09releaseOpenAI made the GPT-5.6 family generally availableOpenAI launched GPT-5.6 Sol, Terra, and Luna across ChatGPT, Codex, and the API, introducing durable capability tiers at three price and performance levels.
Jul 09releaseMeta opened Muse Spark 1.1 and its Model API previewMeta released Muse Spark 1.1 with stronger coding, tool use, and multi-agent orchestration, alongside public-preview developer access through the Meta Model API.
June 2026
Jun 30releaseAnthropic released Claude Sonnet 5Claude Sonnet 5 substantially increased the Sonnet tier's autonomous coding, reasoning, and tool-use capabilities while retaining a faster cost profile than Opus.
Jun 17releaseZ.ai released the open-weight GLM-5.2GLM-5.2 shipped with MIT-licensed weights, a long context window, and a focus on sustained coding and agentic tasks.
Jun 15benchmarkArtificial Analysis reweighted its index toward agentic workIntelligence Index v4.1 upgraded several tasks, including Terminal-Bench 2.1, and shifted more weight toward realistic agentic workloads.
Jun 09releaseAnthropic released Claude Fable 5 for long-running agentsClaude Fable 5 became generally available as Anthropic's highest-capability model for sustained agentic work; Mythos 5 launched with restricted access.
Jun 09siteCodex 5.5 xhigh arena waveNine arena prompts re-run on the Codex side at reasoning xhigh.
May 2026
May 28releaseAnthropic released Claude Opus 4.8 with larger Claude Code workflowsClaude Opus 4.8 arrived with coding and agentic gains, while Claude Code added dynamic workflows for very large-scale problems.
May 28toolingMistral unified work and coding agents under VibeMistral turned Vibe into a cross-surface agent with a web coding mode, a VS Code extension, parallel cloud sessions, and updated CLI controls.
May 19releaseGoogle launched Gemini 3.5 Flash and managed agentsGoogle introduced Gemini 3.5 Flash for high-speed agentic work and added managed agents powered by its Antigravity harness to the Gemini API.
May 14toolingGrok Build entered beta as a terminal coding agentSpaceXAI opened Grok Build beta access with interactive and headless terminal modes, approval controls, and parallel subagents.