Claude Opus 4.8 Arrives to Fanfare — Then Scores 12 Points Behind GPT-5.5 on DeepSWE
Anthropic shipped Opus 4.8 with dynamic agentic workflows, but independent benchmarks show it trailing OpenAI's month-old GPT-5.5 by a wide margin on the coding task suite that matters most to developers.
Anthropic released Claude Opus 4.8 this week, positioning it as a step forward in agentic coding capability with new dynamic workflow features. The launch rhetoric was confident. The benchmark results were not. As @winneravgwin documented, Opus 4.8 scored 58% on DeepSWE — a benchmark increasingly regarded as the most meaningful test of real-world coding agent performance — while OpenAI's GPT-5.5, released a full month earlier, sits at 70%. That's a 12-point gap on the test that Anthropic's own marketing language implicitly claims to dominate.
Unlock the full briefing
Get every story in today's briefing, the full archive, and the daily AI intelligence brief.
All stories today
Full archive
Daily brief
Cancel anytime. Payments powered by Stripe.