GPT 5.2 Scores 95.7% on Single Tasks but Collapses to Under 10% When Steps Are Chained

An Oxford and Lawrence Livermore benchmark exposes a brutal reliability gap in agentic AI — individual steps work fine, but chain them together and performance falls off a cliff.

A benchmark from Oxford and Lawrence Livermore National Laboratory, surfaced by @heygurisingh, found that GPT 5.2 achieved 95.7% accuracy on individual task steps but collapsed to just 9.83% when those steps were chained into multi-step workflows. The finding is a sobering counterpoint to the day's avalanche of agentic product announcements.

Unlock the full briefing

Get every story in today's briefing, the full archive, and the daily AI intelligence brief.

All stories today

Full archive

Daily brief

Cancel anytime. Payments powered by Stripe.