OpenAI's GPT-6 Astra Posts Benchmark Jumps So Large They're Straining Credulity
OpenAI shipped its new flagship model claiming 99.9% on ARC-AGI-3 and state-of-the-art results across math, cyber, and agentic work — numbers so far ahead of the field that even seasoned observers say they can't parse the landscape anymore.
OpenAI has launched GPT-6 Astra, and the early numbers circulating on X describe a capability jump unlike anything in recent memory. According to @paji_a, the model reportedly hit 99.9% on ARC-AGI-3 and 62.7% on the standard variant — a figure he called an "abnormal jump" that has left him unable to make sense of the competitive picture. ARC-AGI has long been treated as a stubborn measure of fluid reasoning precisely because frontier models struggled to crack it. A near-perfect score, if it holds up to independent verification, would represent a discontinuity rather than an increment.
The official framing leans into that ambition. @MiddlechildKE relayed OpenAI's positioning of Astra as "their most intelligent & aligned model yet," one that "can handle any computer task fast" and sets new state-of-the-art marks on math, ARC-AGI, cyber, and professional work. That last category — "pro work" — is the tell. The pitch is no longer about answering questions well. It is about completing the kind of multi-step, tool-using labor that fills a knowledge worker's day.
Get our free daily newsletter
Get this article free — plus the lead story every day — delivered to your inbox.
Want every article and the full archive? Upgrade anytime.
No spam. Unsubscribe anytime.