OpenAI Ships GPT-6 Astra, Gates Its Own Cyber Capabilities Behind a Safety Tripwire It Calls 'Daybreak'

OpenAI's first major leap since GPT-5 posts near-perfect agent and exploit benchmarks — and reportedly triggered an internal 'Critical' cyber threshold that forced the company to gate parts of the model on launch day.

OpenAI released GPT-6 Astra this week, its first significant model step since GPT-5, and the most notable detail isn't the benchmark line — it's the guardrail. According to @venturesintern, Astra posts near-perfect scores on ARC-AGI-3 variants and a claimed 100% on ExploitBench, and crossed what OpenAI internally labels a 'Critical' cyber threshold. The consequence: the company reportedly gated certain capabilities behind a control layer it calls 'Daybreak.' A model that scores 100% on an exploit benchmark is, definitionally, a model that can write working attacks. The fact that OpenAI shipped it anyway — with a switch — is the story.

The rollout moved fast. Early reporting suggested a limited launch, but by the weekend Astra had reached Pro, Enterprise, Business Premium, Work, Codex, and the API, per @marcopapa99. Separately, @ainunnajib relayed that Sam Altman confirmed access extending to Plus and Business tiers, describing rising 'computer-use agent adoption pressure.' In other words, the model didn't stay in a sandbox for long. Within days it was in the hands of paying developers with agent runtimes attached.

Get our free daily newsletter

Get this article free — plus the lead story every day — delivered to your inbox.

Want every article and the full archive? Upgrade anytime.

No spam. Unsubscribe anytime.