Prime Intellect Ran 153 Autonomous AI Research Runs. The Best One Closed 82% of a Human Record.

In the largest open experiment of its kind, frontier models were set loose on real optimizer research — and came within striking distance of a benchmark that took human researchers months to build.

The question of whether AI can meaningfully accelerate its own development stopped being purely theoretical this week. Prime Intellect published results from what it called the largest open experiment on how frontier models conduct AI research: more than 100 autonomous runs across 10-plus models, each tasked with pushing on genuine optimizer research rather than toy problems. As @PrimeIntellect described it, the best runs "closed 82% of the gap to a record built by dozens of humans over months."

The framing matters. This was not a model answering a benchmark question or completing a coding task with a known solution. The agents were operating in a sandbox where the goal was open-ended improvement — the kind of exploratory, iterative work that has historically been the exclusive domain of trained researchers. The 82% figure is striking precisely because the human record it chased was itself the product of sustained, collaborative effort.

Get our free daily newsletter

Get this article free — plus the lead story every day — delivered to your inbox.

Want every article and the full archive? Upgrade anytime.

No spam. Unsubscribe anytime.