Claude Solves 30% of Expert-Stumped Biology Problems in New Anthropic Benchmark

Anthropic's new BioMysteryBench gave Claude 99 open-ended biological research problems that had stumped human experts — and the model cracked roughly a third of them, raising pointed questions about where AI fits in the scientific discovery pipeline.

Anthropic published results from BioMysteryBench, a new evaluation that tests AI on exactly the kind of messy, open-ended research problems that benchmarks typically avoid. As @AnthropicAI detailed on its science blog, the benchmark consists of 99 problems drawn from real biological data — cases where domain experts had attempted analysis and gotten stuck. Claude solved roughly 30% of them.

Unlock the full briefing

Get every story in today's briefing, the full archive, and the daily AI intelligence brief.

All stories today

Full archive

Daily brief

Cancel anytime. Payments powered by Stripe.