Claude Solves 30% of Expert-Stumped Biology Problems in New Anthropic Benchmark
Anthropic's new BioMysteryBench gave Claude 99 open-ended biological research problems that had stumped human experts — and the model cracked roughly a third of them, raising pointed questions about where AI fits in the scientific discovery pipeline.
Anthropic published results from BioMysteryBench, a new evaluation that tests AI on exactly the kind of messy, open-ended research problems that benchmarks typically avoid. As @AnthropicAI detailed on its science blog, the benchmark consists of 99 problems drawn from real biological data — cases where domain experts had attempted analysis and gotten stuck. Claude solved roughly 30% of them.
Unlock the full briefing
Get every story in today's briefing, the full archive, and the daily AI intelligence brief.
All stories today
Full archive
Daily brief
Cancel anytime. Payments powered by Stripe.