Grok 4.20 Posts Record-Low 17% Hallucination Rate on New Benchmark

xAI's latest model scores dramatically lower on hallucination metrics than competitors, potentially carving out a differentiated position as the 'honest' model in a market where confidence often masks confabulation.

xAI's Grok 4.20 achieved a 17% hallucination rate on a new benchmark that prioritizes factual honesty over completeness, as reported by @XFreeze. For comparison, the same benchmark reportedly scored GPT-5.4 at 89% — though the specific benchmark methodology and conditions weren't detailed in the initial report, which warrants caution about direct comparisons.

Unlock the full briefing

Get every story in today's briefing, the full archive, and the daily AI intelligence brief.

All stories today

Full archive

Daily brief

Cancel anytime. Payments powered by Stripe.