Grok 4.20 Posts Record-Low 17% Hallucination Rate on New Benchmark
xAI's latest model scores dramatically lower on hallucination metrics than competitors, potentially carving out a differentiated position as the 'honest' model in a market where confidence often masks confabulation.
xAI's Grok 4.20 achieved a 17% hallucination rate on a new benchmark that prioritizes factual honesty over completeness, as reported by @XFreeze. For comparison, the same benchmark reportedly scored GPT-5.4 at 89% — though the specific benchmark methodology and conditions weren't detailed in the initial report, which warrants caution about direct comparisons.
Unlock the full briefing
Get every story in today's briefing, the full archive, and the daily AI intelligence brief.
All stories today
Full archive
Daily brief
Cancel anytime. Payments powered by Stripe.