A Refusal-Stripped Qwen 3.8 Build Lands the Same Week Open Models Close the Gap

A community MLX build of Qwen 3.8 with safety refusals removed is circulating — reportedly capable of producing malware, fraud, and weapons instructions — even as legitimate optimizations push the same weights to impressive local speeds.

The open-weights community's twin nature was on full display this week. On one side, real engineering progress: @ai_hakase_ reported running Qwen-3.8-27B at 99 tokens per second on a single RTX 3090, using fp8 KV cache and int8 quantization to fit a capable model on consumer hardware. That is the kind of quiet optimization that makes local inference genuinely practical for daily work.

Unlock the full briefing

Get every story in today's briefing, the full archive, and the daily AI intelligence brief.

All stories today

Full archive

Daily brief

Cancel anytime. Payments powered by Stripe.