DeepSeek Ships V4-Flash With Agent Focus — And It Runs Locally the Same Day
DeepSeek's new V4-Flash reportedly outperforms its own Pro-tier model on agent benchmarks, and within hours the community had it running on a single workstation in 3-bit.
DeepSeek launched the public beta of its V4-Flash API on Friday, and the pitch was blunt: the company says it has "massively upgraded" the model's agent capabilities, with benchmark scores that @deepseek_ai claims now "far surpass" the earlier V4-Pro-Preview. The announcement drew outsized attention — roughly 5.6 million views and 23,000 likes — which is notable given that Flash is nominally the smaller, cheaper tier in DeepSeek's lineup. When a company's budget model beats its flagship on the benchmarks that matter most for agents, that reordering is itself the story.
What makes this release different from a typical API drop is the speed of the local-inference follow-through. Within hours, @UnslothAI had published quantized builds, reporting that the "0731" checkpoint of V4-Flash can run losslessly in 4-bit on 168GB of RAM, and in 3-bit on 110GB. Those are large numbers for a laptop but modest for a well-specced workstation — meaning individual developers and small teams can now run a frontier-class agent model without touching a cloud API at all.
Get our free daily newsletter
Get this article free — plus the lead story every day — delivered to your inbox.
Want every article and the full archive? Upgrade anytime.
No spam. Unsubscribe anytime.