DeepSeek's Smaller Flash Model Beats Its Flagship
DeepSeek has released V4.1-Flash, a 552-billion-parameter mixture-of-experts model that the company says outperforms its larger V4-Pro on performance, speed, and cost. Starting September 14, API calls to V4-Pro will be rerouted to V4.1-Flash, cutting output costs by roughly 70%. The model includes built-in image understanding and a significantly reduced memory footprint.
Benchmark results show V4.1-Flash edging out Anthropic's Claude Opus 5 and OpenAI's GPT-5.6 Sol on Terminal-Bench and software engineering tests, though both US models lead on science reasoning. Weights are available on Hugging Face under the MIT license. The release coincides with Anthropic naming DeepSeek in a threat intelligence report over alleged distillation campaigns against Claude.
