2026-09-10
DeepSeek V4.1 Flash Ships — and Retires Its Own Flagship
DeepSeek released V4.1 Flash on September 10: a 552-billion-parameter mixture-of-experts model with native vision that activates only ~8B parameters on input and ~16B on output.
The unusual part is the pricing and the product decision. Weights are on Hugging Face under an MIT licence, off-peak pricing is about $0.15 per million input tokens and $0.60 per million output tokens (cache hits are effectively free at $0.003), and — the eye-catching bit — DeepSeek says Flash beats its own flagship V4 Pro on most benchmarks. So on September 14 it is retiring V4 Pro entirely and routing those requests to Flash at Flash prices.
On the published tables, Flash scores 88.1 on CyberGym and 74.2 on DeepSWE v1.1 — ahead of GPT-5.6 Sol and roughly on par with Claude Opus 5 — at about 1/40th of the price. It is the most aggressive pricing move since GPT-5.6 Luna's 80% cut in July, and this one ships with open weights.