DeepSeek Is Retiring V4 Pro and Will Reroute Its Requests on Sept. 14

DeepSeek moved its V4.1-Flash model into general availability on Sept. 10, 2026, and set a date to stop serving the model above it. From 04:00 UTC on Sept. 14, the company's release note says, every request sent to deepseek-v4-pro will be answered by V4.1-Flash instead and billed at the Flash price, and that will hold until a V4.1-Pro is released.
The company describes V4.1-Flash as the smallest model in a new architecture family, one that reads images directly and was built for faster answers and higher throughput. In the API it is named deepseek-flash. The older names deepseek-v4-flash and deepseek-v4-flash-vision-exp now point to it as well, and both of those models are retired.
DeepSeek's release note puts the model at 552 billion parameters in a mixture-of-experts design, where only part of the network runs on each request: 8 billion parameters to read the input and 16 billion to produce the output. New API prices took effect on Sept. 10, with off-peak rates set at half the peak rates.
The note's claims about how good the model is are the company's own. It says tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed, and total runtime, and that new training methods put its benchmark results ahead of flagship models. It names no testers and gives no scores, and no independent evaluation of V4.1-Flash has been published.
V4-Pro is not an old model. DeepSeek's release index dates its general-availability announcement to Aug. 13, 2026. The company has posted V4.1-Flash's weights and a technical report on Hugging Face, and says the partner tools WorkBuddy and OpenCode already support the model.
