DeepSeek ran V4.1 Flash for 48 hours, then closed the endpoint
Key takeaways
- The DeepSeek V4.1 Flash beta opened on 9 September with the endpoint scheduled to close on 10 September, a 48 hour window
- Multimodal input is native to the base model this time, rather than a vision encoder bolted onto a text model as in V4 Flash
- No model card has been published, so audio support is unconfirmed and testers only exercised text and images
Two days. That was the entire life of the DeepSeek V4.1 Flash beta, which opened to a limited group of testers on 9 September with an endpoint scheduled to go offline on 10 September. No model card was published, and no price.
What testers got in exchange was a look at an architecture DeepSeek describes as rebuilt rather than updated, with multimodal input handled natively instead of grafted on.
What actually changed in DeepSeek V4.1 Flash
The previous generation gave away its seams in the name. V4 Flash processed images through a vision encoder attached to a text base, and shipped as DeepSeek-V4-Flash-Vision-Exp. Vision was an extension, and it read like one.
V4.1 Flash is described as multimodal from the ground up, with text, image and audio handled in a unified way by the base model itself. That distinction is not marketing. A model that learns across modalities during pre-training tends to reason across them better than one that receives a translated image embedding at inference time.
The caveat is worth keeping in view. DeepSeek has published no model card and no modality specification. What testers exercised in practice was text and images. Treat audio as unconfirmed until the company says otherwise.
On speed, developer benchmarks during the window put output generally above 300 tokens per second, with peaks around 507. Fast enough that latency stops being the thing you notice.
Why a 48 hour window is the point
A two day beta is not a soft launch. It is a load test with an audience attached.
DeepSeek gets real traffic patterns against a new architecture, plus a wave of independent benchmark posts, and commits to nothing on pricing or availability. If the numbers hold up, the eventual launch arrives with third party evidence already in circulation. If they do not, the endpoint was always going to close on Thursday.
It also keeps the company in the conversation during a month that has been dense with model news, without spending a launch on it. Four frontier models landed inside 72 hours earlier this month, and the wider release schedule has not slowed since.
The part worth watching
The model card is the document that matters. Until DeepSeek publishes context window, pricing and the actual modality list, the audio claim sits in the same category as every unpublished spec: plausible, unverified.
Watch the gap between this beta and the real release. A short gap suggests the load test went well. A long one suggests it did not. Chinese labs are also moving on governance as well as models, with Alibaba and Cambricon taking PyTorch Foundation board seats this month.