property="og:url" content=https://futuretechnologyhq.com/article/deepseek-v4-1-flash-pricing-stock-crash/>
Future TechnologyFuture Technology
AI

DeepSeek's new model costs a fraction of a cent and just crashed two stock prices

· 2 min read · By Future Technology

Key takeaways

  • DeepSeek V4.1 Flash uses 552 billion parameters but only activates 8 to 16 billion per query
  • MiniMax and Z.ai shares dropped more than 8% on launch day in Hong Kong
  • DeepSeek says it has reduced high-bandwidth memory requirements, rattling Samsung and SK Hynix shares
  • The company ordered Huawei AI chips for a new data centre, bypassing the Nvidia supply chain entirely

552 billion parameters. Eight billion active on the input side, 16 billion on the output. Those are the numbers behind DeepSeek V4.1 Flash, which launched on 10 September and immediately compressed margins across the AI inference market.

The DeepSeek V4.1 Flash pricing is what moved stock prices. MiniMax and Z.ai, two competing Chinese AI firms, saw their Hong Kong-listed shares drop more than 8% the same day. When a model launch triggers a sell-off in rival stocks, the market is telling you something about where inference margins are headed.

How the architecture keeps costs down

V4.1 Flash is a mixture-of-experts model. The full 552 billion parameter network exists, but only a small subset activates for any given query. DeepSeek built an asymmetric split: 8 billion parameters handle input processing and 16 billion handle output generation. The result is performance that DeepSeek claims matches or beats Moonshot's Kimi K3 on standard benchmarks, at a fraction of the compute cost per token.

The pricing gap depends on the benchmark and the use case, but the direction is consistent. Every DeepSeek release pushes the floor lower. For anyone tracking how this affects hardware demand, our earlier analysis covers when RAM prices might start falling as the inference cost curve steepens.

The memory problem nobody expected

On 11 September, Samsung and SK Hynix shares fell more than 3%. The trigger was DeepSeek telling investors it has reduced the amount of high-bandwidth memory its models require.

HBM has been one of the biggest beneficiaries of the AI spending boom. Every major GPU ships with stacks of expensive high-bandwidth memory, and the chip makers building it have been printing money. If DeepSeek's architectural decisions reduce HBM demand, even at the margin, the growth story for memory makers gets harder to sustain.

The Huawei order

Separately, DeepSeek placed a large order for Huawei AI chips to power a new data centre. This is a deliberate step away from the Nvidia supply chain, signalling that Chinese AI labs are building inference capacity on domestic hardware.

This follows the V4.1 Flash beta, which gave early testers a 48-hour window before the public launch. The feedback from that beta appears to have been strong enough to justify the aggressive pricing.

What to watch

The companies most exposed are those whose entire value proposition is access to a capable model rather than the workflow built around it. If frontier-class performance is available at commodity prices, the premium API business gets harder to defend.

The part worth watching is whether OpenAI, Anthropic, and Google respond with their own price cuts or try to differentiate on reliability and tooling instead. So far, each DeepSeek release has triggered a pricing response within weeks. This one will likely be no different.

Some links in this article are affiliate links. We may earn a small commission at no extra cost to you.

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.

More from Future Technology