DeepSeek and Huawei are going after CUDA, the moat Nvidia actually relies on
Key takeaways
- DeepSeek open-sourced six software modules for Huawei's Ascend chips, including a port of its TileLang programming language
- The tools target Ascend 950 accelerators and a 128-chip supernode the two companies tuned together
- The real target is CUDA, the software habit that keeps developers on Nvidia even when rival chips exist
Six software modules. That is what DeepSeek published on Wednesday for Huawei's Ascend chips, and it matters more than another benchmark chart would. The Hangzhou lab announced the release on its WeChat account as a joint effort with Huawei to make AI development outside Nvidia's CUDA ecosystem less painful, according to Quartz.
In plain terms, DeepSeek and Huawei are building a CUDA alternative, and they are giving it away for free.
What DeepSeek and Huawei released
The centrepiece is TileLang, DeepSeek's high-level open-source language for programming chips, now adapted to run on Huawei's Ascend 950 accelerators. The South China Morning Post reports that the new version brings native code generation, automatic scheduling and synchronisation to the Ascend platform.
Around it sit libraries DeepSeek already relies on for its own models. Startup Fortune lists them as DeepGEMM for matrix maths, DeepEP for communication between chips, FlashMLA for long-context attention, TileKernels for standard vector work and DeepSelect for data filtering. The two companies also tuned a "supernode" setup that links 128 Ascend 950 chips, optimising both the compute and the traffic between them.
DeepSeek says TileLang offers "a simpler programming model" than CUDA. That is a claim worth testing rather than repeating, but the intent is clear.
Why CUDA is Nvidia's real moat
Nvidia's lead usually gets described in terms of silicon, yet the stickier advantage is software. CUDA has been the default way to program Nvidia GPUs for nearly two decades, and most AI code, tooling and hiring assumes it. A rival chip can match the raw numbers and still lose, because moving a working codebase means rewriting kernels, retesting everything and retraining people.
We made a similar point in our piece on how AI hardware gets designed: the chip is only half the product. The same logic explains why software support, more than raw speed, decides what runs well when you run a local LLM on a Mac.
That is the part that makes this release interesting. DeepSeek is not asking developers to trust an unfamiliar chip. It is handing them code that already runs its own models, ported to hardware Chinese labs can actually buy.
Export controls set the stage
US export controls block Chinese companies from buying Nvidia's most advanced processors, which has pushed Huawei to widen the Ascend family. Last week Huawei said it would pull its next-generation Ascend 960DT forward to the first quarter of 2027 and laid out a chip roadmap through 2029, Invezz and Quartz report. Rotating chairman Eric Xu claimed Ascend already holds a bigger share of China's AI chip market than Nvidia, though he offered no data to back that up.
The partnership has been building for a while. DeepSeek previewed its V4 model earlier this year adapted to run on Ascend, and DeepSeek says this software release followed Huawei's new processors and supernode systems by about two weeks. Huawei expects those systems to handle model training next year, and every one of those clusters needs serious electricity, a constraint we covered in how much power AI data centres use.
What to watch
Open-source tools only matter if people use them. The signals worth tracking are whether Chinese labs beyond DeepSeek start shipping models trained on Ascend with TileLang, and whether Huawei's supernodes really take on training workloads in 2027.
If both happen, Nvidia's China problem stops being only about export licences. It becomes a question of developer habit, which is much harder to win back.