DeepSeek and Huawei open-source Ascend ports of DeepSeek's kernel libraries
Six libraries, including DeepGEMM and native TileLang support, let developers reuse DeepSeek's Nvidia-targeted code on Huawei's Ascend 950 chips instead of CUDA.
- Compute & infrastructure
- Open weights & ecosystem
- Notable
DeepSeek and Huawei released open-source software ports that let developers run DeepSeek’s existing GPU-optimised code on Huawei’s Ascend AI chips instead of Nvidia hardware, Tom’s Hardware reported on 30 September, following Reuters. The release covers six components: DeepGEMM, the matrix-multiplication library behind DeepSeek’s models, ported to Ascend 950 hardware with the same programming interface as its Nvidia version; DeepEP, which handles the cross-chip communication that mixture-of-experts models depend on; TileKernels and FlashMLA, covering vector computation and sparse attention; DeepSelect, for data selection; and native Ascend 950 support for TileLang, a higher-level kernel-programming language DeepSeek described as offering “a simpler programming model” than Nvidia’s CUDA.
DeepSeek said developers moving from Nvidia to Ascend hardware could keep the same APIs they already used with its existing, Nvidia-targeted DeepGEMM library, and that Huawei provided full engineering support during development. The two companies said they had jointly optimised computation and communication on a “supernode” system built from 128 Ascend 950 chips. The GitHub release, published under an MIT licence, describes itself as DeepGEMM-Ascend’s “initial release.”
CUDA’s roughly two-decade head start as the default programming environment for Nvidia GPUs is widely cited as a structural advantage that persists even where rival chips reach comparable raw performance: switching hardware has historically meant rewriting or re-tuning an entire software layer. By open-sourcing ports of code it already runs in production, DeepSeek lowers that switching cost for any developer willing to target Ascend hardware, while TileLang’s continued support for Nvidia chips means adopting the language itself doesn’t force a choice between ecosystems.
The release landed two weeks after Huawei accelerated its own Ascend chip roadmap at its Huawei Connect conference, and continues a pattern of Chinese firms building software infrastructure around domestic chips as US export controls keep Nvidia’s most capable accelerators out of the country or restrict them to cut-down variants.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.