← All posts

Artificial intelligence

DeepSeek Open-Sourced Six Libraries for Huawei's Ascend Chips. The Software Layer Is the Part Export Controls Never Touched

DeepSeek published a TileLang backend for Huawei's Ascend 950 and five kernel and communication libraries on September 30, developed with Huawei and merged upstream next to the CUDA backend. Export policy restricted the chips; it never restricted CUDA, and CUDA was always the harder thing to replace.

MAI
The GitHub social card for the deepseek-ai/DeepGEMM-Ascend repository, showing the project name and its description as a matrix multiplication kernel library for Huawei Ascend NPUs.

DeepSeek published six open-source software components for Huawei's Ascend AI accelerators on September 30, developed jointly with Huawei and announced on the company's own WeChat channel. The headline item is TileLang, a language for writing compute kernels, which now carries an official Ascend 950 backend. The other five are kernel and communication libraries that mirror, almost one for one, the set DeepSeek has already published for Nvidia GPUs.

That mirroring is the point. This is not a model release. It is DeepSeek stating that the tooling it uses to train its own frontier models now treats Chinese silicon as a first-class target.

What shipped

ComponentWhat it does
TileLang (Ascend 950 backend)Kernel language with native code generation, automatic scheduling and synchronisation for Ascend
DeepGEMM-AscendMatrix-multiplication kernels
TileKernelsVector compute and memory operators
FlashMLASparse attention for long-context inference
DeepEPCross-device expert-parallel communication
DeepSelectEfficient data selection

The Ascend 950 backend went upstream into the tile-ai/tilelang project as pull request #3308, merged on September 29 — 134 commits, credited to contributors SiriusNEO and LeiWang1999. That detail matters more than it appears. This is not a vendor fork parked in a corner of the internet. It landed in the main branch of the language, next to the CUDA backend, and it defines its own dialect for kernel launch rather than borrowing CUDA's threading model.

What the backend exposes is the Ascend machine as it actually is. Kernels can compose Cube matrix operations on the AIC units with Vector operations on the AIV units; memory can be placed explicitly across the UB, L1 and L0 levels; block-scaled MXFP8 and MXFP4 arithmetic has its own paths. An AutoSchedule pass handles dependency-aware instruction pipelining, multi-buffering and the insertion of intra-core and cross-core synchronisation barriers. Device code compiles through the bisheng compiler in Huawei's CANN toolkit, and tensors and streams interoperate with PyTorch's NPU backend. DeepSeek and Huawei also describe a supernode configuration linking 128 Ascend 950 chips, optimised across computation and communication together.

The performance claims are DeepSeek's own, published in the DeepGEMM-Ascend repository rather than verified by anyone outside it. The README reports dense GEMM reaching up to 99.8% of hardware limits, FP8 inference at up to 861 TFLOPS, MegaMoE operations at up to 846.3 TFLOPS, and MQA logits saturating the FIX pipe at 99% utilisation, all on the Ascend 950 series with CANN 9.20. Treat those as vendor numbers until someone reproduces them. The pull request's own evaluation is more modest and more useful: bfloat16 GEMM at parity with or better than PyTorch's existing NPU path, FP8 casts near bandwidth limits.

The control that was never imposed

US export policy has spent three years restricting the sale of chips. It has never restricted CUDA, and CUDA is the harder thing to replace. Nvidia's advantage is not only that its accelerators are fast; it is that twenty years of kernels, libraries, framework integrations and engineer muscle memory assume its programming model. A Chinese lab handed a warehouse of Ascend parts still faces the problem that the software it wants to run was written for something else.

That is the gap this release addresses, and it addresses it from the only position that can. Huawei has been open-sourcing CANN and wiring it into PyTorch, Triton and vLLM for a while, but a chip vendor writing its own tooling is a vendor making a pitch. DeepSeek writing it is the country's most capable model lab certifying that the tooling is good enough for frontier training — and, per the company, every TileLang operator used in its V4-series training now has a high-performance Ascend implementation. DeepSeek ported V4 to Huawei hardware back in April. It is reported to be building a data centre in Inner Mongolia holding roughly 160,000 Ascend accelerators. The libraries published today are the ones that make that number mean something.

What this is not

Six libraries and one compiler backend are not CUDA. The backend targets the Ascend 950 specifically; earlier A2 and A3 parts are still served by a separate community project. The PTO codegen path and virtual machine intrinsics were dropped during upstreaming. The kernels cover what DeepSeek's own workload needs — attention, GEMM, mixture-of-experts communication, data selection — which is a narrow and extremely well-chosen slice, not a general-purpose ecosystem. Anyone whose model does not look like DeepSeek's is largely on their own.

But the direction is the thing to notice. The chips were always going to arrive eventually; yields improve, designs iterate, stockpiles get built. The software layer was the part that looked like it would take a decade of patient, unglamorous work by people with no particular incentive to do it. DeepSeek has an incentive, it has the engineers, and it is doing the work in public under permissive licences where anyone in China — or outside it — can pick it up. Export controls can stop a shipment. They cannot un-merge a pull request.

Sources: Bloomberg: DeepSeek Unveils Huawei AI Chip Tools That May Replace Nvidia's, South China Morning Post: China's DeepSeek open-sources tools to help Huawei chips supplant Nvidia in AI, deepseek-ai/DeepGEMM-Ascend on GitHub, tile-ai/tilelang pull request #3308: Introduce Ascend 950 backend, tile-ai/tilelang-ascend on GitHub

Keep reading