On September 30, 2026, DeepSeek open sourced an entire low-level operator stack for Huawei's Ascend platform: FlashMLA, its attention kernel library, shipped Ascend kernels; TileKernels, the general-purpose operator library, gained an Ascend backend; and two new repositories, the compute library DeepGEMM-Ascend and the communication library DeepEP-Ascend, went live the same day. According to press coverage (AI Tools daily briefing, September 30 issue, relaying DeepSeek's official announcement), the components map one-to-one to the counterparts DeepSeek previously released on the NVIDIA platform, and the Ascend version of TileLang already supports the high-performance implementation of training operators for the DeepSeek V4 model series, with multiple tests approaching hardware limits. All star counts in this article are from a GitHub API snapshot taken on October 1, 2026.
Before this, among leading model labs, only Meituan had publicly moved a training pipeline onto domestic chips (the LongCat 2 China-chip training line, covered in our Meituan LongCat 2 hotspot piece; that article's sourcing stands on its own, and this one follows its own verified fact base). Meituan's story was about how to train; DeepSeek's is more fundamental: what the operator layer looks like when a flagship model runs on Ascend, now open at source-code level. The model weights were open long ago (see our DeepSeek V4.1 hotspot piece); now the layer that computes and communicates those weights on this hardware has been handed over as well.
The Four Repositories at a Glance
First, the lineup. Star counts are from the October 1, 2026 GitHub API snapshot; licenses follow the LICENSE file in each repository root.
| Repository | Stars | License | Role |
|---|---|---|---|
| TileKernels | 1883 | MIT | TileLang DSL operator library, NVIDIA and Ascend backends |
| DeepGEMM-Ascend | 418 | MIT | Matrix multiplication library, Ascend port of DeepGEMM |
| DeepEP-Ascend | 190 | No LICENSE file in root | MoE expert-parallel communication, Ascend implementation of DeepEP |
| FlashMLA | 13015 | MIT | MLA sparse attention kernels, one repo covering both platforms |
The division of labor reads like a map of where the compute goes. TileKernels is the topmost general operator library, written in TileLang, a domain-specific language supporting multiple hardware backends. DeepGEMM-Ascend carries matrix multiplication, the single largest consumer of FLOPs in large models. DeepEP-Ascend handles the expert-parallel all-to-all communication that makes or breaks MoE training and inference. FlashMLA owns attention, the biggest chunk of inference latency.
TileKernels: 1883 Stars, One API for Two Kinds of Chips
TileKernels is the oldest of the four (created April 22, 2026), the only one distributed through pip (package name tile-kernels), and the one that received the pivotal update on September 30. Built on tile-ai/tilelang, it provides dozens of deeply optimized operators: MoE routing with top-k expert selection and scoring; quantization spanning per-token, per-block, and per-channel FP8/FP4 casting and dequantization, plus a fused SwiGLU-plus-quantization op; Engram gating with fused RMSNorm and weight-gradient reduction; manifold hyper-connections (mHC) including Sinkhorn normalization; and RoPE and Rand kernels.
Two lines deserve underlining. First, per the README's own claim, most kernels reach performance close to the hardware's compute or memory bandwidth limits, and all of them are already used in DeepSeek's internal training and inference workloads. Second, the September 30 news item: Huawei Ascend support. Following the NVIDIA path, the kernels now ship a second backend that is selected automatically at runtime, so the same Python APIs run on both NVIDIA GPUs and Huawei NPUs. The README's acknowledgement section explicitly thanks Huawei for its technical support and engineering investment in developing the Ascend backend.
DeepGEMM-Ascend: 418 Stars, a Fully API-Compatible GEMM Library
DeepGEMM-Ascend is the Ascend port of DeepGEMM, created September 29 and released September 30, 2026 (the README news item states initial release with Ascend 950 device support). Its most practical design choice is full API compatibility with DeepGEMM, down to the package name, deep_gemm: install the package and you keep the same APIs and development workflow as DeepGEMM on other platforms. It supports BF16, FP8, and FP4 GEMM, plus MQA logits and MegaMoE operators.
On the engineering side, the library provides a lightweight abstraction over Ascend's matrix multiply-add (MAD) primitives, hiding fractal layouts, alignment constraints, address calculations, and verbose low-level parameters so that GEMM kernels stay concise. It makes extensive use of Ascend-specific techniques such as sparse data loading and coroutine-based pipelining. The README's official claim: peak hardware performance across a wide range of matrix shapes. No specific TFlops figures are given for this repository; the performance statement stands as written in the README, which also frames the code as a reference for extreme performance optimization on Ascend.
Check the dependencies before planning a port: the CANN 9.20 toolkit (providing bin/bisheng and bin/ld.lld), torch_npu, Python 3.10 or higher, compilers with C++20 format support, and a dependency on tilelang used by the HC prenorm kernel. One easily missed detail: the README states plainly that the scaling factor format on Ascend differs from NVIDIA's. Each pair of UE8M0 scaling factors along the K dimension is packed into an int16, and the packed values are stored in MN-major order. API compatibility does not mean data formats flow through untouched; pipelines migrating from the NVIDIA side must adapt this layer themselves.
DeepEP-Ascend: 190 Stars, the Hardest Engineering Problem and the Fuzziest Compliance Story
DeepEP-Ascend is the MoE expert-parallel communication library, created and pushed on September 30, 2026. Training, inference prefill, and decoding all share one EPBuffer interface and the deep_ep package name; the implementation follows DeepEP's EPBuffer V2.5 APIs, with modes and stream behavior per Ascend's characteristics. Beyond EP, the README marks the pipeline-parallel, Bucket-collective, and Engram remote-memory primitives as work in progress. Its Ascend C kernels are compiled at runtime through the DeepJIT submodule (deepseek-ai/DeepJIT, 359 stars, MIT).
The validated stack is spelled out precisely: Ascend 950DT, CANN 9.2.0, Python 3.12, PyTorch 2.13.0+cpu, and torch_npu 2.13.0rc1, plus a PoC HDK configuration. On hardware it requires Ascend 950 (A5) with UBMEM connectivity and UBC_CTP/URMA channels between participating ranks, and an even number of ranks for multi-rank communication; on software, a CANN installation with Ascend C, the Bisheng compiler, HCCL, and HCOMM. The performance test conditions per the README: 16,384 tokens per rank, hidden size 7168, top-6 routing over 256 experts, 32 AI cores and 64 AIVs, FP8 dispatch with row-major scales, BF16 combine, and 10 warmups plus 50 samples. The README presents bandwidth results across EP sizes as charts; the specific throughput figures should be read from the repository itself and are not reproduced here. Hidden size 7168 matches the V4 series model width and top-6 over 256 experts is the classic MoE configuration, so the benchmark shapes follow the real model.
One problem in this repository must be called out on its own: the missing license. As of the October 1, 2026 snapshot, DeepEP-Ascend has no LICENSE file in its root. The GitHub API license field is empty, and the root contents listing contains no LICENSE entry. Compare the rest of the batch: DeepGEMM-Ascend and TileKernels both carry MIT license texts (verified against the LICENSE files), and FlashMLA is MIT as well. Only this one is blank. Without a license file, default copyright protection applies: you may look, but strictly speaking you may not copy, modify, or redistribute. Given that this is the hardest piece of engineering in the batch, treat it as compliance pending observation and wait for an official LICENSE before production introduction.
FlashMLA: 13015 Stars, Where 410 TFlops Counts as 95 Percent
FlashMLA is the most-starred and oldest of the four (created February 21, 2025), and it serves both NVIDIA and Ascend from a single repository. The September 30, 2026 news: release of sparse attention prefill and decoding kernels for the Huawei Ascend 950 NPU. By the README's official figures, prefill reaches 410 TFlops, 95 percent of the hardware peak, and decoding reaches 360 TFlops, 83 percent of the hardware peak. Alongside the kernels, DeepSeek published a deep-dive technical report on the algorithms and optimization techniques (docs/20260930-ascend-prefill-deep-dive.md in the repository, in both Chinese and English).
The NVIDIA-side comparison gives the numbers texture: on B200 with CUDA 13.3, the fused kernel combining Q-norm, Q-RoPE, attention, O-RoPE, and the cast to FP8 reaches 1460 TFlops in prefill and 950 TFlops in decoding. The architectures differ, so the figures cannot be compared head-to-head, but hitting 95 percent of peak on Ascend prefill places the kernel quality in the top tier. This update also optimized the fused kernel's decoding performance by roughly 10 to 15 percent (README official figures), with one caveat stated plainly in the README: the fused kernel currently supports CUDA only.
FlashMLA powers inference of DeepSeek-V4.1 on both NVIDIA GPUs and Huawei Ascend NPUs, covering prefill, decoding with FP8 or FP4 KV caches, the fused kernel, and dense attention prefill and backward kernels. That is where the headline comes from: 410 TFlops counts as only 95 points. Prefill is pressed against the hardware ceiling; decoding, at 83 percent, retains headroom, and that remaining 17 points is the most obvious thing to watch in the next release.
Relation to the NVIDIA-Side Repositories: One-to-One, Cards on the Table
The press framing said one-to-one mapping to the NVIDIA platform. At the repository level: DeepGEMM-Ascend corresponds to DeepGEMM (7907 stars, MIT, same snapshot date); DeepEP-Ascend corresponds to DeepEP (10231 stars, MIT); and FlashMLA is a single repository serving both platforms (13015 stars, MIT). The small star counts on the two Ascend-side repositories are unsurprising, given their September 29 and 30 creation dates; TileKernels, already a dual-platform repository, sits at a healthy 1883. Two background repositories round out the picture: DeepJIT (359 stars, MIT), the submodule of DeepEP-Ascend, and clangd-ascend (14 stars), an IDE support project whose license field reads NOASSERTION.
The strategy is clear from a zoom-out. Open model weights (V4.1), open training and inference operators (this batch), open dual-platform backends (TileKernels and FlashMLA): DeepSeek is currently the lab that has laid out the widest span of its technology stack. For the Ascend ecosystem, these repositories provide flagship-grade reference implementations, from abstracting GEMM over MAD primitives, to EP communication over HCCL/HCOMM and URMA, to pressing sparse attention to 95 percent of peak. On the training-infrastructure line, our Agentic RL frameworks comparison and the MiMo with verl walkthrough cover how frameworks organize training; this article covers the operator layer those frameworks stand on.
The shortcomings deserve equal clarity. First, the validation environments carry high entry costs: DeepEP-Ascend needs a 950-series system with UBMEM connectivity and a specific CANN version, and DeepGEMM-Ascend is bound to CANN 9.20; developers without hardware can only watch. Second, parts of the stack, such as FlashMLA's fused kernel, still support CUDA only. Third, DeepEP-Ascend's missing license must be resolved before any commercial use.
Closing Thoughts
Taken together, the signal of this release outweighs any single number: the operator layer is no longer a black box, Ascend has public flagship-grade implementations for the first time, and the performance claims are stated in hardware-peak percentages that anyone can audit. 410 TFlops counts as only 95 points; the remaining 5, plus the 17 decoding points, are the next steps this stack announces openly. Teams following the Ascend training line should start with the tile-kernels pip package and FlashMLA's deep-dive report; DeepGEMM-Ascend rewards careful reading by operator developers; DeepEP-Ascend is worth watching until its license lands.
FAQ
Q1: What does each repository cover, and do I need all four? A1: TileKernels provides general operators (MoE routing, quantization, Engram gating, mHC, RoPE), DeepGEMM-Ascend handles matrix multiplication, DeepEP-Ascend handles MoE expert-parallel communication, and FlashMLA handles sparse attention. Install by need: for inference, start with FlashMLA plus DeepGEMM-Ascend; add DeepEP-Ascend for training; install tile-kernels for the general operators.
Q2: DeepGEMM-Ascend is fully API-compatible with DeepGEMM. Can code be moved over directly? A2: At the interface level, yes: the package name is deep_gemm either way. But the README states plainly that the scaling factor format on Ascend differs from NVIDIA's: each pair of UE8M0 values along the K dimension is packed into an int16 and stored in MN-major order. Quantization pipelines must adapt this layer themselves, and the environment needs CANN 9.20, torch_npu, and Python 3.10 or higher.
Q3: DeepEP-Ascend has no LICENSE file. Can it be used commercially? A3: Not without care. As of the October 1, 2026 snapshot, the repository root contains no LICENSE file and the GitHub API license field is empty, so default copyright protection applies: strictly, viewing is permitted but copying, modifying, and redistributing are not. The other three repositories in the batch are all MIT. Treat it as compliance pending observation and wait for an official license before production introduction.
Q4: How good is 410 TFlops on the FlashMLA Ascend kernels? A4: These are the README's official figures: on the Ascend 950 NPU, sparse attention prefill reaches 410 TFlops, 95 percent of the hardware peak, and decoding reaches 360 TFlops, 83 percent of the peak. Prefill sits close to the ceiling, while decoding retains headroom. For reference, the fused kernel on B200 reaches 1460 TFlops in prefill and 950 TFlops in decoding, though the architectures differ and the numbers cannot be compared directly.
Q5: What environment is needed to get started? A5: On hardware, an Ascend 950 series NPU; on software, CANN 9.2.0 or higher across the board (CANN 9.20 for DeepGEMM-Ascend) plus torch_npu. TileKernels additionally needs Python 3.12, PyTorch 2.13, and TileLang 0.1.15; the DeepEP-Ascend validated stack is Python 3.12 with PyTorch 2.13.0+cpu and torch_npu 2.13.0rc1, requiring an even number of ranks and UBMEM connectivity. Without Ascend hardware, start with the FlashMLA deep-dive report and each repository's README.