Frontline Hotspot
Frontline Hotspot

DeepSeek Ports the NVIDIA Stack to Ascend: Four Repos, One Day

On 2026-09-30 DeepSeek open-sourced its Huawei Ascend infrastructure components (official basis relayed via AI-tool aggregators): a full kernel stack mapped one-to-one to its NVIDIA-side repos - TileKernels (TileLang-based operator library) gaining an Ascend backend selected automatically at runtime so one Python API runs on both GPU and NPU, DeepGEMM-Ascend (fully API-compatible, BF16/FP8/FP4 GEMM and MegaMoE), DeepEP-Ascend (EPBuffer V2.5-compatible), and FlashMLA Ascend sparse-attention kernels (410 TFlops prefill at 95% of hardware peak, 360 TFlops decoding at 83%, with a deep-dive technical report, README official basis). "Multiple tests near hardware limits" is the official claim. Cold thought: DeepEP-Ascend ships no LICENSE file as of the snapshot, some fused kernels are CUDA-only, and the validated stack centers on Ascend 950 series.

Published October 1, 20269 min read
<!-- deepseek-ascend-opensource-hotspot | hotspot | DeepSeek Ports the NVIDIA Stack to Ascend: Four Repos, One Day -->

On September 30, 2026, DeepSeek did something no other large model company had done before: it moved the entire infrastructure operator stack behind its training and inference from the NVIDIA platform to Huawei's Ascend, and open-sourced all of it in a single day. What appeared over the course of that day includes TileKernels, a deeply optimized operator library built on TileLang; DeepGEMM-Ascend, an Ascend port of the DeepGEMM matrix computation library; DeepEP-Ascend, an Ascend implementation of the DeepEP expert-parallel communication library; and a set of new Ascend attention operators added to FlashMLA, the repository with 13,015 stars.

The official framing, as relayed by AI tools newsletter coverage of DeepSeek's announcement, goes like this: these components correspond one-to-one with the components previously released for the NVIDIA platform, covering matrix computation, distributed communication, sparse attention and other core capabilities, with multiple tests approaching hardware limits. The Ascend version of TileLang, according to the same source, already supports the high-performance implementation of training operators for the DeepSeek V4 series of models. The most eye-catching numbers come from FlashMLA: on Huawei's Ascend 950 NPU, sparse attention prefill reaches 410 TFlops, which is 95 percent of hardware peak, and decoding reaches 360 TFlops, or 83 percent of hardware peak (README, official claim).

One thing needs emphasis up front: the statement that "multiple tests approach hardware limits" is an official claim relayed through official channels, not an independent third-party benchmark. What do these numbers actually mean, and what does the release mean for the domestic silicon ecosystem? This article takes the four components apart one by one, and puts the question marks on the table.

Four Repos in One Day: A Line-by-Line Mapping of the NVIDIA Stack

First, some context on what each piece does. Training and inference for large models can be roughly divided into three layers: at the bottom are operators, the most basic compute units such as matrix multiplication, attention, and quantization; in the middle is communication, since expert-parallel architectures require high-frequency data exchange between devices; on top sit the model and the training framework. DeepSeek had previously open-sourced infrastructure components on the NVIDIA side, including DeepGEMM, DeepEP, and FlashMLA, all under the MIT license. The new Ascend-side components line up almost perfectly:

RepositoryRole in the stackStarsLicense
TileKernelsTileLang-based operator library, Ascend backend added Sept 301,883MIT
DeepGEMM-AscendAscend port of DeepGEMM (7,907 stars on the NVIDIA side)418MIT
DeepEP-AscendAscend implementation of DeepEP (10,231 stars on the NVIDIA side)190No LICENSE file in repo root
FlashMLA13,015-star attention library, Ascend operators added Sept 3013,015MIT

Star counts are from an October 1, 2026 snapshot. There is also a background repository, deepseek-ai/clangd-ascend, with 14 stars and a NOASSERTION license status, which merits only a passing mention.

The first component worth unpacking is TileKernels. The repository was created on April 22, 2026, originally as a deeply optimized operator library built on TileLang. TileLang itself is a multi-backend domain-specific language from the tile-ai/tilelang project: developers describe computation in tile-like abstractions, and the compiler generates efficient implementations for different hardware. On top of it, DeepSeek built dozens of operators: MoE routing with top-k expert selection; quantization operators covering per-token, per-block, and per-channel FP8 and FP4 casts and dequantization, including fused SwiGLU with quantization; Engram gating with fused RMSNorm, forward and backward passes, and weight gradient reduction; manifold hyper-connection (mHC) with Sinkhorn normalization; plus basic operators like RoPE and Rand. Two statements in the official description matter: most operators approach hardware compute throughput or bandwidth limits, and all operators are already used in internal training and inference.

The September 30 news is that TileKernels added a second backend along the NVIDIA route, namely Huawei Ascend. The backend is selected automatically at runtime, and the same Python API runs on both NVIDIA GPUs and Ascend NPUs. On the requirements side, the CUDA backend needs SM90 or SM100 hardware with CUDA 13.1 or later, while the Ascend backend needs an Ascend 950 NPU with CANN 9.2.0 or later; overall the project requires Python 3.12 or later, PyTorch 2.13 or later, and TileLang 0.1.15 or later, with the pip package name tile-kernels. The acknowledgements section thanks Huawei for its technical support and engineering investment in developing the Ascend backend for Tile Kernels, suggesting joint work with deep involvement from Huawei's engineering teams rather than a unilateral port by DeepSeek.

The second component is DeepGEMM-Ascend. Matrix multiplication (GEMM) is the single largest share of compute in large models; every Transformer layer rests on it, and if GEMM cannot saturate the hardware, the whole card is leaking compute. This is DeepSeek's Ascend port of DeepGEMM, fully API-compatible with the original, down to keeping the same package name deep_gemm. It supports GEMM in BF16, FP8, and FP4 precision, plus MQA logits and MegaMoE operators. Two implementation details stand out as Ascend-specific. First, it provides a lightweight abstraction over the Ascend MAD (matrix multiply-accumulate) primitive, hiding fractal layouts, alignment constraints, and address computation. Second, it applies Ascend-specific optimizations such as sparse data loading and coroutine pipelines. The first release on September 30, 2026 supports Ascend 950 devices; the environment requires CANN 9.20 (with bin/bisheng and bin/ld.lld), torch_npu, Python 3.10 or later, a compiler with C++20 format support, and a dependency on tilelang for the HC prenorm kernel. The README's official claim is performance approaching hardware limits across various matrix shapes. Note that no specific TFlops figures are given there, so none should be invented. One easily missed engineering detail: the scaling factor format on the Ascend side differs from NVIDIA, with UE8M0 packed along the K dimension into int16 and stored MN-major (README text). This means the "port" was not copy-paste work; quantization data layouts had to be redesigned for Ascend hardware.

The third component is DeepEP-Ascend. DeepSeek's MoE architecture distributes experts across devices, and tokens must be routed across the network, so communication efficiency directly determines cluster utilization. DeepEP-Ascend implements and follows DeepEP's EPBuffer V2.5 API, with training, inference prefill, and decoding sharing the same EPBuffer interface and the same deep_ep package name, while modes and stream behaviors are implemented according to Ascend characteristics. The officially stated validation stack is specific: Ascend 950DT, CANN 9.2.0, Python 3.12, PyTorch 2.13.0+cpu, torch_npu 2.13.0rc1, plus a PoC HDK configuration. On the hardware side it requires UBMEM connectivity on Ascend 950 (A5) with UBC_CTP and URMA channels, and an even number of ranks; on the software side it depends on CANN components including Ascend C, the Bisheng compiler, HCCL, and HCOMM. The README's performance test conditions are 16,384 tokens per rank, hidden size 7,168, top-6 routing over 256 experts, 32 AI cores and 64 AIVs, Dispatch in FP8 row-major scale format, combine in BF16, with 10 warmup runs and 50 samples. But the actual throughput figures appear only in charts, not as text, so this article quotes no specific throughput values. The repository also depends on a DeepJIT submodule (deepseek-ai/DeepJIT, 359 stars, MIT).

410 TFlops and 95 Percent of Peak: The Official Numbers

The most complete dataset among the four belongs to FlashMLA. DeepSeek's models combine MLA (multi-head latent attention) with sparse attention, making this the most compute-hungry part of inference and the optimization target every hardware vendor loses sleep over. The repository, created on February 21, 2025, was already the star of DeepSeek's open-source matrix (13,015 stars, MIT). On September 30 it released Ascend attention operators: sparse attention prefill and decoding kernels for Huawei's Ascend 950 NPU, with prefill at 410 TFlops, or 95 percent of hardware peak, and decoding at 360 TFlops, or 83 percent of hardware peak (README, official claim). The repository also published a deep-dive report on the operator's algorithms and optimization techniques at docs/20260930-ascend-prefill-deep-dive.md, in both Chinese and English.

For comparison, the NVIDIA side: on B200 with CUDA 13.3, the fused norm-RoPE-attn-RoPE-cast kernel reaches up to 1,460 TFlops on prefill and 950 TFlops on decoding. Two caveats are essential. First, 410 and 1,460 are not directly comparable: the hardware differs, and so do the kernel shapes. The 1,460 figure comes precisely from that fused kernel, and the Ascend side currently has no equivalent fused kernel. Second, both the 95 percent and 83 percent figures are README self-reported numbers; independent reproduction will take time. The README also states that optimization work related to Ascend sped up the decoding fused kernel by roughly 10 to 15 percent (official claim), while explicitly noting that this fused kernel currently supports CUDA only and does not support Ascend.

On the functionality side, FlashMLA now powers DeepSeek-V4.1 inference on both NVIDIA GPUs and Huawei Ascend NPUs (for the open-source details of the V4.1 model itself, see our earlier coverage). It supports prefill and decoding with FP8 and FP4 KV caches, fused kernels that fold small operations around attention, and dense attention prefill and backward operators. In other words, the sparse attention path, the most compute-intensive route in DeepSeek's architecture, now has a corresponding open-source implementation on both hardware platforms.

Why It Matters for the Domestic Silicon Ecosystem

Zoom out, and the weight of this release becomes clear. In public reporting, Meituan had been the only company openly building its LongCat 2 training line on domestic chips, with a detailed technical write-up (see our article on how Meituan moved its LongCat 2 training onto domestic silicon). What Meituan demonstrated is that one company can train large models on domestic chips. DeepSeek's move is different in kind: rather than delivering a trained model or a reproduction report, it open-sourced the entire operator stack with all its engineering details, so the whole industry can keep building on top of it.

The value of "one-to-one correspondence" lies there too. For operator libraries, the hardest part of cross-hardware porting is not producing a version that runs; it is grinding performance close to hardware limits, which requires mastering both the algorithms and the hardware microarchitecture. TileKernels' multi-backend design means the same Python API behaves consistently on both hardware platforms. DeepGEMM-Ascend does not even change the package name, so migration cost for existing code is theoretically minimal. DeepEP-Ascend follows the same EPBuffer V2.5 API. For teams attempting their own domestic-chip adaptations, this is a textbook they can read side by side: which operators need rewriting, which data formats must change (such as the UE8M0 scaling factor layout mentioned above), and which communication channels depend on Ascend's UBMEM, UBC_CTP, and URMA. The answers are all laid out in the repositories.

One layer deeper, TileLang as a multi-backend DSL is emerging as the foundation for this kind of porting work, and the official statement that the Ascend version of TileLang supports DeepSeek V4 series training operators shows it has passed real-world testing. We have previously discussed the competitive landscape of the open-source ecosystem in our flagship open-source model comparison. This release fills the biggest gap in that ecosystem: not model weights, but the infrastructure that makes the weights run. For domestic silicon, the long-standing shortage of operators, toolchains, and engineering patterns has just been partially filled, in open-source form, by a frontier model company.

Three Cold Showers: License, Fused Kernels, and a Narrow Validation Stack

Beyond the excitement, three question marks deserve to be laid out calmly.

First, DeepEP-Ascend, as of the October 1, 2026 snapshot, has no LICENSE file in its repository root, and the GitHub API license field is empty. By contrast, the LICENSE texts of both DeepGEMM-Ascend and TileKernels are MIT. This means teams wanting to adopt DeepEP-Ascend are in a compliance-watching position for now. Under default copyright rules, code without a license cannot be freely used, so the only option is to wait. For a high-profile open-source day, this is the most conspicuous gap.

Second, the fused kernel supports CUDA only. The README states it explicitly: the fused kernels that fold small operations around attention currently do not support Ascend. And the 1,460 TFlops prefill figure on the NVIDIA side comes precisely from that fused kernel. In other words, what has been published for Ascend is the unfused path at 410 TFlops; the extra gains from fusion (roughly 10 to 15 percent on the decoding fused kernel) remain on the CUDA side, with no timeline for an Ascend equivalent.

Third, the validation stack is concentrated on the Ascend 950 series. DeepGEMM-Ascend's first release supports Ascend 950 devices, DeepEP-Ascend's validation environment is Ascend 950DT with CANN 9.2.0, and TileKernels' Ascend backend requires an Ascend 950 NPU with CANN 9.2.0 or later. The public validation across all three repositories does not cover Ascend models before the 950, and there is no public information on whether existing Ascend hardware can run any of this, or how well. Add the fact that DeepEP-Ascend's performance data exists only as charts without text figures, and community reproduction will have to wait for teams with 950-series hardware.

The Takeaway

Four repositories open-sourced in one day, one-to-one correspondence with the NVIDIA stack, Ascend attention operators hitting 95 percent of hardware peak by official account. On September 30, 2026, DeepSeek turned the question "can domestic silicon run large-model infrastructure" from a debate into an engineering fact. But the license gap, the missing fused kernels, and the narrow validation stack are equally real. Three things to watch next: when DeepEP-Ascend adds a LICENSE, when fused kernels arrive on the Ascend side, and whether teams that get 950-series hardware can reproduce those near-peak numbers. The real meaning of open-sourcing this stack is turning one company's engineering experience in training and inference on domestic chips into a public asset. This time, DeepSeek did not just open code; it opened a complete playbook that nobody had published before.

This article is AI-assisted and human-edited. Last updated: 2026-10-01

Related

Frontline Hotspot

Ant open-sources Ming-Image: AI that designs, not draws

In mid-September 2026 Ant Group inclusionAI open-sourced two design-generation models, Ming-Image-0.1-Design and Design-Layer, under the MIT license (repo created 2026-09-17; 2026-09-24 GitHub API snapshot: 91 stars, 6 forks, Python). Both are 6B-class: Design generates UI, dashboards, infographics and posters end to end, while Design-Layer decomposes a flat design into 2 to 9 RGBA transparent layers. Highlights: end-to-end holistic generation, native RGBA VAE, and 8K long structured prompts, with a hardware floor of a single 80 GiB GPU. Facts pinned by this site: MIT is confirmed commercial across four sources (GitHub API and others); Qwen-Image 2.1 uses the non-commercial Qwen Research License, forming the license red line. Leaderboard and Crello figures (top of the UI/UX open-source chart, 4.3x faster, and so on) are uniformly tagged vendor or model-card basis and were not independently retested; unpublished pricing and quotas are not invented.

Sep 25, 20269 min read
Frontline Hotspot

DeepSeek V4.1 Flash Open Weights: The Asymmetric Design

DeepSeek open-sourced V4.1 Flash on 2026-09-10: a 552B-parameter MoE with an asymmetric Causal-Encoder-Decoder design that activates only 8B on input and 16B on output, natively multimodal, with officials citing significant KV Cache compression to cut agent-scenario cost. The API shipped alongside it - just switch the model name to deepseek-flash - and Tencent WorkBuddy, CodeBuddy plus OpenCode have integrated it fully. The model first surfaced on 9-08 as an internal preview build before being promoted on 9-10, a timeline worth noting in itself. This piece breaks down each release claim, argues the real engineering signal is not parameter count but the shift in long-context and agent cost structure implied by the asymmetric design plus 8B input activation, runs the numbers on what KV Cache compression means for accumulated multi-turn trajectories, and closes with cold takes: no published benchmark comparison, an unresolved relationship to its own V4-Flash, and concurrency and pricing still unconfirmed. Note that what shipped is model weights on HuggingFace; there is no dedicated code repository for V4.1 Flash under the official DeepSeek org.

Sep 10, 20269 min read