Open Source
Open Source

410 TFlops, 95 Percent of Peak: DeepSeek Opens Its Ascend Stack

A repo-by-repo teardown of DeepSeek's Ascend kernel stack (2026-10-01 GitHub API snapshot): TileKernels, 1883 stars, MIT (dozens of TileLang-DSL operators - MoE routing, FP8/FP4 quantization, Engram gating, manifold hyper-connections - with an Ascend backend added 09-30 that auto-selects at runtime, acknowledging Huawei's engineering support); DeepGEMM-Ascend, 418 stars, MIT (fully API-compatible with the main repo, Ascend 950 first release, CANN 9.20 + torch_npu + Python 3.10+, Ascend scaling-factor packing differs from NVIDIA); DeepEP-Ascend, 190 stars (EPBuffer V2.5-compatible, validated on 950DT + CANN 9.2.0 + torch_npu 2.13.0rc1, and - as of the snapshot - no LICENSE file in the repo root, so commercial use awaits clarification); FlashMLA, 13015 stars, MIT (Ascend sparse attention at 410 TFlops / 95% peak prefill and 360 TFlops / 83% decoding, versus 1460/950 TFlops on B200; fused kernel is CUDA-only). Cross-referenced against the NVIDIA-side repos (DeepGEMM 7,907 stars, DeepEP 10,231); treat each repo's LICENSE file as the final license authority.

Published October 1, 20269 min read
<!-- deepseek-ascend-stack-resource | open-source | 410 TFlops, 95 Percent of Peak: DeepSeek Opens Its Ascend Stack -->

On September 30, 2026, DeepSeek open sourced an entire low-level operator stack for Huawei's Ascend platform: FlashMLA, its attention kernel library, shipped Ascend kernels; TileKernels, the general-purpose operator library, gained an Ascend backend; and two new repositories, the compute library DeepGEMM-Ascend and the communication library DeepEP-Ascend, went live the same day. According to press coverage (AI Tools daily briefing, September 30 issue, relaying DeepSeek's official announcement), the components map one-to-one to the counterparts DeepSeek previously released on the NVIDIA platform, and the Ascend version of TileLang already supports the high-performance implementation of training operators for the DeepSeek V4 model series, with multiple tests approaching hardware limits. All star counts in this article are from a GitHub API snapshot taken on October 1, 2026.

Before this, among leading model labs, only Meituan had publicly moved a training pipeline onto domestic chips (the LongCat 2 China-chip training line, covered in our Meituan LongCat 2 hotspot piece; that article's sourcing stands on its own, and this one follows its own verified fact base). Meituan's story was about how to train; DeepSeek's is more fundamental: what the operator layer looks like when a flagship model runs on Ascend, now open at source-code level. The model weights were open long ago (see our DeepSeek V4.1 hotspot piece); now the layer that computes and communicates those weights on this hardware has been handed over as well.

The Four Repositories at a Glance

First, the lineup. Star counts are from the October 1, 2026 GitHub API snapshot; licenses follow the LICENSE file in each repository root.

RepositoryStarsLicenseRole
TileKernels1883MITTileLang DSL operator library, NVIDIA and Ascend backends
DeepGEMM-Ascend418MITMatrix multiplication library, Ascend port of DeepGEMM
DeepEP-Ascend190No LICENSE file in rootMoE expert-parallel communication, Ascend implementation of DeepEP
FlashMLA13015MITMLA sparse attention kernels, one repo covering both platforms

The division of labor reads like a map of where the compute goes. TileKernels is the topmost general operator library, written in TileLang, a domain-specific language supporting multiple hardware backends. DeepGEMM-Ascend carries matrix multiplication, the single largest consumer of FLOPs in large models. DeepEP-Ascend handles the expert-parallel all-to-all communication that makes or breaks MoE training and inference. FlashMLA owns attention, the biggest chunk of inference latency.

TileKernels: 1883 Stars, One API for Two Kinds of Chips

TileKernels is the oldest of the four (created April 22, 2026), the only one distributed through pip (package name tile-kernels), and the one that received the pivotal update on September 30. Built on tile-ai/tilelang, it provides dozens of deeply optimized operators: MoE routing with top-k expert selection and scoring; quantization spanning per-token, per-block, and per-channel FP8/FP4 casting and dequantization, plus a fused SwiGLU-plus-quantization op; Engram gating with fused RMSNorm and weight-gradient reduction; manifold hyper-connections (mHC) including Sinkhorn normalization; and RoPE and Rand kernels.

Two lines deserve underlining. First, per the README's own claim, most kernels reach performance close to the hardware's compute or memory bandwidth limits, and all of them are already used in DeepSeek's internal training and inference workloads. Second, the September 30 news item: Huawei Ascend support. Following the NVIDIA path, the kernels now ship a second backend that is selected automatically at runtime, so the same Python APIs run on both NVIDIA GPUs and Huawei NPUs. The README's acknowledgement section explicitly thanks Huawei for its technical support and engineering investment in developing the Ascend backend.

DeepGEMM-Ascend: 418 Stars, a Fully API-Compatible GEMM Library

DeepGEMM-Ascend is the Ascend port of DeepGEMM, created September 29 and released September 30, 2026 (the README news item states initial release with Ascend 950 device support). Its most practical design choice is full API compatibility with DeepGEMM, down to the package name, deep_gemm: install the package and you keep the same APIs and development workflow as DeepGEMM on other platforms. It supports BF16, FP8, and FP4 GEMM, plus MQA logits and MegaMoE operators.

On the engineering side, the library provides a lightweight abstraction over Ascend's matrix multiply-add (MAD) primitives, hiding fractal layouts, alignment constraints, address calculations, and verbose low-level parameters so that GEMM kernels stay concise. It makes extensive use of Ascend-specific techniques such as sparse data loading and coroutine-based pipelining. The README's official claim: peak hardware performance across a wide range of matrix shapes. No specific TFlops figures are given for this repository; the performance statement stands as written in the README, which also frames the code as a reference for extreme performance optimization on Ascend.

Check the dependencies before planning a port: the CANN 9.20 toolkit (providing bin/bisheng and bin/ld.lld), torch_npu, Python 3.10 or higher, compilers with C++20 format support, and a dependency on tilelang used by the HC prenorm kernel. One easily missed detail: the README states plainly that the scaling factor format on Ascend differs from NVIDIA's. Each pair of UE8M0 scaling factors along the K dimension is packed into an int16, and the packed values are stored in MN-major order. API compatibility does not mean data formats flow through untouched; pipelines migrating from the NVIDIA side must adapt this layer themselves.

DeepEP-Ascend: 190 Stars, the Hardest Engineering Problem and the Fuzziest Compliance Story

DeepEP-Ascend is the MoE expert-parallel communication library, created and pushed on September 30, 2026. Training, inference prefill, and decoding all share one EPBuffer interface and the deep_ep package name; the implementation follows DeepEP's EPBuffer V2.5 APIs, with modes and stream behavior per Ascend's characteristics. Beyond EP, the README marks the pipeline-parallel, Bucket-collective, and Engram remote-memory primitives as work in progress. Its Ascend C kernels are compiled at runtime through the DeepJIT submodule (deepseek-ai/DeepJIT, 359 stars, MIT).

The validated stack is spelled out precisely: Ascend 950DT, CANN 9.2.0, Python 3.12, PyTorch 2.13.0+cpu, and torch_npu 2.13.0rc1, plus a PoC HDK configuration. On hardware it requires Ascend 950 (A5) with UBMEM connectivity and UBC_CTP/URMA channels between participating ranks, and an even number of ranks for multi-rank communication; on software, a CANN installation with Ascend C, the Bisheng compiler, HCCL, and HCOMM. The performance test conditions per the README: 16,384 tokens per rank, hidden size 7168, top-6 routing over 256 experts, 32 AI cores and 64 AIVs, FP8 dispatch with row-major scales, BF16 combine, and 10 warmups plus 50 samples. The README presents bandwidth results across EP sizes as charts; the specific throughput figures should be read from the repository itself and are not reproduced here. Hidden size 7168 matches the V4 series model width and top-6 over 256 experts is the classic MoE configuration, so the benchmark shapes follow the real model.

One problem in this repository must be called out on its own: the missing license. As of the October 1, 2026 snapshot, DeepEP-Ascend has no LICENSE file in its root. The GitHub API license field is empty, and the root contents listing contains no LICENSE entry. Compare the rest of the batch: DeepGEMM-Ascend and TileKernels both carry MIT license texts (verified against the LICENSE files), and FlashMLA is MIT as well. Only this one is blank. Without a license file, default copyright protection applies: you may look, but strictly speaking you may not copy, modify, or redistribute. Given that this is the hardest piece of engineering in the batch, treat it as compliance pending observation and wait for an official LICENSE before production introduction.

FlashMLA: 13015 Stars, Where 410 TFlops Counts as 95 Percent

FlashMLA is the most-starred and oldest of the four (created February 21, 2025), and it serves both NVIDIA and Ascend from a single repository. The September 30, 2026 news: release of sparse attention prefill and decoding kernels for the Huawei Ascend 950 NPU. By the README's official figures, prefill reaches 410 TFlops, 95 percent of the hardware peak, and decoding reaches 360 TFlops, 83 percent of the hardware peak. Alongside the kernels, DeepSeek published a deep-dive technical report on the algorithms and optimization techniques (docs/20260930-ascend-prefill-deep-dive.md in the repository, in both Chinese and English).

The NVIDIA-side comparison gives the numbers texture: on B200 with CUDA 13.3, the fused kernel combining Q-norm, Q-RoPE, attention, O-RoPE, and the cast to FP8 reaches 1460 TFlops in prefill and 950 TFlops in decoding. The architectures differ, so the figures cannot be compared head-to-head, but hitting 95 percent of peak on Ascend prefill places the kernel quality in the top tier. This update also optimized the fused kernel's decoding performance by roughly 10 to 15 percent (README official figures), with one caveat stated plainly in the README: the fused kernel currently supports CUDA only.

FlashMLA powers inference of DeepSeek-V4.1 on both NVIDIA GPUs and Huawei Ascend NPUs, covering prefill, decoding with FP8 or FP4 KV caches, the fused kernel, and dense attention prefill and backward kernels. That is where the headline comes from: 410 TFlops counts as only 95 points. Prefill is pressed against the hardware ceiling; decoding, at 83 percent, retains headroom, and that remaining 17 points is the most obvious thing to watch in the next release.

Relation to the NVIDIA-Side Repositories: One-to-One, Cards on the Table

The press framing said one-to-one mapping to the NVIDIA platform. At the repository level: DeepGEMM-Ascend corresponds to DeepGEMM (7907 stars, MIT, same snapshot date); DeepEP-Ascend corresponds to DeepEP (10231 stars, MIT); and FlashMLA is a single repository serving both platforms (13015 stars, MIT). The small star counts on the two Ascend-side repositories are unsurprising, given their September 29 and 30 creation dates; TileKernels, already a dual-platform repository, sits at a healthy 1883. Two background repositories round out the picture: DeepJIT (359 stars, MIT), the submodule of DeepEP-Ascend, and clangd-ascend (14 stars), an IDE support project whose license field reads NOASSERTION.

The strategy is clear from a zoom-out. Open model weights (V4.1), open training and inference operators (this batch), open dual-platform backends (TileKernels and FlashMLA): DeepSeek is currently the lab that has laid out the widest span of its technology stack. For the Ascend ecosystem, these repositories provide flagship-grade reference implementations, from abstracting GEMM over MAD primitives, to EP communication over HCCL/HCOMM and URMA, to pressing sparse attention to 95 percent of peak. On the training-infrastructure line, our Agentic RL frameworks comparison and the MiMo with verl walkthrough cover how frameworks organize training; this article covers the operator layer those frameworks stand on.

The shortcomings deserve equal clarity. First, the validation environments carry high entry costs: DeepEP-Ascend needs a 950-series system with UBMEM connectivity and a specific CANN version, and DeepGEMM-Ascend is bound to CANN 9.20; developers without hardware can only watch. Second, parts of the stack, such as FlashMLA's fused kernel, still support CUDA only. Third, DeepEP-Ascend's missing license must be resolved before any commercial use.

Closing Thoughts

Taken together, the signal of this release outweighs any single number: the operator layer is no longer a black box, Ascend has public flagship-grade implementations for the first time, and the performance claims are stated in hardware-peak percentages that anyone can audit. 410 TFlops counts as only 95 points; the remaining 5, plus the 17 decoding points, are the next steps this stack announces openly. Teams following the Ascend training line should start with the tile-kernels pip package and FlashMLA's deep-dive report; DeepGEMM-Ascend rewards careful reading by operator developers; DeepEP-Ascend is worth watching until its license lands.

FAQ

Q1: What does each repository cover, and do I need all four? A1: TileKernels provides general operators (MoE routing, quantization, Engram gating, mHC, RoPE), DeepGEMM-Ascend handles matrix multiplication, DeepEP-Ascend handles MoE expert-parallel communication, and FlashMLA handles sparse attention. Install by need: for inference, start with FlashMLA plus DeepGEMM-Ascend; add DeepEP-Ascend for training; install tile-kernels for the general operators.

Q2: DeepGEMM-Ascend is fully API-compatible with DeepGEMM. Can code be moved over directly? A2: At the interface level, yes: the package name is deep_gemm either way. But the README states plainly that the scaling factor format on Ascend differs from NVIDIA's: each pair of UE8M0 values along the K dimension is packed into an int16 and stored in MN-major order. Quantization pipelines must adapt this layer themselves, and the environment needs CANN 9.20, torch_npu, and Python 3.10 or higher.

Q3: DeepEP-Ascend has no LICENSE file. Can it be used commercially? A3: Not without care. As of the October 1, 2026 snapshot, the repository root contains no LICENSE file and the GitHub API license field is empty, so default copyright protection applies: strictly, viewing is permitted but copying, modifying, and redistributing are not. The other three repositories in the batch are all MIT. Treat it as compliance pending observation and wait for an official license before production introduction.

Q4: How good is 410 TFlops on the FlashMLA Ascend kernels? A4: These are the README's official figures: on the Ascend 950 NPU, sparse attention prefill reaches 410 TFlops, 95 percent of the hardware peak, and decoding reaches 360 TFlops, 83 percent of the peak. Prefill sits close to the ceiling, while decoding retains headroom. For reference, the fused kernel on B200 reaches 1460 TFlops in prefill and 950 TFlops in decoding, though the architectures differ and the numbers cannot be compared directly.

Q5: What environment is needed to get started? A5: On hardware, an Ascend 950 series NPU; on software, CANN 9.2.0 or higher across the board (CANN 9.20 for DeepGEMM-Ascend) plus torch_npu. TileKernels additionally needs Python 3.12, PyTorch 2.13, and TileLang 0.1.15; the DeepEP-Ascend validated stack is Python 3.12 with PyTorch 2.13.0+cpu and torch_npu 2.13.0rc1, requiring an even number of ranks and UBMEM connectivity. Without Ascend hardware, start with the FlashMLA deep-dive report and each repository's README.

This article is AI-assisted and human-edited. Last updated: 2026-10-01

FAQ

What does each repository cover, and do I need all four?
TileKernels provides general operators (MoE routing, quantization, Engram gating, mHC, RoPE), DeepGEMM-Ascend handles matrix multiplication, DeepEP-Ascend handles MoE expert-parallel communication, and FlashMLA handles sparse attention. Install by need: for inference, start with FlashMLA plus DeepGEMM-Ascend; add DeepEP-Ascend for training; install tile-kernels for the general operators.
DeepGEMM-Ascend is fully API-compatible with DeepGEMM. Can code be moved over directly?
At the interface level, yes: the package name is deep_gemm either way. But the README states plainly that the scaling factor format on Ascend differs from NVIDIA's: each pair of UE8M0 values along the K dimension is packed into an int16 and stored in MN-major order. Quantization pipelines must adapt this layer themselves, and the environment needs CANN 9.20, torch_npu, and Python 3.10 or higher.
DeepEP-Ascend has no LICENSE file. Can it be used commercially?
Not without care. As of the October 1, 2026 snapshot, the repository root contains no LICENSE file and the GitHub API license field is empty, so default copyright protection applies: strictly, viewing is permitted but copying, modifying, and redistributing are not. The other three repositories in the batch are all MIT. Treat it as compliance pending observation and wait for an official license before production introduction.
How good is 410 TFlops on the FlashMLA Ascend kernels?
These are the README's official figures: on the Ascend 950 NPU, sparse attention prefill reaches 410 TFlops, 95 percent of the hardware peak, and decoding reaches 360 TFlops, 83 percent of the peak. Prefill sits close to the ceiling, while decoding retains headroom. For reference, the fused kernel on B200 reaches 1460 TFlops in prefill and 950 TFlops in decoding, though the architectures differ and the numbers cannot be compared directly.
What environment is needed to get started?
On hardware, an Ascend 950 series NPU; on software, CANN 9.2.0 or higher across the board (CANN 9.20 for DeepGEMM-Ascend) plus torch_npu. TileKernels additionally needs Python 3.12, PyTorch 2.13, and TileLang 0.1.15; the DeepEP-Ascend validated stack is Python 3.12 with PyTorch 2.13.0+cpu and torch_npu 2.13.0rc1, requiring an even number of ranks and UBMEM connectivity. Without Ascend hardware, start with the FlashMLA deep-dive report and each repository's README.

Related

Open Source

MiniMax Opens Its Deck: mcode, the Terminal Agent You Can Audit

MiniMax open-sourced mcode, its terminal coding agent: repository MiniMax-AI/minimax-code (1,443 stars, 159 forks, TypeScript, MIT, created 2026-06-01, last push 2026-09-20, per the 2026-09-20 GitHub API), pitched as continuously unlocking model capability through excellent harness design. The core claim: the coding-agent battlefield has moved from the model to the harness, where permissions, sandboxing and auditability decide whether enterprises dare to use it. Three entry points (interactive TUI, headless mcode exec, ACP), BYOK to OpenAI and Anthropic compatible APIs, plus MCP, skills, parallel subagents and AGENTS.md. The vendor reports a 76.7 percent FrontierHarness pass rate at a 4 minute 33 second median. A cold look: 1,443 stars is still early and plugin-ecosystem depth is unproven, but for regulated industries auditability can outweigh a few points of pass rate.

Sep 20, 20268 min read
Open Source

diagram-design: AI diagrams as deliverable static files

The GitHub repo cathrynlavery/diagram-design ranked second on the OpenGithubs weekly momentum chart dated 2026-09-14, gaining 7,208 stars that week; verified on 2026-09-15 it holds 39,807 stars, 2,528 forks, HTML as its main language, an MIT license, created 2026-04-16, last pushed 2026-09-10, with only 44 open issues. It is a diagram skill pack for Agent Skills compatible hosts including Claude Code, Codex, Factory Droid, Pi, GitHub Copilot, Kiro and OpenCode, and the official README claims 39 editorial diagram types, while the weekly chart blurb says 38, a discrepancy this piece resolves in favor of the README. Its output is self-contained HTML with inline SVG: no build step, no JavaScript, no external image dependency, openable offline by double-click, with each type shipping three static variants, minimal light, minimal dark and full-editorial. The design system is what defeats the AI look: a single accent color, one or two focal elements per diagram, 1px hairline borders, no shadows, a 10px border-radius ceiling, and every coordinate and gap divisible by four. It can redraw draw.io, Mermaid and Excalidraw sources into that system through four dials, format, size, detail and audience, emitting a fidelity ledger; it inherits components, relationships, grouping and direction but never source coordinates, palette or fonts. Its tagline is No Mermaid slop, yet it ships a Mermaid import path, a tension worth reading closely. The piece also covers brand onboarding that reads your homepage for palette and font stack, maps them to semantic tokens like paper, ink, muted and accent, checks WCAG AA contrast and emits a fidelity receipt; multi-client profile isolation; and the genuinely serious engineering: CI across three platforms, clipping detected by pixel diffing rather than geometry, plus gates for Sankey conservation, waterfall running totals, treemap area error and label collision, all built to catch diagrams that lie.

Sep 15, 202610 min read
Open Source

God's Eye View: a public-data globe you run locally

The GitHub repo bilawalsidhu/gods-eye-view topped the OpenGithubs weekly momentum chart for the week dated 2026-09-13 (that snapshot records 29,396 stars and +11,455 for the week); verified on 2026-09-14 it had reached 32,399 stars, 6,480 forks, JavaScript, 199 open issues, under the MIT license (read from the repo's LICENSE file - the GitHub API license field reports NOASSERTION, which is wrong here). Its pitch is a spy-satellite simulator in your browser where every source is public and the data is real: a photorealistic 3D globe overlaid with live aircraft, ships, satellites, earthquakes, traffic and public cameras, with hands-free voice control powered by a realtime AI agent; formerly named WorldView, it grew out of a YouTube series with 5M+ views, hit number one on GitHub Trending daily and weekly in August 2026, and landed at number 8 on Product Hunt that day. Two install paths: one click with Pinokio 8.2+, or a terminal run on Node 24.x/26.x with npm ci, npm run doctor and npm run dev (localhost:4173), keyless out of the box via Esri imagery plus keyless terrain with OSM as fallback. This piece maps the capability surface and the privacy and compliance boundary, and stresses what it is not: traffic is simulated along real roads, and CCTV poses and rocket trajectories are coarse estimates. It also contrasts its MIT license with the same-batch LingBot-World 2.0, which is CC BY-NC-SA 4.0 and non-commercial.

Sep 14, 202610 min read