Field SOP
Field SOP

No Ascend Hardware? Three Routes into DeepSeek's Ascend Stack

A hands-on SOP for DeepSeek's Ascend kernel stack along three routes of rising hardware demands: (1) the no-NPU reading route - clone the four repos, diff the ports against the NVIDIA-side originals, study the FlashMLA deep-dive technical report (docs/20260930-ascend-prefill-deep-dive.md), and note CUDA users can stay on the main repos since the fused kernel remains CUDA-only; (2) Ascend deployment - TileKernels installs with pip install tile-kernels (Python 3.12+, PyTorch 2.13+, TileLang 0.1.15+, Ascend 950 NPU + CANN 9.2.0+) or from source via pip install -e ".[dev]", while DeepGEMM-Ascend builds from source (git clone --recursive, ./develop.sh, pip install . --no-build-isolation on CANN 9.20); (3) DeepEP-Ascend for collective communication (git clone plus the DeepJIT submodule, ASCEND_HOME_PATH setup), honestly flagged with its missing LICENSE file and validated-stack boundary (950DT, CANN 9.2.0, torch_npu 2.13.0rc1). Includes a hardware/software requirements table, common pitfalls (scaling-factor packing, even-rank requirement, recursive submodules) and a pre-flight checklist; all commands and versions come verbatim from each repo's README.

Published October 1, 202610 min read
<!-- deepseek-ascend-stack-sop | sop | No Ascend Hardware? Three Routes into DeepSeek's Ascend Stack -->

On 2026-09-30, DeepSeek open-sourced its infrastructure stack for the Huawei Ascend platform, releasing four cloneable projects: the operator library TileKernels, the matrix computation library DeepGEMM-Ascend, the communication library DeepEP-Ascend, and the Ascend edition of FlashMLA. Per the official announcement (relayed by AI Tools Daily Digest, issue 09-30), the components map one-to-one onto their NVIDIA counterparts DeepGEMM, DeepEP and FlashMLA, covering matrix computation, communication and sparse attention, with many tests approaching hardware limits.

For most engineers the real barrier is hardware: do you have an Ascend 950? This SOP splits onboarding into three routes by threshold. Without Ascend hardware, take the code-reading route and still absorb the algorithms and trade-offs; with a card, start with TileKernels via pip; for multi-card expert parallelism, move up to DeepEP-Ascend, checking its license status and validated-stack boundary first. Every command and version below comes verbatim from the repository READMEs. One scoping note: our DeepSeek V4.1 Flash open-source hotspot covers the model release itself, and our DeepSeek V4.1 Flash integration SOP covers API-layer integration, so this article handles only the operator stack without overlap.

What Is in the Stack: Four Repositories, Four Layers

First, build the mental map. The four repositories do not overlap; remember them as operators, matrix, communication and attention:

  • TileKernels: dozens of highly optimized kernels implemented in TileLang, a domain-specific language supporting multiple hardware backends. It covers MoE routing (top-k expert selection), quantization (per-token, per-block and per-channel FP8/FP4 casting and dequantization, plus fused SwiGLU with quantization), Engram gating with fused RMSNorm, manifold hyper-connections mHC with Sinkhorn normalization, RoPE and random number kernels. Per the README's own description, most kernels perform close to the hardware's compute or bandwidth limits, and all already serve internal training and inference. On 2026-09-30 the project added Ascend support: a second backend ships and is selected automatically at runtime, so the same Python APIs run on both NVIDIA GPUs and Ascend NPUs.
  • DeepGEMM-Ascend: a port of DeepGEMM to Ascend, fully API-compatible with the original and sharing the same package name deep_gemm. It supports BF16, FP8, FP4 GEMM, MQA logits and MegaMoE. It wraps the Ascend MAD (matrix multiply-add) primitives in a lightweight abstraction that hides fractal layouts, alignment constraints and address calculations, and applies Ascend-specific optimizations such as sparse data loading and coroutine-based pipelining. The initial release on 2026.09.30 supports Ascend 950 devices.
  • DeepEP-Ascend: a high-performance communication library for training and inference on Ascend NPUs. It provides expert-parallel all-to-all operations for MoE dispatch and combine, including FP8 dispatch and deferred epilogues, plus primitives for pipeline parallelism, Bucket collectives and Engram remote memory access (some still in progress). Its public buffer APIs align with NVIDIA DeepEP; training, prefill and decoding share the same EPBuffer interface and deep_ep package name, following DeepEP's EPBuffer V2.5 APIs.
  • FlashMLA: DeepSeek's optimized attention kernel library, powering DeepSeek-V4.1 inference on both NVIDIA GPUs and Ascend NPUs. The 2026.09.30 release added Ascend sparse attention prefill and decoding kernels which, per the README, reach 410 TFlops during prefill (95 percent of hardware peak) and 360 TFlops during decoding (83 percent), plus a deep-dive technical report.

The NVIDIA-side counterparts DeepGEMM, DeepEP and FlashMLA are all MIT-licensed mature repositories, giving you a complete reference frame for side-by-side reading. On domestic-chip training, our Meituan LongCat 2 hotspot previously noted that Meituan was the only company publicly training on domestic chips; DeepSeek has now put an entire operator stack on the table, a different order of magnitude.

Threshold Table: What Needs Hardware, What Only Needs Code Reading

The three routes differ sharply in requirements. Check the table before committing:

RouteHardware thresholdSoftware thresholdWhat you get
Code readingNo Ascend hardware neededNoneKernel implementation details, porting differences, deep-dive reports
TileKernels and DeepGEMM-AscendAscend 950 NPU (TileKernels); DeepGEMM-Ascend developed and validated on the Ascend 950 seriesTileKernels: Python 3.12+, PyTorch 2.13+, TileLang 0.1.15+, CANN 9.2.0+; DeepGEMM-Ascend: CANN 9.20, torch_npu, Python 3.10+, C++20-capable compilerSingle-card operators: quantization, routing, GEMM
DeepEP-AscendAscend 950 (A5) with UBMEM connectivity and UBC_CTP/URMA channels; an even number of ranks for multi-rank communicationValidated stack: Ascend 950DT, CANN 9.2.0, Python 3.12, PyTorch 2.13.0+cpu, torch_npu 2.13.0rc1, plus PoC HDK configurationExpert-parallel all-to-all communication

Note the layering: TileKernels is the only repository with a pip package; DeepGEMM-Ascend installs from source only; DeepEP-Ascend has the highest bar because it handles communication between cards. CUDA users need not wait at all: TileKernels' CUDA backend requires an SM90 or SM100 GPU with CUDA 13.1+, and the NVIDIA-side mainline repositories work directly.

Route One: Reading Code Without Ascend Hardware

No Ascend card? You can still learn most of this stack, because kernel implementations, performance data and deep-dive reports are all public. Proceed in three steps.

Step one, clone the repositories; two carry submodules, so the clone commands matter (see pitfalls):

bash
# TileKernels: the TileLang operator library
git clone https://github.com/deepseek-ai/TileKernels.git

# DeepGEMM-Ascend: submodules must be cloned too
git clone --recursive https://github.com/deepseek-ai/DeepGEMM-Ascend.git

# DeepEP-Ascend: the communication library
git clone https://github.com/deepseek-ai/DeepEP-Ascend.git

# FlashMLA: attention kernels, also carries submodules
git clone https://github.com/deepseek-ai/FlashMLA.git flash-mla
cd flash-mla
git submodule update --init --recursive

Step two, read the Ascend deep-dive report at docs/20260930-ascend-prefill-deep-dive.md inside the FlashMLA repository, in both Chinese and English, covering the algorithms and optimization techniques behind the Ascend prefill kernels. It is the rare piece of this release that hands you methodology directly: how sparse attention prefill reaches 95 percent of hardware peak on Ascend 950, beyond what code reading alone reveals.

Step three, compare against the NVIDIA-side mainline repositories to understand porting differences. Points deserving focused attention: DeepGEMM-Ascend abstracts the Ascend MAD primitives and applies sparse data loading and coroutine-based pipelining; the scaling factor format on Ascend differs from NVIDIA's (detailed in the pitfalls section); DeepEP-Ascend's Ascend C kernels run on the HCCL/HCOMM, UBMEM and URMA communication stack and compile at runtime through DeepJIT, a completely different compilation model from the CUDA side.

CUDA users can simply use the NVIDIA-side mainline repositories. One caveat: FlashMLA's fused kernel combining Q-norm, Q-RoPE, attention, O-RoPE and the FP8 cast supports CUDA only; Ascend is not supported, per the README. Also note the 2026.09.30 release carries a breaking change: it removed support for the Hopper architecture and earlier models (DeepSeek V3, V3.2, V4.0) and changed the FP8/FP4 KV cache format; to run older models, switch to the old commit the README links to.

Route Two: With Ascend Hardware, Start with TileKernels and DeepGEMM-Ascend

Given an Ascend 950, the smoothest starting point is TileKernels, the only repository shipped as a pip package:

bash
# Install a release version
pip install tile-kernels

# Or install a local development version from source
pip install -e ".[dev]"

Verify with the repository's pytest suite; the benchmark flag adds benchmarks on top of correctness checks:

bash
python -m pytest tests/quant/test_per_token_cast.py -n 4
python -m pytest tests/quant/test_per_token_cast.py --run-benchmark

Requirements restated: Python 3.12 or higher, PyTorch 2.13 or higher, TileLang 0.1.15 or higher; the Ascend backend needs an Ascend 950 NPU with CANN 9.2.0 or higher. Since the backend is selected automatically at runtime, the same Python APIs run on an NVIDIA machine (SM90 or SM100, CUDA 13.1+) unchanged.

Next, DeepGEMM-Ascend, which requires a source install. Confirm the environment first: the CANN 9.20 toolkit (providing bin/bisheng and bin/ld.lld), torch_npu, Python 3.10 or higher, C++20-capable compilers and standard libraries, plus the tilelang dependency used by the HC prenorm kernel. The workflow:

bash
# Submodule must be cloned
git clone --recursive https://github.com/deepseek-ai/DeepGEMM-Ascend.git
cd DeepGEMM-Ascend

# Link some essential includes and build the C++ extension
cat develop.sh
./develop.sh

pip install . --no-build-isolation

After installation, usage matches DeepGEMM exactly: the package name is deep_gemm, and the kernel interfaces follow DeepGEMM's documentation. The library also exposes utility functions such as set_num_sms, set_npu_arch and transform_sf_into_required_layout, and JIT behavior is controlled via environment variables like DG_JIT_DEBUG and DG_JIT_CACHE_DIR (cache defaults to $HOME/.dj).

Route Three: DeepEP-Ascend, Together with Its Compliance Notes

DeepEP-Ascend is the highest-threshold layer. Settle three things before installing.

First, hardware and the validated-stack boundary. It requires Ascend 950 (A5) with UBMEM connectivity and UBC_CTP/URMA channels between participating ranks, and an even number of ranks for multi-rank communication. The README is explicit about the validated stack: Ascend 950DT, CANN 9.2.0, Python 3.12, PyTorch 2.13.0+cpu and torch_npu 2.13.0rc1, with a manually configured PoC HDK. Kernel support on other Ascend generations or CANN versions is not established, so stepping off this stack is uncharted territory. On firmware, Huawei's Q3 commercial HDK for the Atlas 850E is planned for public availability in mid-October 2026 (around October 15, a vendor release plan) and should become the recommended public deployment baseline; the README's performance numbers came from a PoC HDK supplied to DeepSeek with manual configuration, not that unreleased commercial version.

Second, the license status. Stated plainly: as of the 2026-10-01 snapshot, the DeepEP-Ascend repository root has no LICENSE file (the GitHub API license field is empty, and the root directory listing contains no LICENSE). This contrasts sharply with the MIT-licensed DeepGEMM-Ascend, TileKernels and FlashMLA. For personal code reading it hardly matters, but commercial or compliance-sensitive teams should wait for an official license before adopting; do not assume it is MIT.

Third, the installation flow. Activate the CANN environment, install matching PyTorch and torch_npu packages, set ASCEND_HOME_PATH (or use ASCEND_TOOLKIT_HOME), then run:

bash
git clone https://github.com/deepseek-ai/DeepEP-Ascend.git
cd DeepEP-Ascend

# Fetch the required DeepJIT submodule over HTTPS.
git submodule update --init --recursive third-party/deep_jit

python -m pip install --no-build-isolation .

The installation builds the host extension; device kernels compile on first use, so keep the CANN toolkit and Bisheng compiler available at runtime. For development, bash develop.sh builds and links the extension into the source tree. To reproduce the README's EP performance cases, set these variables first:

bash
export EP_AVOID_RECORD_STREAM=1
export TASK_QUEUE_ENABLE=0
export HCCL_IF_BASE_PORT=48361

python tests/ep/test_ep.py --num-processes 8 --num-ai-cores 32 \
    --num-tokens 16384 --hidden 7168 --num-topk 6 --num-experts 256 \
    --dispatch-dtype fp8 --test-first-only

The README's test conditions are 16,384 tokens per rank, hidden size 7168, top-6 routing over 256 experts, 32 AI cores with 64 AIVs, FP8 dispatch with row-major scales, BF16 combine, and 10 warmups plus 50 samples, with sustained dispatch bandwidth reaching roughly 90 to 95 percent of the physical payload limit for EP sizes up to 32; larger EP sizes and combine remain under optimization. For exact numbers, consult the repository's charts rather than secondhand citations.

Common Pitfalls: Four Frequent Failure Points

Pitfall one, the scaling factor format. Ascend differs from NVIDIA: each pair of UE8M0 scaling factors along the K dimension is packed into an int16, and the packed values are stored in MN-major order for optimal hardware efficiency. Porting from the NVIDIA side or migrating layouts will hit this; the transform_sf_into_required_layout utility exists precisely to convert layouts.

Pitfall two, even ranks. Multi-rank communication requires an even number of ranks, and the reference tests spawn eight processes. An odd rank count is unsupported usage per the README, so do not report it as a bug.

Pitfall three, incomplete submodules. DeepGEMM-Ascend requires git clone --recursive; DeepEP-Ascend depends on the DeepJIT submodule and needs a separate git submodule update --init --recursive third-party/deep_jit; FlashMLA also needs submodule update --init --recursive. Missing submodules typically surface as build failures about missing headers or dependencies.

Pitfall four, never drop --no-build-isolation. DeepEP-Ascend's documented command is python -m pip install --no-build-isolation ., and DeepGEMM-Ascend's is pip install . --no-build-isolation. These C++ extension projects reference the environment's torch or torch_npu during the build, and isolated build environments fail directly.

Self-Check Checklist

Run through this list before and after hands-on work:

  • Environment match: Ascend hardware model (950 or 950DT), CANN version (9.2.0 or 9.20), Python, PyTorch and torch_npu versions against each README's requirements.
  • Submodules: DeepGEMM-Ascend cloned with --recursive; DeepEP-Ascend's third-party/deep_jit fetched separately; FlashMLA with submodule update --init --recursive.
  • Environment variables: ASCEND_HOME_PATH (or ASCEND_TOOLKIT_HOME) points to the active CANN installation; EP_AVOID_RECORD_STREAM set before DeepEP tests.
  • Build flags: --no-build-isolation present in install commands; develop.sh run first for DeepGEMM-Ascend.
  • Compliance: DeepEP-Ascend has no LICENSE file, evaluate separately before commercial use; the other three repositories are MIT.
  • Version expectations: FlashMLA 2026.09.30 is incompatible with older versions (KV cache format change); switch commits for older models; trust DeepEP-Ascend only inside its validated stack.
  • Data layout: Ascend-side scaling factors are UE8M0 pairs packed into int16 in MN-major order, unlike NVIDIA; convert before migrating.

FAQ

Q1: How far can I get without Ascend hardware?

More than seventy percent. All four repositories publish their kernel implementations, and FlashMLA ships a deep-dive report on algorithms and optimizations. The NVIDIA-side mainline repositories map one-to-one onto the Ascend ones, so side-by-side reading reveals porting differences. Only hands-on performance tuning needs actual hardware.

Q2: I have an Ascend 950. Which repository should I install first?

Start with TileKernels, the only repository with a pip package: pip install tile-kernels is one command, then verify with pytest. After that, install DeepGEMM-Ascend from source, and only then consider DeepEP-Ascend, which additionally demands multi-card setup, UBMEM connectivity and even ranks.

Q3: The build fails after cloning. What are the most common causes?

Three frequent ones: incomplete submodules (DeepGEMM-Ascend and FlashMLA need recursive clones, DeepEP-Ascend needs DeepJIT fetched separately); the CANN environment not activated or ASCEND_HOME_PATH unset; and install commands missing --no-build-isolation. Cross-checking against this article's verbatim commands usually locates the problem.

Q4: How good is Ascend-side performance, actually?

Per the README's official figures: FlashMLA's Ascend kernels reach 410 TFlops during prefill (95 percent of hardware peak) and 360 TFlops during decoding (83 percent); DeepGEMM-Ascend approaches hardware limit performance across a wide range of matrix shapes. These are official self-reported numbers, and reproduction should follow the repository's test scripts.

Q5: Can DeepEP-Ascend be used in production today?

Hold off for now. First, the repository root had no LICENSE file as of the 2026-10-01 snapshot, so commercial compliance remains an open question. Second, the validated-stack boundary is narrow (the 950DT plus CANN 9.2.0 combination), and stepping outside it has no official backing. Third, the recommended deployment baseline awaits Huawei's Q3 commercial HDK release. Fine for learning and testing; reconsider production once the license and commercial firmware land.

This article is AI-assisted and human-edited. Last updated: 2026-10-01

FAQ

How far can I get without Ascend hardware?
More than seventy percent. All four repositories publish their kernel implementations, and FlashMLA ships a deep-dive report on algorithms and optimizations. The NVIDIA-side mainline repositories map one-to-one onto the Ascend ones, so side-by-side reading reveals porting differences. Only hands-on performance tuning needs actual hardware.
I have an Ascend 950. Which repository should I install first?
Start with TileKernels, the only repository with a pip package: pip install tile-kernels is one command, then verify with pytest. After that, install DeepGEMM-Ascend from source, and only then consider DeepEP-Ascend, which additionally demands multi-card setup, UBMEM connectivity and even ranks.
The build fails after cloning. What are the most common causes?
Three frequent ones: incomplete submodules (DeepGEMM-Ascend and FlashMLA need recursive clones, DeepEP-Ascend needs DeepJIT fetched separately); the CANN environment not activated or ASCEND_HOME_PATH unset; and install commands missing --no-build-isolation. Cross-checking against this article's verbatim commands usually locates the problem.
How good is Ascend-side performance, actually?
Per the README's official figures: FlashMLA's Ascend kernels reach 410 TFlops during prefill (95 percent of hardware peak) and 360 TFlops during decoding (83 percent); DeepGEMM-Ascend approaches hardware limit performance across a wide range of matrix shapes. These are official self-reported numbers, and reproduction should follow the repository's test scripts.
Can DeepEP-Ascend be used in production today?
Hold off for now. First, the repository root had no LICENSE file as of the 2026-10-01 snapshot, so commercial compliance remains an open question. Second, the validated-stack boundary is narrow (the 950DT plus CANN 9.2.0 combination), and stepping outside it has no official backing. Third, the recommended deployment baseline awaits Huawei's Q3 commercial HDK release. Fine for learning and testing; reconsider production once the license and commercial firmware land.

Related

Field SOP

Change One Line of base_url: 5M Free Tokens for Your Agent

A hands-on SOP for wiring LongCat-2.5-Preview into your existing agent toolchain along three routes of rising effort: (1) web-first (sign in at longcat.ai for chat, image upload and simple agent tasks); (2) OpenAI-protocol access (point base_url at https://api.longcat.ai/openai with model ID LongCat-2.5-Preview - existing OpenAI SDKs migrate with zero code changes); (3) Anthropic-protocol access for Claude Code (three env vars: ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN, ANTHROPIC_MODEL), with Codex, OpenClaw, OpenCode and Kilo Code following the same base_url-plus-model-name swap. Includes usage tips for the 1M context and 128K output, the 5M-free-tokens offer (relayed basis), and Preview-stage caveats (no public benchmarks, unreleased weights, interface may change); validate price, latency and success rate on a small traffic slice before production.

Sep 29, 202610 min read
Field SOP

MiMo-V2.6 Integration SOP: Desktop, API and Local Weights

A hands-on SOP for wiring MiMo-V2.6 into your workflow along three routes of rising effort: (1) the MiMo Desktop client (install, sign in or plug in your own API key, switch MiMo-V2.6-Pro/Flash in the model list, UltraSpeed mode, screenshot-feedback iteration); (2) the MiMo open-platform API (create an app for a key, pass model name, messages, tools and multimodal inputs; for agent tasks, wire environment logs, test pass rates, screenshots and verifier feedback into an execute-check-correct loop; validate price, latency and success rate on a small traffic slice before production); (3) local weights (download from the HF collection collections/XiaomiMiMo/mimo-v26; parameter counts are unpublished, so hardware floors defer to the model cards; for RL reproduction, go through the five environment scripts in XiaomiMiMo/verl). Includes a pitfall table and a pre-launch checklist; API pricing defers to the platform documentation.

Sep 27, 202610 min read
Field SOP

One Line Change, Half the Cost: GPT-6 Sol/Luna Migration SOP

A hands-on SOP for migrating existing GPT-5.6 calls to GPT-6 Sol/Luna: with OpenAI-compatible access, swapping the model name (gpt-6-sol / gpt-6-luna) will most likely run — but a name swap alone captures none of the migration's other half. You must restructure caching (stable prefix first, explicit breakpoints around the invariant region) to fully capture the 90% discount, and assign reasoning effort per turn (low for simple steps, high for critical ones), leaning on the official basis that mid-conversation re-tiering preserves the cache — save where you can, spend where you must. Includes a regression-comparison workload step, the availability red line that Free/Go accounts can only use Luna, and a decision method for drawing the Luna/Sol boundary by task difficulty. Request fields and effort values defer to the official documentation as the final basis.

Sep 27, 202610 min read