Frontline Hotspot
Frontline Hotspot

Behind Domestic AI Chips Breaking 40% Share: Usable but Far from Good, Training Is Still the Deep End

In 2025 domestic AI accelerator cards shipped 1.65M units in China breaking 40% share, but don't get carried away. Breaks down two routes (Hygon DCU compatible with ROCm/HIP vs Huawei Ascend/Cambricon native NPU+CANN), training is the deep end, compute centers use heterogeneous deployment, three reminders to avoid label-only compute stocks.

Published July 25, 20264 min read
<!-- domestic-ai-chip-hotspot | hotspot | Domestic AI Chip Reality Check -->

Don't get carried away by "Nvidia's share plummets"-the real situation of domestic AI chips is harsher than the numbers.

A recent report shows that in 2025, domestic AI accelerator cards shipped about 1.65 million units in China, with share stably above 40% for the first time. Nvidia's absolute dominance seems, for the first time, genuinely shaken.

But numbers easily go to your head. After the excitement, let's calmly look at the real situation and the traps everywhere.

Two Routes: The Compatible Camp and the Native Camp

Domestic AI chips are usable now, but far from good. Currently they mainly rest on two routes.

One is the compatible route, like Hygon. Its DCU (Deep Computing Unit) fully docks with AMD's ROCm/HIP ecosystem via its own DTK software platform. The programming model and library interfaces are made to look very much like Nvidia CUDA-developers can run with a few code changes, extremely low migration cost. This is popular among research institutions and industry users who need both AI and traditional scientific computing. This route doesn't compete on hardware-it competes on the smoothness of replacement.

The other is the native route, like Huawei Ascend and Cambricon-dedicated AI accelerator chips (NPU/ASIC), fully abandoning graphics rendering, optimizing only for neural-network compute. Huawei's CANN ecosystem has iterated fast in recent years and can already run hundred-billion-parameter large models-like China Telecom's Xingchen model-through full-pipeline training on a pure-domestic 10,000-card cluster. This isn't paper talk-it's real combat.

The Fatal Problem: Training Is the Deep End

But the fatal problem is exactly in "real combat."

The real test is in training. Inference scenes many domestic chips can already compete, but training is the deep end. A hundred-billion-parameter model training run often takes months and tens of millions of yuan. If the chip cluster is unstable, one accidental interruption can wash millions of investment down the drain.

So even as domestic-chip share rises, many compute centers still use "heterogeneous deployment": Nvidia A100/H100 to guarantee core-model training stability, while domestic chips accumulate experience in fine-tuning and inference. The commercial victory of domestic chips is, for now, more at the order-and-revenue level, not yet a full performance-and-ecosystem crush. Customers ultimately buy not the cold "PetaFLOPS" spec, but stable, efficient productivity.

Three Reminders

First, stay away from so-called "domestic compute" stocks that only slap on labels but have no actual million-level shipments and no big-internet-company repeat orders.

Second, in fields not under embargo, use whoever has the best price-performance-that's an iron business law. Force domestication, and if you hit a stability pit, no one backs you up.

Third, the eventual winner isn't necessarily the one with the strongest single-card compute, but-like Huawei in telecom back in the day-a "full-stack solution" company that connects bottom-layer chips, software stack, cluster scheduling, and final delivery service. For example, a company that can link compute centers across regions and chips, unified scheduling, making compute usable like water and electricity-that's the real "infrastructure" player.

Conclusion: From "Does It Exist" to "Is It Good"

The domestic-AI-chip battle has moved from "does it exist" to "is it good." It's not just hardware substitution-it's a long rebuild of ecosystem and trust. Beyond the excitement, we need to stay clear-eyed.

FAQ

Can domestic AI chips replace Nvidia now?
Inference scenes many domestic chips can compete, but training is the deep end. A hundred-billion-parameter training run costs tens of millions; one unstable cluster interruption wastes millions. Compute centers mostly use heterogeneous deployment (Nvidia for training stability + domestic for fine-tuning/inference). Commercial victory is at the order level, not yet full performance/ecosystem dominance.
What is the difference between Hygon and Huawei Ascend?
Hygon takes the compatible route-DCU docks with AMD ROCm/HIP, interfaces like CUDA, low migration cost, suited for research/industry users needing both AI and scientific computing; Huawei Ascend/Cambricon take the native route-NPU/ASIC optimized for neural networks, CANN ecosystem already supports hundred-billion-model 10k-card training.
How to spot fake domestic-compute stocks?
Two things: actual million-level shipments + big-internet-company repeat orders. Those that only slap a "domestic compute" label with no shipments and no repeats-stay away. The eventual winner is a full-stack company that connects chip + software stack + cluster scheduling + delivery service.

Related