China Is Now Copying Something Harder Than Nvidia's GPUs: CUDA
DeepSeek and Huawei have open-sourced compute, communication and programming tools for Ascend. The contest with Nvidia is moving beyond chip speed to how easily developers can switch.

Harder than building a GPU
AI developers do not work on a GPU's transistors directly. Their work runs through compilers, libraries, frameworks, kernels and communication software. Nvidia launched CUDA in 2006, and the edge it has built over roughly 20 years lies in that entire development environment.
That is why Nvidia's moat is better described as Hardware × Software × Developers. If a chip is fast but hard to program, and moving an existing PyTorch workload takes heavy engineering effort, total cost of ownership rises. The reverse also holds. If hardware is somewhat weaker but the software is easy to use and the whole cluster runs efficiently, a buyer's decision can change.
- 2006: CUDA first released, followed by about 20 years of developer ecosystem building
- 300+: CUDA-based libraries, according to Nvidia
- 5M+: developers in the CUDA ecosystem, according to Nvidia
Three gaps DeepSeek filled
DeepGEMM-Ascend is a kernel library that runs the core matrix operations of AI efficiently on Huawei's Ascend chips. It supports BF16, FP8, FP4 and other formats, and keeps API compatibility with the original DeepGEMM. The point is that developers can keep their existing workflow while thinking less about Ascend's low-level hardware complexity.
DeepEP-Ascend is the more strategic piece. It handles communication among multiple NPUs, including expert-parallel traffic for mixture-of-experts (MoE) models. For very large AI models, performance and cost depend less on the FLOPS of a single chip than on how efficiently dozens, hundreds or thousands of accelerators are tied together.
TileLang's support for Ascend 950 lowers the barrier to entry. It is designed so developers can write high-performance kernels in a high-level, Python-like syntax while the tool handles Ascend code generation, scheduling and synchronization.
The three layers map to compute (DeepGEMM-Ascend, efficiency), communication (DeepEP-Ascend, chip-to-chip links) and programming (TileLang, ease of development).
China's answer is not only a better chip
Huawei's recent strategy makes the picture clearer. Built on UnifiedBus, the company links large numbers of NPUs in a single system. Over the longer term, it has laid out a roadmap to connect multiple SuperPoDs into clusters of up to 1 million NPUs. That figure is Huawei's plan. It does not mean the real performance and economics have been verified.
That raises the value of communication software such as DeepEP-Ascend. The more accelerators are linked, the more latency, bandwidth and memory movement become bottlenecks. Owning thousands of chips does not help if effective compute time is low, and the economics break down. This is why the unit of competition in AI semiconductors is shifting from the chip to the effective throughput of whole racks and clusters.
Why DeepSeek, not Huawei alone
When a hardware vendor builds only its own tools, the result tends to be a closed, supplier-centered ecosystem. DeepSeek sits where real users do, training and running large models. That matters because the Ascend versions of DeepGEMM and DeepEP try to match existing APIs and development flows as closely as possible.
Here, open source works as a distribution strategy. Outside developers use the tools, find bugs and add optimizations, which draws in more developers: a developer flywheel. What China is trying to copy is less the CUDA code itself than the network effect CUDA has built over the past 20 years.
What investors should watch
The key question is whether these tools draw real developer adoption, and whether Ascend clusters deliver competitive effective throughput and economics. Huawei's 1-million-NPU roadmap is a plan, not a verified result.
Insight Times Editorial Desk





