US export controls are driving China's AI chipmakers to build a self-reliant ecosystem, sparking a strategic debate between versatile GPUs and specialized ASICs. While companies like Moore Threads champion general-purpose GPUs for flexibility, giants such as Huawei, Cambricon, and Alibaba are heavily investing in ASICs (NPUs, TPUs, PPUs) for superior efficiency and cost-effectiveness. Analysts predict ASICs will see wider commercial adoption due to lower deployment costs, though the distinction between the two technologies is blurring as models evolve.
Under the weight of sustained US export controls on advanced semiconductors, China’s AI chipmakers are battling to forge a self-reliant silicon ecosystem capable of breaking Nvidia’s stranglehold on the market. At the centre of this rivalry is a fundamental design debate: Should the country rely on the versatile graphics processing unit (GPU) or pivot to the highly specialised application-specific integrated circuit (ASIC)? The fight is no longer about finding a single Nvidia clone; it is about building a domestic ecosystem of chips that can reliably support top Chinese AI models from the likes of DeepSeek and Alibaba Group Holding. As competition heats up among major domestic players like Huawei Technologies, Cambricon Technologies and Moore Threads, the South China Morning Post breaks down the differences between these two paths – and which one is poised to dictate the future of Chinese AI. GPU: Who is China’s Nvidia? The GPU was originally engineered to render video game graphics. Nvidia popularised the term in the 1990s with its GeForce 256, marketed as “the world’s first GPU”. However, the GPU’s biggest breakout moment came years later when it was thrust into the AI spotlight: Nvidia recognised that a GPU’s ability to handle many calculations simultaneously made it ideal for powering complex AI networks. What’s more, a GPU is highly versatile and programmable: instead of being hard-wired to do just one specific task, its software can be rewritten over and over again. This inherent flexibility allows AI developers to quickly adapt their software code to the fast-evolving architectures of large language models. In China, start-ups like Moore Threads and Biren Technology are the leading contenders in the general-purpose GPU (GPGPU) space, alongside Shanghai-based Enflame and Iluvatar CoreX. Moore Threads, founded in 2020 by Nvidia’s former China executive Zhang Jianzhong, is leading the domestic GPU charge. Intimately familiar with what made its US rival successful, the company has dedicated itself to general-purpose chips, such as its top-spec MTT S5000 series. Why are tech giants rushing towards ASICs? If the GPU is a versatile, multitalented worker, an ASIC is a specialised machine built to do exactly one job at lightning speed. Instead of being a general-purpose computer chip, an ASIC is custom-designed for a single, narrow task. Because they do not waste energy or space on features they do not need, ASICs are faster at running AI maths. Depending on their mathematical focus, three different streams are gaining traction in the domestic market: Neural processing units (NPUs): chips built specifically to mimic and power brain-like AI networks known as neural networks. Tensor processing units (TPUs): hyper-efficient chips pioneered by Google to crunch large blocks of data in parallel. Parallel processing units (PPUs): a custom-designed variation created and used by Alibaba. Huawei Technologies is betting big on its Ascend NPU series, including the widely deployed 910C and the upcoming 950. Similarly, Cambricon leans heavily into ASICs and domain-specific architectures with its Siyuan 590 and 690 series. By maximising hardware efficiency just for AI, these chips offer domestic companies a targeted way to bypass US hardware restrictions. Meanwhile, Alibaba is doubling down on the PPU path through its semiconductor unit, T-Head. At its annual cloud computing summit last week, Alibaba launched the Zhenwu M890 PPU. The company claims the processor delivers three times the performance of its predecessor, the Zhenwu 810E, making it exceptionally well-suited for complex “agentic AI workloads” – tasks where AI acts like an autonomous digital assistant. Alibaba owns the South China Morning Post. Chinese firms are also racing to build home-grown alternatives to Google’s TPU design, which uses less power and processes data faster than traditional set-ups. Start-ups like Zhonghao Xinying have already pushed their own versions into mass production. This mirrors a global trend where tech giants like Google prefer to build their own custom silicon rather than relying entirely on buying expensive, generic GPUs. However, as AI models become more complex, the boundaries between custom ASICs and flexible GPUs are “becoming increasingly blurry”, said semiconductor industry analyst Zhang Haijun. Currently, because these custom chips are far cheaper to run when deploying finished AI models to the public, the ASIC path is expected to see wider commercial adoption as Chinese tech giants build out their own cloud networks, according to Zhang. Which path will win the race for China’s AI advancement? The choice between ASIC and GPU ultimately comes down to an enterprise’s specific workload and engineering maturity, said Su Lian Jye, chief analyst at Omdia. For enterprises with robust AI engineering capabilities and a clear vision for their AI roadmap, ASICs offer stronger performance and a better cost-to-performance ratio, Su said, adding that those with mixed workloads still tend to rely on Nvidia GPUs and general purpose GPUs that offer greater flexibility and easier porting. For now, market momentum and volumes favour the ASIC heavyweights. According to a May 8 Morgan Stanley report led by analyst Charlie Chan, Huawei was projected to capture a 62 per cent share of the domestic AI accelerator market in 2026, followed by Cambricon at 14 per cent. Among big tech firms building proprietary chips, Baidu and Alibaba are expected to stand out, each capturing about 5 per cent of the domestic market. Performance metrics appeared to justify this shift towards custom silicon. Morgan Stanley data showed that Huawei’s Ascend 950 cards and Cambricon’s Siyuan 690 can outperform Nvidia’s H20 – the most powerful chip Nvidia is allowed to sell to China – by 50 to 150 per cent. This was measured in “tokens per second”, a metric that essentially counts how many words or other data AI can process and spit out in a single second. For China’s highly commercialised market, which focuses on deploying AI apps to millions of users, this high speed reflects that domestic hardware and software are working hand in hand. Crucially, China’s semiconductor race is no longer just about silicon. To truly break the lock-in of Nvidia’s ubiquitous CUDA platform, domestic players are racing to mature their own software stacks – led by Huawei’s CANN and Moore Threads’ MUSA.