Chinese AI companies are facing significant compute shortages for advanced inference tasks due to restricted access to Nvidia chips. This bottleneck is pushing firms to optimize software and devise hybrid hardware strategies to meet soaring demand.
- Nvidia H200 chip rental costs surge amid rising AI demand
- Domestic chips handle low-quality inference; high-quality uses Nvidia hardware
- Hybrid chip setups enable cost and efficiency improvements
What happened
Chinese AI companies are grappling with a shortage of high-end Nvidia chips crucial for complex AI inference tasks such as coding, forcing them to optimize software and adapt workloads to available domestic hardware. While training AI models demands premium chips, inference can partially run on less advanced processors, but this only supports lower-tier tasks.
The rapid increase in AI token usage — basic units processed in responses — has intensified the crunch, with China’s daily token calls exceeding 140 trillion by March 2026. This surge has made Nvidia chips both scarce and expensive, with rental prices for the H200 model surpassing 100,000 yuan ($14,857), nearly doubling in under a year.
Why it matters
The inability of Chinese domestic chips to meet performance demands for high-quality AI inference limits commercial opportunities and could slow AI progress in the region. High-tier AI applications, which generate significant revenue, still hinge on Nvidia’s hardware, representing a critical dependency amid export restrictions and geopolitical tensions.
If current trends continue, the required scale of high-end inference chips in China could become unrealistic, with estimates suggesting a need for hundreds of thousands to potentially hundreds of millions of units by 2030. This hardware gap underlines the strategic urgency for China to innovate around chip scarcity or risk falling behind in AI deployment.
What to watch next
Chinese firms are adopting ‘heterogeneous’ chip architectures combining Nvidia GPUs with domestic processors like Huawei’s Ascend 910B to optimize task efficiency and control costs. Startups such as Approaching.AI and Moonshot are pioneering software that enables seamless data exchange and distributed processing across mixed chip environments, which may serve as a blueprint for scaling AI deployment under supply constraints.
Industry observers will be closely monitoring if these hybrid solutions can sustainably reduce costs and improve AI performance to satisfy growing demand, as well as how chip supply dynamics evolve with changing geopolitics. Advances in software-driven inference optimization and domestic chip development remain crucial indicators for China’s AI competitiveness trajectory going forward.