Alibaba Cloud has made available the Lingjun Zhenwu M890 supernode for hourly rental exclusively in North China's Inner Mongolia, delivering large-scale AI model hosting via a 64-GPU tightly interconnected cluster. While slower than global leaders, it provides distinctive advantages in latency, domestic hardware, and cloud integration under Chinese regulatory constraints.

  • Supernode composed of 64 domestic GPUs with 144GB HBM memory and custom interconnect
  • Supports AI models exceeding 2 trillion parameters, enabling advanced commercial AI workloads
  • Access limited to Inner Mongolia region, emphasizing regulatory and geographic restrictions

Infrastructure signal

Alibaba’s new supernode offering represents a milestone in Chinese cloud infrastructure by tightly integrating 64 homegrown GPU cards into a single scalable instance with very high inter-chip bandwidth. The platform’s design enables direct treatment of the cluster as one unified chip, significantly reducing data transfer overhead common in multi-GPU deployments. This technical architecture supports hosting AI models of unprecedented size domestically, reinforcing China’s goal of self-reliance in advanced computing hardware.

The use of T-Head’s Zhenwu M890 processor with 144GB of HBM memory and Alibaba’s proprietary ICNSwitch 1.0 fabric showcases the evolution of China’s semiconductor and cloud ecosystem. Although not matching global top-tier multi-GPU setups from Nvidia or AMD in raw speed, the supernode benefits from lower regional latency and tight cloud integration. The upgrade path through 2028 promises even greater performance and memory enhancements, aligning with growing computational demands for trillion-parameter AI models.

Developer impact

For AI developers and enterprise teams within China, especially in Inner Mongolia, this hourly rental supernode provides a flexible and scalable compute resource without the need for capital expenditure on physical infrastructure. This model enables rapid iteration and deployment of large language models that were previously constrained by hardware accessibility and geographic availability. The ability to run up to 10 trillion parameter mixture-of-experts architectures on a single instance opens new frontiers in generative AI capabilities.

However, limitations in geographic availability restrict external developer access, which constrains innovation and partnership potential beyond the domestic market. The lack of public pricing or comprehensive benchmarking data forces teams to evaluate cost-effectiveness cautiously. Nonetheless, Alibaba’s offering simplifies the developer workflow by delivering a turnkey large-scale compute environment with integrated memory and networking, allowing teams to focus on model optimization and deployment rather than infrastructure assembly.

What teams should watch

Teams involved in AI model development, cloud platform strategy, and infrastructure deployment should monitor Alibaba’s expansion roadmap, especially as it evolves toward the 2027 Zhenwu V900 and 2028 J900 iterations with substantial performance gains. Staying abreast of these advancements is critical for projects targeting truly large-scale, low-latency AI workloads within China’s regulated cloud ecosystem. The regional lock on Inner Mongolia access signals ongoing geopolitical and regulatory dynamics affecting hardware supply chain and cloud sovereignty.

Additionally, product and platform teams should evaluate implications on API design, data locality, and observability tools to optimize for Alibaba’s supernode environment. Enterprises planning multi-region or international cloud strategies might anticipate gradual availability expansion or comparable domestic alternatives. Observing how Alibaba integrates this capability with broader cloud services, including databases and AI model hosting, will be essential for shaping future deployment architectures.

Source assisted: This briefing began from a discovered source item from TechRadar. Open the original source.
How SignalDesk reports: feeds and outside sources are used for discovery. Public briefings are edited to add context, buyer relevance and attribution before they are published. Read the standards

Related briefings