Microsoft is adopting AMD’s Helios rack design, incorporating advanced GPU, CPU, and data processing units to boost Azure's AI and HPC workloads while improving efficiency and lowering infrastructure costs.

  • Helios racks combine AMD Instinct GPUs, Venice CPUs, and Pensando DPUs for optimized AI workloads
  • New Azure VMs focus on AI inference, dataset preparation, and enhanced HPC simulations
  • Liquid cooling and UALoE networking improve efficiency and observability in Azure data centers

Infrastructure signal

Microsoft’s incorporation of AMD’s Helios rack design marks a significant shift towards specialized, modular hardware optimized for AI workloads in Azure’s data centers. Each rack features 72 Instinct MI455 GPUs with a new CDNA 5 architecture, supported by Venice-series Epyc CPUs produced on advanced 2nm TSMC technology, and Pensando data processing units designed to handle encryption and storage tasks more efficiently than CPUs alone. This triad of components, organized in wide, liquid-cooled trays, facilitates enhanced thermal management and higher hardware density than traditional rack designs.

The adoption of the Helios racks also leverages an open network protocol called UALoE to enable high-speed, flexible inter-component communication. AMD’s partnership with Microsoft underscores a focus on full-stack optimization, aiming to streamline data center operations by offloading infrastructure management functions to DPUs, thus freeing up CPU resources for customer-facing workloads. These design choices are expected to improve cloud cost efficiency by decreasing CPU load and increasing overall system throughput.

Developer impact

From a developer workflow perspective, the launch of new Azure virtual machine instances powered by these Helios racks enables specialized compute options tailored to specific AI and HPC needs. The ND MI455X v7 series targets AI inference workloads, improving the execution of agent-based and search applications. In parallel, the HDv2 series optimizes CPU-heavy AI data preparation processes, offering up to 500 Epyc Vulcan cores, along with substantial memory and flash storage to accommodate large datasets.

Additionally, the HXv2 instance family enhances performance for electronic design automation (EDA) and scientific simulations, with per-core clock speeds exceeding 5 GHz and a 50% increase in cache over prior generation hardware. Developers focusing on high-performance computing workloads will benefit from improved core density, enhanced memory configurations, and high-bandwidth InfiniBand connectivity, facilitating large-scale MPI execution and more efficient parallelization of compute-intensive tasks.

What teams should watch

Teams managing cloud infrastructure, AI platform deployment, and observability should monitor the rollout of Helios racks closely due to their impact on cost optimization and reliability. The integration of Pensando DPUs shifts infrastructure tasks away from CPUs, leading to more predictable system performance and potentially lower operational costs. Liquid cooling adoption requires adjustments to data center environment monitoring and maintenance protocols to ensure optimal thermal conditions.

Developers and platform engineers should also be aware of the new open networking standard UALoE implemented in Helios, as it may require updates to network management and diagnostics tooling. Finally, those overseeing database services and API endpoints should assess how the increased compute capabilities and specialized hardware can be leveraged to support advanced AI features, accelerate inference times, and improve overall application responsiveness within Azure’s evolving cloud ecosystem.

Source assisted: This briefing began from a discovered source item from SiliconANGLE. Open the original source.
How SignalDesk reports: feeds and outside sources are used for discovery. Public briefings are edited to add context, buyer relevance and attribution before they are published. Read the standards

Related briefings