IBM announced a $240 million infrastructure agreement with AI startup Together AI to provision a cluster of advanced HGX B300 systems optimized for large-scale inference workloads. This marks a significant investment in AI-optimized cloud infrastructure ahead of deployment in early 2027.
- IBM to deploy HGX B300 clusters with Nvidia Blackwell Ultra GPUs in IBM Cloud by H1 2027
- Together AI platform processes approximately 400 trillion tokens monthly for AI inference
- New deal enables scalable, token-priced inference services using open-source AI models
Infrastructure signal
IBM is set to roll out a substantial cluster comprising HGX B300 systems designed expressly for AI workloads, embedding eight Nvidia Blackwell Ultra GPUs each. These GPUs provide up to 15 petaflops of computational power, enabling high-throughput processing critical for deep learning tasks. The architecture incorporates advanced networking components like ConnectX-8 SuperNICs and BlueField-3 DPUs to optimize data flow and latency across the cluster.
This hardware suite, integrated into IBM Cloud, signals a major increase in specialized AI infrastructure offerings. It supports workloads requiring intensive memory and storage configurations, including multiple terabytes of NVMe flash and dual 56-core CPUs, highlighting IBM's focus on balancing GPU-centric AI acceleration with robust CPU and memory resources. The large-scale deployment will underpin Together AI's public cloud inference services, marking a shift toward customized, high-performance AI cloud platforms.
Developer impact
For developers leveraging Together AI's cloud platform, the new infrastructure deal means enhanced availability and efficiency in running AI inference workloads, particularly those based on open-source models. The platform supports multiple inference modalities, including container-based media generation and serverless execution, with token-based pricing models designed to optimize cost management during large-scale usage.
This upgraded infrastructure will enable faster iteration cycles by reducing latency and improving throughput, facilitating model fine-tuning and deployment at scale. Developers can expect more streamlined workflows around managing inference services, with improved reliability from enterprise-grade hardware and network optimizations. The diverse service options—from dedicated machines to batch processing environments—offer flexibility in balancing performance needs against budget constraints.
What teams should watch
Cloud infrastructure teams should closely monitor the rollout of the HGX B300 cluster on IBM Cloud to understand how its integration of high-end Nvidia GPUs and advanced networking components alters deployment practices and operational monitoring. Observability tools may need updates to effectively capture GPU-centric metrics and data movement patterns specific to this architecture, impacting performance tuning and capacity planning.
Platform and API teams will want to track changes in Together AI’s service offerings, notably the introduction of token-priced inference throughput and container-based services optimized for media generation. These services could shift demand patterns and resource consumption, influencing API rate limits, service reliability SLAs, and cost management strategies. Additionally, database and storage strategies may evolve to support large intermediate datasets generated by these AI workloads, requiring closer alignment between infrastructure and application layers.