StorageReview.com

IBM and Together AI Put $240M Into a Dedicated HGX B300 Inference Cluster on IBM Cloud

AI  ◇  Enterprise

IBM has announced a multi-year, $240 million agreement with Together AI to deploy a dedicated NVIDIA HGX B300-based AI inference cluster on IBM Cloud. Together AI plans to use the environment to deliver inference services for open-source models, expanding its AI Native Cloud platform for enterprise customers.

The deployment is positioned as IBM Cloud’s first dedicated large-scale inference cluster built around NVIDIA HGX B300 systems, following the Grace Blackwell capacity IBM added through CoreWeave last year. It will also use NVIDIA Spectrum-X Ethernet networking, creating an AI factory architecture intended to support high-throughput, low-latency inference workloads. IBM says the deployment is built to deliver 30x more AI factory output compared to prior generations. However, workload-level performance will depend on model architecture, precision, batch size, and serving configuration.

NVIDIA Blackwell Ultra rack render of the type IBM Cloud will deploy for the Together AI HGX B300 inference cluster

Together AI offers infrastructure and software services spanning inference, model training, fine-tuning, and agentic AI workflows. The company reports that its inference platform now serves 400 trillion tokens monthly. The new IBM Cloud deployment is intended to add GPU capacity for production inference while improving performance and token economics for organizations deploying open-weight and open-source models at scale.

“Together AI is proud to lead the way in bringing production inference to market with NVIDIA’s latest AI infrastructure on IBM Cloud,” said Vipul Ved Prakash, CEO at Together AI. “Working alongside IBM with NVIDIA accelerates our mission to make advanced AI broadly accessible through open source and to empower builders with the infrastructure and platform capabilities they need to build the future.”

The collaboration brings together IBM Cloud’s enterprise infrastructure, NVIDIA’s HGX B300 compute systems and Spectrum-X Ethernet fabric, and Together AI’s inference platform. The resulting stack is designed to support organizations that need to deploy and operate AI services across cloud and hybrid environments, particularly for workloads where throughput, response time, and infrastructure utilization directly affect operating costs.

“Enterprises are in a race to adopt agentic AI at scale to drive real business outcomes,” said Alan Peacock, general manager of IBM Cloud. “IBM and NVIDIA are delivering scalable, economical, enterprise-grade AI infrastructure that can help Together AI accelerate innovation for the next generation of AI infrastructure.”

NVIDIA HGX B300 system board of the type IBM Cloud will deploy for the Together AI inference cluster

Together AI selected IBM and NVIDIA based on their GPU infrastructure roadmaps and their ability to provide capacity at the pace needed for AI service expansion. The company raised an $800 million Series C at an $8.3 billion valuation in July to expand its AI Native Cloud platform.

IBM characterized the agreement as part of its broader work with NVIDIA across AI infrastructure and software. The companies have also referenced joint work around GPU-native data analytics, unstructured data extraction, hybrid infrastructure, and consulting services. For IBM Cloud customers, the Together AI deployment adds another route to production-grade inference infrastructure built on NVIDIA’s latest HGX platform and high-performance Ethernet networking.

Engage with StorageReview

Newsletter | YouTube | Podcast iTunes/Spotify | Instagram | Twitter | TikTok | RSS Feed

Harold Fritts

I have been in the tech industry since IBM created Selectric. My background, though, is writing. So I decided to get out of the pre-sales biz and return to my roots, doing a bit of writing but still being involved in technology.