IBM, Together AI ink $240M deal for NVIDIA B300 inference cluster
Multi-year agreement targets enterprise open-source AI deployment with first large-scale HGX B300 system on IBM Cloud, expected Q1 2027.
IBM and Together AI have signed a multi-year $240 million agreement to build what the companies describe as the first large-scale inference cluster using NVIDIA HGX B300 systems on IBM Cloud, with expected availability in the first quarter of 2027.
The deployment will leverage NVIDIA's HGX B300 systems paired with Spectrum-X Ethernet networking. According to NVIDIA, this infrastructure is designed to deliver 30 times more AI factory output compared to previous generations. Together AI will use the cluster to provide open-source model inference services to enterprise customers.
Why it matters
As enterprises race to deploy AI at production scale, infrastructure economics become critical. This deal signals a strategic bet that open-source models can compete with proprietary alternatives on performance while offering better token economics. Together AI's reported 400 trillion tokens served monthly demonstrates real enterprise demand for alternatives to closed-model providers, and the $240 million commitment suggests IBM sees infrastructure-as-a-service for AI inference as a growth market worth significant capital investment.
Together AI's enterprise push
Together AI recently closed an $800 million Series C funding round at an $8.3 billion valuation. The company, founded in 2022, has built its business on the premise that open-source models are essential for AI's future. Its platform spans inference, training, fine-tuning, and agentic workflows.
Vipul Ved Prakash, CEO of Together AI, said enterprises want frontier-model performance without closed-model pricing, which requires fast and reliable infrastructure at scale. The company selected IBM and NVIDIA based on their product roadmaps and ability to deliver GPU capacity at the pace required for rapid AI scaling and competitive token costs.
Infrastructure collaboration expands
The agreement represents the latest development in an ongoing collaboration between IBM and NVIDIA. The companies have been working together on GPU-native data analytics, unstructured data extraction, on-premises and cloud infrastructure, and consulting services designed to help organizations operationalize AI at scale.
Alan Peacock, General Manager of IBM Cloud, framed the partnership as addressing enterprise demand for agentic AI at scale. Dion Harris, Senior Director of HPC and AI Infrastructure Solutions at NVIDIA, characterized AI factories as becoming essential enterprise infrastructure comparable to electricity and telecommunications.
The hybrid environment on IBM Cloud powered by NVIDIA GPUs and Together AI's inference platform is intended to give organizations a foundation for building, deploying, and scaling AI systems. IBM Cloud's enterprise-grade capabilities combined with Together AI's open-source focus aim to make advanced AI infrastructure more accessible to developers and enterprises globally.
These details were first reported by IBM in a press release issued August 11, 2026.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
