IBM and Together AI have signed a $240 million multi-year agreement to build a large-scale AI inference cluster on IBM Cloud. Scheduled to go online in early 2027, the U.S.-based cluster will feature about 2,000 Nvidia Blackwell 300 chips on HGX B300 systems. The partnership aims to scale production-grade inference for open-source models like DeepSeek and Kimi, helping enterprises bypass the high costs of closed models amid severe global GPU capacity constraints.
IBM-Together AI $240M infrastructure deal
- ▪The $240 million partnership pairs IBM's enterprise cloud infrastructure with Together AI's GPU-accelerated platform to deliver high-performance model training and inference.
- ▪IBM and Together AI signed a multi-year, $240 million agreement to build a dedicated, large-scale Nvidia-powered AI inference cluster on IBM Cloud.
Nvidia HGX B300 cluster deployment
- ▪The Nvidia HGX B300 cluster on IBM Cloud is expected to come online and become available in the first half of 2027, with initial deployment targeted for the first quarter of 2027.
- ▪The initial IBM Cloud cluster will feature approximately 2,000 Nvidia Blackwell 300 chips and will be physically located in the United States.
- ▪The new IBM Cloud cluster will deploy Nvidia HGX B300 systems, which are motherboards equipped with eight Blackwell Ultra chips and wired with Nvidia's Spectrum-X Ethernet networking.
Together AI inference platform scale
- ▪Together AI processes approximately 400 trillion tokens per month for its customers across its platform.
- ▪Together AI raised $800 million in a Series C funding round in July 2026, valuing the San Francisco-based startup at $8.3 billion.
- ▪Together AI offers five distinct inference services, including Provisioned Throughput, serverless infrastructure, dedicated machines, batch processing, and a service optimized for media generation in containers.
Open-source AI enterprise adoption
- ▪Together AI's platform allows companies to run workloads on open models such as DeepSeek, MiniMax, and Kimi at lower costs than closed systems.
- ▪Together AI will use the IBM Cloud cluster to run production-grade inference workloads on open-source AI models for its enterprise customers.
- ▪Open-source AI models are gaining traction among businesses seeking to reduce AI costs and address cybersecurity concerns associated with closed models from Anthropic, OpenAI, and Meta.
AI infrastructure supply constraints
- ▪Together AI's chief revenue officer Kai Mak stated that the upcoming IBM Cloud cluster is expected to be fully sold out two to three months before it is ready for service.
- ▪AI infrastructure is in short supply due to power, datacenter capacity, and supply chain constraints, forcing service providers to rent compute from competing cloud providers.
Story comments
Loading comments…