IBM struck a multi-year pact worth $240 million with Together AI. The deal sets up a large cluster of advanced Nvidia systems on IBM Cloud dedicated to running open models at scale. Availability won’t come until the first quarter of 2027. Yet the agreement already signals how the battle for enterprise AI workloads has shifted from raw training power to efficient, cost-effective serving.
Together AI will operate the infrastructure to deliver inference services. The setup relies on Nvidia HGX B300 systems paired with Spectrum-X Ethernet networking. Nvidia claims the platform can produce 30 times more output from AI factories than earlier generations. IBM’s announcement positions the cluster as the first of its kind for large-scale inference on its cloud.
The initial build includes roughly 2,000 Nvidia Blackwell 300 chips. It will sit in a U.S. data center. Executives at Together AI expect strong demand. “We think this will be sold out at least two to three months ahead of time,” Kai Mak, the company’s chief revenue officer, told Reuters. “We’ll have full offtake well before it’s ready for service.”
Such confidence stems from Together AI’s rapid growth. The San Francisco startup, founded in 2022, now serves more than 400 trillion tokens each month. Its platform supports inference, training, fine-tuning and agentic workflows built on open-source models. Customers gain access to options such as DeepSeek, MiniMax and Kimi. These deliver capable performance at lower costs than proprietary alternatives.
Last month the company closed an $800 million Series C round. That financing valued it at $8.3 billion. Momentum like this explains why Together AI sought dedicated capacity. “Together AI selected IBM with Nvidia because of their innovative product roadmaps and their ability to deliver GPU capacity at the pace required for rapid AI scaling and lowest token cost,” the joint statement noted.
Vipul Ved Prakash, Together AI’s CEO, put the appeal in simple terms. “Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale.” He added that the cluster would let the firm “bring production-grade inference to more companies, faster.” The remark came directly from IBM’s press materials.
On IBM’s side the partnership fits a deliberate strategy. The company largely sat out the frenzy to build massive training clusters. Instead it has focused on hybrid cloud strengths, enterprise trust and integration with existing systems. Alan Peacock, general manager of IBM Cloud, highlighted the shift. “Enterprises are in a race to adopt agentic AI at scale to drive real business outcomes,” he said. “IBM and NVIDIA are delivering scalable, economical, enterprise-grade AI infrastructure that can help Together AI accelerate innovation for the next generation of AI infrastructure.”
Dion Harris, senior director for HPC and AI infrastructure solutions at Nvidia, framed the bigger picture. “AI factories are becoming essential enterprise infrastructure—like electricity and telecommunications—turning compute and data into intelligence.” The combination of HGX B300 systems and Spectrum-X networking on IBM Cloud, he argued, would give organizations the performance, efficiency and scale needed for real-time services.
This transaction marks more than a one-off hardware sale. It forms part of a wider IBM-Nvidia alliance that now spans GPU-native data analytics, unstructured data extraction, on-premises deployments, cloud resources and consulting. Recent progress in those areas aims to help companies move AI from experiments into daily operations. IBM has also committed $5 billion through its Project Lightwell initiative with Red Hat to advance secure open-source AI tooling.
For Together AI the move carries risks and rewards. The firm already rents capacity from multiple clouds and specialized providers. It maintains an OpenAI-compatible API and stays flexible on underlying silicon. Yet committing to a large dedicated cluster on IBM Cloud ties it more closely to one partner’s roadmap. Supply constraints on advanced GPUs have eased in some spots but remain tight for the latest generations. IBM’s ability to allocate capacity when Together AI needed it proved decisive.
Analysts see inference as the next major arena. Training grabs headlines and venture dollars. Serving models generates the steady revenue. Token economics decide winners. Lower cost per token, predictable latency and strong security matter most to corporate buyers. Open weights models can win market share only when their total cost of ownership beats closed APIs. Reserved capacity deals like this one help lock in those economics.
Enterprises have grown cautious. Many want frontier-level results without vendor lock-in or surprise bills. They also demand compliance, data sovereignty and integration with legacy systems. IBM’s long track record in regulated industries gives Together AI instant credibility in boardrooms where pure startups might struggle. At the same time Together AI’s focus on open models and developer tools brings fresh energy to IBM’s offerings.
The cluster won’t power up for another 18 months. Construction, integration and testing take time. Yet forward bookings already point to robust interest. If the systems deliver on promised efficiency gains, the arrangement could expand. Both companies left room for that outcome in their statements.
Broader market forces support the thesis. Demand for AI compute continues to outstrip supply in many segments. Hyperscalers and specialized clouds scramble to add capacity. Nvidia’s latest chips command premiums. In this setting a $240 million commitment represents a sizable but targeted investment. It buys Together AI guaranteed access. It gives IBM a high-profile showcase for its cloud in the inference market.
Questions remain about long-term economics. Power consumption for these systems runs high. The HGX B300 nodes draw significant electricity. Networking at scale adds complexity. Real-world token costs will depend on utilization rates, software optimizations and future hardware improvements. Together AI’s own research into efficiency and scalability will play a role in keeping prices competitive.
Still, the partnership underscores a maturing industry. Startups that once chased raw scale now hunt for reliable, enterprise-ready infrastructure. Incumbents that once focused on proprietary stacks now court open-source communities. The result may be faster adoption across sectors that previously hesitated. Financial services, healthcare and government organizations in particular prize the combination of performance, openness and trusted operations.
IBM and Together AI have placed their wager. The cluster, when it arrives, will test whether open-source inference at enterprise scale can deliver the economics and reliability that corporations require. Early signs suggest strong appetite. The real proof will come in production workloads running in 2027 and beyond.