Together AI's IBM compute capacity expected to sell out pre-launch
The open-source AI platform says enterprise demand for Nvidia-based infrastructure is outpacing supply as companies seek cost control and customization.

Together AI anticipates its newly announced compute capacity will be fully reserved before it even goes live, according to the company's Chief Revenue Officer Kai Mak. The platform expects all Nvidia-based infrastructure from its $240 million deal with IBM to be "pre-sold well before it comes online in January," Mak told Fierce.
The projection underscores accelerating enterprise adoption of open-source AI models as organizations look beyond proprietary alternatives from major cloud providers.
Why it matters
The pre-sale trajectory signals a fundamental shift in how enterprises are approaching AI infrastructure decisions. Rather than defaulting to closed models from hyperscalers, companies are actively seeking alternatives that offer cost advantages, customization capabilities, and data sovereignty—even when it means committing capacity months in advance.
Cost drives initial interest, control sustains it
While lower token costs initially attract enterprises to open models, the value proposition extends beyond economics. According to Mak, organizations are prioritizing the ability to customize and optimize AI systems for specific performance and latency requirements. Data control and avoiding vendor lock-in have emerged as equally important considerations.
The company is seeing strong uptake of models including GLM, Kimi, MiniMax, Nvidia's Nemotron family, DeepSeek, Google's Gemma, and Alibaba's Qwen family. Together AI also serves a growing segment of customer-owned models that start with open foundations but incorporate extensive proprietary training data.
Matching models to workloads
One of the most significant opportunities for reducing inference costs involves aligning model capability with task complexity. Mak noted that most enterprise workloads don't require the largest frontier models, using the example that "you don't need a god to compose your emails."
Routing routine tasks to smaller open models while reserving more powerful systems for complex problems can reduce costs by a factor of 10 without degrading user experience. This approach requires education, as users often default to higher-capability models and consume more tokens than necessary.
Beyond model selection, Mak pointed to advances in harness engineering, Nvidia chip improvements, kernel optimization, quantization, attention mechanisms, and speculative decoding as additional levers for driving down per-token costs.
Hybrid strategies emerging
Despite the momentum behind open-source AI, Mak doesn't expect proprietary models to disappear. Instead, he anticipates enterprises will adopt hybrid strategies that deploy open, custom, and closed models based on economic and technical fit for specific use cases.
The IBM partnership was chosen specifically to combine sought-after Nvidia compute with the "reliability, security and operational maturity enterprises expect from IBM Cloud," according to Mak.
These details were first reported by Fierce.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call