When Together AI closed its Series C, the number that made the headlines was 800 million dollars. Six weeks later came the number that actually explains it: 240 million, the value of the multi-year agreement the company signed with IBM on August 11 to stand up a dedicated cluster of roughly 2,000 Nvidia Blackwell chips on IBM Cloud. Not for training. For inference.

That sequence — raise enormous growth capital, then convert it immediately into contracted serving capacity — is the defining pattern of AI venture investment in 2026, and Together AI is now one of its cleanest examples.

A note on timing: the Series C was announced on July 1, 2026, not in August, as several funding roundups have reported. The August news is the IBM cluster.

The round

Together AI raised 800 million dollars at an 8.3 billion dollar post-money valuation in a round led by Aramco Ventures, with participation from Vista Equity Partners, General Catalyst, Emergence Capital, Nvidia, Salesforce Ventures, March Capital, Pegatron and S Ventures, the investment arm of SentinelOne. The valuation is more than double the 3.3 billion the San Francisco company carried roughly 16 months earlier.

The equity is only part of the package. Alongside the round, investors committed more than 500 megawatts of compute capacity to Together AI, to be capitalized separately from the raise — a structure that has become standard among the so-called neoclouds, where the balance-sheet cost of GPUs, power and land is simply too large to fund out of a venture round. The company says it expects its infrastructure footprint to grow roughly 50-fold over the next five years.

The business underneath is real, if young. Together AI reported more than 1.15 billion dollars in annual bookings in its most recent quarter and says it now serves about 400 trillion tokens a month, largely running open-weight models — DeepSeek, Nemotron, MiniMax, Kimi — for customers who want frontier-class output without frontier-class per-token pricing.

“Intelligence is becoming a foundational resource for the modern economy, every bit as essential as electricity, bandwidth or capital,” said Vipul Ved Prakash, co-founder and chief executive of Together AI, in announcing the round.

The IBM deal

The August 11 agreement puts a concrete price tag on that thesis. IBM will deploy Nvidia HGX B300 systems — Blackwell-generation accelerators linked with Nvidia Spectrum-X Ethernet networking — on IBM Cloud, with Together AI as the anchor tenant. IBM describes it as the first dedicated, large-scale inference cluster on IBM Cloud built on HGX B300. General availability is expected in the first quarter of 2027.

“Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale,” Prakash said. “This cluster lets us bring production-grade inference to more companies, faster, and it's a big step in our push to make open-source AI the obvious choice for enterprises.”

Alan Peacock, general manager of IBM Cloud, framed it from the incumbent's side: “Enterprises are in a race to adopt agentic AI at scale to drive real business outcomes.” For IBM, the deal is a quiet entry into the neocloud supply chain — renting enterprise-grade capacity to a startup that competes with hyperscalers IBM also partners with.

Why this matters

The training-versus-inference split is no longer a technical distinction. It is a capital allocation decision, and the money has picked a side.

Consider the 90-day window around Together AI's round. Fireworks AI raised 1.505 billion dollars in a Series D on July 16 at a 17.5 billion dollar post-money valuation, led by Atreides Management, Index Ventures and TCV, on the back of a claimed one-billion-dollar annualized revenue run rate and more than 40 trillion tokens served daily. Baseten raised roughly 1.5 billion in the same stretch. Three inference companies, north of 3.8 billion dollars, a single quarter.

None of that capital is going toward training a frontier model. It is going toward chips, power, networking and the deeply unglamorous work of driving down cost per token. That is a bet that the durable margin in AI sits in serving rather than in model weights — and, implicitly, a bet that open-weight models are now good enough that the serving layer, not the model, becomes the product.

It also reflects a market that has grown extraordinarily top-heavy. AI startups took in a reported 407 billion dollars in the first half of 2026, more than all of 2025 combined, with OpenAI and Anthropic alone accounting for roughly 217 billion of that total. For everyone else, the viable path is not to out-train the labs. It is to sell them, and their customers, capacity.

The risk is just as legible. Inference is a commodity business with a hard physical cost floor, and the same 500-megawatt commitments that make growth possible make a demand slowdown ruinously expensive. Nvidia's dual role — investor in both Together AI and Fireworks, and supplier of the hardware both companies buy — adds a circularity that has drawn steady scrutiny across the sector.

What to watch

Three things. Whether Together AI's 1.15 billion in bookings converts into recognized revenue at a gross margin that justifies an 8.3 billion valuation. Whether the IBM cluster actually lands in the first quarter of 2027 or slips — power and supply timelines have been this industry's most reliable source of disappointment. And whether enterprise appetite for open-weight inference holds, or whether the frontier labs simply cut prices far enough to make the arbitrage disappear.

The 800 million was the easy part. The 500 megawatts is the test.

“Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale.”
— Vipul Ved Prakash, Co-founder and CEO, Together AI
$800M
Series C
$8.3B
Post-money valuation
$240M
IBM Cloud inference deal
400T
Tokens served per month