4 SEP 2026 — Equinix has announced Inference Exchange, built with Nvidia and Together AI, to run more than 200 open models on Nvidia hardware inside its data centres and connect them to customers through Equinix Fabric. It launches in the first quarter of 2027. The 2 September announcement is an architecture and three logos, roughly four months before anyone can buy it.

What is being offered

The three pieces are Nvidia's validated Enterprise Reference Architectures, Together AI's inference platform, and Equinix's own estate of data centres and its interconnection fabric. The pitch is that enterprises move inference to where their data already sits rather than sending the data to a cloud region.

It was unveiled at Equinix Horizon, the company's first customer and partner event. No pricing, no launch markets and no capacity figures have been published.

Q1 2027Stated availability
200+Open models via Together AI's platform
3Companies in the stack
NonePrices, launch markets or capacity disclosed

What Equinix is now selling

Equinix has historically sold floor space, power and cross-connects, and let customers decide what to put in the rack. Assembling somebody else's silicon with somebody else's serving software and selling the result as a service is a different business with different margins and a different competitor set.

That competitor set is the awkward part. A managed inference service sold from a data centre competes with the hyperscalers, and the hyperscalers are Equinix's largest tenants. Equinix connects to them through the same fabric this product runs on.

The reconciliation is presumably that this targets workloads which will not move to a public cloud for data-residency or latency reasons. That is a real segment, and a smaller one. Whether it holds as the product grows is the strategic question the announcement does not address.

Equinix is not building models here, nor buying accelerators to resell as raw capacity. It is packaging two partners' products and charging for the location. That keeps its own capital commitment modest and leaves the differentiation in somebody else's hands.

Why inference sits differently from training

Training is a batch job that runs wherever the cheapest large block of compute happens to be. Distance costs nothing, because nobody is waiting on a response.

Inference is a request-response transaction inside an application, so it inherits every latency and locality constraint that application has. It also touches production data on every call, which is what makes residency rules bite: an inference call carrying customer records to another jurisdiction is a transfer, whatever the compute economics say.

That is an argument for distributing inference to where the data lives, and it is stronger in this region than in most. We wrote this week about grids that cannot absorb the load being requested of them. Inference in existing colocation facilities is one of the few paths that does not require a new interconnection queue position.

Four months is a long time in this market

The gap between announcement and availability is the detail worth holding on to. Between now and the first quarter of 2027 the model landscape will turn over at least once, Nvidia will ship a new generation, and the price of hosted inference will fall again as it has every quarter.

Announcing early is rational for Equinix, which needs enterprises to keep it in a 2027 procurement cycle instead of committing to a hyperscaler this year. The enterprise is the one being asked to hold a slot open for a product with no published price.

The 200-model figure will also age. Open model catalogues grow and churn, so the number describes Together AI's platform today, not a commitment about what will be served in fifteen months. A useful version would name which models are guaranteed at launch. This one does not.

Nvidia is on both sides of this

Nvidia supplies the reference architectures here, and this week it agreed to buy Hugging Face, the repository most of those 200 open models are distributed from. It also licenses NVLink to rivals so their chips run on its fabric.

Nvidia therefore occupies three positions at once. It makes the accelerators, it will own the place the models are published, and it writes the architecture enterprises are told to copy. None of those moves is objectionable alone. Together they leave very few moments at which an alternative accelerator would be reached for without somebody deliberately reaching.

For a buyer, the practical consequence is narrow. Ask what a validated reference architecture validates against, and whether the validation survives substituting a different accelerator.

What would make this concrete

Three disclosures would turn this from an architecture into an offer. A price per token or per GPU-hour, which a customer could hold against a hyperscaler quote. A list of launch metros, since the whole premise rests on being near the data and that is a question about cities. And committed capacity, because 200 models available in principle is not capacity reserved for a customer's peak.

The regional test is narrower still. Equinix operates in Singapore and across the region, and whether Inference Exchange appears in those metros at launch or later determines whether this is relevant to a Southeast Asian enterprise in 2027 or in 2028.

Until then, three companies with the assets to build this have described what they intend to build, at an event one of them hosted, and left out everything a buyer would need to judge it.