StorageReview.com

Equinix Inference Exchange Brings NVIDIA Compute and 200+ Open Models Closer to Enterprise Data

AI  ◇  Enterprise

Equinix has expanded its partnership with NVIDIA and entered a new collaboration with Together AI to launch Equinix Inference Exchange. Designed as a distributed AI inference architecture for enterprise deployments, the platform aims to shift compute workloads closer to core data repositories, end users, and operational applications. Announced alongside Equinix Fabric One at the Equinix Horizon event, the initiative combines NVIDIA enterprise hardware architectures, Together AI’s open-source model platform, and Equinix’s global interconnection infrastructure to address latency, data sovereignty, and networking bottlenecks.

Equinix Inference Exchange graphic showing interconnected Equinix data centers across the globe, bringing inference closer to where data, users, and applications live

Addressing Distributed Inference and Data Gravity

As enterprise AI initiatives transition from prototype validation to high-throughput production, inference placement dictates overall cost, response latency, and compliance posture. Moving model execution closer to data sources reduces the cost of backhauling data across public clouds while providing deterministic latency profiles. Managing distributed infrastructure across disparate environments, however, introduces operational overhead and networking complexity.

Equinix addresses this operational hurdle by leveraging its footprint of more than 280 data centers across 77 metros, 230 cloud on-ramps, and a network of over 10,500 interconnected enterprises. With eight of the top ten AI model providers and nine of the top ten AI clouds already operating within Equinix facilities, the platform functions as a neutral aggregation layer for distributed AI pipelines.

Architecture and Stack Integration

The Equinix Inference Exchange stack is structured into three distinct layers across compute, software, and physical infrastructure:

Equinix manages the base physical and interconnect foundation. This includes power provisioning, advanced liquid cooling capabilities, day-two facility operations, and private interconnectivity to networks, public clouds, and SaaS platforms via Equinix Fabric.

NVIDIA provides the compute and reference validation layer, supplying enterprise AI infrastructure engineered for high token throughput and reduced inference costs per query.

Together AI provides the inference-serving software layer that supports over 200 open-source models. The platform supports both shared multitenant environments for general compute efficiency and dedicated single-tenant topologies for workloads that require isolated capacity.

Together AI logo; the company supplies the inference-serving software layer of Equinix Inference Exchange

By routing traffic over Equinix Fabric, the platform reduces time-to-first-token (TTFT) metrics for regional users while establishing private peering paths between on-premises storage and model endpoints.

Enterprise Deployment Targets and Availability

The platform targets three primary deployment topologies. For metro-edge inference, organizations can run models in local metro hubs, delivering low-latency API responses while maintaining centralized management and private interconnects. For open-source model migration, the architecture provides a low-friction path for shifting from closed, proprietary APIs to open-source foundation models, reducing vendor lock-in while leveraging existing network fabrics. For sovereign AI and compliance mandates, the distributed layout enables enterprises in regulated sectors to keep inference execution and data processing within specific geographic boundaries, helping meet strict data residency and governance requirements.

Equinix Fabric One diagram showing a connectivity request declared through an MCP client or API and composed into a network spanning Chicago and Ashburn metros

Leadership from Equinix, NVIDIA, and Together AI noted that the integration establishes a vendor-neutral, high-throughput inference fabric that balances compute performance with operational agility.

Equinix Inference Exchange is scheduled to be generally available to enterprises in the first quarter of 2027.

Engage with StorageReview

Newsletter | YouTube | Podcast iTunes/Spotify | Instagram | Twitter | TikTok | RSS Feed

Harold Fritts

I have been in the tech industry since IBM created Selectric. My background, though, is writing. So I decided to get out of the pre-sales biz and return to my roots, doing a bit of writing but still being involved in technology.