Equinix has announced the Equinix Inference Exchange, a new platform designed to facilitate distributed AI inference for enterprise deployments. Developed in collaboration with NVIDIA and Together AI, the service aims to move compute workloads closer to core data repositories and end users. The platform is scheduled for general availability in the first quarter of 2027.
The architecture integrates NVIDIA enterprise hardware with Together AI’s software layer, which supports over 200 open-source models. By utilizing Equinix’s global interconnection infrastructure, the exchange seeks to address networking bottlenecks and latency issues. It provides a neutral aggregation layer for distributed AI pipelines, allowing enterprises to manage inference placement across various geographic boundaries to meet data residency and compliance requirements.
The platform supports both shared multitenant environments and dedicated single-tenant topologies. By routing traffic over Equinix Fabric, the system is intended to reduce time-to-first-token metrics for regional users while maintaining private peering paths between on-premises storage and model endpoints. This approach is designed to help organizations transition from proprietary APIs to open-source foundation models.
Separately, NVIDIA has released the Personal AI Router, or PAIR, a virtual inference router for local networks. The tool is designed to distribute independent inference requests across multiple systems to alleviate bottlenecks in multi-agent workflows. It functions as a proxy for existing interfaces like Ollama and LM Studio, meaning users do not need to modify their agent harnesses to utilize the additional compute capacity.
NVIDIA PAIR supports a range of hardware, including GeForce RTX 20 Series GPUs and newer, RTX PRO workstation GPUs, DGX Spark, and Apple M4+ silicon. The software handles secure pairing through mDNS discovery and MTLS encryption, while performing live scheduling based on node readiness and GPU utilization. In demonstrations, a three-device cluster using PAIR completed multi-agent workloads significantly faster than a single system.
