d-Matrix Partners With NVIDIA to Integrate AI Inference Chips Into NVLink Fusion Infrastructure

Photo courtesy of NVIDIA

d-Matrix Partners With NVIDIA to Integrate AI Inference Chips Into NVLink Fusion Infrastructure

AI chipmaker d-Matrix is expanding its collaboration with NVIDIA through a multi-year product roadmap that will integrate its next-generation inference processors into NVIDIA’s rack-scale AI infrastructure, targeting growing demand for faster and more energy-efficient AI inference.

The companies said d-Matrix will connect its next-generation Raptor XPUs to NVIDIA’s AI infrastructure platform using NVLink Fusion, giving the startup access to the networking, rack architecture, power, cooling and supply-chain infrastructure already used across NVIDIA’s AI ecosystem.

The first phase will integrate Raptor XPUs into the NVIDIA MGX rack architecture, alongside technologies including NVIDIA Vera CPUs, BlueField-4 DPUs, ConnectX-9 SuperNICs and Spectrum-X Ethernet networking. Initial availability of Raptor XPUs integrated into NVIDIA MGX racks is expected in the fourth quarter of 2027, according to d-Matrix.

NVIDIA Opens Infrastructure to Specialized AI Chips

NVLink Fusion allows companies developing custom CPUs and AI accelerators to connect their processors with NVIDIA’s rack-scale infrastructure rather than building an entirely separate system around their chips.

For d-Matrix, that provides a route for its specialized inference processors to operate within the broader NVIDIA AI platform.

NVIDIA said the approach allows data centers to use a common rack architecture while supporting different types of processors, including GPUs, CPUs and XPUs. That can reduce the additional engineering, validation and infrastructure work required to deploy custom AI silicon at scale.

“NVLink Fusion gives partners like d-Matrix a path to integrate seamlessly with NVIDIA compute platforms,” NVIDIA CEO Jensen Huang said in the announcement.

d-Matrix is also working with Astera Labs on connectivity solutions designed to support high-throughput data movement within the system.

Targeting the Growing AI Inference Market

Unlike AI training, which is used to build and refine models, inference is the process of running trained models to generate responses and perform tasks for users.

Demand for inference computing is increasing as AI applications move into large-scale commercial deployment, particularly with the growth of AI agents, coding assistants, real-time chatbots and voice applications.

The d-Matrix MGX system is being designed specifically for workloads where response latency is critical.

Under the architecture described by the companies, operators could divide inference workloads between NVIDIA GPUs and d-Matrix XPUs. For AI coding applications, for example, NVIDIA GPUs could process the compute-intensive prefill stage while d-Matrix accelerators handle the latency-sensitive decode stage that generates tokens.

Raptor Brings New Memory Architecture

Raptor is the successor to d-Matrix’s Corsair XPU, which is already in production. The new processor extends the company’s memory-centric architecture by combining a DRAM memory chip with an SRAM compute chip in a vertically integrated package.

The architecture is designed to reduce the distance data must travel between memory and computation, addressing one of the bottlenecks affecting AI inference performance and energy consumption.

d-Matrix said Raptor is expected to tape out before the end of 2026 and is currently being evaluated by AI hyperscalers and frontier AI labs. The technology is backed by more than 100 patents, according to the company.

The partnership reflects a broader shift in AI infrastructure toward heterogeneous computing, where specialized accelerators can operate alongside GPUs rather than relying on a single processor architecture for every stage of an AI workload.

For NVIDIA, NVLink Fusion extends its infrastructure ecosystem to third-party silicon. For d-Matrix, the collaboration provides a path to deploy its inference-focused processors using NVIDIA’s existing rack-scale architecture and data-center ecosystem.

Facebook
Twiter
LinkedIn
Picture of Newsroom

Newsroom

More News