AI inference chipmaker d-Matrix said on September 10 that its next-generation Raptor accelerator chips will connect to Nvidia's NVLink Fusion interconnect, giving the startup's custom silicon a direct route into Nvidia's rack-scale data centre architecture, according to a Nvidia announcement.
Custom chips inside Nvidia's rack
NVLink Fusion is Nvidia's interconnect technology for linking non-Nvidia processors, known as XPUs, into its NVLink scale-up network, Spectrum-X scale-out networking and MGX rack architecture. d-Matrix's Raptor chips were designed from the ground up around it, according to Nvidia, and are expected by the end of 2027, according to The Register.
Nvidia has signed similar licensing arrangements with Qualcomm, Arm, Marvell, Amazon, Fujitsu and MediaTek, and has invested $3.5 billion in MediaTek and $2 billion in Marvell alongside those deals, according to a report by The Register. The outlet described the arrangement as a "walled garden" that gives chipmakers fast access to a proven rack ecosystem while tying them to Nvidia's surrounding hardware, including Vera CPUs, NVSwitch, and BlueField and ConnectX networking components.
Betting on memory bandwidth
An NVL144-class rack holding up to 144 Raptor accelerators would carry a combined 2.3 terabytes of memory capacity and 7.2 petabytes per second of aggregate bandwidth, according to The Register, which reported that each Raptor card carries 32 gigabytes of 3D-stacked DRAM with 100 terabytes per second of memory bandwidth.
"Demand for inference is soaring, but capital, time and energy remain finite. With NVLink Fusion and MGX, we can integrate our Raptor XPUs into a broadly deployed, liquid-cooled architecture, giving customers a faster, lower-risk path to deploy and scale ultralow-latency inference," said d-Matrix cofounder and chief executive Sid Sheth in the Nvidia announcement.
Memory bandwidth, rather than raw compute, has become the main bottleneck for AI inference at scale, and The Register calculated that where a trillion-parameter model might need more than 2,000 Groq 3 LPUs at eight-bit precision, a single d-Matrix system would need about 64, or 32 at four-bit precision.
Fitting into Nvidia's wider AI factory
Nvidia said the Raptor-based racks are designed to operate alongside its own Vera Rubin NVL72 GPU systems for what it calls disaggregated inference, splitting the separate stages of processing a model's request across specialised hardware. Nvidia presents its wider platform, which it lists as including the Vera Rubin NVL72, Groq 3 LPX, Vera BlueField-4 STX storage and Spectrum-6 SPX Ethernet networking, as a set of options for running any AI workload with, in its words, the best performance per watt and lowest cost per token.
Why it matters
The deal illustrates how Nvidia is extending its dominance beyond GPUs into the surrounding data-centre infrastructure that AI labs, hyperscalers and neoclouds depend on, even as it opens its network to competing chip architectures. For d-Matrix, the arrangement offers a faster, lower-risk path to market than building an independent rack ecosystem from scratch, though commercial deployment remains more than a year away.