Thursday, September 10, 2026
Tech Beat
Sep 10, 2026, 1:00 PMAI Hardware

d-Matrix Connects Raptor AI Inference XPUs to NVIDIA NVLink Fusion

d-Matrix will connect Raptor XPUs to NVIDIA NVLink Fusion and MGX, targeting faster, lower-risk rack-scale deployment for ultralow-latency AI inference.

Listen to this briefingAudio briefing

Summary

d-Matrix announced on September 10, 2026, that its next-generation Raptor inference XPUs will use NVIDIA NVLink Fusion, linking them to NVLink scale-up networking, Spectrum-X scale-out networking, MGX racks and the wider NVIDIA AI platform. It also plans to integrate Vera CPUs, ConnectX-9 SuperNICs, BlueField-4 DPUs and Spectrum-X Ethernet. Its liquid-cooled racks could operate beside Vera Rubin NVL72 GPU systems for disaggregated, ultralow-latency inference within one AI factory architecture.

Sixth-generation NVLink provides 3 TB/s of all-to-all bandwidth per XPU, while NVIDIA claims three times lower XPU-to-XPU latency and 10 times higher packet rates than off-the-shelf Ethernet. NVLink Fusion supports Arm, x86 and RISC-V CPUs alongside NVIDIA GPUs. The broader platform includes Vera Rubin NVL72, Groq 3 LPX, the Vera CPU rack, Vera BlueField-4 STX storage and Spectrum-6 SPX Ethernet, spanning compute, networking, storage, security, power, cooling and software.

NVLink Fusion supplies the interconnect, validated racks, software and supply chain for semi-custom AI factories, letting d-Matrix focus on silicon while avoiding separate rack designs. Partners include AWS, Arm, Intel, Fujitsu, SiFive, Alchip, Astera Labs, GUC, Marvell, MediaTek, Samsung, Cadence, Synopsys, Ayar Labs and Lightmatter. CEO Sid Sheth said on September 9 that rising inference demand is constrained by capital, time and energy; integration and large-scale deployment remain planned rather than completed.

Positives

  • Sixth-generation NVLink delivers 3 TB/s of all-to-all bandwidth per XPU for Raptor deployments.
  • NVIDIA claims three times lower XPU latency and 10 times higher packet rates than standard Ethernet.
  • MGX gives d-Matrix validated, liquid-cooled rack designs, power infrastructure, cooling and an established supply chain.
  • One common rack architecture can support GPUs, CPUs and XPUs without separate processor-specific rack designs.
  • Raptor racks can work beside Vera Rubin NVL72 systems, enabling specialized and GPU-based disaggregated inference in unified AI factories.

Risks & concerns

  • Raptor integration and rack-scale deployment remain future plans, with no completion date disclosed.
  • Capital, time and energy remain finite even as inference demand rises, CEO Sid Sheth said.
  • The stated latency and packet-rate advantages are NVIDIA platform claims rather than disclosed d-Matrix deployment results.
  • Custom XPU deployment still involves complex networking, interface integration, validation, rack certification, power, cooling and software requirements.
Primary sourceNVIDIA Bloghttps://blogs.nvidia.com/blog/d-matrix-nvlink-fusion/
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

AI HardwareAug 27

NVIDIA Vera CPU Begins Shipping at Scale for AI Agents

A metal doughnut awakens atop stacked coins, symbolizing OpenAI’s costly effort to make AI hardware feel alive. AI HardwareAug 7

OpenAI ChatGPT Smart Speaker Could Cost $400, Launch in 2027

A polished metal donut forms an oversized house key, symbolizing OpenAI’s costly route into smart homes. AI HardwareAug 6

OpenAI’s Donut-Shaped ChatGPT Speaker Could Cost Up to $400