d-Matrix to integrate chips into Nvidia servers using NVLink
The AI inference startup will use Nvidia's chip-linking technology to deploy its Raptor processors in data-center systems by 2027.
d-Matrix adopts Nvidia interconnect for AI inference chips
AI chip startup d-Matrix announced it will use Nvidia's NVLink Fusion technology to integrate its processors directly into Nvidia's data-center infrastructure, according to Reuters. The move positions the Santa Clara-based company to address the growing demand for AI inference workloads as the industry shifts from model training to deployment.
The startup's new Raptor chips will connect to Nvidia server racks through NVLink Fusion, a chip-linking technology that provides specialized connectors and memory to enable third-party AI processors to operate within Nvidia's broader system architecture. d-Matrix expects to complete the final design stage for Raptor by the end of this year, with Nvidia-compatible racks becoming available in 2027.
Why it matters
This partnership highlights a strategic opening in Nvidia's ecosystem. While Nvidia's graphics processors dominate the lucrative AI training market, the company is creating pathways for specialized chips to handle inference—the process of running trained models in production. For enterprises, this could mean more cost-effective options for deploying AI applications that require rapid response times, such as coding assistants and voice agents, without abandoning their Nvidia infrastructure investments.
Targeting speed-critical applications
d-Matrix is focusing on AI services where latency is critical. The combined systems are designed for applications including chatbots, coding assistants, and voice agents—use cases where milliseconds matter for user experience. The startup is also working with connectivity firm Astera Labs to optimize data flow across the integrated system.
The company did not disclose financial terms of the collaboration with Nvidia.
Backed by Microsoft, valued at $2 billion
d-Matrix has raised substantial funding since its founding, including a $110 million round in 2023 that included backing from Microsoft. The startup shipped its first AI chip in November 2024 and was valued at $2 billion when it raised $450 million last year.
The inference market represents a different value proposition than training. While training workloads require massive computational power concentrated in relatively few locations, inference happens at scale across countless deployments. Specialized chips optimized for this task could offer better performance-per-watt or cost advantages compared to repurposing training hardware.
Details were first reported by Reuters, with reporting by Anhata Rooprai in Bengaluru and Stephen Nellis in San Francisco.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
