A Different Route Around the Memory Wall
At Hot Chips 2026, d-Matrix described Raptor as a 3D-DRAM accelerator built for generative AI inference, and the timing is not accidental. As of 02 September 2026, inference hardware is under pressure from three directions at once: larger models, tighter latency targets, and the rising cost and power budget of feeding accelerators with enough memory bandwidth. Raptor’s pitch is simple but technically ambitious: instead of placing high-bandwidth memory beside the compute die, d-Matrix stacks compute directly on top of a custom DRAM die to shorten the data path.
Article contains affiliate links, commission may be earned.
The Key Specs d-Matrix Has Discussed
The reported Raptor configuration uses a TSMC 4nm compute die bonded face-to-face to a custom DRAM die at a 36-micron pitch. Tom’s Hardware reported the card at 32GB of memory and around 100 TB/s of bandwidth, with the vertical interface measured at 0.37 pJ/bit. That is notably lower than the roughly 2.4 pJ/bit figure cited for moving data into an HBM4 base die. These are architecture and silicon-level claims rather than independent application benchmarks, so they should be read as product-overview figures, not a review verdict.
View NVIDIA DGX Spark Personal AI Desktop Supercomputer on partner website
Why Stacking DRAM Changes the Conversation
AI inference often spends a lot of energy and time moving model data rather than doing math. HBM has become the standard answer for bandwidth-hungry accelerators, but it still sits next to the logic die and depends on advanced packaging, interposers, and memory PHY overhead. Raptor’s approach moves the memory path vertically. By bonding logic and DRAM face-to-face, d-Matrix is trying to reduce the energy cost per bit while giving the compute die a much wider connection to nearby memory. In practical terms, the idea is to keep inference engines fed more consistently without simply adding more external memory bandwidth.
See Crucial T705 13 GB/s NVMe SSD price
Built for Generative Inference, Not General-Purpose GPUs
Raptor should not be viewed as a direct clone of a GPU with a different memory stack. d-Matrix has positioned its architecture around generative inference, where low latency, predictable throughput, and power efficiency matter heavily in production deployments. The 32GB-per-card figure also matters here: it suggests Raptor is aimed at serving and scaling inference workloads through system design, rather than trying to fit every large model entirely on a single accelerator card. The interesting part is the trade-off: less conventional memory packaging, but potentially much denser access between compute and DRAM.
Buy ASUS GeForce RTX 5090 here
What Still Needs Watching
The open question is how Raptor performs once it moves beyond architecture slides and reported silicon characteristics into real deployment. Memory bandwidth is crucial, but production inference also depends on software support, model compatibility, interconnects, scheduling, thermals, yield, and total system cost. d-Matrix’s concept is compelling because it attacks the memory wall directly, but the broader market will be watching for shipped systems, customer workloads, and reproducible performance data. For now, Raptor stands out as one of the more concrete examples of how AI accelerator design is shifting from just more compute toward smarter data movement.
View Lenovo Legion Pro 7i laptop with Core Ultra 9 and NVIDIA RTX 5090 24GB on partner website
Comments
No comments yet. Be the first to share your thoughts.