At Hot Chips 2026, Santa Clara startup d-Matrix unveiled Raptor, a 3D DRAM accelerator designed to bypass the AI memory wall. By bonding a TSMC 4nm logic die directly onto a custom DRAM die using vertical microbumps at a 36-micron pitch, Raptor achieves 100 TB/s of bandwidth from 32GB of capacity. This architecture slashes interface energy to 0.37 pJ/bit—roughly six times lower than HBM3—directly targeting the power-hungry autoregressive decode phase of frontier-class models.
Raptor stacked DRAM architecture
- ▪The d-Matrix Raptor architecture features a TSMC 4nm compute logic die bonded face-to-face on top of a custom-designed DRAM die.
- ▪At Hot Chips 2026, d-Matrix unveiled Raptor, a 3D DRAM accelerator for generative AI inference that delivers 100 TB/s of memory bandwidth from 32GB of capacity per card.
- ▪An ISCA 2026 paper co-written by d-Matrix and the University of British Columbia projects that the Raptor architecture delivers approximately 4.7 times higher throughput per card than HBM-based designs.
Memory wall in inference
- ▪In AI inference at scale, the memory wall is the primary driver of high costs, power consumption, and user concurrency constraints in GPU clusters running frontier models.
- ▪The memory wall, first described in a 1995 paper by Wulf and McKee, represents the widening performance gap between processor speed and memory bandwidth.
HBM interposer energy cost
- ▪The physical interface driving signals across a centimeter-scale silicon interposer in High Bandwidth Memory systems consumes approximately 2 to 3 picojoules per bit.
- ▪Pushing a hypothetical HBM4 system to 100 terabytes per second of bandwidth would consume nearly 2 kilowatts of power for the memory interface alone.
Vertical microbump integration
- ▪The vertical microbump interface in the d-Matrix Raptor operates at an energy cost of 0.37 picojoules per bit, which is roughly six times lower than HBM3 across an interposer.
- ▪The d-Matrix Raptor connects its logic and DRAM dies using sub-millimeter vertical microbumps at a 36-micron pitch, bypassing the need for a silicon interposer.
- ▪The logic-on-top orientation of the d-Matrix Raptor allows a liquid-cooling cold plate to contact the logic die directly, protecting the thermally sensitive DRAM below.
Stream blocking bandwidth optimization
- ▪To prevent overfetch, d-Matrix's stream blocking technique aggregates four 96-byte DRAM accesses to assemble exactly three 128-byte flits for Raptor's 256 tensor engines.
- ▪The d-Matrix Raptor uses a stream flipping technique to compare each flit to the previous one and invert when it reduces transitions, saving 20% in I/O power.
Autoregressive decode bottleneck
- ▪A workload of 64 simultaneous users at a one million-token context window generates approximately 935 gigabytes of key-value cache in addition to model weights.
- ▪During the autoregressive decode phase of large language models, compute cores sit largely idle because the chip must read massive model weights and KV cache from memory on every step.
Story comments
Loading comments…