4hi HBM samples started showing up on suppliers’ radar right at the start of 3Q26. AI chip makers, NVIDIA included, are evaluating every possible configuration, including lower stack counts, to work around the ongoing HBM shortage for their next-generation compute platforms.
There’s a real use case for 4hi HBM at the edge, especially in autonomous driving and robotics. What’s less clear is whether it goes any further than that. Suppliers have made no commitment to mass-produce the spec, the base die math works against a shorter stack rather than for it, and the resulting price per Gb weakens the case for adoption in some applications.
Where 4hi Actually Makes Sense: Edge AI
Edge AI players, not data center players, are the ones showing real interest in 4hi HBM. TrendForce observes that in certain autonomous driving and robotics applications, DRAM bandwidth has gradually become a performance bottleneck for computing platforms, prompting HBM to be included in architecture evaluations.
There’s precedent for this. SK hynix supplied HBM2E to Waymo’s autonomous driving platform in 2024 and obtained automotive qualification, indicating that robotics and automotive applications have already become one of the extension areas for HBM usage.
That said, these applications are still early. Relevant specifications remain under evaluation. As a result, actual demand volumes are limited, and the short-term contribution of these segments to overall HBM demand is expected to remain very low.
Why 16GB Is a Hard Sell for the Data Center
Industry discussion around 4hi HBM is really a discussion about HBM5. Based on an HBM5 mono-die density of 32Gb, a single HBM5 4hi cube will offer a capacity of just 16GB. However, its I/O speed and bandwidth will be further enhanced compared to HBM4e.
That capacity-for-bandwidth trade-off is exactly what’s driving the split in interest. TrendForce assesses that HBM5 4hi samples are primarily being positioned as a memory architecture candidate for certain clients’ Edge AI applications. However, the likelihood of this specification being adopted in mainstream data center computing platforms is relatively low due to the following bottlenecks:


