A New Memory Layer Between HBM and SSDs
As of August 18, 2026, High Bandwidth Flash, or HBF, is best understood as an emerging server-memory specification rather than a shipping GPU feature. Sandisk and SK hynix announced the first HBF technical specification through the Open Compute Project in early August 2026, positioning it as an open framework for AI inference systems that need more memory close to compute. The core idea is simple: HBM is fast but capacity-limited and expensive, while SSDs are spacious but sit too far away in the memory hierarchy for many latency-sensitive inference workloads. HBF aims to land in the gap between them, using NAND-based stacks with a much wider, faster connection than conventional storage. SK hynix describes it as a new memory layer between HBM and SSDs, not as a direct HBM replacement. (news.skhynix.com)
Article contains affiliate links, commission may be earned.
The Early HBF Spec: 512GB Packages and Three Bandwidth Grades
The first specification outlines up to 512GB per HBF package, based on either 8-high or 16-high NAND die stacks. Bandwidth is split into three grades, ranging from roughly 0.4TB/s to 3.0TB/s, giving system builders a ladder of possible implementations rather than one fixed part. The spec also includes guidance for interfaces, electrical behavior, reliability, packaging, and software I/O. A major technical point is the use of UCIe, the Universal Chiplet Interconnect Express standard, which is meant to help HBF connect with processors such as GPUs, CPUs, and other accelerator designs. That does not mean every AI GPU will suddenly gain terabytes of HBF, but it does give vendors a more defined path for experimenting with near-compute flash memory. (news.skhynix.com)
See SANDISK Optimus GX PRO 8100 PCIe 5 SSD price
Why AI Inference Is the Target
HBF is aimed mainly at AI inference, where large models and long-context workloads can strain accelerator memory capacity. Training still leans heavily on the lowest-latency, highest-bandwidth memory available, which keeps HBM central. Inference, however, often has a different pressure point: keeping larger model weights and active data closer to the accelerator without constantly falling back to slower system memory or storage. In that role, HBF could act like a high-capacity escape hatch. A future accelerator package could, in theory, combine HBM for the hottest data with HBF for a larger nearby pool. Sandisk has previously described HBF as NAND-based technology designed for AI inferencing, with its first-generation product target listed as 1.6TB/s read bandwidth and 512GB total capacity in a 16-die stack, though real product behavior will depend on final implementations. (documents.sandisk.com)
View NVIDIA DGX Spark Personal AI Desktop Supercomputer on partner website
The Catch: Latency, Ecosystem Support, and Timing
The big trade-off is that NAND-based HBF is not expected to match HBM on latency, even if the top bandwidth grade looks impressive on paper. That makes workload placement important: HBF would likely be most useful for data that benefits from high capacity and strong sequential or parallel read bandwidth, while the most latency-sensitive operations stay in HBM. Adoption is also still uncertain. Sandisk and SK hynix are primary contributors to the OCP specification, and the companies say Google and Tenstorrent have joined the consortium, but broad accelerator-vendor commitment is still the major milestone to watch. In short, HBF is not a magic memory upgrade for today’s GPUs. It is a serious architectural proposal for future AI servers, and its importance will grow if inference models keep outgrowing the memory that can practically be placed beside compute. (investor.sandisk.com)
Buy ASUS GeForce RTX 5090 here
Comments
No comments yet. Be the first to share your thoughts.