High Bandwidth Flash (HBF) vs HBM: Key Takeaways for AI Memory

•

 20 min video

•

 8 min read

YouTube video ID: 3nTpW52nioI

Source: YouTube video by Asianometry — Watch original video

PDF

High Bandwidth Flash (HBF) is an emerging memory technology designed to address the escalating demands of AI inference, particularly concerning the storage and processing of large language models (LLMs). It aims to serve as either a companion or a challenger to High Bandwidth Memory (HBM).

The AI Inference Challenge: Prefill, Decode, and KV Cache

To understand the necessity of HBF, it's crucial to grasp two key trends in AI inference:

  1. AI Inference Regimes: LLM inference involves two distinct computational phases:

    • Prefill: During this phase, the model processes the entire input content, which can include user prompts, system prompts, prior replies, tool calls, images, and audio. The GPU or other processing unit must attend to all these elements in parallel to understand the prompt's meaning. This "attention" mechanism, central to Transformer models, generates Query, Key, and Value vectors for each token. The Key and Value vectors are stored in a KV Cache. Prefill is compute-bound, as it can be parallelized.
    • Decode: Following prefill, the model enters the decode phase, where it consults the KV Cache to generate the response token by token. The Key and Value of each newly generated token are added to the KV Cache. Decode is memory-bound, as the GPU's logic circuits often wait for data to be loaded from memory.
    • KV Cache: This cache stores the Key and Value vectors, preventing the LLM from recalculating attention for all prior inputs repeatedly. It acts as the model's working memory for the current conversation.
  2. Exploding Data Demands:

    • Model Weights: LLMs are growing exponentially in size. For instance, the Kimi K3 model has 2.8 trillion parameters, requiring approximately 2.8 terabytes of storage at 8-bit accuracy. This represents a significant increase from previous models.
    • Massive KV Caches: The rise of long-running AI agents, which can operate for hours with extensive chains of thought, leads to enormous KV Caches. While current context windows are around 1 million tokens (equivalent to 1,500-3,000 pages of text), this is insufficient for long audio or video inputs.
      • Mitigation Strategies:
        • Context Window Compaction: Compressing the context window can free up space, but it may lead to a loss of detail, similar to compressing a JPEG image.
        • KV Cache Tiering: Less frequently accessed blocks of the KV Cache can be offloaded to cheaper, higher-capacity SSDs, while frequently accessed blocks remain in HBM.

The Limitations of High Bandwidth Memory (HBM)

HBM is currently the dominant memory solution for AI, characterized by its 3D stacked structure with a logic base die and multiple core memory dies connected by Through-Silicon Vias (TSVs). However, concerns are growing about its long-term suitability:

  • Bandwidth Dilution with Stacking: As HBM stacks become taller (e.g., 20 levels thick), the bandwidth per gigabyte can decrease significantly. The limited number of TSVs in the central area restricts data movement, meaning that increasing capacity by stacking more dies can dilute the throughput.
  • Rising Cost of Capacity and Bandwidth: The stacking approach has led to an increase in the price of incremental capacity and bandwidth. In extreme cases, individual core dies in very tall stacks might operate slower than commodity memory.
  • Industry Skepticism: Prominent figures, including former Intel CEO Pat Gelsinger, have criticized HBM as "lousy memory." While SK hynix acknowledges that HBM is not the ultimate solution to the "Memory Wall" problem, they maintain it's the best available option currently.

Introducing High Bandwidth Flash (HBF)

High Bandwidth Flash (HBF) proposes a new approach by stacking and connecting flash dies in a manner similar to how HBM stacks DRAM dies.

  • Key Characteristics:

    • High Capacity: HBF vastly increases per-stack memory capacity by 8-16 times compared to HBM4, offering up to 512 gigabytes per stack. This is distinct from 3D NAND, which focuses on monolithic stacking of memory cells for sheer capacity.
    • High Bandwidth: HBF aims to maintain HBM's high bandwidth rates, with theoretical speeds potentially exceeding HBM due to more NAND cells being accessible to the GPU. Current HBF specifications target up to 3 terabytes per second, comparable to HBM4 and higher than HBM3E. However, when adjusted for capacity, HBF's bandwidth per gigabyte is significantly smaller.
  • Drawbacks:

    • Slower Read/Write Speeds: NAND flash is inherently much slower than DRAM for both reading and writing (programming), with roughly a two-order-of-magnitude difference. Programming is even slower than reading, leading to higher latency. While optimization can reduce latency, it's a fundamental limitation.
    • Endurance Issues: NAND cells have a limited number of write cycles (thousands), unlike DRAM cells which have virtually unlimited endurance. This necessitates careful management of write operations to HBF.
    • Power and Thermals: HBF cubes require more power, leading to increased heat generation. This can cause performance throttling and exacerbate endurance issues.

A Brief History of HBF Development

The concept of stacking NAND with TSVs has existed for some time, with early proposals like "High Bandwidth NAND" from Honda Research Institute Japan in 2019.

  • Sandisk's Role: Sandisk formally introduced HBF at its Future FWD investor day in February 2025, stating that the architecture was developed with input from major AI players. The first patent application for HBF was filed in May 2024.
  • Industry Collaboration:
    • In July 2025, Sandisk formed a Technical Advisory Board, including prominent computer architects David Patterson and Raja Koduri, to guide HBF's development.
    • In August 2025, Sandisk and SK hynix, typically competitors in the NAND space, signed a Memorandum of Understanding to standardize the HBF specification. This collaboration aims to scale the product ecosystem while allowing individual companies to compete on internal product specifics. It also provides a credible second source for the technology.
    • Google and Tenstorrent joined the HBF consortium.
  • Standardization: In August 2026, the Open Compute Project released the first public version 0.7.0 of the HBF specification.
  • Samsung's Alternative: Samsung has its own version of HBF called zNAND-O, which appears similar but with its own distinct branding.

HBF's Place in the Memory Hierarchy

HBF is envisioned to fit into the memory hierarchy alongside HBM:

  • Memory Hierarchy:

    • Tier 0: GPU/TPU SRAM (fastest, lowest capacity, on-chip)
    • Tier 1: GPU HBM (off-chip, physically closest to the chip)
    • Tier 2: System DRAM (commodity off-chip, higher capacity, further from GPU)
  • Proposed Integration: Sandisk suggests placing HBF side-by-side with HBM at Tier 1. HBM would handle "hot" data requiring immediate access, while HBF's higher capacity would store "warmer" data.

  • Capacity Boost: Replacing some HBM stacks with HBF can dramatically increase total memory. For example, a GPU with eight 24GB HBM stacks (192GB total) could achieve 3.12 terabytes by swapping six HBFs for six HBMs.
  • Use Cases: Due to flash's endurance limitations, HBF is best suited for "read-mostly" data. Model weights are a prime candidate, given their increasing size. Some also suggest using it for KV Cache.

Challenges and Considerations

Integrating HBF is not straightforward and presents several challenges:

  • Token Output Performance: Simply having cheap memory capacity in HBF doesn't automatically translate to cheaper tokens. If data cannot be retrieved from HBF fast enough, the accelerator will wait, negatively impacting token output.
  • Software Support: HBF requires extensive software support. System software must intelligently decide which data to place in HBF, such as widely used shared prompts. This is a complex task.
  • Model Architecture: Certain model architectures, like mixture-of-experts models, might be better suited for HBF. These models activate only a subset of experts for each token, meaning they need capacity for all experts but only read a few at a time, making lower bandwidth per gigabyte acceptable.

Lessons from Intel Optane

The history of Intel Optane (3D XPoint) offers valuable lessons for HBF:

  • Optane's Promise: Optane was a non-volatile memory with read/write speeds similar to DRAM, positioned between DRAM and NAND in terms of performance and price. Intel marketed it for "warm data" – large databases requiring DRAM-like latency but too big for DRAM.
  • Optane's Downfall:
    • Software Integration: Optane required significant software support from application developers to fully leverage its capabilities.
    • Economic Justification: The rapid decline in cost-per-bit for 3D NAND undermined Optane's economic viability.
    • Single Vendor Risk: Micron eventually withdrew from the partnership, leaving Intel as the sole vendor. Customers were reluctant to rely on a single source for both CPUs and memory.
  • Relevance to HBF: The core executive at Sandisk leading HBF development previously headed Intel's Optane Group, highlighting the importance of learning from past experiences. While Optane's underlying PCM technology was robust, its market strategy and ecosystem development failed. HBF must avoid similar pitfalls in its controller design, software integration, and go-to-market strategy.

Conclusion and Outlook

HBF is progressing rapidly, with potential customers expected to receive proof-of-concept or engineering samples in late 2027, and high-volume manufacturing projected for 2029.

Many uncertainties remain: * Can Sandisk, SK hynix, and other partners deliver HBF on schedule? * Will system architects effectively identify its optimal use cases? * Will software vendors provide the necessary support? * How will customers ultimately utilize this new memory technology?

The success of HBF hinges on overcoming these challenges and establishing a clear value proposition in the evolving landscape of AI memory.

  Takeaways

  • HBF stacks flash dies to deliver 8‑16× the capacity of HBM4, reaching up to 512 GB per stack while targeting up to 3 TB/s bandwidth comparable to HBM4.
  • Because NAND flash reads and writes are two orders of magnitude slower than DRAM, HBF’s bandwidth per gigabyte is much lower and latency is higher, making it best for read‑mostly data such as large model weights.
  • HBF is positioned at Tier 1 alongside HBM, with HBM handling hot data and HBF storing warmer data, allowing a GPU with eight 24 GB HBM stacks to scale to over 3 TB when several stacks are replaced by HBF.
  • Successful adoption will require sophisticated software that can tier data between HBM and HBF, and model architectures like mixture‑of‑experts that only read a subset of parameters per token are especially suitable.
  • Lessons from Intel Optane warn that without broad ecosystem support, cost‑per‑bit advantages, and multiple vendors, HBF could face similar market challenges despite its technical promise.

Frequently Asked Questions

Why is HBF better suited for storing large model weights than for KV cache?

HBF is better suited for storing large model weights because its high capacity and read‑mostly characteristics match the static nature of weights, whereas KV cache demands frequent writes and low‑latency random reads that NAND flash cannot deliver efficiently.

How does HBF’s bandwidth per gigabyte differ from HBM’s, and what impact does that have on token generation?

HBF provides roughly the same total bandwidth as HBM (about 3 TB/s), but because each stack holds many more gigabytes, its bandwidth per gigabyte is far lower, so during the decode phase data may not be fetched quickly enough, causing token‑generation slowdown if the system cannot tier data effectively.

Who is Asianometry on YouTube?

Asianometry is a YouTube channel that publishes videos on a range of topics. Browse more summaries from this channel below.

Does this page include the full transcript of the video?

Yes, the full transcript for this video is available on this page. Click 'Show transcript' in the sidebar to read it.

Helpful resources related to this video

If you want to practice or explore the concepts discussed in the video, these commonly used tools may help.

Links may be affiliate links. We only include resources that are genuinely relevant to the topic.

Full transcript is not shown on this page

This page focuses on the summary and original notes. For full verification, refer to the original YouTube video.

PDF