longbridgelongbridge
  • Platform Features
    Features
    Investment ProductsPrivate Wealth ManagementTrading ToolsMarket Data ServicesAnalysis ToolsNews ServicesFor Developers
    Account Types
    For IndividualsFor Institutions
  • Café
longbridge
© 2026 Longbridge|Terms of ServicePrivacy Policy

WDCX

WDCX
17.2700.17%( +0.030 )

LongbridgeAI
D
Dolphin Research

8 hours ago

From Sidekick to Star: How AI inference can rewrite NAND's destiny?

LongbridgeAII'm LongbridgeAI, I can summarize articles.
SandiskDeep Research

NAND has long been a 'technology-cursed' industry: as Dolphin Research noted in the earlier piece 'Born to Overproduce: How Does SanDisk Hold 80% GPM?', pre-AI NAND was a textbook commodity biz. Capacity was ruthlessly cleared by the market, tech leadership failed to command premiums, and the best a leader could do was 'lose less' as price cuts absorbed all volume gains.

The root cause lies in supply being too prolific: by stacking vertically on a single wafer and shrinking laterally, from BiCS5 to BiCS11 over five gens, SanDisk lifted bit output per wafer by 54% each gen, implying ~27% CAGR.

In other words, even without a cent of CapEx, bit capacity inflates nearly 30% per year automatically. The 3D transition delivered a step-function jump, pushing supply far ahead of demand and making asset-light NAND even harsher than asset-heavy DRAM through downcycles.

AI’s breakout is rewriting the NAND script. In this report, Dolphin Research focuses on how the AI wave is structurally reshaping downstream demand for NAND.

Full text below:

1. How exactly is AI structurally reshaping NAND’s downstream demand?

Before AI, core demand owners in consumer cycles — smartphones, PCs, and traditional servers — grew steadily with little upside surprise. Lacking explosive drivers, the amplitude of the cycle was effectively dictated by chaotic supply expansions and sharp contractions.

Growth in demand came from three steady levers: device-level capacity upgrades, EM penetration, and SSD replacing HDD stock. The pace was gradual but stable.

Stable demand paired with mechanically rising supply kept NAND trapped in the classic 'silicon cycle': glut and price collapse → cuts by OEMs while low prices spur per-box capacity upgrades → rebalance and price recovery → OEMs chase high margins with fresh CapEx → overcapacity again.

As long as NAND remains a peripheral component, this loop doesn’t break. The only exit is a use case that is less price-sensitive and directly binds to GPU deployment — which is exactly the breaker role the AI wave has begun to play.

Phase 2 (post-2025): Structural boom driven by AI inference

2026 marks a watershed in NAND history: data centers’ share of total NAND demand jumps from ~30% in 2025 to ~45% in 2026. This surpasses smartphones for the first time, with phones shrinking to ~23% over the same period, making DCs the largest end-market base.

The demand driver shifts from price-sensitive, inventory-whipsawed consumer electronics to cloud providers that decide based on TCO and ROI, with far less price sensitivity.

Clouds no longer ask 'are bits cheap this quarter', but 'can we secure enough high-perf eSSDs on time to ensure the next three years of GPU build-outs don’t stall'. The pivot from short-term price squeezing to long-term assured supply underpins OEMs’ push for NBM long-term agreements.

As the base shifts, the AI focus moves from training to inference, fundamentally altering NAND’s role:

In early AI, NAND acted as a peripheral data store: training corpora, model weights, and crash-safe checkpoints were staged near GPUs for quick fetches. These datasets are highly substitutable — if lost, one can reload from a data lake at the cost of bandwidth and minor power, without re-computing on costly GPUs.

But inference is high-concurrency and continuous 'dynamic consumption' — storage demand tightly binds to 'AI users × usage frequency × context length'. More critically, NAND now stores compute intermediates for the first time.

The massive KV Cache from long-context inference is not a static import but the 'working memory' generated by GPUs after heavy compute. Once lost, there is no central backup and it must be recomputed with expensive GPU cycles and power.

Storage value can now be directly translated into 'compute and power'. As SanDisk puts it, SSDs have effectively become a 'token battery': tokens represent work already done; rather than discard and recompute, 'charge' the disk by storing KV Cache, saving clouds power and GPU cycles.

Kioxia’s outlook corroborates this: DC NAND demand rises from 286 EB in 2025 to 1,686 EB by 2031 (34% CAGR), with sharp divergence internally:

AI inference: surges from 86 EB in 2025 to 1,251 EB by 2031, a 56% CAGR. Its share of DC NAND climbs from 30% to 74%.

AI training and traditional loads: training CAGR is only 11% over the same period, while CPU-centric cloud workloads grow 14%.

Inference grows 5x faster than training — this NAND super-cycle is primarily driven by inference demand.

Main NAND applications in AI data centers:

① KV Cache offload — NAND becomes the GPU’s context store

During inference, GPUs compute and maintain the KV Cache for each token in real time. In early designs, this cache was treated as a disposable: once a dialogue ends or the GPU frees HBM for the next user, the KV Cache is purged.

When a user follows up, or many users hit the same long document, the system must re-feed the context and rerun the costly prefill from scratch. That prefill step reads all inputs at once, computes token relationships via attention, and regenerates the KV Cache.

Prefill is the most compute- and power-intensive stage of inference.

Historically, discarding the cache made sense for two reasons:

a. Recompute was cheaper than storing: in the GPT-3 era with 2K tokens, KV Caches were tiny and a GPU could recompute in fractions of a second. Persisting to SSD required CPU and DRAM bounce buffers and PCIe 'roundtrips', often taking longer than recompute.

b. HBM is small and expensive: scarce HBM had to be reserved for active tasks and couldn’t hold inactive users’ history.

The long-context era invalidates this math: from 2K tokens in 2020’s GPT-3 to 1M+ in 2026 mainstream models, context expanded 500x in six years. Recomputing a 1M-token KV Cache keeps pricey GPUs fully loaded for seconds, incurring real compute and power costs.

This drove a reclassification of KV Cache: NVIDIA’s ICMS reframes it from a disposable scratchpad to 'AI-native persistent data'.

Now that ephemeral becomes persistent, does the economics work?

a. Capacity constraints: compression can’t keep up with KV growth

First, HBM capacity growth lags the surge in KV Cache.

Total KV Cache = context length × concurrent sessions × per-token footprint. The first two inputs are expanding rapidly.

Context length is doubling yearly or faster: mainstream models raced from 4K and 128K to 1M+, scaling KV linearly. Meanwhile, concurrency and iterative inference cycles are exploding under three architectures:

a. RAG: single queries are augmented with massive retrieved documents, pushing tokens far above the original question and inflating KV per task.

b. Agentic Search: AI shifts from Q&A to autonomous 'plan-call-evaluate-retrieve', lifting tokens per task from hundreds to millions.

c. Multi-Agent: parallel agents each generate a KV Cache, multiplying storage needs with agent count.

Vendors are doubling context annually, while HBM scaling is hitting limits: per-card capacity doubles only every 2-3 years, and benefits are slowing.

Assess overflow at the cabinet-level HBM pool. Take Rubin NVL72’s 21 TB HBM pool as an example.

Deploying a 2.8T-parameter flagship MoE, weights in FP8 consume ~2.8 TB; subtract activations, paging, and framework overhead (~12%, ~2.5 TB), leaving ~15.4 TB pooled headroom for KV Cache.

At no compression, 100K tokens are ~40 GB. Even with cutting-edge compression like Google TurboQuant (16-bit down to 3-bit; ~5–6x lossless), 1M tokens still occupy 67–80 GB.

This means a ~$5 mn cabinet can only serve 193–230 simultaneous 1M-token sessions at the limit. For enterprise cloud, that implies low ROI on compute monetization.

Push to 5M tokens or increase concurrency and the cabinet HBM pool quickly saturates. Offloading KV from HBM progressively to DRAM and then to NAND becomes inevitable.

b. Cost matters — and NAND’s price is decisively attractive

Even if HBM capacity barely suffices, using it to retain KV long-term is uneconomic: HBM4 is ~$16/GB vs. eSSD at ~$0.3–0.4/GB. That is a 40–50x delta.

Parking 'warm/cold' KV that’s revisited infrequently on HBM is gross misallocation. NVIDIA’s new tiering addresses this: G1 HBM → G2 system DRAM → G3 local SSD (NAND) → G4 shared network storage (NAND/HDD), with NVIDIA Dynamo orchestrating transparent KV migration.

NVIDIA estimates multi-agent inference can require up to 16 TB KV Cache per GPU module. Rubin adds an ICMS/CMX context tier between G3 local SSD and G4 shared storage, dubbed G3.5, because G3 capacity is limited and lacks cross-node sharing for large-scale reuse.

CMX at ~16 TB per layer is nearly 60x a single card’s 288 GB HBM — NAND’s low cost enables large capacity without breaking budgets.

c. Can NAND meet latency targets?

The offload decision ultimately hinges on TTFT — first-token latency. Offload only makes commercial sense when 'read from SSD' is faster and cheaper than 'recompute on GPU'.

GPU Direct Storage (GDS) is the key enabler: it bypasses CPU and DRAM, creating a direct high-speed path via PCIe/RDMA between eSSDs and GPU HBM.

Public tests show GDS plus compression can improve KV restore latency by 10x, making ultralow-latency context delivery practical.

KV Cache data types for NAND: 'shared memory' with high reuse

Even with capacity and cost hurdles cleared, offloading KV to NAND requires two conditions:

Condition 1: it must be shared-type cache

a. Ephemeral drafts (token-by-token): HBM holds these

During decode, the AI rereads the entire context for each token. Forcing SSD writes here would choke throughput due to bandwidth limits, and fragmented high-frequency appends would wear out enterprise SSDs in weeks.

These draft buffers must not hit disk and must remain in HBM.

b. Shared caches (static, batched): store in NAND

Only after prefill completes and the user or agent pauses do we get a complete context that can be written in large sequential blocks. Such 'big-block, sequential, low-frequency' data maps perfectly to NAND and can be placed there.

Condition 2: high reuse is required to pay for itself

Beyond feasibility, economics matter: offload net benefit = reuse count × (recompute cost − read cost) − (write cost + storage occupancy cost).

GPU recompute burns power and cycles, while SSD reads are cheap. Reusing just 1–2 times can break even, but SSD capacity is scarce and only high-reuse contexts deserve the space.

This boils down to cache hit rate: the higher the hit rate, the closer marginal service cost gets to zero. When a new request arrives, the system fetches matching prefixes and skips the expensive prefill.

As Manus notes, tokens served from cache are ~10x cheaper than misses. General LLMs show 40%–70% hit rates (MiniMax 71%, Z.ai GLM5 ~40%), while multi-step agent flows reach 95%+ persistently (DeepSeek measured 98.7%), delivering 'write once, read thousands' economics.

In practice, longer dialogues or user switches cause interruptions and free HBM by dropping current KV Cache. If sessions are likely to resume, store them in NAND and reload on restore.

Typical offload scenarios include: user reading pauses, cross-epoch history recall, long agent suspensions, massive concurrent sharing, and enterprise KB retrievals.

KV Cache tiering: AI inference lifts not just one node but the full span from high-perf flash (pSLC/TLC) to large 'warm-cold' stores (QLC), along with DRAM and HBM. The entire enterprise storage stack gets pulled up.

Division of labor across tiers is clear:

DRAM (G2) buffers 'hot spillover' from active sessions: it cushions HBM and holds overflow KV within the same active session. This is fast and avoids stalls but is volatile.

NAND (G3–G4) anchors 'persistent reuse' of high-value assets: once past the power-loss threshold, data moves into non-volatile flash for persistent reuse and sinks down tiers by medium characteristics:

a. G3 direct-attached (pSLC/perf TLC): holds recently swapped contexts likely to be reawakened.

b. G3.5 local network (durable TLC): holds System Prompts and agent states for high-frequency cross-node sharing within a cluster.

c. G4 remote archive (large QLC/HDD): holds massive, very low-frequency RAG KBs and historical dialogue archives.

② Staging: the TLC and pSLC hard base

SanDisk estimates staging — 'standby' data waiting for HBM calls — will account for 40% of DC AI NAND by 2030. This is the largest single slice and is fully served by TLC (incl. pSLC).

Think of it this way: data pulled from the lake lands first in high-speed NVMe SSD buffers near the GPU motherboard (G3 direct-attached SSD). Like a prep tray by the cutting board — you don’t run back to the cold room for every slice.

a. Training-side staging: bursty overwrite

Key tasks are periodic checkpoints and dataset buffering. For trillion-parameter models, single checkpoints hit TB scale and roll every tens of minutes.

Even with async checkpointing that avoids stalling GPUs, such heavy dumps still consume significant bandwidth. The longer it takes, the more it disrupts the compute pipeline.

Training-side staging thus faces two constraints: withstanding frequent overwrites and minimizing write windows.

b. Inference-side staging: read-heavy concurrency

Three read-dominant tasks stand out:

Cold starts and elastic scaling: unlike training nodes that run continuously, inference fleets scale with diurnal peaks, and memory is volatile. New nodes rely on strong sequential read bandwidth from local SSDs to 'instant-load' weights for 'seconds-to-serve' starts, repeated dozens of times per day.

Multi-model/version residency: production nodes often host 'family buckets' of tens of TB across sizes and precisions. With 288 GB HBM and limited host memory, only 4–16 TB direct SSDs can stage the full toolkit.

Letting rarely used models hog expensive HBM is misallocation; swapping via SSD is optimal.

Under such extreme I/O, several media fall out:

HBM & DRAM: too costly and scarce, and DRAM is volatile, undermining checkpoint recovery. These must be reserved for active KV.

HDD: throughput cannot handle bursty, large writes, and size/power constraints keep it out of GPU trays.

QLC: endurance is ~1/10 of TLC (QLC ~1,000 cycles vs. TLC ~10,000). Staging is the write-heaviest layer; QLC would wear out quickly — a hard constraint, not a cost trade-off.

TLC fits staging, yet in some cases even TLC falls short — hence pSLC:

pSLC treats TLC/QLC cells as SLC by storing 1 bit per cell instead of 3–4. This yields far higher endurance and lower write latency, at the cost of sacrificing ~2/3 effective capacity.

The direct implication: achieving the same effective capacity as standard SSDs requires 3–4x wafer output. The core trigger is high-frequency vector search in RAG and agentic search.

New search patterns force NAND to run in high-speed pSLC mode:

a. Extremely fine access granularity: AI often compares tiny embeddings (hundreds of bytes). Standard TLC reads full 4 KB pages even for a few bytes, wasting bandwidth.

b. Massive concurrency: agentic flows spawn thousands of dictionary-like lookups per task. Standard TLC queues saturate under high concurrency.

c. Read-while-write pressure: agents update indices and states while reading. This 'read and jot notes' pattern wears out TLC fast, while pSLC endures several times longer.

Still, the next-gen direct-attached layer to serve this — NVIDIA Storage-Next (G2.6, between G2.5 CXL memory and G3 local SSD) — is being co-defined by NVIDIA and vendors like Kioxia.

Its role is to add a GPU-direct pSLC tier between DDR and local SSD, delivering DRAM-like IOPS and 512B random granularity while retaining NAND’s low cost and high capacity.

③ Fast Data Lake: QLC’s main battlefield

If GPU-attached SSDs are the prep tray by the cutting board, they must both read and write. The G4 remote Fast Data Lake is the 'central pantry' of AI DCs, read-dominant: it holds all the 'inventory' — training corpora, multimodal datasets, model weights across versions, RAG KBs, and inference logs — at PB scale.

Per SanDisk’s 2030 split, Fast Data Lake is ~25% of AI DC NAND demand and is almost entirely QLC-based.

Core question: can QLC replace HDD? We assess along these axes:

\

a. Hot data — QLC firmly in the lead

Data frequently consumed by CPUs/GPUs — active training sets and RAG-ready KBs — requires very high throughput. HDD’s millisecond seek vs. AI’s microsecond needs is a two-order-of-magnitude gap.

HDD is physically out; QLC’s capacity and bandwidth make it the definitive replacement here.

b. Warm/cold data — HDD holds the head table

For 'store-for-reference' and archival data, QLC inclusion depends on two factors:

① Supply-side disturbance: short-term 'pseudo substitution'

HDD will face a short-term 300–400 EB supply gap. Some clouds are backfilling with eSSDs, creating the appearance of accelerated substitution.

But HDD can also raise output via techniques akin to SSD scaling. As HDD supply returns, pure capacity demand temporarily served by eSSD should revert, weakening the substitution thesis.

② Economics: QLC fails the math after price rises

QLC is structurally bid up by hot-tier demand. Today, HDD’s $/GB advantage has stretched to 20–25x, making QLC uneconomic for cold archives.

Only after ultra-dense, low-cost QLC products arrive could substitution odds improve.

Thus QLC vs. HDD is not a one-way trend but splits three ways:

a. On performance-critical hot paths, substitution is done and irreversible.

b. On warm/cold archives, large price gaps block normal replacement; current 'shares' reflect HDD shortages and may revert.

c. Long-term wins hinge on ultra-dense QLC crossing the TCO line of 2–3x vs. HDD via space and energy compression.

SanDisk disclosed that ultra-high-capacity QLC saw first revenue in FQ4'26 (Apr–Jun 2026), with early ramp in FQ1'27. QLC’s share of bit shipments is expected to rise from ~20% in 2025 to 40% by FY26-end, driven by hyperscale uptake of Stargate (ultra-dense QLC).

④ HBF: upside option on demand

In traditional AI tiers, HBM near GPUs offers sub-µs latency and very high bandwidth but is constrained by DRAM scaling, resulting in small stacks and high costs. Dropping to NVMe SSD introduces a severe latency and bandwidth gap.

NAND vendors are pushing HBF (High Bandwidth Flash) — architecturally akin to HBM: multi-layer stacks, TSVs, and co-packaging with XPUs, but with NAND dies instead of DRAM.

SanDisk’s HBF aims to stack 16 NAND layers, sustaining HBM-class bandwidth while boosting capacity 8x. See: From HBM Caps to NAND Scale: NVIDIA vs. Memory Suppliers.

HBF and HBM are complementary, not substitutes:

HBF’s strengths are capacity and cost, but weaknesses are clear: microsecond latency limits responsiveness for fine-grained random reads, and write endurance caps usage. It fits 'read-mostly' roles, not heavy write hot paths.

HBF remains in engineering: pilot lines in 2H26 with commercialization in 2027, while thermal standards near hot GPUs remain under definition.

Division of labor emerges:

HBM focuses on 'performance first', the compulsory choice for flagship GPUs. It serves latency-critical prefill, high-frequency active KV, and symmetric bandwidth demands in training.

HBF is 'capacity first', handling sequential-read-heavy inference, long-context and MoE cases where full models must fit per node, and 'read-mostly' warm KV offloads. This enables mid/low-end GPUs or future edge devices to run very large models.

Thus, Dolphin Research views HBF as complementary to HBM — HBM for latency-sensitive hot data, HBF for capacity-hungry warm data. Adoption trajectories for HBF still warrant monitoring.

SanDisk defines four deployment modes; see 'AI Inference Boom: Can SanDisk Rise from the Ashes?'. We see the highest near-term odds for ③ additive (HBF+HBM mix) and ④ decoupled (pooled) modes.

The additive path is simplest: keep HBM and add HBF. GPU vendors avoid major redesigns; HBM keeps hot roles while HBF holds weights and cold KV.

SK hynix’s new H3 architecture blends HBM and HBF, with volume in 2027 and customer adoption in 2028, making it the least resisted path today.

The decoupled path pools HBF away from GPUs, like external disk trays connected over the fabric. This aligns with NVIDIA’s rack-level CMX storage pooling, but network upgrades are needed, making 2028+ a more realistic window.

By vendor roadmaps, HBF is still in standards coalition phase and not yet in GPU vendors’ (esp. NVIDIA’s) official architectures.

SanDisk and SK hynix are the two substantive drivers: they formed a standards alliance and built pilot lines in early 2026, ~6 months ahead of plan. Chips are targeted by end-2026, controllers in 2027, and volume thereafter.

Google, NVIDIA, and AMD could adopt as customers no earlier than 2028, but adoption ≠ definition. For HBF to enter GPU architecture stacks, it must pass standards, reference designs, and ecosystem integration.

Others: Samsung has patents but no public prototypes or timelines; Kioxia is pursuing XL-FLASH (G2.6 Storage-Next), not directly competing; Micron has no public plan.

Net-net, while HBF timelines are accelerating, the structure remains 'memory vendors are pushing; GPU vendors have yet to board'. True divergence hinges on 2027 samples and whether leading customers formalize it in 2028 designs.

Summary:

AI’s reshaping of NAND is more than a level shift up the demand curve; it could break the 'silicon cycle'. With compute deployment and inference loads taking the lead, NAND moves from 'peripheral' to a must-have on the compute path.

In the 'training-to-inference' and long-context era, storing KV Cache in NAND preserves expensive GPU cycles and power. SSDs are, to a degree, 'token batteries'.

SanDisk estimates DC AI NAND shipments reach 1.2 ZB in 2030 (~3x vs. 2026), with three core workloads clearly split:

a. Staging (40%): ultra-high-intensity I/O buffers near GPUs, constrained by overwrite endurance and high concurrency, monopolized by high-perf TLC (and pSLC).

b. KV Cache offload (35%): new in the inference era, absorbing HBM spillover via TLC/QLC laddering.

c. Fast Data Lake (25%): remote central store, dominated by large QLC.

By NAND type, TLC is 66% and QLC 34%: staging’s write intensity naturally excludes QLC, cementing TLC’s base. QLC focuses on data lakes and warm KV offload.

These workloads evolve in opposite directions: lakes pursue 'cheaper' via denser QLC, while staging and hot paths pursue 'faster and tougher' via TLC and pSLC, trading ~2/3 capacity for ~10x endurance.

How long can this price upcycle fly? From today’s valuation base, does post-rebirth SanDisk still have upside? Dolphin Research will unpack this in the next piece — stay tuned.

<End>

Risk disclosure and statements: Dolphin Research Disclaimer and General Disclosures

Updates:

I. Dolphin Research AI storage series

DRAM industry:

'From Compute to Memory: Why HBM Stands Out'

'From HBM Caps to NAND Scale: NVIDIA vs. Memory Suppliers'

Stocks:

SanDisk (Part I): 'AI Inference Boom: Can SanDisk Rise from the Ashes?'

SanDisk (Part II): 'Born to Overproduce: How Does SanDisk Hold 80% GPM?'

II. Dolphin Research AI DC interconnect series

CPO industry

'AI’s Hyper-Connected Era: Toward Light?' 'Copper Never Left: CPO — Real Opportunity or Mirage?'

Networking architectures

NVIDIA networking: 'AI-era DC Interconnects: Beyond Single Chips, Teaming via Nets — Is There a China Window?'

AMD vs. NVIDIA: 'AMD Arm-Wrestles NVIDIA — Is Helios Ready?'

Google networking: 'Challenging NVIDIA’s Hegemony: Google’s 'Optical' Edge'

Stocks

Lumentum: 'From Optical Veteran to 'All-Purpose Water Seller': Lumentum’s Edge'

III. Dolphin Research AI XPU series

Agents and CPUs: 'Muse Goes Viral — Is This CPU’s ChatGPT Moment?'

…More on the Dolphin Research site.

Sandisk

Sandisk

USSNDK

Micron Tech

Micron Tech

USMU

Roundhill Memory ETF

Roundhill Memory ETF

USDRAM

Western Digital

Western Digital

USWDC

Apple

Apple

USAAPL

The copyright of this article belongs to the original author/organization.

The views expressed herein are solely those of the author and do not reflect the stance of the platform. The content is intended for investment reference purposes only and shall not be considered as investment advice. Please contact us if you have any questions or suggestions regarding the content services provided by the platform.