Unlocking KV Cache Efficiency with HDDs

Unlocking KV Cache Efficiency with HDDs

How you can improve your bottom line by using HDDs in your KV cache infrastructure

With the speed of AI development, enterprises are learning on the fly. The world is now transitioning from training to generative inference and agentic AI, and this change to workflows exerts pressure on infrastructure in novel ways.

Unlike web search, for example, the user experience is frequently not simply prompt->response. Rather, the user issues a prompt, then gets a response. Based on that response, the user issues clarifying or additional prompts to refine the output or expand it based on what was generated. In the simplest case, an LLM user may treat their interaction with AI as a conversation, where questions require the context of previous prompts and responses to be retained, or else the output to a question down the line won’t make sense.

If the LLM is stateless, meaning each prompt is a distinct event, this becomes a problem. Each time a follow-up question is asked, the LLM must return to the first prompt and recompute every prompt/response/follow-up between the initial prompt and the current prompt. This consumes valuable time and consumes expensive GPU cycles. A Hugging Face performance benchmark found a specific task took 1 minute, 1 second using standard inference on an NVIDIA® T4 GPU.

To avoid this resource consumption, LLMs instead use key-value (KV) caching to retain those stored prompts and responses. Instead of recomputing, they can simply be read from cache, avoiding those expensive GPU cycles. That same Hugging Face benchmark showed that with KV caching enabled, the task took only 11.7 seconds, enabling a 5.21x throughput improvement on the GPU.

This performance advantage translates into real dollars and cents. Whether owned or rented, GPUs are very expensive assets. A task not utilizing KV cache can cost more than 5 times as much to run than a task utilizing KV cache. And more importantly, a very expensive GPU can perform more than 5 times as much work per hour when making use of KV cache.

As inference and agentic AI get more powerful and take on more complex tasks, the complexity of prompts grows rapidly. This is a problem, as Andrew Baker of Capitec Bank explains, because the effect of computation grows quadratically with prompt length; but when KV caching is in play, it grows linearly. Linear growth problems are a lot easier to solve economically than quadratic growth problems—especially when quadratic growth problems are affecting expensive assets like GPUs.

So, the takeaway is simple: You need KV caching for AI inference. KV caching turns quadratic problems into linear problems. But there’s another issue … KV caching is only effective when the data you need is stored in the KV cache!

The impact of KV cache misses: Direct and opportunity cost

Every cache miss results in recomputation. It’s that simple. The value of KV caching is to avoid spending expensive GPU assets recomputing what it’s already computed.

And that is both a waste of time and money. A task which could have been completed in 11.7s with KV cache now takes 1:01 because of a cache miss. This has two costs:

  • Direct cost: This is simple. Whether it’s an owned GPU asset or a rented GPU asset, your task now takes 49.3s longer than it could have. You’re paying for that time. You can measure this directly in P&L.
  • Opportunity cost: That 49.3s of extra GPU time could have been spent doing something more productive than recomputing something it had already computed. You miss the value of what that GPU could have been doing if only it had been able to read the data from cache.

Avoiding the expense of KV cache misses is therefore a valuable contributor to the bottom line.

KV cache capacity up, cache miss frequency down

How do you avoid cache misses? Oddly, the answer is simple. The best way to avoid KV cache misses is to have a bigger KV cache. This allows you to store more context, so it is easily readable to avoid GPU recomputation. But calling something simple doesn’t necessarily mean it’s easy—or free.

KV cache has historically lived in on-GPU high-bandwidth memory (HBM) with spillover into AI server DRAM. These can be impossible to expand, and difficult and expensive even when you’re able to do so. To get beyond the limits of memory, the next option is to use storage in the KV cache stack. This gives four possible levels of KV cache:

  1. HBM: This is typically fixed and resides within the GPU. It is generally not expandable.
  2. DRAM: Hosted as CPU memory inside the AI server, this offers high performance but is generally subject to expansion limits based on server architecture and very expensive.
  3. Flash: This moves further down the performance envelope from memory, offering sub-millisecond (ms) latency, but is also much less expensive—about an order of magnitude less expensive than DRAM.
  4. HDD: Further down the performance envelope, with latencies in the ms to second range depending on workload, but also about an order of magnitude less expensive per TB than flash.

What is needed is a KV cache solution that achieves high capacity, cost efficiency, and performance high enough to provide a quality of service (QoS) that satisfies customers.

How much storage performance does KV cache actually need?

Here is where it gets interesting. It’s easy to assume that KV cache must use the highest performance memory or storage technology to be worthwhile. But it overlooks a simple question: what is the alternative?

The GPU is the most expensive asset in the AI hardware stack. A cache miss entails asking the GPU to repeat work it has already completed. We’ve highlighted that doing so has both direct costs and opportunity costs. But it’s important to think about it quantitatively, not just qualitatively.

  • Time: We can return to the comparison between a task taking 61 seconds without KV cache vs 11.7 with caching enabled. Conceivably, this means that a retrieval time from KV cache of anything up to 49.3 seconds speeds up the inference exercise. So just as a simple back of the napkin exercise, we can see that we’re not dealing with microsecond or even millisecond latency requirements.
  • Cost: A modern cloud GPU can rent anywhere between $1 and $10/hour. Compare that to a bank of 13 hard disk drives (HDDs) running a 10:3 erasure coding scheme for data durability. At worst-case active power consumption, each drive will consume a little under 10W of power. Even if we assume a 2:1 ratio between HDD power consumption and all power used for ancillary tasks and cooling, that equates to 260W, or 0.26kW. At an average US per-kWh price of 14.48 cents, those 13 drives together will cost less than 4 cents an hour in power to operate. This is more than an order of magnitude difference.

In the direct comparison of time, one can simply conclude that time-saving KV cache, even if it’s a second or two having the data within KV cache, is inherently a good thing. But when adding in the actual cost of GPU vs storage, one can easily make the argument that a storage technology that slowed inference completion would still be worthwhile, because of the incredible cost advantages conferred by avoiding GPU recomputation.

So, performance is important, but even more important is having a high cache hit rate. Because the penalty for data not residing in cache, in both time and dollars, is dramatic. The answer isn’t that faster is automatically better. Rather, it’s that as long as the storage is fast enough, that’s sufficient because the value comes from having the data resident in KV cache rather than from raw performance.

Benchmark results tell the HDD performance story

WD set out to answer the question of whether HDDs can keep up in our technical brief KV Cache Offload Across Storage Tiers: A Controlled Benchmark on an 8x NVIDIA® B200 Inference Host.

For this paper, WD tested combinations of tiered KV caching. This included large main caching tiers of both SSD and HDD configurations. In addition, it included configurations of a staging tier in front of the main capacity: no tier, a 30GiB CPU DRAM staging tier on the host server, or a 15.36TB SSD within the host server for staging.

When no staging tier was included, the SSD configuration outperformed the HDD configurations. The SSDs exhibited a higher cache hit rate, a shorter time to first token (TTFT), and a higher average number of decode tokens/second. However, when either staging tier was enabled, the HDD configurations matched the SSD configuration in all three metrics of hit rate, TTFT, and decode tokens/second. All three with the staging tier also outperformed either SSD or HDD with no staging tier in TTFT and decode tokens/second.

If the goal is to maximize KV cache capacity and increase the potential to have high cache hit rates, the testing strongly implies that a tiered KV cache comprising local DRAM or SSD in front of bulk capacity on HDD is a blend of performance and cost-efficiency.

The upshot: KV cache capacity and hit rate improve AI infrastructure

As AI increasingly transitions from training workloads to inference and agentic workloads, the importance of KV caching is growing. As we’ve shown, the benefits of KV caching are contingent on the data being within the cache, and cache misses carry a massive penalty.

The implication is simple: increasing the KV cache capacity maximizes the likelihood that any given workload will hit, rather than miss, the cache. HDDs are the most cost-effective way to deploy bulk capacity, and WD’s testing suggests that HDDs are capable of serving the KV cache results within the performance windows AI applications require. By deploying large-capacity KV cache, GPU assets can be saved for the highest-value usage rather than repeating already-completed work.

Explore KV Cache Offload Across Storage Tiers