Don’t Confuse Flash-y with Foundational

Don’t Confuse Flash-y with Foundational

This week’s Future of Memory Storage event is a place where flash gets most of the attention. NAND density, controller innovation, and interface performance are all moving quickly. For the right workloads, flash is indispensable.

But if you are building AI infrastructure today, whether you are an AI lab, a neocloud, or a sovereign cloud, the harder question is not simply which storage medium is fastest. The question is whether the architecture you choose this year will still work when the data you are generating today is petabytes, even exabytes, larger.

That is where many designs fail. Not immediately, and not always visibly. They fail as an economics problem. They fail when retained data grows faster than expected. They fail when a system that looked clean at a few petabytes becomes painful at hundreds of petabytes. And at that point, fixing the architecture is much harder, and way more expensive, than designing it correctly from the beginning.

My view is simple: AI storage will be won by those who deliver economically scalable capacity. Customers are not choosing storage based on recording physics. They are choosing confidence, the confidence to scale AI with capacity, reliability, and minimal operational disruption.

That is the lens I use when I think about flash and HDD. Not as rivals, and not as a religious debate. They are different tools in a system. The work is to match the medium to the workload, and to do it early enough that the economics hold up at scale.

What works at small scale breaks at AI scale

In year one, an all-flash architecture can look reasonable. The datasets are still manageable. GPU procurement gets most of the attention. Latency feels like the only variable that matters.

The problem is that AI does not stay small. Models create data. Inference creates data. Customer interactions create data. Compliance and observability create data. All of it retained for training future models. Simple AI data compounds. Over time, the system becomes less about a single performance tier and more about the full lifecycle of data—how it is created, retained, retrieved, protected, and reused.

This is why so much cloud storage capacity still runs on HDDs today. It is not because the industry has failed to modernize. It is because most storage capacity at scale is bulk, sequential, persistent, and cost-sensitive. It does not need microsecond access. It needs to be reliable, dense, efficient, and affordable for years.

A flash-everywhere design can be attractive in a demo. It can even work in the first phase of deployment. But it becomes the most expensive way to store the data that AI keeps accumulating. At AI scale, that is not a small optimization problem. It becomes a design flaw.

Customers don’t build AI with a single storage tier

One of the biggest misconceptions in the market is that AI infrastructure can be reduced to a single storage technology. Successful customers don’t build AI with a single storage tier. They optimize across performance, economics, power efficiency, and the full data lifecycle.

AI did not simplify storage. It made the storage stack deeper.

Traditional cloud storage often separated hot data from cool and warm data. AI adds more pressure across the whole stack. Some workloads need extreme performance. Some need fast retrieval. Some need retention at enormous scale. Some need to stay online because the next model, the next audit, or the next customer interaction may depend on them.

A practical AI storage architecture looks more like this:

  • Model weights and GPU overflow need very high performance and low latency. Flash is the right answer there.
  • KV cache and session context create a new performance-sensitive tier. This is where flash innovation is important, and it is why FMS is focused on the right problem this week.
  • Vector databases, embeddings, and RAG workloads need fast lookup over indexes that often point back to a much larger source corpus.
  • Bulk storage is the largest tier by far: training corpora, logs, checkpoints, compliance records, synthetic data, millions of user contexts, and inference output. This is where HDD economics matter most.

The important point is not that one medium wins. The important point is that the architecture has to recognize what each tier is structurally suited to do. Flash handles the moment. HDDs handle the lifetime. A system that confuses those roles may be simple to explain, but it will be unequivocally expensive to operate.

At scale, economics becomes architecture

Below a certain scale, cost can look like a line item. At AI scale, cost becomes architecture. It determines how much data you can afford to retain, how much data you can bring back into the next training run, how many users you can serve while keeping prices affordable and the business profitable, and whether the system can keep growing without consuming the rest of the budget.

This is not just a media price discussion. It is a capital allocation discussion. Every dollar spent storing long-lived, bulk data on a premium performance tier is a dollar not available for compute, networking, power, people, or the next generation of infrastructure.

That matters because AI data does not behave like a small application dataset. Inference is a continuous write. Session state is created, cooled, and retained. Logs and compliance records may need to live for years. Synthetic data gets curated back into training. The data flywheel only works if the storage foundation is economically sound.

This is where engineering realism matters. The right question is not, “Can flash store it?” Of course it can. The right question is, “Can customers afford to store all of it on flash, operate it for years, and still build a sustainable AI business?” For many bulk workloads, the answer is no.

Resilience is not something you add later

At hyperscale, failure is not an exception. Failure comes from software, hardware, networks, power, and storage. It is the steady state. Somewhere in the fleet, something is always failing. The job of the architecture is not to pretend failure can be eliminated. The job is to absorb it, isolate it, recover from it, and keep the customer’s system running.

That is true across media types. But each medium brings different physics into the operating model. HDDs wear mechanically and predictably in well-managed fleets. Flash wears by use; every write, garbage-collection cycle, and wear-leveling operation consumes endurance. Those differences matter when a system runs continuously for years.

This is why I do not think about resilience as a feature that gets added after the architecture is selected. It is a design choice made up front. If the workload is long-lived, bulk, and continuously written, then the medium has to match that reality. Quality, reliability, power, serviceability, and customer operations all have to be considered together.

Customers do not want a storage supplier to win a benchmark and then introduce operational risk. They want innovation that does not break their business. That means smooth transitions, predictable behavior, and technology that proves itself in real customer environments.

You cannot scale AI without scaling data efficiently

AI improves when it has more useful data. But that only helps if the data can be stored, managed, protected, and accessed economically. Otherwise, the data you cannot afford to retain becomes the data you cannot learn from.

This is especially important because AI generates data at every stage. Training creates checkpoints and artifacts. Fine-tuning creates versions. Inference creates session history, logs, and outputs. Compliance creates retained records. The next model may depend on some of that data being available again, even if it is no longer hot.

For AI labs, this becomes a compounding retention problem. For neoclouds, it becomes a margin problem. For sovereign clouds, it becomes a compliance and durability problem. The details are different, but the architecture principle is the same: the storage foundation must be built for the economics of the workload, not just the excitement of the moment.

This is why I often describe storage as the foundation of the AI data economy. It is not the most visible part of the AI stack, but if it is wrong, everything built on top of it becomes harder to sustain.

The stack keeps expanding

This is not a simple “add more HDD” argument. Customers are not building AI infrastructure with one tier. They are optimizing across performance, cost, power, density, availability, and operational complexity.

That is why the future is a broader portfolio of fit-for-purpose technologies. Some workloads need flash. Some need HDD capacity. Some need more throughput while preserving HDD-class economics. Some need lower-power capacity at the colder end of the stack. Some need archival solutions for data that outlives the workload that created it.

The design principle is consistent across all of these tiers: match the medium to the economics and behavior of the workload. Do not default to the fastest medium because it feels safe. Do not default to the lowest-cost medium when the workload needs performance. Engineer the system deliberately, tier by tier.

That is how customers reduce risk. It is also how they preserve choice. A good architecture gives customers a way to grow without forcing a disruptive redesign every time the dataset doubles.

The architecture question

Flash will get attention this week, and it makes sense to discuss. There are important AI workloads where flash is the right answer. But customers do not build AI around a benchmark. They build AI around outcomes. They need the confidence that their infrastructure can continue scaling as data volumes compound year after year.

But the question that will determine whether an AI lab, neocloud, or sovereign cloud is still economically sound in three years is bigger than flash performance. It is whether the bulk tier was designed in from the beginning, or discovered later as a budget problem after the data had already compounded.

If you design this wrong, it fails at scale and, quite frankly, breaks your business.

If you design it right—tiered by workload, grounded in economics, built for resilience, and operated in diligently—it becomes a foundation. Not a flash-y foundation. A real one. The kind that lets the next model, the next application, and the next generation of AI infrastructure keep growing.

That is the work in front of us. It is not about flash versus HDD. It is about building storage architectures that work in the real world, for real customers, at the scale AI is now demanding.