StorageAI: Giving Storage a Seat at AI's Table
In the last few articles in this series, I've obsessed over memory — CXL, SOCAMM, High Bandwidth Flash. All of that is about getting bits closer to compute, faster. But there's a layer sitting quietly behind the memory conversation that rarely gets the same spotlight: storage. And it turns out storage has been an unspoken bottleneck all along.
So today, let's talk about StorageAI, SNIA's open standards project aimed squarely at fixing how storage talks to AI accelerators.
Why Storage Needed Its Own AI Moment
We've spent a lot of time in this series on compute and memory. But there's a third pillar that decides whether a GPU is actually busy or just expensively idle: how fast and how directly data gets from storage into accelerator memory.
SNIA frames the problem plainly. AI exposed urgent gaps in data services, particularly in how storage interacts with compute. Data pipelines are inefficient, with data round-tripping between storage and compute in ways that waste both power and performance. And when data isn't in the right place at the right time, accelerators like GPUs sit idle, which, given what a GPU cluster costs to run, is an expensive way to wait.
AI data today typically sits on separate, siloed storage networks, forcing workloads across inefficient network boundaries, and every stage of the AI pipeline - ingestion, preprocessing, training, checkpointing, inference, has its own demanding storage profile that legacy architectures were never built for. Add in the fact that models and datasets keep growing, and moving data is quickly becoming the dominant cost, not just an inconvenience. Without shared standards, every organization ends up reinventing the same plumbing instead of spending that engineering effort on the AI itself.
This is really an extension of the same memory-wall story I've been writing about all along, just one hop further down the stack. HBM and HBF solve for what sits right next to the GPU die. CXL and SOCAMM solve for the memory tier around the server. StorageAI is SNIA's attempt to solve for the tier after that — the storage systems and data services that feed everything upstream.
What Is StorageAI?
SNIA announced StorageAI on August 4, 2025, describing it as an open standards project for efficient data services related to AI workloads, built on industry-standard, non-proprietary, vendor-neutral approaches. SNIA itself isn't a new name here; it's a not-for-profit that's been developing storage standards for more than 25 years, with efforts like SMI-S, the NVMe Management Interface, Swordfish, and CDMI to its name.
SNIA Chair Dr. J Metz framed the original rationale simply: the demands of AI require a holistic view of the data pipeline, from storage and memory to networking and processing, and no single company can solve that alone.
At launch, early press coverage (Blocks & Files among others) grouped SNIA's initial technical scope into six areas: AiSIO (Accelerator-Initiated Storage I/O), CNM (Compute-Near-Memory), FDO (Flexible Data Placement), GDB (GPU Direct Bypass), NVMP (the NVM Programming Model), and SDXI (the Smart Data Accelerator Interface). Eight months on, the StorageAI Chair's own keynote lays out a fuller and more current picture of the technical focus areas the community is actually working:
Smart Data Accelerator Interface (SDXI): standardized, processor-agnostic memory-to-memory data movement that skips redundant software layers
Accelerator-Initiated I/O & GPU Direct Access: letting GPUs pull data straight from storage without routing through the CPU
RDMA & High-Performance Fabrics: running File, Object, Block, and KV storage protocols over RDMA and emerging fabrics like Ultra Ethernet
StorageAI Infrastructure & Data Discovery: extending Swordfish, Redfish, and CDMI so AI agents can discover, provision, and manage storage on their own
Agentic AI Optimization: work on the Model Context Protocol (MCP) and Agent-to-Agent (A2A) communication, turning storage into an active participant in AI workflows rather than a passive repository
Data Resilience, Lifecycle & Sanitization: tackling newer problems like ephemeral data, rack-scale failure domains, machine unlearning, and the “right to be forgotten”
Power & Efficiency: measuring and optimizing the energy footprint of storage under AI-specific workloads
Computational Storage & Memory Technologies: pushing compute closer to data via computational storage, CXL-attached memory, and persistent memory
That's a meaningfully broader remit than the launch-day summary suggested; the agentic AI and data-sanitization items in particular weren't part of the original public framing, and they tell you this project is trying to stay ahead of where AI workloads are heading, not just where they are today.
Where Storage Actually Touches the AI Pipeline
One of the more useful things in Duquette's keynote is a straightforward map of where storage sits inside a typical AI pipeline. Underneath the familiar stages: ingestion, data prep, training, evaluation, serving, monitoring, there are six concrete storage touchpoints: the raw data lake, a feature store, checkpoints during training, a model registry, the KV cache used during inference, and logs and metrics for monitoring.
Each of those touchpoints already leans on a different mix of the same underlying protocols - NVMe and NVMe-oF, RDMA, GPUDirect Storage, CXL, SDXI, S3 and object storage, parallel file systems, Swordfish and Redfish. Seeing it laid out this way makes the case for StorageAI more concrete: it's not one bottleneck to fix, it's six of them, all leaning on overlapping infrastructure.
The USP: Coordination, Not Reinvention
What I find genuinely useful about StorageAI's approach is that it isn't trying to invent an entirely new protocol stack from a blank page. It's coordinating and aligning specifications that already exist across SNIA and its partner organizations: UEC, NVM Express, OCP, OFA, DMTF, SPEC, and others, into one coherent story for the AI data path.
That's a deliberate choice, and a sensible one. AI infrastructure already has plenty of overlapping standards bodies (I've written about several of them in this series). Rather than adding one more competing spec to the pile, StorageAI positions itself as connective tissue: a vendor-neutral framework for coordinating data services rather than a single new thing to adopt.
Betting on a Pattern That's Repeated Before
Duquette's keynote leans on a history lesson, and it's a fair one. Storage standards have shown up at the base of essentially every major computing shift: PATA/IDE and SCSI underpinned personal computing and early enterprise servers in the 1980s; Fibre Channel, iSCSI, NFS, and USB enabled networked storage and the modern data center in the 1990s; SATA, SAS, and Advanced Format helped carry in cloud computing and the SaaS economy in the 2000s; and NVMe, NVMe-oF, FDP, and ZNS underwrote the flash era that powers today's big data and hyperscale platforms.
The framing StorageAI is pitching is that its own moment - StorageAI, SDXI, CXL, and MCP working together - is the next entry in that same sequence, this time aimed at AI-optimized, increasingly autonomous storage. Duquette's closing line for the keynote sums up the pitch: “a rising tide floats all boats - but only if we build the harbor together.”
It's a compelling pattern, and it's historically accurate as far as it goes. Whether StorageAI actually repeats it is a separate question from whether the pattern exists — more on that below.
Ecosystem and Momentum
The founding roster that signed on at launch included AMD, Cisco, DDN, Dell, IBM, Intel, KIOXIA, Microchip, Micron, NetApp, Pure Storage, Samsung, Seagate, Solidigm, and WEKA - a genuinely broad cross-section of compute, memory, and storage vendors.
Since then, the participant list on SNIA's StorageAI page has grown considerably, now including names like Arm, Broadcom, HPE, Hitachi Vantara, Huawei, Lenovo, Marvell, Supermicro, Western Digital, and Nutanix, alongside research participants such as Los Alamos National Laboratory and university partners. The community hosted its own dedicated event, SDC: StorageAI, held in Denver on April 29, 2026. StorageAI also had a visible presence at SC25's Open Standards Pavilion alongside other alliance partners.
Outlook
Zoom out across this series, and a pattern emerges: CXL, SOCAMM, and HBF have been chipping away at the memory hierarchy. ESUN has been doing the same for scale-up networking. StorageAI is SNIA's attempt to do it for the last mile, the path between storage and the accelerator itself so that all that memory and network innovation isn't wasted waiting on data that hasn't arrived yet.
If SNIA can bring the same multi-decade patience to StorageAI that it brought to NVMe and Swordfish and if the education push actually gets vendors and developers moving in the same direction this could quietly become the plumbing everyone assumes and nobody thinks about, which for an infrastructure standard is about the highest compliment there is. The harbor Duquette described still has to get built, and whether that happens without Nvidia's participation is the open question I'll be watching.
I might have missed some pieces of the StorageAI story — it's moving fast. Please comment and let me know what I've missed; I'll add to this article as the picture fills in.