Enterprise

NVIDIA Opens cuFile APIs, Pushes Storage-as-Memory for AI

At FMS 2026, the chip giant unveiled open-source GPU storage access and a new industry initiative to eliminate bottlenecks as AI agents demand microsecond data retrieval.

Omega Editorial· August 4, 2026· 3 min read

NVIDIA Opens cuFile APIs, Pushes Storage-as-Memory for AI

As AI models consume datasets too large to fit in system memory, NVIDIA is redefining the boundary between storage and RAM — turning drives into an extension of memory that GPUs can access in microseconds rather than milliseconds.

At the Future of Memory and Storage conference this week, the company announced it is open-sourcing cuFile, the application programming interface that lets GPUs read from and write to storage directly without routing requests through CPUs. The move aims to make GPU-accelerated storage interoperable across hardware and software platforms, according to details first reported by NVIDIA.

Why it matters

AI agents and large-context models are generating thousands of concurrent storage requests that traditional architectures can't serve fast enough. By collapsing the performance gap between memory and storage, enterprises can run more capable AI workloads without proportionally scaling expensive high-bandwidth memory — a shift that changes both the economics and the technical ceiling of AI infrastructure.

Direct GPU storage access becomes open standard

cuFile is a component of NVIDIA GPUDirect Storage. It uses hundreds of thousands of GPU threads and high-bandwidth memory to enable data access from storage in microseconds. The APIs are now hosted as open source with Google, Intel, NVIDIA, and Meta serving as inaugural maintainers.

The open-source release supports interoperability between GPUs and diverse storage systems while maintaining Linux security protocols. NVIDIA positions the move as foundational for AI-powered cybersecurity, where defenses need storage access at speeds matching threat detection algorithms. The technology aligns with the newly formed Open Secure AI Alliance.

Storage-Next initiative tackles AI bottlenecks

NVIDIA also launched Storage-Next, an industry initiative bringing together more than 40 storage and flash vendors — including DDN, KIOXIA, and Micron — to define how GPU-driven storage should behave and codify those practices into open standards.

Central to the effort is SCADA (scaled, accelerated data access), a framework that allows massively parallel GPUs to pull only necessary data directly from storage into their own memory. DDN is integrating SCADA with Infinia, its AI-native data intelligence platform designed to eliminate storage bottlenecks at scale.

"AI success will be defined not by how much infrastructure organizations own, but by how productively they use it," said Sven Oehme, chief technology officer at DDN.

Vera CPU delivers 3x throughput for data services

Storage systems serving AI workloads must continuously encrypt, compress, verify, and reconstruct data — operations that become bottlenecks when thousands of agents access storage simultaneously. Benchmarks highlighted in an NVIDIA technical blog show the NVIDIA Vera CPU, part of NVIDIA Vera BlueField-4 STX, delivers up to 3.21 times higher throughput than x86 CPUs in two-stage compression and encryption pipelines.

The Vera BlueField-4 STX storage processor is part of a modular, rack-scale foundation powered by the NVIDIA Vera Rubin platform and NVIDIA Spectrum-X Ethernet networking. The system uses the unified NVIDIA DOCA security stack to enable continuous policy enforcement in the AI data path.

Security by design

NVIDIA SCADA addresses the security risk inherent in letting applications talk directly to drives. The framework splits the job into two components: user-level application parts that need speed stay outside the trusted computing base, while a separate privileged component configures protected access between the application and approved storage at setup, adhering to standard Linux security protocols.

NVIDIA CMX Context Memory Storage provides an AI-native context tier for long-context, multi-turn, agentic AI inference, built on the STX platform.

The announcements were made at FMS, running August 4-6, 2026, in Santa Clara, California, as reported by NVIDIA.

#gpu storage#cufile#nvidia bluefield#ai infrastructure#open source#storage-next

This is an original analysis by the Omega editorial team. Source reporting: AI Watch.

Want systems like this working for your business?

Book a Call

More in Enterprise

Enterprise· 2 min read

Palantir Stock Jumps 15% on Strong AI Revenue Growth

US commercial revenue climbed 40% year-over-year as enterprise customers accelerate artificial intelligence deployments.

Via AI Watch · Aug 4, 2026
Enterprise· 4 min read

AI Agents Need Management, Not Just Prompts, Companies Learn

As autonomous AI systems take on work independently, employees must shift from using AI tools to actively supervising them—a capability gap most organizations haven't addressed.

Via AI Watch · Aug 4, 2026
Enterprise· 4 min read

AI Productivity Gains Are Creating Faster Burnout, Not Free Time

New research shows employees save two hours daily with AI tools, but organizations are converting those gains into higher output expectations rather than breathing room.

Via AI Watch · Aug 4, 2026