Enterprise

NVIDIA Opens cuFile APIs, Pushes Storage-as-Memory for AI

At FMS 2026, the chip giant unveiled open-source GPU storage access and a new industry initiative to eliminate bottlenecks as AI agents demand microsecond data retrieval.

Omega Editorial· August 4, 2026· 3 min read

NVIDIA Opens cuFile APIs, Pushes Storage-as-Memory for AI

As AI models consume datasets too large to fit in system memory, NVIDIA is redefining the boundary between storage and RAM — turning drives into an extension of memory that GPUs can access in microseconds rather than milliseconds.

At the Future of Memory and Storage conference this week, the company announced it is open-sourcing cuFile, the application programming interface that lets GPUs read from and write to storage directly without routing requests through CPUs. The move aims to make GPU-accelerated storage interoperable across hardware and software platforms, according to details first reported by NVIDIA.

Why it matters

AI agents and large-context models are generating thousands of concurrent storage requests that traditional architectures can't serve fast enough. By collapsing the performance gap between memory and storage, enterprises can run more capable AI workloads without proportionally scaling expensive high-bandwidth memory — a shift that changes both the economics and the technical ceiling of AI infrastructure.

Direct GPU storage access becomes open standard

cuFile is a component of NVIDIA GPUDirect Storage. It uses hundreds of thousands of GPU threads and high-bandwidth memory to enable data access from storage in microseconds. The APIs are now hosted as open source with Google, Intel, NVIDIA, and Meta serving as inaugural maintainers.

The open-source release supports interoperability between GPUs and diverse storage systems while maintaining Linux security protocols. NVIDIA positions the move as foundational for AI-powered cybersecurity, where defenses need storage access at speeds matching threat detection algorithms. The technology aligns with the newly formed Open Secure AI Alliance.

Storage-Next initiative tackles AI bottlenecks

NVIDIA also launched Storage-Next, an industry initiative bringing together more than 40 storage and flash vendors — including DDN, KIOXIA, and Micron — to define how GPU-driven storage should behave and codify those practices into open standards.

Central to the effort is SCADA (scaled, accelerated data access), a framework that allows massively parallel GPUs to pull only necessary data directly from storage into their own memory. DDN is integrating SCADA with Infinia, its AI-native data intelligence platform designed to eliminate storage bottlenecks at scale.

"AI success will be defined not by how much infrastructure organizations own, but by how productively they use it," said Sven Oehme, chief technology officer at DDN.

Vera CPU delivers 3x throughput for data services

Storage systems serving AI workloads must continuously encrypt, compress, verify, and reconstruct data — operations that become bottlenecks when thousands of agents access storage simultaneously. Benchmarks highlighted in an NVIDIA technical blog show the NVIDIA Vera CPU, part of NVIDIA Vera BlueField-4 STX, delivers up to 3.21 times higher throughput than x86 CPUs in two-stage compression and encryption pipelines.

The Vera BlueField-4 STX storage processor is part of a modular, rack-scale foundation powered by the NVIDIA Vera Rubin platform and NVIDIA Spectrum-X Ethernet networking. The system uses the unified NVIDIA DOCA security stack to enable continuous policy enforcement in the AI data path.

Security by design

NVIDIA SCADA addresses the security risk inherent in letting applications talk directly to drives. The framework splits the job into two components: user-level application parts that need speed stay outside the trusted computing base, while a separate privileged component configures protected access between the application and approved storage at setup, adhering to standard Linux security protocols.

NVIDIA CMX Context Memory Storage provides an AI-native context tier for long-context, multi-turn, agentic AI inference, built on the STX platform.

The announcements were made at FMS, running August 4-6, 2026, in Santa Clara, California, as reported by NVIDIA.

#gpu storage#cufile#nvidia bluefield#ai infrastructure#open source#storage-next

This is an original analysis by the Omega editorial team. Source reporting: AI Watch.

Want systems like this working for your business?

Book a Call

More in Enterprise

Enterprise· 3 min read

UN launches AI-ready data platform with Google to fix agent accuracy

New system addresses dismal 21% accuracy rate when large language models answer questions about global development statistics.

Via AI Watch · Sep 17, 2026
Enterprise· 3 min read

OpenAI Tests Sponsored AI Agents That Chat for Advertisers

The company's new ad format lets users converse with business-backed bots after clicking ChatGPT ads, with HubSpot and Shopify as launch partners.

Via AI Watch · Sep 17, 2026
Enterprise· 2 min read

CHOP Uses AI to Build Patient-Specific Heart Models in Seconds

Pediatric surgeons now rehearse complex cardiac procedures on virtual replicas generated from medical imaging, reducing planning time from hours to seconds.

Via AI Watch · Sep 17, 2026