ScaleFlux has introduced an AI-optimized SSD platform designed for NVIDIA CMX and other inference architectures that use SSDs as a shared KV-cache tier beyond GPU HBM and host DRAM. The platform combines high-endurance SSD hardware, Flexible Data Placement (FDP) support, and workload telemetry intended to improve data placement, reduce write amplification, and extend effective endurance in high-churn inference environments.
The platform addresses three storage challenges with offloaded KV cache: characterizing real-world workload behavior, separating data by lifecycle, and sustaining heavy write activity without using excess flash capacity to absorb writes.
Long-context inference, shared-prefix reuse, agentic application flows, and retained idle sessions increase the volume of reusable runtime state stored outside GPU memory. Unlike conventional enterprise workloads, KV-cache blocks may be written frequently, retained for varying periods, reactivated after inactivity, and invalidated asynchronously across sessions, workers, and tenants. These patterns increase garbage collection activity and write amplification when data with incompatible lifecycles share the same flash blocks.
ScaleFlux’s high-endurance architecture is designed to deliver 7 to more than 10 effective drive writes per day at five years for KV cache workloads, with the company noting that effective endurance depends on workload characteristics, FDP utilization, and device configuration. Higher effective endurance reduces the raw flash capacity operators must deploy purely to absorb write traffic, leaving more installed capacity available to hold active KV cache and other AI runtime state. ScaleFlux calls that overhead the “endurance tax,” and reducing it is the platform’s core economic argument.
The SSD platform supports over 200 FDP write streams per drive. This lets inference software group data by lifecycle, session, tenant, shared-prefix classification, ownership, or reuse behavior before placing it on flash media. The goal is to write data with similar invalidation patterns together, reducing internal data movement during garbage collection and limiting interference across data classes.
In preliminary controlled testing, ScaleFlux measured more than a twofold reduction in write amplification using lifecycle-aware FDP placement compared with a baseline placement configuration. The company notes that actual results depend on workload characteristics, lifecycle classification, software integration, and device configuration.
“AI inference infrastructure needs SSDs that provide more than additional capacity,” said Hao Zhong, CEO and co-founder of ScaleFlux. “Infrastructure teams need to understand how KV workloads affect the drive, separate data according to lifecycle, and sustain high write rates without deploying excess capacity simply to dilute writes. ScaleFlux brings workload intelligence, scalable FDP placement, and 7-10+ effective DWPD together in one AI-optimized SSD platform.”
At the software and telemetry layer, ScaleFlux Context-Insight SSD shows how KV-cache policies affect SSD operation. The platform captures latency, queue depth, throughput, request-size distribution, data age, write-to-first-read intervals, read reuse, NAND write volume, garbage collection movement, and write amplification.
Context-Insight can operate in SSD-only mode for initial workload analysis without changes to upper software layers. With deeper integration, it correlates SSD telemetry with application metadata, including session IDs, worker or tenant identifiers, shared-prefix IDs, KV-block ownership, lifecycle state, and key-to-block mappings. This lets operators associate latency, endurance consumption, and write amplification with specific workload classes instead of treating the SSD as an opaque shared resource.
ScaleFlux positions the platform as a complement to NVIDIA’s recently announced CMX Context Memory Storage Platform, which provides a shared, pod-level context tier for high-speed KV cache access and reuse. The pitch is aimed at AI factory operators: CMX handles the context tier, while ScaleFlux addresses the endurance, data placement, and write amplification challenges specific to the underlying SSDs.
“As AI inference systems extend KV cache beyond GPU Memory and DRAM, understanding the behavior and requirements of the SSD tier becomes increasingly important,” said Jason Hardy, vice president of storage technology at NVIDIA. “Our engagement with ScaleFlux is helping characterize how KV cache offload affects storage requirements for latency, endurance, and write amplification, contributing to the broader storage ecosystem around NVIDIA CMX.”
The company is developing a trace-driven simulator that models KV-cache movement across GPU HBM, host memory, and SSD tiers. The simulator generates replayable SSD traces to evaluate placement, eviction, and lifecycle-grouping policies under controlled conditions.
ScaleFlux plans to showcase the platform at FMS, covering Context-Insight workload analysis, KV metadata correlation, lifecycle-aware FDP placement, write amplification reduction, and high-endurance operation for write-intensive KV cache workloads. The company has a substantial presence at the show: ScaleFlux said in July that its experts would lead seven presentations there, including a keynote co-delivered with NVIDIA on memory solutions for scaling the AI data pipeline.
The platform announcement caps a busy stretch of silicon news. Two days earlier, ScaleFlux unveiled two PCIe Gen6 parts it will introduce at the same show: the FC6116 NVMe SSD controller and the MC600 CXL 3.2 Type 3 memory controller. ScaleFlux rates the FC6116 at up to 28 GB/s sequential read and 25 GB/s sequential write, up to 7 million 4K random read IOPS and more than 1 million sustained 4K random write IOPS, under 9W active controller power, with support for TLC, QLC, and SLC NAND up to 256TB across E1.S/L, E3.S/L, and U.2/3. The MC600 draws under 9W typical in a Gen6 x8 configuration and handles quad-channel DDR5 or dual-channel DDR4 with up to 2TB of DDR5, a dual-generation capability ScaleFlux positions as a way to carry existing DDR4 into a CXL deployment. Both begin sampling with key customers in Q4 2026.




Amazon