StorageReview.com

MLPerf Storage v3.0: 877 GiB/s Checkpoints, a Cloud First, and a Leaderboard Turned Over

AI  ◇  Enterprise

MLCommons published the MLPerf Storage v3.0 results today, and the round changes not only what the benchmark is but also who leads it. For the first time, the only audited AI storage benchmark covers the full pipeline: training throughput, checkpointing, and two new inference workloads, a vector database test and a KV cache test. Nineteen organizations submitted 143 results, eleven of them first-time submitters, including Microsoft Azure, NVIDIA, Everpure, and a wave of AI storage startups most readers will not recognize. Just as notable is that DDN, Huawei, Hammerspace, and Lightbits, the names that defined the v2.0 leaderboard, did not submit this round.

One comparability note up front. v3.0 moved its simulated accelerators from H100 to B200, which raises the per-accelerator bandwidth bar, so v3.0 numbers cannot be compared against v2.0 results. These are MLCommons-audited vendor submissions; StorageReview did not conduct or independently verify this testing.

Checkpointing Is the Headline Event

MLCommons added checkpointing as a first-class workload because it has become a verified bottleneck: frontier-scale training jobs write enormous amounts of state at regular intervals, and every second spent flushing a checkpoint is a second of idle accelerator time. The largest number in the round belongs here. Everpure (formerly Pure Storage), in its first MLPerf Storage appearance, submitted FlashBlade//EXA results that scale almost linearly with data node count: 327.6 GiB/s of checkpoint write at 10 data nodes, 484.1 GiB/s at 15, 655.6 GiB/s at 20, and 877.5 GiB/s write with 833.0 GiB/s read at 30 data nodes. Everpure has spent the past year making large throughput claims for //EXA; this is the first time an audited result backs the direction of those claims.

Bar chart of MLPerf Storage v3.0 checkpoint write scaling for Everpure FlashBlade//EXA, from 327.6 GiB/s at 10 data nodes to 877.5 GiB/s at 30

The second-largest checkpoint number is arguably the bigger industry story. Azure Managed Lustre, the first hyperscale cloud storage service ever submitted to the benchmark, posted 642.2 GiB/s of checkpoint write from a 4,096 TiB deployment with 128 clients. A managed cloud file service putting up numbers in the same conversation as purpose-built AI storage arrays would have been hard to picture two rounds ago.

Checkpointing (Top Submissions) System Write B/W (GiB/s)
Everpure FlashBlade//EXA, 30 data nodes 877.5
Azure Managed Lustre, 4,096TiB, 128 clients 642.2
YanRongTech F9000X, 8 clients 313.7
Suzhou Zishan Longlin Multiple configurations 313.7
TuringData F9200, 8 clients 307.1
UBIX UbiPower18000, 3 storage nodes 293.8

Training: Names You Likely Don’t Know

The 3D U-Net training leaderboard belongs to newcomers. YanRongTech’s F9000X sustained 543.9 GiB/s of read bandwidth feeding 99 simulated B200 accelerators, the highest training throughput of the round. TuringData, a Singapore AI infrastructure company making its first submission, landed 4 GiB/s behind at 541.5 GiB/s while feeding 96 simulated B200s from just three storage nodes, and the same three-node F9200 cluster also submitted Checkpoint-70B at 539.9 GiB/s read and 307.1 GiB/s write. UBIX’s UbiPower18000, three storage nodes with 16 x 15.36TB NVMe drives and quad NDR400 networking, hit 454.1 GiB/s. HPE was the only established enterprise array vendor to submit training results, with the K3000 at 345.8 GiB/s.

3D U-Net Training (Top Submissions) Read B/W (GiB/s) Simulated B200 Accelerators
YanRongTech F9000X 543.9 99
TuringData F9200 541.5 96
UBIX UbiPower18000 454.1 80
Suzhou Zishan Longlin 433.4 80
FarmGPU 406.7 75
Azure Managed Lustre 379.1 70

Inference Joins the Benchmark

The two new workloads acknowledge where storage demand is actually growing. The KV cache test measures how quickly a storage system can page the attention state in and out for long-context LLM serving, a pattern we examined in depth in our Dell and Solidigm KV cache offload work. Everpure again posted the round’s largest number, 1,623.4 GiB/s of KV cache read on a 51-host FlashBlade//EXA configuration, with UBIX (443.7 GiB/s) and FarmGPU (443.6 GiB/s) leading the standard-scale submissions. On the vector database side, TTA’s Seahorse system topped the query throughput chart at 57,620 queries per second.

The other structural change is protocol. v3.0 added S3 object storage support, and roughly one-sixth of submissions used it. NVIDIA accounted for most of that with 20 AIStore submissions spanning bare-metal configurations on OCI, AWS, and GCP, topping out at 133.3 GiB/s of checkpoint write. The absolute numbers trail the parallel file systems, but object storage running training and checkpoint workloads within the benchmark marks territory that POSIX file systems had to themselves a round ago.

MLCommons also began tracking power and rack-space efficiency as first-class metrics this round.

Everpure FlashBlade//EXA

Everpure FlashBlade//EXA system held the lead across the 405B and 1.25T parameter model checkpointing categories as well as Key-Value (KV) cache retrieval benchmarks, highlighting the platform’s ability to maintain high throughput during data-intensive artificial intelligence training and inference workloads.

Everpure FlashBlade//EXA architecture diagram showing a compute cluster connected to a metadata core and data nodes over RDMA

Everpure’s submitted configuration paired 30 FlashBlade//EXA blades (120 DirectFlash Modules handling metadata) with 30 Linux/NVMe data nodes for roughly 866 TB of usable capacity in a single file system, splitting metadata over pNFS/TCP and bulk data over NFSv3 on RDMA. In the 1.25T parameter checkpointing simulation with 1,024 client accelerators, the 30-node configuration sustained 877.52 GiB/s of write bandwidth over a 17.74-second operation, alongside 588.28 GiB/s of read bandwidth completed in 28.99 seconds. On the KV cache inference test it reported 85,736 tokens/sec on a storage-only Llama 3.1 8B run, 67,642 tokens/sec on an 8B run combining storage and system memory, and 33,403 tokens/sec on a 70B storage-only run.

What the No-Shows Mean

A benchmark round is evidence of two things: what the submitters can do and what the absentees chose not to show. VAST Data, WEKA, and DDN, the vendors winning the largest neocloud and AI factory deployments, all sat v3.0 out, as did v2.0 submitters Huawei, Hammerspace, and Lightbits. Vendors skip rounds for mundane reasons, engineering cycles, release timing, and benchmark politics; a skipped round is not evidence that a system is slow. But it does mean the audited record and the deployment record now point to different vendor lists, and buyers have to weigh both. Our Best Storage Arrays page tracks exactly that split, labeling what is audited and what is deployment evidence, and has been updated with the full v3.0 results.

The v3.0 results are public today on the MLCommons storage benchmark page.

Engage with StorageReview

Newsletter | YouTube | Podcast iTunes/Spotify | Instagram | Twitter | TikTok | RSS Feed

Brian Beeler

Brian is located in Cincinnati, Ohio and is the chief analyst and President of StorageReview.com.