---
description: MLPerf Inference v6.1 adds End-to-End RAG and Edge Agentic tests, debuts Vera Rubin NVL72, and logs a 512-GPU MI355X run. 30 submitters.
title: "MLPerf Inference v6.1: 5.7x Per-Accelerator Gains, a 512-GPU Run, and Vera Rubin&#039;s First Peer-Reviewed Numbers"
image: https://www.storagereview.com/wp-content/uploads/2026/04/StorageReview-Supermicro-H14-AMD-MI350X-1.jpg
---

[![Storage Review]()](https://storagereview.com/)

---

[Facebook](https://www.facebook.com/@storagereview) [X (Twitter)](https://twitter.twitter.com/storagereview) [Instagram](https://www.instagram.com/storagereview/) [LinkedIn](https://www.linkedin.com/company/storagereview-com) [YouTube](https://www.youtube.com/user/storagereview?sub_confirmation=1) [Email](mailto:info@storagereview.com) [Spotify](https://open.spotify.com/show/1y6VnznABhHeOSMOmbDTz0) [Reddit](https://www.reddit.com/r/StorageReview/) [Discord](https://discord.gg/TwMHb4azdC) [RSS Feed](https://www.storagereview.com/rss.xml) [TikTok](https://www.tiktok.com/@storagereview)

[Facebook](https://www.facebook.com/@storagereview) [X (Twitter)](https://twitter.com/storagereview) [Instagram](https://www.instagram.com/storagereview/) [LinkedIn](https://www.linkedin.com/company/storagereview-com) [YouTube](https://www.youtube.com/user/storagereview?sub_confirmation=1) [Email](mailto:info@storagereview.com) [Spotify](https://open.spotify.com/show/1y6VnznABhHeOSMOmbDTz0) [Reddit](https://www.reddit.com/r/StorageReview/) [Discord](https://discord.gg/TwMHb4azdC) [RSS Feed](https://www.storagereview.com/rss.xml) [TikTok](https://www.tiktok.com/@storagereview)

[![StorageReview.com](https://www.storagereview.com/wp-content/uploads/2026/07/Storage-Review-2x.png)](https://www.storagereview.com)

≡ Menu

- [Home](https://www.storagereview.com/)
- [Storage Reviews](https://www.storagereview.com/review)
  - [Consumer Reviews](https://www.storagereview.com/consumer)
  - [Enterprise Reviews](https://www.storagereview.com/enterprise)
  - [Ubiquiti Reviews](https://www.storagereview.com/best/ubiquiti-reviews)
- [SR Merch](https://store.storagereview.com)
- [Leaderboards](https://www.storagereview.com/best)
  - [Best Storage Arrays](https://www.storagereview.com/best/storage-arrays)
  - [Best Enterprise SSDs](https://www.storagereview.com/best/enterprise-ssds)
  - [Best Servers](https://www.storagereview.com/best/servers)
  - [Best Desktops for Local AI](https://www.storagereview.com/best/desktops-local-ai)
  - [Best Laptops for Local AI](https://www.storagereview.com/best/laptops-local-ai)
  - [Best Local LLM Tools](https://www.storagereview.com/best/local-llm-tools)
  - [Agentic AI Hardware](https://www.storagereview.com/best/agentic-ai-hardware)
  - [Best SSDs & Hard Drives](https://www.storagereview.com/best_drives)
  - [Best Portable SSDs](https://www.storagereview.com/best/portable-ssds)
  - [Best Business Laptops](https://www.storagereview.com/best/business-laptops)
  - [Best Mobile Workstations](https://www.storagereview.com/best/mobile-workstations)
  - [Best Desktop Workstations](https://www.storagereview.com/best/desktop-workstations)
  - [Best Laptop Battery Life](https://www.storagereview.com/best/laptop-battery-life)
- [Storage Reference Guide](https://www.storagereview.com/storage-reference-guide)
- [About SR](https://www.storagereview.com/about-storagereview)
  - [StorageReview.com Sweepstakes Rules and Regulations](https://www.storagereview.com/storagereview-com-sweepstakes-rules-and-regulations)

Search

[Home](https://www.storagereview.com/) » [News](https://www.storagereview.com/news) » MLPerf Inference v6.1: 5.7x Per-Accelerator Gains, a 512-GPU Run, and Vera Rubin’s First Peer-Reviewed Numbers

# MLPerf Inference v6.1: 5.7x Per-Accelerator Gains, a 512-GPU Run, and Vera Rubin’s First Peer-Reviewed Numbers

by Harold Fritts on September 16, 2026

[AI](https://www.storagereview.com/enterprise/ai)  ◇  [Enterprise](https://www.storagereview.com/enterprise)

MLCommons has published MLPerf Inference v6.1, and the round sets a participation record with 30 submitting organizations and 486 datacenter and edge results. Two new tests join the suite: an End-to-End RAG pipeline for the datacenter and an Edge Agentic Inference benchmark for single-user devices, and the results carry the first peer-reviewed numbers for NVIDIA’s Vera Rubin NVL72, AMD’s Instinct MI350P, Intel’s Arc Pro B70, and AMD’s Ryzen AI Max+ 395. On the pace of improvement, MLCommons says the best per-accelerator DeepSeek-R1 result in the server scenario is 5.7x better than in v5.1 a year ago, and the best VLM result improved 2.99x in the six months since v6.0.

![Supermicro H14 server with AMD Instinct MI350X GPUs, front view showing the fan wall and NVMe drive bays](https://www.storagereview.com/wp-content/uploads/2026/04/StorageReview-Supermicro-H14-AMD-MI350X-1.jpg)

## Two New Tests: End-to-End RAG and Edge Agentic Inference

The End-to-End RAG benchmark measures a complete question-answering pipeline, several models and a vector database working together. The reference implementation runs four models together: gpt-oss-120B handles query decomposition, sufficiency checking, and answer generation; gpt-oss-20B grades retrieved documents; e5-base-v2 produces embeddings; and ColBERTv2 reranks passages. The corpus comprises 107,484 passages, chunked from 2,515 HTML files; the questions are 824 multi-hop tasks from Google’s FRAMES dataset; and each task can loop through up to 5 retrieval rounds before the pipeline decides it has enough evidence. Two metrics come out: documents per second for building the FAISS HNSW vector database, and tasks per second for answering questions against it. A Llama 3.1-8B judge scores the final answers against a 97% accuracy target, and the judging isn’t timed.

The Edge Agentic Inference benchmark targets the coding-assistant pattern that has moved onto workstations and desktop AI boxes. The model is Qwen3.6-27B with thinking off, run as a Q4\_K\_M GGUF under llama.cpp in the reference, with a 32K context window served per turn. The performance workload is a recorded replay of 20 agentic coding trajectories drawn from SWE-bench Verified, totaling 1,007 turns, driven in a single stream with one request in flight, the way a developer on a laptop runs an agent. The reported metric is mean latency per turn, with time-to-first-token and time-per-output-token distributions alongside, and accuracy is gated separately by BFCL v4 at 97% of the reference score. MLCommons adapted the framework from its upcoming MLPerf Agentic datacenter benchmark.

“We added the End-to-end RAG test because it’s clear that query-answering has evolved beyond simply an LLM trained on a corpus; stakeholders need to understand the real-world performance of the types of multi-step, multi-component pipelines that are being built today,” said Miro Hodak, MLPerf Inference working group co-chair. “Likewise, we added the Edge Agentic Inference test because complex inference systems with agentic properties are increasingly hosted on edge computing devices, creating a new set of performance challenges our customers face today.”

The round also extends speculative decoding support, previously limited to DeepSeek-R1, to the GPT-OSS benchmark in the interactive scenario, and defines a new interactive scenario for the VLM test with responses targeted at about 1.5 seconds.

## New Silicon From the Desktop to the Rack

NVIDIA’s Vera Rubin NVL72 makes its MLPerf debut in the preview category, submitted by NVIDIA and by Nebius on its VR200 NVL72. NVIDIA reports up to 2.5x higher token throughput than the GB300 NVL72 on DeepSeek-R1 across offline, server, and interactive scenarios using TensorRT-LLM, and up to 3.7x on the Qwen3 vision-language model using vLLM with NVIDIA Dynamo. Those are NVIDIA’s comparisons against its own prior generation, and the preview designation means the platform is expected to be commercially available by the next round.

AMD expanded its MLPerf Inference 6.1 submission to six model families across language, reasoning, text-to-video, and recommendation tasks, deploying the Instinct MI355X, MI350X, and the new MI350P PCIe card. The 512-GPU Crusoe cluster built on the MI355X is covered below, along with AMD’s own breakdown of the round.

Intel’s Arc Pro B70 shows up in a four-GPU node with 128GB of combined VRAM that Intel used for Llama 3.1-8B, Llama 2-70B, gpt-oss-120B, Whisper, and the new E2E-RAG test; Intel reports gpt-oss-120B improved 36% in server and 27% in offline over v6.0 on the same hardware. On the CPU side, Intel says Xeon 6980P Llama 3.1-8B server throughput rose 2.4x from v6.0 on identical silicon, a software-only gain.

The Ryzen AI Max+ 395 appears through Atlas Inference, a first-time submitter that ran the new Edge Agentic workload on both an NVIDIA DGX Spark and an AMD Strix Halo desktop with the same engine and quantization recipe. Atlas reports 20.1 tokens per second on the DGX Spark, completing all 1,007 turns in under 64 minutes, and 19.63 tokens per second on Strix Halo. That’s a narrower gap than we measured between the two platforms with off-the-shelf runtimes in our [Ryzen AI Halo](https://www.storagereview.com/review/amd-ryzen-ai-halo-review-a-dual-os-200b-parameter-desktop-takes-on-the-dgx-spark) and  [DGX Spark reviews](https://www.storagereview.com/review/nvidia-dgx-spark-review-the-ai-appliance-bringing-datacenter-capabilities-to-desktops), and it’s the kind of result the new benchmark is designed to surface.

![AMD Ryzen AI Max+ 395 Strix Halo mainboard from the Ryzen AI Halo desktop with the SoC and LPDDR5X packages exposed]()

## Bigger, More Diverse, and More Distributed

Multi-node submissions hit a record this round, up from zero in v4.0, and three stand out. Crusoe, another first-time submitter, ran the largest system in MLPerf Inference history with AMD: 512 Instinct MI355X GPUs across 64 nodes on a standard RoCE Ethernet fabric, submitted for gpt-oss-120b and DeepSeek-R1. AMD reports 5.75 million tokens per second in the offline scenario and 5.39 million in server on gpt-oss-120b from that cluster, and 2.90 million offline and 2.41 million server on DeepSeek-R1, which AMD calls the highest aggregate token throughput in MLPerf history, with throughput scaling near-linearly from 1 to 64 nodes. The gpt-oss-120b run served the model as 512 independent single-GPU replicas in native MXFP4; DeepSeek-R1 used SGLang with one eight-GPU replica per node. AMD separately cites a 72-GPU GPT-OSS-120B submission at 95% scaling efficiency and 1 million tokens per second, the same headline it hit on [MI355X in v6.0](https://www.storagereview.com/news/amd-instinct-mi355x-achieves-mlperf-inference-v6-0-gains-with-over-1-million-tokens-per-second-and-supports-scalable-rocm-stack).

Cisco submitted the benchmark’s first cross-vendor heterogeneous system, pooling eight NVIDIA H200 and eight AMD Instinct MI350X GPUs into a single inference pool over a Cisco Silicon One G200 fabric, the same mixed-accelerator approach it’s selling through its [Secure AI Factory](https://www.storagereview.com/news/cisco-secure-ai-factory-adds-supermicro-rack-scale-compute-with-vera-rubin-nvl72-support-and-october-availability). The geographically distributed entry came from MangoBoost with Dell: four sites on two continents, hosted by MangoBoost, Dell, TensorWave, and Microsoft Azure, spanning the Pacific and running as one endpoint at what MangoBoost reports as 97% scaling efficiency. MangoBoost also claims the first prefill/decode-disaggregated results on AMD Instinct GPUs.

Elsewhere in the results, CoreWeave reports 1.16 million tokens per second aggregate on GB300 NVL72 with per-GPU offline throughput up 17% since v6.0, HPE cites 8,500 tokens per second per GPU on DeepSeek-R1 across two Compute XD690 servers with Blackwell Ultra, and Google focused its submission on DeepSeek-R1, citing the industry’s shift to large mixture-of-experts models. gpt-oss-120b drew 112 submissions, the most of any MLPerf workload.

## AMD’s Two CDNA 4 Form Factors: OAM and Dual-Slot PCIe

AMD’s CDNA 4 architecture is available in two physical form factors to meet specific data center power and cooling constraints. The flagship Instinct MI355X targets high-density compute nodes using an OAM form factor on an OCP Universal Baseboard (UBB 2.0) platform, providing 256 compute units, 288 GB of HBM3E memory, and 8 TB/s of aggregate memory bandwidth. It delivers up to 10.1 PFLOPS of peak theoretical MXFP4 and MXFP6 matrix compute. The newly introduced Instinct MI350P adapts that same CDNA 4 silicon into a dual-slot PCIe 5.0 add-in card for standard enterprise chassis, housing 128 compute units, 144 GB of HBM3E memory, 4 TB/s of bandwidth, and delivering up to 4.6 PFLOPS of peak theoretical MXFP4 and MXFP6 matrix performance.

![AMD slide comparing the Instinct MI355X OAM GPU (256 CUs, 288GB HBM3E, 8TB/s, 10.1 PFLOPS MXFP4) with the Instinct MI350P dual-slot PCIe card (128 CUs, 144GB HBM3E, 4TB/s, 4.6 PFLOPS MXFP4)]()

## Software-Driven Generational Uplift on Identical Silicon

Software optimizations in AMD ROCm v7 yielded measurable throughput gains on identical MI355X hardware during a single MLPerf cycle. Testing on an eight-GPU MI355X node showed GPT-OSS-120B throughput rose by 28% in the Offline scenario and 38% in the Server scenario, while Wan-2.2 SingleStream performance improved by 70%. At cluster scale, 72 MI355X accelerators in MLPerf 6.1 achieved higher aggregate GPT-OSS-120B throughput than a 94-GPU configuration reported in round 6.0. In competitive comparisons published by AMD, the eight-GPU MI355X led submitted results against the NVIDIA B200 and B300 on GPT-OSS-120B, while the 72-GPU cluster led NVIDIA’s GB200 submission.

In its first MLPerf round, the dual-slot MI350P submitted across five Closed workloads, posting leading results against selected submissions of the NVIDIA RTX PRO 6000 Server Edition and H200 NVL. For deployments sensitive to power and cooling budgets, an eight-GPU MI350X system maintained approximately 80% of the MI355X platform’s benchmark performance across GPT-OSS, Llama, Wan, and DLRM workloads, while the MI355X carries a 40% higher rated TDP. A commissioned study by Signal65 reported that these throughput numbers translated to lower operational cost per document and higher token output per dollar within fixed latency limits.

![AMD slide showing an eight-GPU Instinct MI350X platform retaining about 80% of MI355X performance across GPT-OSS, Llama, Wan, and DLRM workloads with the MI355X carrying a 40% higher rated TDP]()

## AMD Partner Results and the Korea-to-US Cluster

AMD says comparable MI355X submissions from seven partners, Dell Technologies, Oracle, Hewlett Packard Enterprise, Supermicro, MangoBoost, Crusoe, and MiTAC, averaged within 4% of its reference system results, with some runs matching or slightly exceeding them.

![AMD slide on the first 32-GPU heterogeneous MLPerf inference submission by Dell and MangoBoost: 16 Instinct MI300X GPUs in Korea and 16 MI355X GPUs in the US serving GPT-OSS-120B at 285,454 offline and 253,501 server tokens per second]()

The MangoBoost and Dell entry noted above is the benchmark’s first heterogeneous 32-GPU serving configuration, bridging 16 previous-generation Instinct MI300X GPUs in Korea with 16 Instinct MI355X GPUs in the United States into a unified GPT-OSS-120B endpoint. The split-cluster configuration delivered 285,454 Offline tokens per second and 253,501 Server tokens per second. The submissions ran on ROCm 7; AMD points to the [ROCm 10 release](https://www.storagereview.com/news/amd-rocm-10-arrives-with-rocm-ai-ga-hyperloom-agents-amd-skills-and-a-claimed-3-3x-inference-lift), with vLLM, SGLang, and ROCm.AI profiling tools, as the path forward ahead of the [HBM4-based MI400 Series](https://www.storagereview.com/news/amd-mi455x-and-helios-432gb-hbm4-72-gpu-racks-and-a-real-answer-to-vera-rubin) and the MI500 generation that follows.

## The Harness That Replaces MLPerf Inference in the Datacenter

Sixteen of the 30 submitters used MLPerf’s API-centric harness this round, up from a single open-division submitter in v6.0. The harness runs a true client/server setup over standard APIs against a hosted endpoint, which is how datacenter inference is deployed, and it already carries the new Edge Agentic test, VLM-Interactive, gpt-oss, DeepSeek-R1, Llama 3.1-8B, and text-to-video. It’s the foundation of MLPerf Endpoints, which opens on-demand rolling submissions in October 2026 and replaces MLPerf Inference as the datacenter benchmark in 2027, with normalized results and expanded agentic workloads planned for Endpoints v1.0.

“Moving forward, MLPerf Endpoints will replace Inference in our family of benchmarks for the datacenter, and the quick uptake of our API-centric harness will contribute to making that transition seamless,” said David Kanter, head of MLPerf. The six first-time submitters this round are Atlas Inference, Crusoe, Orrick Industries, ScitiX, VibeHPC, and individual contributor Naeem Khoshnevis of Harvard’s Kempner Institute, who submitted a single-H200 Llama 3.1-8B result.

The inference round follows the [MLPerf Storage v3.0 results](https://www.storagereview.com/news/mlperf-storage-v3-0-877-gib-s-checkpoints-a-cloud-first-and-a-leaderboard-turned-over) published two weeks ago, and full datacenter and edge tables, along with submitter supplementals, are available on the MLCommons results pages.

### [MLPerf Inference v6.1 Datacenter Results](https://mlcommons.org/benchmarks/inference-datacenter/)

### [MLPerf Inference v6.1 Edge Results](https://mlcommons.org/benchmarks/inference-edge/)

**Engage with StorageReview**

[Newsletter](https://www.storagereview.com/storage_newsletter) |  [YouTube](https://www.youtube.com/user/StorageReview "Opens in a new window") | Podcast  [iTunes](https://podcasts.apple.com/gb/podcast/storagereview-com-storage-reviews/id1060681115 "Opens in a new window")/ [Spotify](https://open.spotify.com/show/1y6VnznABhHeOSMOmbDTz0 "Opens in a new window") |  [Instagram](https://www.instagram.com/storagereview/ "Opens in a new window") |  [Twitter](https://twitter.com/storagereview "Opens in a new window") |  [TikTok](https://www.tiktok.com/@storagereview? "Opens in a new window") |  [RSS Feed](https://www.storagereview.com/rss.xml)

![]()

### Harold Fritts

I have been in the tech industry since IBM created Selectric. My background, though, is writing. So I decided to get out of the pre-sales biz and return to my roots, doing a bit of writing but still being involved in technology.

Previous post: [Micron Shows off 512GB DDR5 RDIMM: 12TB per Dual-Socket Server at 9,200 MT/s, Volume Production in 2H 2027](https://www.storagereview.com/news/micron-shows-a-512gb-ddr5-rdimm-12tb-per-dual-socket-server-at-9200-mt-s-volume-production-in-2h-2027)

Next post: [CoreWeave Deploys Multi-Rack Vera Rubin NVL72 Cluster: Hundreds of Rubin GPUs, 1.6 Tb/s Per GPU, and a No-Fee Archive Tier](https://www.storagereview.com/news/coreweave-brings-up-a-multi-rack-vera-rubin-nvl72-cluster-hundreds-of-rubin-gpus-1-6-tb-s-per-gpu-and-a-no-fee-archive-tier)

Trusted Vendors

Products and solutions from our affiliate partners:

- [  
  ![Ubiquiti Logo](https://www.storagereview.com/wp-content/uploads/2024/11/2-1.webp)  
   Ubiquiti](https://store.ui.com/us/en?a_aid=StorageReview "Opens in a new window")
- [  
  ![Newegg Logo](https://www.storagereview.com/wp-content/uploads/2025/11/newegg.webp)  
   Newegg](https://click.linksynergy.com/fs-bin/click?id=g5terNMDhz0&offerid=1207190.18&subid=0&type=4 "Opens in a new window")
- [  
  ![Amazon Logo](https://www.storagereview.com/wp-content/uploads/2025/11/amazon.webp) Amazon](https://amzn.to/4fLJUEW "Opens in a new window")

Newsletter

Subscribe to the StorageReview newsletter to stay up to date on the latest news and reviews. We promise no spam!

1  

Leave this field empty if you’re human:

## Advertisement

Content Categories

[Facebook](https://www.facebook.com/@storagereview) [X](https://twitter.com/storagereview) [Instagram](https://www.instagram.com/storagereview/) [LinkedIn](https://www.linkedin.com/company/storagereview-com) [YouTube](https://www.youtube.com/user/storagereview?sub_confirmation=1) [Email](mailto:info@storagereview.com) [Spotify](https://open.spotify.com/show/1y6VnznABhHeOSMOmbDTz0) [Reddit](https://www.reddit.com/r/StorageReview/) [Discord](https://discord.gg/TwMHb4azdC) [RSS](https://www.storagereview.com/rss.xml) [TikTok](https://www.tiktok.com/@storagereview)

Copyright © 1998-2025 Flying Pig Ventures, LLC Cincinnati, Ohio. All rights reserved.

Manage your privacy

To provide the best experiences, we and our partners use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us and our partners to process personal data such as browsing behavior or unique IDs on this site and show (non-) personalized ads. Not consenting or withdrawing consent, may adversely affect certain features and functions.

Click below to consent to the above or make granular choices. Your choices will be applied to this site only. You can change your settings at any time, including withdrawing your consent, by using the toggles on the Cookie Policy, or by clicking on the manage consent button at the bottom of the screen.

Functional Functional Always active

The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.

Preferences Preferences

The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.

Statistics Statistics

The technical storage or access that is used exclusively for statistical purposes. The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.

Marketing Marketing

The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.

Statistics

Marketing

Features

Always active

Always active

- Manage options
- Manage services
- Manage {vendor\_count} vendors
- [Read more about these purposes](https://cookiedatabase.org/tcf/purposes/)

Accept Deny Manage options Save preferences Manage options

- {title}
- {title}
- {title}

Manage your privacy

To provide the best experiences, we use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us to process data such as browsing behavior or unique IDs on this site. Not consenting or withdrawing consent, may adversely affect certain features and functions.

Functional Functional Always active

The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.

Preferences Preferences

The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.

Statistics Statistics

The technical storage or access that is used exclusively for statistical purposes. The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.

Marketing Marketing

The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.

Statistics

Marketing

Features

Always active

Always active

- Manage options
- Manage services
- Manage {vendor\_count} vendors
- [Read more about these purposes](https://cookiedatabase.org/tcf/purposes/)

Accept Deny Manage options Save preferences Manage options

- {title}
- {title}
- {title}

Manage consent Manage consent

```json
{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/www.storagereview.com\/news\/mlperf-inference-v6-1-5-7x-per-accelerator-gains-a-512-gpu-run-and-vera-rubins-first-peer-reviewed-numbers","url":"https:\/\/www.storagereview.com\/news\/mlperf-inference-v6-1-5-7x-per-accelerator-gains-a-512-gpu-run-and-vera-rubins-first-peer-reviewed-numbers","name":"MLPerf Inference v6.1: 5.7x Per-Accelerator Gains, a 512-GPU Run, and Vera Rubin's First Peer-Reviewed Numbers - StorageReview.com","isPartOf":{"@id":"https:\/\/www.storagereview.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.storagereview.com\/news\/mlperf-inference-v6-1-5-7x-per-accelerator-gains-a-512-gpu-run-and-vera-rubins-first-peer-reviewed-numbers#primaryimage"},"image":{"@id":"https:\/\/www.storagereview.com\/news\/mlperf-inference-v6-1-5-7x-per-accelerator-gains-a-512-gpu-run-and-vera-rubins-first-peer-reviewed-numbers#primaryimage"},"thumbnailUrl":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2026\/04\/StorageReview-Supermicro-H14-AMD-MI350X-1.jpg","datePublished":"2026-09-16T15:00:00+00:00","description":"MLPerf Inference v6.1 adds End-to-End RAG and Edge Agentic tests, debuts Vera Rubin NVL72, and logs a 512-GPU MI355X run. 30 submitters.","breadcrumb":{"@id":"https:\/\/www.storagereview.com\/news\/mlperf-inference-v6-1-5-7x-per-accelerator-gains-a-512-gpu-run-and-vera-rubins-first-peer-reviewed-numbers#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.storagereview.com\/news\/mlperf-inference-v6-1-5-7x-per-accelerator-gains-a-512-gpu-run-and-vera-rubins-first-peer-reviewed-numbers"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.storagereview.com\/news\/mlperf-inference-v6-1-5-7x-per-accelerator-gains-a-512-gpu-run-and-vera-rubins-first-peer-reviewed-numbers#primaryimage","url":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2026\/04\/StorageReview-Supermicro-H14-AMD-MI350X-1.jpg","contentUrl":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2026\/04\/StorageReview-Supermicro-H14-AMD-MI350X-1.jpg","width":1500,"height":1099,"caption":"Supermicro H14 server with AMD Instinct MI350X GPUs, front view showing the fan wall and NVMe drive bays"},{"@type":"BreadcrumbList","@id":"https:\/\/www.storagereview.com\/news\/mlperf-inference-v6-1-5-7x-per-accelerator-gains-a-512-gpu-run-and-vera-rubins-first-peer-reviewed-numbers#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.storagereview.com\/"},{"@type":"ListItem","position":2,"name":"News","item":"https:\/\/www.storagereview.com\/news"},{"@type":"ListItem","position":3,"name":"MLPerf Inference v6.1: 5.7x Per-Accelerator Gains, a 512-GPU Run, and Vera Rubin&#8217;s First Peer-Reviewed Numbers"}]},{"@type":"WebSite","@id":"https:\/\/www.storagereview.com\/#website","url":"https:\/\/www.storagereview.com\/","name":"StorageReview.com","description":"StorageReview.com is a leading provider of news and reviews throughout the entire IT stack - from the datacenter to the edge, and all points in between.","publisher":{"@id":"https:\/\/www.storagereview.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.storagereview.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/www.storagereview.com\/#organization","name":"StorageReview.com","url":"https:\/\/www.storagereview.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.storagereview.com\/#\/schema\/logo\/image\/","url":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2020\/02\/Storage-Reviews-2-2.png","contentUrl":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2020\/02\/Storage-Reviews-2-2.png","width":344,"height":61,"caption":"StorageReview.com"},"image":{"@id":"https:\/\/www.storagereview.com\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/x.com\/storagereview","http:\/\/youtube.com\/user\/storagereview"]}]}
```
