---
description: AMD MLPerft Inference 6.0 gains indicate significant advancement with ROCm software at the center of the gains.
title: AMD Instinct MI355X Achieves MLPerf Inference v6.0 Gains with Over 1 Million Tokens per Second and Supports Scalable ROCm Stack
image: https://www.storagereview.com/wp-content/uploads/2026/04/Storagereview-AMD-MLPerf-tokens-per-sec.jpg
---

 

[](//storagereview.com/) 

---

[ ](https://www.facebook.com/@storagereview) [ ](https://twitter.com/storagereview) [ ](https://www.instagram.com/storagereview/) [ ](https://www.linkedin.com/company/storagereview-com) [ ](https://www.youtube.com/user/storagereview?sub%5Fconfirmation=1) [ ](mailto:info@storagereview.com) [ ](https://open.spotify.com/show/1y6VnznABhHeOSMOmbDTz0) [ ](https://www.reddit.com/r/StorageReview/) [ ](https://discord.gg/TwMHb4azdC) [ ](https://www.storagereview.com/rss.xml) [ ](https://www.tiktok.com/@storagereview) 

[](https://www.storagereview.com) 

≡ Menu 
* [Home](https://www.storagereview.com/)
* [Storage Reviews](https://www.storagereview.com/review)  
  * [Consumer Reviews](https://www.storagereview.com/consumer)
  * [Enterprise Reviews](https://www.storagereview.com/enterprise)
  * [Ubiquiti Reviews](https://www.storagereview.com/best/ubiquiti-reviews)
* [SR Merch](https://store.storagereview.com)
* [Leaderboards](https://www.storagereview.com/best)  
  * [Best Storage Arrays](https://www.storagereview.com/best/storage-arrays)
  * [Best Enterprise SSDs](https://www.storagereview.com/best/enterprise-ssds)
  * [Best Servers](https://www.storagereview.com/best/servers)
  * [Best Desktops for Local AI](https://www.storagereview.com/best/desktops-local-ai)
  * [Best Laptops for Local AI](https://www.storagereview.com/best/laptops-local-ai)
  * [Best Local LLM Tools](https://www.storagereview.com/best/local-llm-tools)
  * [Agentic AI Hardware](https://www.storagereview.com/best/agentic-ai-hardware)
  * [Best SSDs & Hard Drives](https://www.storagereview.com/best%5Fdrives)
  * [Best Portable SSDs](https://www.storagereview.com/best/portable-ssds)
  * [Best Business Laptops](https://www.storagereview.com/best/business-laptops)
  * [Best Mobile Workstations](https://www.storagereview.com/best/mobile-workstations)
  * [Best Desktop Workstations](https://www.storagereview.com/best/desktop-workstations)
  * [Best Laptop Battery Life](https://www.storagereview.com/best/laptop-battery-life)
* [Storage Reference Guide](https://www.storagereview.com/storage-reference-guide)
* [About SR](https://www.storagereview.com/about-storagereview)  
  * [StorageReview.com Sweepstakes Rules and Regulations](https://www.storagereview.com/storagereview-com-sweepstakes-rules-and-regulations)

Search 

[Home](https://www.storagereview.com/) » [News](https://www.storagereview.com/news) » AMD Instinct MI355X Achieves MLPerf Inference v6.0 Gains with Over 1 Million Tokens per Second and Supports Scalable ROCm Stack

# AMD Instinct MI355X Achieves MLPerf Inference v6.0 Gains with Over 1 Million Tokens per Second and Supports Scalable ROCm Stack

by Harold Fritts on April 2, 2026 

[AI](https://www.storagereview.com/enterprise/ai) ◇ [Enterprise](https://www.storagereview.com/enterprise) 

AMD has released its MLPerf Inference v6.0 results, positioning the [Instinct MI355X GPU](https://www.storagereview.com/news/from-mi350-to-mi500-amds-bold-ai-accelerator-roadmap-through-2027) as a scalable inference platform across single-node, multinode, and heterogeneous deployments. The submission extends beyond incremental gains by adding new workloads, demonstrating cluster-scale throughput exceeding 1 million tokens per second, and validating reproducibility across a growing partner ecosystem.

### CDNA 4 Architecture Targets High-Capacity Inference

The Instinct MI355X GPU is based on [AMD’s CDNA 4](https://www.storagereview.com/news/supermicro-expands-air-cooled-solutions-and-unveils-new-ai-factory-clusters) architecture built on a TSMC 3nm | 6nm FinFET process (it uses a dual-process chiplet design: the compute dies (XCDs) use TSMC’s 3nm node while the I/O dies use 6nm), integrating 185 billion transistors and supporting FP4 and FP6 data formats. It’s worth noting that this is across the entire multi-chiplet package, not a monolithic die. Each GPU includes up to 288GB of HBM3E memory, enabling support for models up to 520 billion parameters on a single device. AMD positions this combination of compute density and memory capacity as critical to large-model inference without excessive model partitioning.

The platform is available in UBB8 configurations with both air-cooled and direct liquid-cooled options, aligning with data center deployment requirements.

### Multinode Throughput Surpasses 1 Million Tokens per Second

A key result from this round is AMD surpassing 1 million tokens per second at the cluster scale. Using Instinct MI355X GPUs, AMD achieved this threshold on Llama 2 70B in both Server and Offline scenarios, and on GPT-OSS-120B in Offline.

![AMD MLPerf 1M tokens per second graphic]() 

These results reflect a shift toward evaluating inference performance at the cluster level rather than per accelerator. Aggregate throughput and time-to-serve are increasingly used to determine production readiness for large-scale AI deployments.

AMD also demonstrated efficient scaling. On Llama 2 70B, a configuration of 11 nodes and 87 GPUs reached over 1 million tokens per second across Offline, Server, and Interactive scenarios, with scale-out efficiency ranging from 93% to 98%. On GPT-OSS-120B, a 12-node, 94-GPU cluster achieved similar throughput with over 90% scaling efficiency. These results indicate that performance gains translate effectively as deployments expand beyond a single system.

### Generational Gains and Competitive Single-Node Performance

AMD reported a 3.1x performance increase on Llama 2 70B Server compared to the prior Instinct MI325X generation, reaching 100,282 tokens per second. The improvement reflects both architectural changes and ROCm software optimizations. Offline scores improved by 4.4x and Server scores improved by 4.8x compared to prior rounds. These gains are primarily attributed to FP4 quantization.

![AMD Inference results vs previous gen graphic]() 

In single-node comparisons, MI355X demonstrated competitive positioning against NVIDIA platforms. On Llama 2 70B, AMD matched NVIDIA B200 in Offline throughput, reached near parity in Server performance, and exceeded Interactive performance. Against Nvidia’s B300, AMD’s GPU delivers 92% in offline mode, 93% in server mode, and exceeds it with 104% in interactive mode.

### First-Time Model Enablement Expands Coverage

MLPerf Inference v6.0 includes several new workloads, and AMD used this round to demonstrate rapid model enablement. GPT-OSS-120B, a mixture-of-experts model, was introduced for the first time and achieved competitive results compared to NVIDIA systems across both Offline and Server scenarios.

AMD also submitted results for Wan-2.2 text-to-video generation, marking its entry into multimodal and generative video inference. While the official submission focused on Single Stream latency, the results were competitive with those of existing platforms. Post-submission tuning further improved performance, indicating headroom for optimization as software matures.

These additions highlight AMD’s focus on expanding beyond traditional LLM benchmarks to support emerging AI workloads.

### ROCm Software Enables Scaling and Heterogeneous Inference

AMD attributes much of the performance and scalability to its ROCm software stack. Enhancements include optimized FP4 execution, improved GPU-to-GPU communication for distributed inference, and support for dynamic workload distribution across heterogeneous environments.

![AMD MLPerf inference results instinct mI355x graphic ]() 

The initial MLPerf heterogeneous submission was developed using three AMD Instinct GPU models: MI300X, MI325X, and MI355X. Submitted by Dell and MangoBoost, the configuration achieved 141,521 tokens per second on Llama 2 70B Server and 151,843 tokens per second on Llama 2 70B Offline.

Worth noting, the AMD Instinct MI355X platform was located in Dell’s lab in the United States, while the Instinct MI300X and MI325X platforms were in Korea. This demonstrates the capability to coordinate systems across different geographic locations.

### Ecosystem Growth and Reproducibility

AMD’s partner ecosystem expanded in this MLPerf round, with nine companies submitting results across multiple Instinct GPU generations. Participating vendors included Cisco, Dell, Giga Computing, HPE, MangoBoost, MiTAC, Oracle, Supermicro, and Red Hat.

Partner submissions closely matched AMD’s internal results, typically within 4% and, in some cases, within 1%. This consistency indicates that performance is reproducible across OEM and cloud platforms, reducing deployment risk and improving confidence in real-world outcomes.

**Engage with StorageReview**

[Newsletter](https://www.storagereview.com/storage%5Fnewsletter) | [YouTube](https://www.youtube.com/user/StorageReview "Opens in a new window") | Podcast [iTunes](https://podcasts.apple.com/gb/podcast/storagereview-com-storage-reviews/id1060681115 "Opens in a new window")/[Spotify](https://open.spotify.com/show/1y6VnznABhHeOSMOmbDTz0 "Opens in a new window") | [Instagram](https://www.instagram.com/storagereview/ "Opens in a new window") | [Twitter](https://twitter.com/storagereview "Opens in a new window") | [TikTok](https://www.tiktok.com/@storagereview? "Opens in a new window") | [RSS Feed](https://www.storagereview.com/rss.xml)

![]()

### Harold Fritts

I have been in the tech industry since IBM created Selectric. My background, though, is writing. So I decided to get out of the pre-sales biz and return to my roots, doing a bit of writing but still being involved in technology. 

Previous post: [Dell Technologies Enhances PowerProtect Portfolio for Improved Cyber Resilience](https://www.storagereview.com/news/dell-technologies-enhances-powerprotect-portfolio-for-improved-cyber-resilience)

Next post: [NerdioCon ’26 Focuses on Azure Virtual Desktop, Windows 365, and AI Convergence](https://www.storagereview.com/news/nerdiocon-26-focuses-on-azure-virtual-desktop-windows-365-and-ai-convergence)

Trusted Vendors

Products and solutions from our affiliate partners:

* [ Ubiquiti ](https://store.ui.com/us/en?a%5Faid=StorageReview "Opens in a new window")
* [ Newegg ](https://click.linksynergy.com/fs-bin/click?id=g5terNMDhz0&offerid=1207190.18&subid=0&type=4 "Opens in a new window")
* [  Amazon ](https://amzn.to/4fLJUEW "Opens in a new window")

Newsletter

Subscribe to the StorageReview newsletter to stay up to date on the latest news and reviews. We promise no spam!

1  

Leave this field empty if you’re human: 

## Advertisement

Content Categories

[ Facebook ](https://www.facebook.com/@storagereview) [ X ](https://twitter.com/storagereview) [ Instagram ](https://www.instagram.com/storagereview/) [ LinkedIn ](https://www.linkedin.com/company/storagereview-com) [ YouTube ](https://www.youtube.com/user/storagereview?sub%5Fconfirmation=1) [ Email ](mailto:info@storagereview.com) [ Spotify ](https://open.spotify.com/show/1y6VnznABhHeOSMOmbDTz0) [ Reddit ](https://www.reddit.com/r/StorageReview/) [ Discord ](https://discord.gg/TwMHb4azdC) [ RSS ](https://www.storagereview.com/rss.xml) [ TikTok ](https://www.tiktok.com/@storagereview) 

Copyright © 1998-2025 Flying Pig Ventures, LLC Cincinnati, Ohio. All rights reserved.

Manage your privacy

To provide the best experiences, we and our partners use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us and our partners to process personal data such as browsing behavior or unique IDs on this site and show (non-) personalized ads. Not consenting or withdrawing consent, may adversely affect certain features and functions.

Click below to consent to the above or make granular choices. Your choices will be applied to this site only. You can change your settings at any time, including withdrawing your consent, by using the toggles on the Cookie Policy, or by clicking on the manage consent button at the bottom of the screen.

Functional Functional Always active 

The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network. 

Preferences Preferences 

The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user. 

Statistics Statistics 

The technical storage or access that is used exclusively for statistical purposes. The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you. 

Marketing Marketing 

The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes. 

Statistics

Marketing

Features

Always active

Always active

* [Manage options](#)
* [Manage services](#)
* [Manage {vendor\_count} vendors](#)
* [Read more about these purposes](https://cookiedatabase.org/tcf/purposes/)

Accept Deny Manage options Save preferences [Manage options](#) 

* [{title}](#)
* [{title}](#)
* [{title}](#)

Manage your privacy

To provide the best experiences, we use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us to process data such as browsing behavior or unique IDs on this site. Not consenting or withdrawing consent, may adversely affect certain features and functions.

Functional Functional Always active 

The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network. 

Preferences Preferences 

The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user. 

Statistics Statistics 

The technical storage or access that is used exclusively for statistical purposes. The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you. 

Marketing Marketing 

The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes. 

Statistics

Marketing

Features

Always active

Always active

* [Manage options](#)
* [Manage services](#)
* [Manage {vendor\_count} vendors](#)
* [Read more about these purposes](https://cookiedatabase.org/tcf/purposes/)

Accept Deny Manage options Save preferences [Manage options](#) 

* [{title}](#)
* [{title}](#)
* [{title}](#)

Manage consent Manage consent

```json
{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/www.storagereview.com\/news\/amd-instinct-mi355x-achieves-mlperf-inference-v6-0-gains-with-over-1-million-tokens-per-second-and-supports-scalable-rocm-stack","url":"https:\/\/www.storagereview.com\/news\/amd-instinct-mi355x-achieves-mlperf-inference-v6-0-gains-with-over-1-million-tokens-per-second-and-supports-scalable-rocm-stack","name":"AMD Instinct MI355X Achieves MLPerf Inference v6.0 Gains with Over 1 Million Tokens per Second and Supports Scalable ROCm Stack - StorageReview.com","isPartOf":{"@id":"https:\/\/www.storagereview.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.storagereview.com\/news\/amd-instinct-mi355x-achieves-mlperf-inference-v6-0-gains-with-over-1-million-tokens-per-second-and-supports-scalable-rocm-stack#primaryimage"},"image":{"@id":"https:\/\/www.storagereview.com\/news\/amd-instinct-mi355x-achieves-mlperf-inference-v6-0-gains-with-over-1-million-tokens-per-second-and-supports-scalable-rocm-stack#primaryimage"},"thumbnailUrl":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2026\/04\/Storagereview-AMD-MLPerf-tokens-per-sec.jpg","datePublished":"2026-04-02T11:50:31+00:00","description":"AMD MLPerft Inference 6.0 gains indicate significant advancement with ROCm software at the center of the gains.","breadcrumb":{"@id":"https:\/\/www.storagereview.com\/news\/amd-instinct-mi355x-achieves-mlperf-inference-v6-0-gains-with-over-1-million-tokens-per-second-and-supports-scalable-rocm-stack#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.storagereview.com\/news\/amd-instinct-mi355x-achieves-mlperf-inference-v6-0-gains-with-over-1-million-tokens-per-second-and-supports-scalable-rocm-stack"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.storagereview.com\/news\/amd-instinct-mi355x-achieves-mlperf-inference-v6-0-gains-with-over-1-million-tokens-per-second-and-supports-scalable-rocm-stack#primaryimage","url":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2026\/04\/Storagereview-AMD-MLPerf-tokens-per-sec.jpg","contentUrl":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2026\/04\/Storagereview-AMD-MLPerf-tokens-per-sec.jpg","width":960,"height":540},{"@type":"BreadcrumbList","@id":"https:\/\/www.storagereview.com\/news\/amd-instinct-mi355x-achieves-mlperf-inference-v6-0-gains-with-over-1-million-tokens-per-second-and-supports-scalable-rocm-stack#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.storagereview.com\/"},{"@type":"ListItem","position":2,"name":"News","item":"https:\/\/www.storagereview.com\/news"},{"@type":"ListItem","position":3,"name":"AMD Instinct MI355X Achieves MLPerf Inference v6.0 Gains with Over 1 Million Tokens per Second and Supports Scalable ROCm Stack"}]},{"@type":"WebSite","@id":"https:\/\/www.storagereview.com\/#website","url":"https:\/\/www.storagereview.com\/","name":"StorageReview.com","description":"StorageReview.com is a leading provider of news and reviews throughout the entire IT stack - from the datacenter to the edge, and all points in between.","publisher":{"@id":"https:\/\/www.storagereview.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.storagereview.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/www.storagereview.com\/#organization","name":"StorageReview.com","url":"https:\/\/www.storagereview.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.storagereview.com\/#\/schema\/logo\/image\/","url":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2020\/02\/Storage-Reviews-2-2.png","contentUrl":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2020\/02\/Storage-Reviews-2-2.png","width":344,"height":61,"caption":"StorageReview.com"},"image":{"@id":"https:\/\/www.storagereview.com\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/x.com\/storagereview","http:\/\/youtube.com\/user\/storagereview"]}]}
```
