AMD ROCm 10.1, released October 5, puts its weight on the path between storage and the GPU. Checkpoint files, key-value caches, and model weights now routinely outgrow accelerator memory, and when they do, the time spent moving data starts to govern iteration time more than the math does. The release builds out AMD Infinity Storage, the direct-storage stack that arrived in ROCm 7.14, by wiring storage access into the HIP runtime and adding NUMA-aware host memory, so data reaches the GPU without a detour through CPU memory.
Direct GPU Storage Pipelines and NUMA-Aware Memory
The hipFILE library inside AMD Infinity Storage gains an asynchronous fast-path backend. Read and write requests run directly on a HIP stream, skipping the host-memory staging step the traditional path requires, and AMD says the result is lower latency for checkpoint engines and KV-cache offload while sustained bandwidth holds up. The release notes name the calls: hipFileReadAsync() and hipFileWriteAsync() now support the fast path.
A batch I/O API submits multiple file requests at once and spreads them across an internal pool of worker threads, which keeps the storage queues full when a training job is pulling a large dataset. Telemetry grows with it: hipFileGetStatsL1, L2, and L3 return progressively more detailed I/O statistics, and those counters feed the ROCm profilers so an operator can tell whether a pipeline is limited by compute, memory, or data delivery.
On two-socket hosts, a buffer allocated on the wrong socket pays an inter-socket penalty on every transfer. ROCm 10.1 extends the HIP virtual memory management APIs with hipMemLocationTypeHostNuma and hipMemLocationTypeHostNumaCurrent, so host memory can be placed on a specific NUMA node through the full create, reserve, map, and teardown lifecycle. Libraries such as RCCL can build NUMA-local, zero-copy communication buffers on top of that and keep host-to-device traffic off the socket interconnect.
ROCm CLI v1.0.0 and AI Agent Workflows
ROCm CLI reaches v1.0.0 as a single binary for Linux, Windows, and WSL2 that installs, configures, and runs local AI workloads on AMD GPUs. It detects the flash-attn and amd-aiter wheels that vLLM needs, supports the updated ROCm10 package layout, and adds a flag for fully non-interactive installs in scripted or CI environments. A full-screen TUI dashboard shows live GPU telemetry next to model serving and a local chat window for checking that an endpoint answers.
AMD Skills, the packaged integrations for coding agents such as Claude Code and Codex, picks up two more. quark-install sets up AMD Quark with a matching PyTorch build, and quark-torch-llm-ptq runs post-training quantization on PyTorch and Hugging Face LLMs with schemes such as FP8 and INT4.
Profiling and Telemetry Modernization
ROCprofiler-SDK adds kernel replay, in beta. GPU hardware exposes a limited number of performance counters per dispatch, so collecting a full counter set used to mean rerunning the whole application once per counter group. Kernel replay re-executes each dispatch in place, resetting GPU memory between passes, and gathers the full set in a single run. The SDK also lets profiling start and stop at any point in a run for PyTorch Kineto and Triton Proton, and ships a PC sampling agent skill that walks a coding assistant through rocprofv3 PC sampling to find compute stalls.
ROCm Compute Profiler’s roofline report is now an interactive HTML page, so an engineer can isolate a single kernel and switch the arithmetic-intensity axis between cache levels. PC sampling reports per kernel, down to sample and stall counts on individual instructions, and a new guide covers profiling vLLM. ROCm Systems Profiler adds Linux tracing for Gorgon Point 1, 2, and 3 APUs (gfx1150, gfx1152, and gfx1153), and ROCm Optiq now visualizes rocprofv3 profiles alongside Systems Profiler traces and Compute Profiler analysis.
AMD SMI (amd-smi) takes over from ROCm SMI, which is removed from the standard build in this release. It covers AMD GPUs, CPUs, and APUs on Linux and Windows, bare metal and virtualized, with a CLI, a C library (amd_smi_lib), and Python, Go, and Rust bindings. Container awareness extends to containerd, CRI-O, Podman, LXC, and LXD, including Kubernetes pods, so a cluster administrator can match GPU processes and utilization to the container running them.
Core Compiler Updates, Libraries, and Runtime Support
The toolchain moves from LLVM 23 to LLVM 24, and the clang_major macro now reports 24, which anyone pinning compiler versions or linking against libLLVM will want to check before upgrading. Incremental builds get faster through caching of the partitions that device-code link-time optimization creates, so unchanged partitions are reused on the next build.
Composable Kernel speeds up attention on Radeon GPUs and adds more quantized matrix-multiply types to its dispatcher. MIGraphX swaps its backend compiler from rocMLIR to rocMLIRTriton, which AMD says improves inference performance and often shortens model compile times while keeping the compiler in step with upstream Triton. rocSPARSE runs triangular solves directly on matrices stored in the ELLPACK (ELL) format, with no conversion step first.
hipThreads, new under the ROCm Core SDK threading libraries, brings standard C++ threading concepts to AMD GPUs on Linux and Windows. The point is to accelerate existing CPU-threaded code on the GPU without a full rewrite into HIP kernels and synchronization primitives.
Windows Subsystem for Linux 2 enters tech preview, with standard Linux packages installing in the guest and GPU paravirtualization handing work to the Windows host driver. The preview covers .deb packages and Python wheels but not RPMs, and profiling, debugging, and KFD-dependent tooling don’t work inside WSL2 yet. On the virtualization side, KVM SR-IOV is now supported with Ubuntu 26.04 LTS as both host and guest, paired with version 9.3.0.K of AMD’s GIM virtualization driver.
The supported hardware list adds the Radeon AI PRO R9600 alongside the Instinct MI355X, MI350X, and MI350P, the MI300 and MI200 series, the rest of the Radeon AI PRO R9700 family, and Ryzen AI processors.




Amazon