Our current AI system reviews center on inference serving in the lab, with open-weight models from the Llama, Qwen, DeepSeek, and gpt-oss families running through vLLM or SGLang and measured on output throughput and first-token latency. Coverage runs from the edge to the full GPU node: the NVIDIA RTX PRO 4500 Server Edition delivered a median 2.7x the output tokens per second of the L4 it was built to replace inside an HPE ProLiant DL145 Gen11, for about 105W more at the wall; the Dell, GIGABYTE, and HP versions of DGX Spark ran distributed inference as two-node clusters over the 200 Gb fabric; and Supermicro JumpStart gave us hands-on time with an NVIDIA HGX B200 and an AMD Instinct MI350X system. Storage gets the same treatment because the data path decides how busy those GPUs stay, so we measure GPUDirect Storage throughput with GDSIO and test KV cache offload to flash for long-context and agentic workloads.
If you are sizing hardware for AI, the Best Servers page names our current AI inference and edge picks on the data from these reviews, Best Enterprise SSDs covers the drives that keep the GPUs fed, and Best Storage Arrays covers the AI storage platforms behind training clusters and neoclouds.