We’ve taken a look at a couple of GB300 DGX Stations now, and recently had two of them physically in our lab at the same time: the ASUS ExpertCenter Pro ET900N G3 and the HP ZGX Fury AI Station. With both towers sitting next to each other, we wanted to connect them and see what a second GB300 could add to the performance of a single machine for more coordinated deskside AI workloads.
Growing a desktop setup one box at a time isn’t a new idea; NVIDIA pitched this explicitly with the GB10 platform. One of the things that made the DGX Spark so special was the ability to start with one system and connect more as your workloads grew. The GB300 starts with considerably more performance, but running larger models or serving more users can tax even one of the fastest desktop machines we have tested. That is what we were getting at in our original MSI review: if your workloads justify one of these Stations, it is easy to see how they could eventually justify several.
For many buyers, cost will determine when that expansion becomes practical. Understanding how well two Stations work together now gives someone starting with one a clearer idea of what they could gain from expanding later.
Networking on the DGX Station
Before getting into the cluster, let’s recap the networking capabilities we covered in our MSI review. The ConnectX-8 SuperNIC does more than provide the network connection. Its integrated PCIe switch also handles part of the Station’s internal I/O, connecting to the Grace as well as the 2x PCIe Gen 6 x4 storage slots.
For the actual networking, the system exposes two QSFP112 ports rated at 400Gb/s each, providing 800Gb/s of aggregate bandwidth to connect Stations together or attach them to a larger network.
Compared with the DGX Spark’s implementation, which we covered in our original DGX Spark review, the Spark’s 400Gb/s-class ConnectX-7 was more capable than the GB10’s PCIe connectivity could fully use. The controller sat behind two separate PCIe 5.0 x4 links, limiting useful aggregate network throughput to about 200Gb/s. Its two physical ports also appeared as four logical Ethernet interfaces in Linux. Getting the full bandwidth required mapping those interfaces back to the connected cage, configuring both logical interfaces behind it, and making sure RDMA traffic could use both PCIe paths.
The Station gives us a simpler arrangement, with two 400Gb/s ports that can each serve as an independent network rail.
Pushing four times the Spark’s usable bandwidth through a pair of QSFP cages also raises the cooling load, and during our testing, every unit handled it. The ConnectX-8 itself is liquid-cooled, and some vendors extend that cooling to the QSFP cages like MSI, while other systems use heatsinks and dedicated fans over the cages, as we saw in our ASUS review. That cooling is especially useful for optical modules, which generate more heat than passive DACs. Both approaches have worked well in our testing, and we have not run into cooling issues with the networking.
Connecting Two Stations
Clustering two DGX Stations together was simple: we connected the ASUS and HP directly using two 400G direct attach copper (DAC) cables between their matching QSFP112 ports, giving us two independent 400Gb/s connections.
On the software side, NVIDIA provides a step-by-step guide for clustering the systems, and the instructions were clear enough that we simply gave the link to Claude Code and let it handle the setup. It configured the network between the two Stations, ran NCCL bandwidth tests, and gave us a report of the results.
The guide is a numbered playbook run from a third machine that reaches both Stations over SSH. The ports are cabled as matching rails, QSFP0 to QSFP0 and QSFP1 to QSFP1, and each rail becomes its own point-to-point RoCEv2 link: the scripts bring up the two ConnectX-8 interfaces at a 9,000-byte MTU, give each rail a private /24 of its own, and apply the RoCE and GPUDirect runtime settings.
A validation pass then confirms that both links trained at 400G, that jumbo pings cross each rail, and that the two RDMA devices, mlx5_0 and mlx5_1, report active, and an ib_write_bw run on each rail proves the fabric moves RDMA traffic and not just pings. After that, NCCL is pointed at both rails with NCCL_IB_HCA=mlx5_0,mlx5_1, and the job launches like any other two-node run. The results below came out of that sequence.
The first table is from nccl-tests, the collectives a distributed inference or training job issues, run across the two B300s with both rails in use. With one GPU per Station, every collective comes down to the two GPUs exchanging data over the pair of 400G links, so the bus bandwidth column reads directly against the 100GB/s (800Gb/s) the two cables provide. Peak is the best result across message sizes; the 16 GiB column is the largest message tested, where bandwidth has settled.
| Operation | Peak size | Peak bus bandwidth | Bus bandwidth at 16 GiB |
|---|---|---|---|
| All-reduce | 1 GiB | 92.242 GB/s | 91.346 GB/s |
| All-gather | 16 GiB | 87.893 GB/s | 87.893 GB/s |
| Reduce-scatter | 2 GiB | 93.930 GB/s | 92.905 GB/s |
| Broadcast | 16 GiB | 97.975 GB/s | 97.975 GB/s |
| Reduce | 16 GiB | 97.978 GB/s | 97.978 GB/s |
| All-to-all | 8 GiB | 80.522 GB/s | 80.512 GB/s |
| Send/receive, unidirectional | 16 GiB | 97.953 GB/s | 97.953 GB/s |
| Send/receive, bidirectional | 16 GiB | 83.748 GB/s per direction | 167.496 GB/s aggregate |
Broadcast, reduce, and one-way send/receive all reached 98GB/s, 98% of what the two cables can carry. The patterns with data moving in both directions at once landed lower, 92GB/s for all-reduce and 80GB/s for all-to-all. The bidirectional send/receive result is the reference point for those: with both GPUs transmitting at the same time, each direction settled at 84GB/s, 167GB/s in total across the link.
The second table drops below NCCL to the raw RDMA path, using ib_write_bw one rail at a time. The GPUDirect rows write 1 MiB transfers from one B300’s HBM into the other’s, with the NIC reading and writing GPU memory directly; both rails hit 392Gb/s, 98% of line rate. The host-memory rows are the guide’s basic smoke test, a single queue pair sourcing from Grace memory, which NVIDIA’s playbook rates at 220 to 230Gb/s on a 400G rail; we saw 228Gb/s. The latency row is an 8-byte write: 1.46 µs on average.
| Test | Bandwidth | Link utilization |
|---|---|---|
| GPUDirect RDMA, GPU memory, ib_write_bw, 1 MiB transfers | ||
| Rail 0 | 392.15 Gb/s | 98.04% |
| Rail 1 | 392.15 Gb/s | 98.04% |
| Host-memory RDMA, both rails | ||
| RDMA write, one-way | 228.15 Gb/s | |
| RDMA read, one-way | 228.23 Gb/s | |
| RDMA send, one-way | 228.16 Gb/s | |
| RDMA write, bidirectional | 447.10 Gb/s aggregate | |
| RDMA write latency, 8 bytes | 1.46 µs average; 1.63 and 1.66 µs 99th percentile by rail | |
How We Tested the Cluster
Time with the pair was limited, so we focused on inference, the most popular AI workload today. We plan to expand this work to larger clusters. For this article, we concentrated on larger models and compared a single Station with two Stations using tensor parallelism (TP) and prefill/decode (PD) disaggregation.
Tensor parallelism splits the weight tensors and computations within the model across multiple GPUs. In our two-Station configuration, or TP2, each B300 holds a portion of the model, and both work on each request. That gives us more high-bandwidth memory (HBM) in which to place the weights, along with more compute to run them. It also requires the GPUs to exchange intermediate results and synchronize through NCCL throughout inference, so communication overhead becomes part of the cost of running the model.
Disaggregated inference assigns the prefill and decode stages of a request to separate workers. Prefill processes the input prompt and builds the key/value (KV) cache, the attention state the model reuses during generation. Decode then uses that state to generate output tokens one at a time. The two stages have different processing requirements: prefill is compute-heavy, whereas decode is memory bandwidth intensive. In larger setups, disaggregating inference like this allows the operator to independently scale these workers according to demand, as our KV cache offload paper showed.
In a single-node setup, when both phases share a worker, vLLM can interleave prefill work for incoming prompts with the decode work of requests already in progress. That prompt processing can increase the time between output tokens. In our PD setup, one Station handles prefill and transfers the request’s KV cache to the other Station, which handles decode. Each worker has access to its own copy of the model, including any data placed in Grace memory. The decode worker no longer has to schedule new prompt processing alongside token generation, speeding up inference throughput.
For our testing, the model’s memory requirements drove which approach we used. For the largest checkpoints, where a single Station had to offload a substantial amount of model weights to Grace memory, TP gave us the best performance. Spreading the weights across two GPUs allowed more of the model to stay in HBM and reduced the offload penalty. For models that fit reasonably well on one Station, either entirely in HBM or with limited offload to Grace, we used PD disaggregation. That included DeepSeek V4.1 Flash, with its Engram lookup tables placed in Grace DRAM.
To maintain comparability with our older coverage, we retained our legacy methodology with fixed input sequence lengths (ISL) and output sequence lengths (OSL). The equal workload uses 512 input tokens and 512 output tokens, while the prefill-heavy workload uses 8,192 input tokens and 1,024 output tokens. We vary the number of concurrent requests while keeping those sequence lengths fixed.
The single-Station baseline in each chart is the ASUS ET900N G3 from our review, swept to 32 streams; the two-Station runs go to 256. Each point is a single pass, so a dip at one concurrency level is reported as collected.
Inference Configuration
| Model | Cluster mode | Software | Speculative decoding | Grace memory |
|---|---|---|---|---|
| DeepSeek V4.1 Flash (FP8) | Prefill/decode disaggregation, full copy per Station | vllm/vllm-openai:deepseekv41-flash-0909 | Model’s own drafter, fixed 5 tokens | Two FP8 Engram tables (188.84GiB) pinned in Grace memory plus 131.97GB of expert weights offloaded, on each Station |
| GLM-5.3 NVFP4 | Tensor parallel (TP2) | NVIDIA vLLM 26.08 NGC container (vLLM 0.27.1) | MTP, fixed 5-token draft | None on the cluster; 215GB of expert weights offloaded on a single Station |
| GLM-5.3 Flash NVFP4 | Prefill/decode disaggregation, ASUS prefill, HP decode | vLLM 0.28.1 RC1 | MTP, fixed 5-token draft | None |
| GLM-5.2 NVFP4 | Tensor parallel (TP2) | NVIDIA vLLM 26.08 NGC container (vLLM 0.27.1) | MTP, fixed 3-token draft | None on the cluster; about 215GB of expert weights offloaded on a single Station |
| MiniMax-M3 NVFP4 | Tensor parallel (TP2) | NVIDIA vLLM 26.08 NGC container (vLLM 0.27.1) | EAGLE3, fixed 3-token draft | None on the cluster; a slice of experts offloaded on a single Station |
Inference Performance
DeepSeek V4.1 Flash
DeepSeek V4.1 Flash ran disaggregated, one Station handling prefill and the other decode, with each Station holding a full copy of the model: the two FP8 Engram tables, 188.84GiB in total, pinned in Grace memory and a further 131.97GB of expert weights offloaded to Grace. That makes it the one configuration in the set with Grace memory in play on both machines. It is also the one model here without a single-Station line, because the ASUS review ran the earlier V4 Flash checkpoint and the two aren’t directly comparable.
On the two-Station cluster, the native FP8 model ran from 441 output tokens per second at one stream to 3,328 at 32, 5,141 at 128, and 5,248 at 256, the highest output throughput of any model in this set, with the curve flattening only at the top of the sweep. For reference, the single ASUS Station served V4 Flash at 1,766 output tokens per second at 32 streams in our review.
On the prefill-heavy profile, the cluster ran from 382 output tokens per second at one stream to 1,466 at 32 and 1,796 at 256, and because each request carries 8,192 input tokens, total throughput reached 16,163 tokens per second at the top of the sweep. With the Engram tables and 131.97GB of expert weights sitting in Grace memory on each Station, the curves stay smooth across the whole sweep regardless.
GLM-5.3
GLM-5.3 ran in TP2 with no Grace offload. On a single Station, it offloads 215GB of expert weights to Grace memory, and the cluster result follows GLM-5.2 through most of the sweep before it gets uneven. One Station ran from 44 output tokens per second at one stream to 188 at 32 on the equal workload.
The two-Station cluster ran from 208 at one stream to 1,301 at 16, between 4.7x and 9.2x the single tower, fell to 853 at 32, and then resumed climbing to 1,652 at 64, 3,051 at 128, and 5,018 at 256. The 32-stream point is in the data as collected; batching decisions and kernel selection both change as concurrency rises, and a tensor-parallel deployment adds a collective on every layer that has to complete before the next step, so the curve isn’t a straight line, but the ceiling at 256 streams is 27x what one Station reached at 32.
On the prefill-heavy profile, the single Station stopped at 16 streams, by design: an offloaded model flattens early, and there’s little to learn from pushing concurrency further. It moved from 46 output tokens per second at one stream to 126 at 16. The cluster moved from 194 to 785 over the same range, 6.2x at the top, then settled between 573 and 611 from 32 streams through 256, a long-prompt ceiling near 600 output tokens per second that is still about five times the single Station’s best.
GLM-5.3 Flash
GLM-5.3 Flash is small enough for each Station to hold a full copy, so it ran disaggregated, with the ASUS handling prefill and the HP handling decode. Flash has a small active parameter count, so on a single Station it already runs fast: 353 output tokens per second at one stream and 828 at 32 on the equal workload, with an uneven middle (858 at four streams, 489 at eight) that we’re reporting as collected.
On the disaggregated pair, the equal workload ran from 362 output tokens per second at one stream to 2,721 at 32, 3.3x the single tower, and on to 3,256 at 256 streams.
At one and two streams, the pair matches a single Station almost exactly, which is what you’d expect when the decode Station is doing the same work alone; the gain arrives once there are enough requests in flight for the prefill Station to stay busy and for the decode Station to benefit from never pausing to prefill.
On the prefill-heavy profile, the single Station climbed from 344 output tokens per second at one stream to 1,216 at 16 and held 1,219 at 32, flat from 16 streams on. The prefill/decode pair kept climbing, from 329 at one stream to 2,448 at 32, twice the single Station where it had leveled off, and on to 3,042 at 256. Long prompts are the case this split is built for, because the prefill Station absorbs the 8,192-token prompts while the decode Station’s KV cache and compute stay on generation.
GLM-5.2
GLM-5.2 ran in TP2 with no Grace offload on the cluster, and it’s the model that makes the strongest case for a second Station. The NVFP4 checkpoint is about 433GB, so on one GB300 roughly 215GB of expert weights live in Grace memory, and every decode step that touches them crosses NVLink-C2C to LPDDR5X at 396GB/s. Split across two Stations, each B300 holds about half the model in HBM3e.
On the equal workload, the single ET900N G3 delivered 35 output tokens per second on one stream and 139 at 32. The two-Station cluster started at 149 tokens per second on one stream, 4.3x the single tower, reached 1,525 at 32 streams, 11x, and kept climbing past the single Station’s sweep to 2,316 at 64 and 3,141 at 256, with a dip to 2,012 at 128. A model that one Station serves at conversational speed for a handful of users becomes a model two Stations can serve to a department.
The prefill-heavy profile tells the same story at a lower height, with one Station moving from 37 output tokens per second at one stream to 120 at 32, and the cluster moving from 141 to 677 over the same range, 3.8x at the bottom of the sweep and 5.6x at the top, then flattening at 710, 725, and 740 through 256 streams.
Long prompts put more of the work in prefill, where both configurations are compute-bound, so the offload penalty that the second Station removes is a smaller share of the total than it is on short prompts. The gain runs well past 2x on both profiles because the second Station takes a 433GB model out of LPDDR5X and puts all of it in HBM, and that matters more here than the second B300’s compute.
MiniMax-M3
MiniMax-M3 ran in TP2 with no Grace offload. On one Station it sits right at the HBM boundary, with a slice of its experts in Grace memory, and in the ASUS review it reached 1,041 output tokens per second at 32 streams on the equal workload with EAGLE3.
The cluster tracks the single Station closely through 16 streams (1,208 against 940) and then pulls away: 2,776 at 32, 2.7x the single tower, 3,610 at 64, a dip to 1,912 at 128, and 3,498 at 256. Two Stations add less here than they do on GLM-5.2 because one Station already keeps most of this model in HBM, so the second GPU’s contribution is closer to plain compute scaling, with KV cache headroom on top.
The prefill-heavy profile is where MiniMax-M3 ran out of room on one Station: 8,192-token prompts consume the remaining KV cache quickly, so the single tower peaked at 282 output tokens per second at two streams, held 271 at four, and fell to 103 at eight, where the sweep ended because there was nothing left to measure.
With the model split across two Stations, the same points read 365, 522, and 736, a 30% gain at two streams where both configurations have cache to spare and 7.1x at eight, where one has run out. The cluster then continued to 909 at 16, 1,506 at 32, 1,724 at 64, 1,366 at 128, and 1,801 at 256, a range the single Station can’t reach on this profile at all.
Conclusion
Two GB300 DGX Stations cabled together over their own ConnectX-8 ports do what NVIDIA’s documentation says they should; the setup was painless, and this is the first throughput data we’ve seen for the dual-system configuration.
The second Station earns its keep on the models that overflow 252GB of HBM3e: GLM-5.2 goes from 139 output tokens per second at 32 streams on one Station to 1,525 on two, and on to 3,141 at 256 streams, because the split takes a 433GB model out of LPDDR5X; GLM-5.3 reaches 5,018 at 256 streams from a single-Station ceiling of 188; MiniMax-M3 on long prompts goes from a tower that slows at eight streams to a pair that is still climbing at 256.
On the two models that fit in one Station, GLM-5.3 Flash and DeepSeek V4.1 Flash, the better way to use the second box is to give each Station a whole copy, split prefill from decode, and let the decode worker generate without pausing; on GLM-5.3 Flash, that doubled long-prompt throughput at 32 streams and tripled short-prompt throughput, and DeepSeek V4.1 Flash posted the highest output throughput in the set at 5,248 tokens per second.
Our view of the product hasn’t moved since the MSI XpertStation WS300 review: this is an interesting machine for a specific market, and if you find a need for one, you may find a need for one more.
The best way to run them is as fungible infrastructure: a pool that a team schedules work onto and re-carves as projects change, and NVIDIA owns the pieces to make that happen. It bought Run:ai for scheduling, Brev for developer environments, and Lepton for a compute marketplace, and any of the three paired with a rack of Stations would make the product more compelling if you have access.
The two-box configuration scaled well; the best gain we measured was GLM-5.3, from 188 output tokens per second on a single Station to 5,018 across the pair (a 27x gain), and at matched concurrency GLM-5.2 turned in 11x at 32 streams, both on models that a single Station can only serve by spilling weights into Grace memory. That is a GB300 result over 2x 400G, and it makes us curious what a Vera Rubin Station will look like when NVIDIA rolls one out, with the next generation of HBM and NVLink-C2C.
Where this fits hasn’t changed: two Stations on a desk, without a rack or a data center behind them, now serve the models that used to justify a GPU server and the facilities that come with it, so a team that can’t get a slot in the data center, or doesn’t want to wait for one, has a path to frontier-class inference that sits next to them. NVIDIA has set the bar for that category, price aside.




Amazon