Deskside AI systems built on NVIDIA’s GB10 have one job, and they do it well: put 128GB of unified memory and a petaflop of compute on a desk close enough to the person using it that iteration is instant. What they don’t have is a connection to the storage an organization already runs. Every DGX Spark and OEM variant ships with a single short M.2 slot, which caps local capacity around 4TB and puts every dataset, model, and checkpoint on that machine outside the governance, access controls, backup, and audit trail that live in the data center. Multiply that by a lab, a department, or a campus, and the result is dozens of isolated 4TB islands, each holding its own copy of the same models and datasets, each one only as secure as the desk it sits on.
This project takes the storage problem back to the data center. We built a DGX Spark shared storage host from a Dell PowerEdge R770AP, eight Solidigm D5-P5336 QLC SSDs in Graid SupremeRAID RAID 5, Broadcom 400GbE networking, and Tuxera Fusion serving NFS over RDMA, then attached eight Sparks to its 100GbE ports. Every Spark on the fabric mounts the same governed, parity-protected storage pool with one line; a single copy of each model and dataset serves all of them, no workstation needs more storage than it ships with, and the Sparks themselves can live in the rack next to the array where they’re easier to secure, cool, and share. The point of this build is where the data lives, and any NVMe server with fast NICs could play the part. Oregon State University, which is exploring a large-scale Spark lab for AI learning and wants to give students access to research datasets measured in tens of terabytes, shared its thinking with us as we designed the build, and for data at that scale, shared storage was the approach that fit this evaluation.
Key Takeaways
- One array, eight Sparks, no client software: Eight DGX Sparks from five OEMs mount NFS/RDMA shares from a single Solidigm D5-P5336 QLC array behind Graid SupremeRAID and Tuxera Fusion, over the 100GbE ports they already have, with a one-line mount on stock Ubuntu.
- Hosted models load like local ones: A 78GB Qwen3.5-122B-A10B-NVFP4 model loaded from the shared array in 269.6 seconds on one Spark, within 13% of the 239.6 seconds the same Spark needed from its internal NVMe, and eight Sparks loading the same single copy at once finished between 232.8 and 326.5 seconds.
- One copy replaces a silo on every desk: Models and datasets on governed, parity-protected data center storage replace per-workstation copies on 4TB M.2 drives, so utilization goes up, copy data goes down, and no Spark needs a storage upgrade or a place on a desk. For a large-scale AI learning lab like the one Oregon State is exploring, where research datasets run to tens of terabytes, this is the approach that fit the use case in our evaluation.
- The drives set the ceiling, and the network has headroom: With all eight Sparks reading at once, the eight-drive RAID 5 array topped out near 25GiB/s, about 214Gb/s, which used 27% of the 800Gb/s the two Broadcom BCM57608 NICs in the host can carry. One Spark reading from the array at 10.8GiB/s matches the fastest internal M.2 we’ve measured in a GB10 system.
- Governance comes with the protocol: NFS with RDMA puts per-user folders, shared datasets, and hosted models in one namespace with Active Directory, Kerberos, and LDAP hooks, which is what lets the data stay under data center control while the Sparks move into the rack.
The Problem With an AI Island on Every Desk
The DGX Spark and its OEM variants from Dell, GIGABYTE, HP, Acer, and ASUS all share the GB10 platform and the same storage ceiling. In our original DGX Spark review, storage was the platform’s clearest limitation. The internal slot takes short M.2 drives, so the 8TB client SSDs that ship in the 2280 form factor don’t physically fit. The Luisuantech GP Spark we reviewed in August answers that with a four-bay M.2 enclosure hanging off one 100GbE port, which is a clean fix for one Spark on one desk and, at lab scale, a bigger silo on every desk.
Capacity is half the problem; the other half is what happens to data once it lands on a desk-side machine. A 78GB model pulled down to a Spark is a copy nobody in IT can see. A research dataset staged onto local flash is a copy outside the backup schedule, outside the access policy, and outside institutional audits. Ten Sparks in a lab means ten copies of the same weights consuming 780GB of flash that could be used for something else, and ten places a dataset can walk out the door with the machine. The issue gets worse as the numbers grow, be it GB10 systems or any desktop AI infrastructure. Every organization that has managed a fleet of laptops knows this pattern, and the answer was the same then: keep the data on shared storage in the data center, and let the endpoint be an endpoint.
Thankfully, NVIDIA designed the Spark with ports to make that work. Every unit carries two QSFP56 cages on an integrated ConnectX-7, with a platform ceiling of 200Gb of usable bandwidth behind a pair of PCIe Gen5 x4 links. In our DGX Spark cluster review, we described the split-role configuration in which one cage goes to a peer Spark and the other to storage. This project uses one cage per Spark for storage at 100GbE, leaving the second free for clustering or additional bandwidth. At 100GbE, each Spark has about 12.5GB/s of theoretical bandwidth to the storage host, which exceeds the bandwidth provided by the internal M.2 slot on most of these systems.
Once the data lives on the other end of that link, the Spark’s role changes. It becomes a compute endpoint with a boot drive and local cache drive, and capacity is no longer a per-workstation purchase. A model or dataset is written once and read by every client, so storage utilization increases and the copy count decreases. Data protection moves to a parity-protected array in a server chassis, and access control, encryption at rest, backup, and auditing all occur where they already do for everything else. And because the Spark now needs nothing but power and a 100GbE link, there’s no reason it has to sit on a desk at all. Racked next to the storage host, it’s cooler, quieter, physically secured, and available to whoever needs it next.
The Storage Host
The storage host is a Dell PowerEdge R770AP, the same 2U Xeon 6 platform we reviewed earlier this year, and we used it because it was free in the lab when the project started. Any NVMe server with room for fast NICs and a small GPU could fill the role, and the R770AP is one way to build it, even if a little overbuilt. Ours runs two Xeon 6 6978P processors with 1.5TB of DDR5 across 24 DIMMs, uses one of its two banks of eight Gen5 NVMe bays, and hangs the drives and the Graid GPU off CPU 0, with the two 400GbE adapters on CPU 1. The server runs Ubuntu 24.04.
| Bestanddeel | Storage Host Configuration |
|---|---|
| Server | Dell PowerEdge R770AP, 2U, 16x 2.5-inch Gen5 NVMe bays |
| CPU | 2x Intel Xeon 6 6978P |
| Geheugen | 1.5TB DDR5 (24x 64GB) |
| Opslag | 8x Solidigm D5-P5336 61.44TB U.2 PCIe 4.0 QLC (491.52TB raw) |
| RAID | Graid SupremeRAID, RAID 5 across all eight drives, NVIDIA RTX A2000 GPU on CPU 0 |
| File System | XFS on the SupremeRAID virtual drive |
| file Server | Tuxera Fusion (Fusion NFS, private preview build), NFS 4.1 over RDMA (RPC-RDMA on port 20049), user-mode |
| Netwerken | 2x Broadcom BCM57608 (Thor 2) single-port 400GbE, PCIe Gen5 x16 on CPU 1, RoCEv2 |
| Stap over voor slechts | Dell PowerSwitch Z9864F-ON, 64x 800GbE OSFP112, Enterprise SONiC |
| Client Links | 800G optic broken out to 8x100GbE, one link per Spark, four Sparks per storage NIC on separate VLANs |
| Klanten | 2x Dell Pro Max with GB10, 2x GIGABYTE AI TOP ATOM, 2x HP ZGX Nano G1n, Acer Veriton GN100, ASUS Ascent GX10 |
| OS | Ubuntu LTS 24.04 |
Why QLC, and Why the D5-P5336
In our conversations with Oregon State about the Spark lab concept, capacity was the first requirement, because the research datasets involved are large. The job of this server is to hold everything a room full of Sparks might want, and to serve it in a form that can be read by many clients at once. That’s a read-heavy, large-block, high-capacity profile, and it’s the profile QLC was built for. The D5-P5336 family runs from 7.68TB to 122.88TB; the 61.44TB U.2 drives used here are rated by Solidigm at up to 7,000MB/s sequential read, 3,000MB/s sequential write, 1.005M random 4K read IOPS, and 65.2PBW of endurance, or 0.58 drive writes per day over five years, on 192-layer QLC NAND with a 16KB indirection unit.
The eight 61.44TB drives fill half of the R770AP’s front bays, and a production build could pick any capacity point in the family and any number of bays. Even so, eight drives in RAID 5 provide about 430TB of usable space with single-drive fault tolerance. That’s more than 100 times what a Spark’s M.2 slot holds and enough for a department’s model library, the shared datasets students work from, and per-user home directories. Filling the other eight bays doubles that without touching the network design, and the sizing numbers later in this review show why the drives are the part to grow.
The 16KB indirection unit matters for this workload because Solidigm’s larger IU trades small random-write efficiency for density and cost, and for a share that serves models and datasets, small random writes are the minority case. Model loading is large sequential reads. Dataset streaming is large sequential reads. Checkpoints and student files are written once and read often. The division of labor that falls out of this is the right one for a Spark lab: the shared QLC array holds everything that’s read by many, and the Spark’s own M.2 handles the scratch and swap traffic that’s written by one.
RAID on a GPU: Graid SupremeRAID
NVMe RAID on eight Gen4 drives, each capable of reading at 7,000MB/s, is the point at which a conventional hardware RAID card becomes the bottleneck because every byte must pass through the card’s PCIe slot. Software RAID faces a different set of resource challenges. Graid SupremeRAID offloads parity calculations to a GPU and keeps the data path on the host’s PCIe lanes, which is why we’ve seen it scale to more than 180GB/s in our SupremeRAID AE testing with 16 drives. For this host, we ran SupremeRAID on an NVIDIA RTX A2000, a small workstation GPU.
The eight D5-P5336 drives were secure-erased and sequentially filled before being handed to Graid, then configured as a single RAID 5 array. The resulting virtual drive was formatted with XFS and served by Tuxera. RAID 5 is the right level for a share like this: a drive failure keeps the array online, rebuilds happen without taking user data offline, and the capacity penalty is one drive in eight. RAID 6 or RAID 10 would be reasonable choices for a production deployment that values rebuild-window protection over capacity, and SupremeRAID supports both.
The Network: Broadcom 400GbE and an 800G Core
The storage side of the fabric is two Broadcom BCM57608 Thor 2 adapters in the R770AP, each a single-port 400GbE NIC on a PCIe Gen5 x16 interface. Broadcom positions the BCM57608 as the industry’s lowest-power 400G NIC, and its feature set makes it a good fit: RoCEv2 with enhanced DCQCN congestion control, adaptive routing, secure boot with a silicon root of trust, and inline kTLS encryption offload. RoCEv2 carries NFS over RDMA in this design, and congestion control prevents multiple clients from stepping on each other when they all start pulling at once. During the build, the bnxt_en driver reported that the optic modules on the 400G ports were reaching their 70C warning threshold under sustained load. We addressed this with a small fan-profile offset on the R770AP, and the NICs ran without incident thereafter.
Both NICs land on a Dell PowerSwitch Z9864F-AAN, our lab’s 800G core switch, running Enterprise SONiC. The Z9864F-ON carries 64 OSFP112 ports at 800GbE and breaks any of them out to 2x400G, 4x200G, or 8x100G. One 800G optic broken out to 8x100G connects the eight Sparks, one link each, to the same switch that terminates the two 400G storage links. Each 400G NIC sits on its own VLAN and serves four Sparks, so the storage host presents two RDMA targets, and the load is split evenly across them by design. In practice, this is a full-rate topology: 8x100G of client bandwidth on one side, 2x400G of storage bandwidth on the other, and a switch with 102.4Tbps of capacity in the middle.
We validated the fabric before any storage software touched it. Server-to-server tests between the R770AP and a second host over the 400G links ran at line rate in both directions, and RDMA traffic from the Sparks confirmed that the 100G breakout ports operated at full speed. Once Tuxera’s engineers had the NFS shares up, iperf3 from the Sparks to their respective NICs showed a clean 99Gbps sustained with four parallel streams and zero retransmits.
Why NFS, and Why Tuxera
With Graid presenting a fast block device, the storage host could serve the Sparks as either NVMe-oF block targets or file shares. We tested the block path in the GP Spark review, and it works well for one client. For a lab of many Sparks, we wanted something else. Block volumes are one-to-one; each user would need a LUN, and the moment two users want the same dataset, you’re copying it, which is the problem this whole design exists to remove. NFS is the middle of the Venn diagram: a single namespace holds shared datasets, hosted models, and per-user folders. Every Spark mount uses the NFS client already in Ubuntu, and the Kubernetes and Slurm tooling a university would put on top of a Spark lab all speak NFS.
Fusion NFS is the NFS half of Tuxera Fusion, the multiprotocol file server that covers SMB and NFS on one user-mode, multithreaded core; we first looked at Fusion SMB in a Solidigm and Tuxera media workflow project in 2024. Tuxera claims up to 22.7GB/s over a single 200GbE connection with RDMA, above kernel NFSD and Ganesha in the company’s own testing, and the product supports NFS 4.1, active-active scale-out clustering, transparent failover, and Active Directory, Kerberos, and LDAP integration. That last set is the governance hook: the same directory that controls who can log in to Spark also controls what they can see on the share. One caveat applies to every Tuxera number in this article: the build we ran is a private preview of Fusion NFS that Tuxera says has not been fully optimized for performance, and the software isn’t available for commercial sale yet. General availability is planned for November, after SC’26.
Tuxera installed Fusion on the R770AP, formatted the Graid volume with XFS, and configured the server with NFS on the standard port and RPC-RDMA on port 20049, leaving SMB listeners out since this project is NFS-only. The initial thread configuration was four protocol threads, each with eight threads for transport transmit and receive, VFS data, and VFS metadata, which Tuxera describes as a starting point for tuning. Directory services were deliberately left out of the lab build; a university deployment would layer its own identity services on top. On the Spark side, mounting the share is one command:
sudo mount -t nfs -o proto=rdma,port=20049 <storage-nic>:/mnt/shares/smbnfs1/<folder> /localnfs
That one line is the entire client integration, with no driver, no initiator configuration, and no vendor software on the Spark. A user who logs into a different Spark tomorrow runs the same line and sees the same files.
Proof of Configuration: Eight Sparks on One Data Set
Before any of the model work, we ran a sanity check on each Spark individually: a 1M sequential fio read with 16 jobs and a queue depth of 16 against its own folder on the array. The eight units returned between 10.0 and 11.2GiB/s, saturating the 100GbE link on each. We then ran a sweep to prove that all eight could hold their share of the load. Each run keeps a fixed 16 threads at a queue depth of 16 in aggregate and spreads them across one, two, four, and eight Sparks, over 1M sequential, 64K random, 16K sequential, and 4K random workloads, with read and write, direct I/O, 60 seconds per test. Every Spark writes to its own folder on the array, four per storage NIC, and there’s no QoS anywhere in the stack. With all eight loaded at once, per-unit 1M sequential reads ranged from 3.09 to 3.13GiB/s, 64K random reads from 6,648 to 6,779 IOPS, and 4K random writes from 2,136 to 2,155 IOPS, across five OEMs on two VLANs with no fairness enforcement. Every Spark reached the array, every Spark got an even share, and the configuration was ready for the workload that matters.
| Spark (eight clients, 2 threads x QD16 each) | Storage NIC | 1M sequentieel lezen | 64K willekeurig lezen IOPS | 4K willekeurig lezen IOPS | 4K willekeurig schrijven IOPS |
|---|---|---|---|---|---|
| Dell Pro Max with GB10 (1) | 1 | 3,176 MiB/s | 6,648 | 10,000 | 2,150 |
| GIGABYTE AI TOP ATOM (1) | 1 | 3,195 MiB/s | 6,779 | 10,300 | 2,142 |
| HP ZGX Nano G1n (1) | 1 | 3,201 MiB/s | 6,696 | 10,100 | 2,136 |
| Acer Veriton GN100 | 1 | 3,194 MiB/s | 6,670 | 10,100 | 2,149 |
| Dell Pro Max with GB10 (2) | 2 | 3,184 MiB/s | 6,694 | 10,200 | 2,155 |
| GIGABYTE AI TOP ATOM (2) | 2 | 3,187 MiB/s | 6,667 | 10,100 | 2,141 |
| HP ZGX Nano G1n (2) | 2 | 3,189 MiB/s | 6,678 | 10,100 | 2,151 |
| ASUS Ascent GX10 | 2 | 3,163 MiB/s | 6,734 | 10,300 | 2,145 |
Sizing the Host: Two NICs and Plenty of Headroom
The same sweep answers the question a lab or an enterprise has to settle before buying: where is the ceiling, and which part of the host sets it? The four panels below show read throughput across all four workloads as the fixed load scales from one Spark to eight; the write side is covered in the text and table further down.
Aggregate 1M sequential read climbed from 10.8GiB/s on one Spark to 21.3GiB/s on two, 24.6GiB/s on four, and 24.9GiB/s on eight. In network terms, that’s 93Gb/s, 183Gb/s, 212Gb/s, and 214Gb/s against the 800Gb/s the two Broadcom 400GbE NICs can carry, so at full eight-client load the array used 27% of the host’s network capacity. The plateau near 25GiB/s belongs to the eight-drive RAID 5 QLC array; the NICs, the switch, and the Sparks’ own 100GbE links never became the limit.
Two server-side NICs already cover eight Sparks at line rate, and they would cover the other eight bays in the R770AP as well: doubling the drive count, whether as a second array or a wider RAID set, puts more read bandwidth behind the same two ports, with room to spare. The same holds for I/O. 4K random reads scaled from 32.5K IOPS on one client to a plateau of about 81K IOPS at four clients and beyond, with per-I/O latency falling from 7.9ms to 3.1ms as each client’s queue got shallower, and 64K random reads went from 1,637 to 3,349MiB/s, with the network well short of its limit in both cases.
The other comparison that matters is against the storage a Spark already has. In our individual GB10 reviews, GDSIO 1M reads from the internal M.2 topped out near 11.2GiB/s on the Acer Veriton GN100 and GIGABYTE AI TOP ATOM. The Dell Pro Max with GB10 and HP ZGX Nano G1n reached around 5.5GiB/s, while the ASUS Ascent GX10 approached 5GiB/s. By comparison, the Luisuantech GP Spark’s external NVMe-oF box delivered 9.5GB/s over the same 100GbE port. A single Spark reading from the shared QLC array at 10.8GiB/s matches the fastest internal drives in that group and roughly doubles the slowest, over one cable, from data every other Spark can read too. The tools differ: GDSIO on the internal drives and fio over NFS here, so this is a range comparison.
The full sweep is below, with aggregate figures for all four workloads at each client count. Every read row climbs with client count until the array runs out of headroom, and every write row stays flat from one Spark to eight, because the RAID 5 parity path holds the array at a fixed write ceiling no matter how many clients share it. For a share built to be read by many and written once at ingest, that’s the profile the design is built for.
| Aggregate (16 threads total) | 1 Spark | 2 Vonken | 4 Vonken | 8 Vonken |
|---|---|---|---|---|
| 1M sequentieel lezen | 10.8 GiB/s | 21.3 GiB/s | 24.6 GiB/s | 24.9 GiB/s |
| 1M sequentieel schrijven | 2.4 GiB/s | 2.4 GiB/s | 2.4 GiB/s | 2.4 GiB/s |
| 64K willekeurig lezen | 1,637 MiB/s | 2,754 MiB/s | 2,964 MiB/s | 3,349 MiB/s |
| Willekeurig schrijven in 64K | 722 MiB/s | 719 MiB/s | 723 MiB/s | 703 MiB/s |
| 16K sequentieel lezen | 466 MiB/s | 833 MiB/s | 890 MiB/s | 1,005 MiB/s |
| 16K sequentieel schrijven | 57 MiB/s | 60 MiB/s | 62 MiB/s | 28 MiB/s |
| 4K willekeurig lezen | 32.5K IOPS | 62.3K IOPS | 82.5K IOPS | 81.2K IOPS |
| Willekeurig schrijven in 4K | 17.8K IOPS | 17.7K IOPS | 17.5K IOPS | 17.2K IOPS |
Hosting Models on Shared Storage
The fio sweep proves that the array and the network are more than capable. The next test that matters for the use case is whether a model living on the shared array loads as well as a model living on the Spark’s own NVMe, because that’s the workflow the whole design depends on: one copy, hosted centrally, loaded by anyone. We picked Qwen3.5-122B-A10B in NVFP4, a 122-billion-parameter mixture-of-experts model that takes 78GB on disk and fits in a single Spark’s 128GB of unified memory. We measured the time from launching the inference server to the server reporting healthy, along with the average storage bandwidth over that window, in three scenarios: one Spark loading from its internal NVMe, the same Spark loading over NFS/RDMA from the shared array, and all eight Sparks loading the same single copy from the array at the same time.
From internal NVMe, the Acer Veriton GN100 reached a healthy server in 239.6 seconds at an average of 0.373GB/s. From the shared array over NFS/RDMA, the same Spark took 269.6 seconds at 0.315GB/s, about 30 seconds, or 13%, longer. The averages show that at 78GB over four minutes, model loading on a Spark is not storage-bound in either case. The GB10’s own processing of the weights sets the pace, and the storage spends most of the window waiting.
The bandwidth trace shows that pattern: local NVMe reads arrive in bursts that peak near 5GB/s with long gaps between them; the single NFS client shows the same burst shape at slightly lower peaks. The eight-Spark aggregate line shows the array under concurrent demand. With all eight pulling the same 78GB at once, the array delivered an aggregate peak of 11GB/s in the first minute, then settled between roughly 4 and 8GB/s for the next two minutes while the Sparks worked through their weights, averaging 2.08GB/s across the full window. Against the 25GiB/s the fio sweep showed the array can deliver, eight Sparks loading a model at once used a fraction of what’s available.
Per-unit load times in the eight-way test ranged from 232.8 seconds on the ASUS Ascent GX10 to 326.5 seconds on one of the GIGABYTE AI TOP ATOM units, with the two HP ZGX Nano G1n systems at 235.3 and 235.4 seconds and the two Dell Pro Max with GB10 units at 245.2 and 245.6 seconds. Five of the eight Sparks loading concurrently from the shared array beat the single Spark loading alone over NFS/RDMA, and three of them beat the local NVMe run. The spread isn’t a vendor story; with no QoS in the stack, which unit ends up waiting on the array at any moment shifts from run to run, and the two units on each storage NIC split evenly across the fast and slow ends of the chart.
For a lab, the result is that it can keep one copy of every model on the array, and any Spark, or every Spark, loads it in about the time it would take to load a local copy, with no download step, no local capacity consumed, and no copy of the weights sitting on a desk.
The Oregon State Use Case
Oregon State’s College of Earth, Ocean, and Atmospheric Sciences has been a repeat collaborator in our AI infrastructure work, from real-time ocean research aboard a research vessel naar AI-assisted academic assessment with Metrum AI. Chris Sullivan, OSU’s Director of Research and Academic Computing, and his architect, Thomas Olson, are exploring a large-scale Spark lab for AI learning that could scale to as many as 100 units, and their thinking about storage informed our design in two specific ways.
First, the storage has to be independent of the Sparks. OSU considered a file system stretched across the Sparks themselves and rejected it for a classroom because adding or removing a single node would disrupt the storage service for everyone. A central array that any Spark mounts means a Spark can be added, removed, reimaged, or handed to a different student without affecting anyone else’s data. Second, storage has to follow the user. We describe the OSU requirement as a beefed-up version of file shares for AI purposes, where students store persistent information as they move between systems. With NFS, that’s a folder per student; log into any Spark, mount the share, and your work is there.
The datasets are why this matters more for OSU than it would for a home-directory server. The College of Earth, Ocean, and Atmospheric Sciences wants to make its research data, oceanographic and atmospheric collections measured in tens of terabytes, available to students for their own research. There’s no version of that on deskside storage: a 4TB M.2 slot can’t hold a single one of those collections, and even a lab full of external enclosures would mean buying and filling a copy for every desk and then keeping every copy current, which no department could afford. Hosting the data once on the array and mounting it over 100GbE is the approach that fit that requirement in our evaluation. The same array hosts the model library, so students spend their session running Qwen, Llama, or any other model, with no download step. The same principle applies to an enterprise deploying desktop AI systems as developer workstations: mount the production dataset over a fast network and let the system work on it in place, which removes the need to stand up a second data lake so desktop machines have something to read. The data stays in the data center, under the access controls, encryption, and backup regime that already exist there, and the Spark holds nothing that matters when it walks out the door.
Conclusie
Deskside AI hardware is good at a lot of things, but it’s not great at being part of an organization’s data estate, and the single M.2 slot in every GB10 system makes sure of that: whatever a user needs ends up copied onto local flash, invisible to IT, unprotected, and duplicated on every other desk that needs the same thing. For a single Spark, an external local storage box is a reasonable answer. For a lab, a department, or an enterprise deploying these by the dozens, every desk with its own data is a silo, and the answer is the one the data center arrived at years ago: put the capacity behind a fast network, govern it once, and share it. The Sparks already have high-speed networking; this project shows what happens when you give it somewhere to go, and it makes the case that the Sparks themselves belong in the rack beside it.
Eight Solidigm D5-P5336 QLC drives in Graid RAID 5 in a Dell server gave eight Sparks a single parity-protected dataset. Two Broadcom 400GbE NICs and a Dell Z9864F-ON with an 800G breakout gave eight Sparks 100GbE each. Tuxera’s Fusion NFS over RDMA made that capacity a one-line mount on stock Ubuntu with directory-service hooks for access control. Under a fixed load spread across all eight, the array delivered 24.9GiB/s of aggregate sequential read, split evenly with no QoS, and one Spark alone saturated its link at 10.8GiB/s, as fast as the best internal drive in any GB10 system we’ve tested; that ceiling used 27% of what the two server-side NICs can carry, so the network is sized for a fuller server than this one. The model test confirmed the workflow: a 78GB model hosted once on the shared array loads on one Spark within 13% of local NVMe, and on eight Sparks at once, with five of the eight finishing faster than the solo NFS run.
QLC flash makes these economics work because the workload is capacity-heavy and read-heavy, the D5-P5336’s strong suit, and a single server can hold what a hundred Sparks could never carry locally, with one copy of everything and every copy under the data center’s control. Oregon State has production research storage for the work that can’t wait. The concept Chris Sullivan and Thomas Olson are exploring is an educational tier the college doesn’t have today: a governed, shared pool sized for a rack or more of Sparks, built from a server and a handful of high-capacity drives, that could put tens of terabytes of real oceanographic and atmospheric data in front of a large-scale AI learning lab on the same fabric its Sparks use.
Dit rapport is gesponsord door Solidigm. Alle standpunten en meningen in dit rapport zijn gebaseerd op onze onbevooroordeelde kijk op het (de) product(en) in kwestie.
References to commercial products in this article are for informational purposes and do not constitute an endorsement by Oregon State University.





Amazon