The NVIDIA RTX PRO 4500 Blackwell Server Edition is the card NVIDIA built primarily to replace the L4 in servers that live outside the data center: single-slot, PCIe Gen5, passively cooled, 165W, with 32GB of GDDR7 and the full Blackwell feature set, including FP4 Tensor Cores and Multi-Instance GPU. HPE sent us a ProLiant DL145 Gen11, the short-depth edge server it pairs with this GPU in its Nemotron edge solution, and we ran the same vLLM serving sweep on the RTX PRO 4500 and on an L4 in the same chassis. This RTX PRO 4500 edge review is about that one comparison: what a site that already runs an L4 in an edge box gets by swapping in the 4500, in tokens, in latency, and in power consumption.
Across two small BF16 models at practical concurrency, the RTX PRO 4500 delivers 2.7x the L4’s output throughput, holds first-token latency under a second at loads where the L4 has already fallen into a queue, and raises the DL145’s whole-server draw from about 196W to about 301W to do it. Per GPU watt, the gain is 1.3 to 1.4x. Per card, it’s the difference between serving 32 to 64 concurrent chat sessions and serving 256. Per box, the DL145 takes up to three L4s, and three at 32 sessions each land close to one 4500 on throughput, so the 4500’s box-level case rests on per-user speed, latency, a single 32GB memory pool, and FP4.
This is an edge server look at one card against its predecessor, on the Metrum AI® Bench Platform, with the models that fit comfortably in 24GB and 32GB. The RTX PRO 4500 Server Edition also belongs in a wider comparison against the rest of NVIDIA’s inference lineup, and we have that testing underway for a separate piece. For this one, the L4 is the most relevant comparison, and we’ve added a single RTX PRO 6000 Blackwell dataset only to show what NVFP4 on a 26B model looks like when there’s a bigger card to measure against. The RTX PRO 6000 isn’t supported in the HPE edge server.
Key Takeaways
- 2.7x the L4 in the same server: Across 18 input and output combinations on Qwen3.5-4B and Gemma 4 E4B, the RTX PRO 4500 delivered a median 2.7x the L4’s output tokens per second at 8 to 32 concurrent requests, and 3.0x to 3.2x at 256, with both cards in the same HPE ProLiant DL145 Gen11.
- Latency is the bigger gap: The 4500 generates at 12 to 15ms per token against 35 to 40ms on the L4, and at 256 concurrent requests its median time to first token stayed at 264ms to 650ms on 256-token prompts, while the L4’s climbed to 9 and 39 seconds.
- 105W more at the wall, 2.7x the tokens: The DL145 drew about 196W serving on the L4 and about 301W on the 4500 at 32 concurrent. Output tokens per system watt went from 2.7 to 5.1 on Qwen3.5-4B and from 2.9 to 5.6 on Gemma 4 E4B.
- Per-user speed holds to 256 sessions on chat-shaped prompts: Dividing output by concurrency, the 4500 still gives every user 11 to 22 tokens per second at 256 concurrent on 256-token or longer answers to short and medium prompts, where the L4 has dropped to 3.7 to 8.7. The L4 falls under a 10 tokens per second floor between 32 and 256 sessions on most shapes and between 8 and 16 on the prefill-heavy 1,024/32 shape; the 4500 stays above it at 256 except on short-answer and 1,024-token-prompt shapes.
- 32GB and FP4 open a model class the L4 can’t run natively: On Gemma 4 26B-A4B in NVFP4 with an FP8 KV cache, the 4500 produced 1,613 output tokens per second at 32 concurrent at 165W, about half of what an RTX PRO 6000 Blackwell managed in a workstation while drawing 266 to 471W at 8 to 32 concurrent. At 256 concurrent, the gap widens to about 2.5x.
What Is the RTX PRO 4500 Server Edition Designed For?
NVIDIA announced the RTX PRO 4500 Blackwell Server Edition at GTC 2026 as the entry tier of its RTX PRO server family and the successor to the L4. It shares a name with the RTX PRO 4500 Blackwell workstation card, and the two share silicon and a 10,496 CUDA core count, but they’re different products. The Server Edition drops the display outputs and the fan, holds board power to 165W, and adds the data center feature set: MIG with up to two 16GB instances, confidential computing, a third NVENC and NVDEC, and vGPU support. NVIDIA rates its memory at 800GB/s and its FP4 Tensor peak at 1.6 PFLOPS with sparsity, and describes the card as a “power-efficient Blackwell GPU for enterprise AI” aimed at “mainstream data center and edge servers.” We covered this in our HPE GTC 2026 coverage.
NVIDIA also changed the form factor: the L4 we reviewed in 2024 is a half-height, half-length, 72W card with 24GB of GDDR6 at 300GB/s, while the RTX PRO 4500 Server Edition is a full-height, full-length single-slot board at 165W, so it still takes one slot but needs a full-length riser and a 16-pin power feed. The DL145 Gen11’s standard slot is full-height, half-length, and HPE configured our unit with the full-length riser and power cable. On paper, the 4500 brings 2.7x the L4’s memory bandwidth, 33% more memory, and 1.7x the FP8 Tensor peak at 2.3x the power, plus an FP4 path the L4 has no hardware for.
| Specification | RTX PRO 4500 Blackwell Server Edition | NVIDIA L4 |
|---|---|---|
| Platform Overview | ||
| Architecture | Blackwell | Ada Lovelace |
| CUDA cores | 10,496 | 7,424 |
| Memory | 32GB GDDR7 with ECC | 24GB GDDR6 |
| Memory bandwidth | 800GB/s | 300GB/s |
| Performance (NVIDIA peak figures, with sparsity) | ||
| FP32 | 51 TFLOPS | 30.3 TFLOPS |
| FP8 Tensor | 811 TFLOPS | 485 TFLOPS |
| FP4 Tensor | 1.6 PFLOPS | Not supported |
| Power and Form Factor | ||
| Board power | 165W | 72W |
| Form factor | Single slot, full height, full length, passive | Single slot, half height, half length, passive |
| Power connector | 1x PCIe CEM5 16-pin | Slot powered |
| Interface | PCIe 5.0 x16 | PCIe 4.0 x16 |
| Features | ||
| MIG | Up to 2 instances at 16GB | No |
| NVENC / NVDEC | 3 / 3 | 2 / 4 |
| Confidential computing | Capable | No |
The Edge Server: HPE ProLiant DL145 Gen11
HPE positions the DL145 Gen11 as the server for sites that don’t have a server room. Its solution brief for the DL145 and NVIDIA Nemotron describes a wall-mountable, near-silent box rated from -5C to 55C for cabinets with little airflow, dusty locations, and hot or cold sites, with iLO and Compute Ops Management for onboarding hundreds of locations that have only power and Ethernet. HPE suggests retail, healthcare, finance, manufacturing, and energy as the target industries, and the workloads it lists are the ones this review measures: small language models, RAG, and vector search running locally so latency-sensitive and regulated data never leaves the site. We reviewed the DL145 Gen11 when it launched, so we won’t repeat the platform tour here.
The unit HPE sent for this testing is configured with a single AMD EPYC 8534P, a 64-core, 128-thread Siena part, 96GB of DDR5-4800 in six 16GB DIMMs, and the RTX PRO 4500 Blackwell Server Edition on the full-length riser. The chassis has a two-bay U.3 backplane; our unit shipped with SATA-only backplane cabling, so we booted from a 6.4TB Solidigm D7-P5620 NVMe SSD on a PCIe sled instead, which has no bearing on the GPU results. The L4 comparison used the same exact server configuration. The 4500 is the card HPE ships in this platform for its Nemotron edge solution; the L4 is what a DL145 owner from the last two years most likely has in that slot today.
Two more platform numbers feed the analysis below: the DL145 reports whole-server power through iLO, and the Bench Platform logged it alongside the GPU’s own board power on every run, so we can talk about watts at the wall and watts at the card. And the server’s air path kept both passive cards inside their limits: the 4500 peaked at 84C average GPU temperature during the longest runs and the L4 at 86C, both under sustained full-power inference load in a 2U short-depth chassis.
How We Tested
All inference data in this review was collected with the Metrum AI Bench Platform (version 3.9.2, recorded in the run files under its former name, Metrum Insights), running vLLM 0.24.0 as the serving engine. The Bench Platform runs a fixed grid of scenarios against each system, records throughput, latency percentiles, GPU telemetry, and host power for every scenario, and exports the results as a single dataset per job. We’ve written about Metrum AI’s work with Oregon State before, and Dell used the same platform for its R770AP technical brief; this is the first time we’ve used it to drive our own GPU comparison end to end. Metrum AI’s engineers and our lab ran the jobs together, with the Bench Platform doing the collection and our team setting the hardware, the models, and the questions.
Metrum AI also publishes an open-source companion, the Metrum AI Bench CLI, which readers can use to run the same style of tests on their own systems. Metrum AI is a registered trademark, and Metrum AI Bench CLI and Metrum AI Bench Platform are trademarks of Metrum AI, Inc. Use of Metrum AI Bench does not imply endorsement by Metrum AI.
The grid is nine prompt shapes: every combination of 32, 256, and 1,024 input tokens with 32, 256, and 1,024 maximum output tokens, each run at 1, 8, 16, 32, and 256 concurrent requests, for 45 scenarios per model per GPU. The L4’s Gemma 4 run also included 64 and 128 concurrent requests, which turned out to be useful for finding its ceiling. Each scenario sends 15 requests per stream from 1 through 32 concurrent, and 1,280 requests at 256. The 4500’s Gemma 4 E4B job was run three times; scenario-to-scenario variance between repeats was under 2%, and we report the median. The L4 and 4500 BF16 jobs and the two Gemma 4 26B jobs, one on each card, come to 378 scenario runs, and across them there was one failed request in one repeat of one 4500 scenario, which we mention only for disclosure.
Three models cover the test: Qwen3.5-4B and Gemma 4 E4B, both served in BF16 with a BF16 KV cache, are the two the L4 and the 4500 both ran, and they’re the size of model an edge deployment on a 24GB card is realistically serving today. Gemma 4 26B-A4B, a 26B mixture-of-experts model with about 4B active parameters, ran in NVFP4 with an FP8 KV cache on the 4500 only, with an RTX PRO 6000 Blackwell Workstation Edition (600W) in a Dell Pro Max Tower T2 as a reference point. The L4 has no native FP4 compute and no room for that model at BF16, so it sits out that part of the test.
Three notes on reading the charts: we start the throughput plots at 8 concurrent requests. The single-request scenarios are short enough (15 requests, a few seconds each) that the first request of each single-user Qwen scenario on the 4500, which hit a 37 to 39 second first-token delay, skews the run’s throughput figure, so for the single-user case we report per-token latency and response time, which aren’t affected. GPU power is a sampled average, so we only compute tokens per watt from runs long enough for the average to settle, which in practice means 16 concurrent and up. And at 256 concurrent with long prompts, both cards run out of KV cache and vLLM queues requests; those points are still real throughput numbers, but the time-to-first-token figures at that load describe a queue, and we treat 32 concurrent as the working range for latency-sensitive use.
Qwen3.5-4B
The grid below is the whole Qwen3.5-4B comparison at a glance: output tokens per second on both cards at every prompt shape, from 8 to 256 concurrent requests. The same curve shape repeats across all nine panels, with the 4500 opening at about 2.5 to 2.8x the L4 at 8 concurrent, scaling a little faster through 32, and landing at 2.4x to 3.2x at 256.
Taking the 256-in, 256-out shape as the chat-sized reference: at 8 concurrent, the L4 produced 192 output tokens per second and the 4500 529 (2.76x); at 32 concurrent, 532 against 1,529 (2.87x); at 256 concurrent, 935 against 2,821 (3.02x). The longest shape, 1,024 tokens each way, tracks it almost exactly: 499 against 1,445 at 32 concurrent and 715 against 2,181 at 256. The narrowest gap in the whole grid is the prefill-heavy 1,024-in, 32-out case, where the 4500 holds a steady 2.4x from 8 concurrent onward, because that workload is almost entirely prompt processing and both cards are compute-bound on it from the start. Qwen3.5-4B rises steadily on that shape on both cards. The one dip at 16 concurrent in the BF16 grids is the 4500 running Gemma 4 E4B on 1,024/32, which showed up in all three repeats; we haven’t isolated the cause.
For a single user, the throughput ratio understates the difference, because what a person sees is how fast the words appear. On the 256/256 shape, the 4500 generated at a median 12.4ms per token, the L4 at 35.0ms, so a 256-token answer took about 3.2 seconds on the 4500 and 9.0 seconds on the L4. At 1,024 tokens each way, it was 13.1 seconds against 36.6 seconds. First-token latency for that lone user was 43ms on the 4500 and 81ms on the L4 for the shorter prompt, 99ms against 219ms for the longer one.
Under load, the latency gap widens: at 32 concurrent on the 256/256 shape, the L4’s median time to first token was 884ms and the 4500’s was 386ms. At 256 concurrent, the L4 went to 39 seconds, while the 4500 held at 650ms. On 1,024/1,024, the L4 was already over a second at 32 concurrent (1,022ms) and reached 255 seconds at 256, which is a request sitting in vLLM’s queue for four minutes waiting for the KV cache to free up. The 4500 crossed the one-second line on that shape only at 256 concurrent, at 12.9 seconds. In the chart below, the L4’s lines leave the plot area between 32 and 256; the 4500’s stay under a second on the shorter shape all the way out.
Another way to read the same data is per user: output tokens per second divided by the number of concurrent requests, which is the rate each person in the queue sees, plotted in the grid below for every shape against a 10 tokens per second floor that serves as a common threshold for interactive use. At 32 concurrent, the 4500 delivers 44 to 53 tokens per second per user on every shape except 1,024-in, 32-out, where it’s 11, while the L4 sits at 10 to 18 and has already dropped to 4.6 on that same shape. At 256 concurrent the 4500 stays above the floor on the four chat-shaped combinations (14.1 on 32/256, 12.7 on 32/1,024, 11.0 on 256/256, and 11.6 on 256/1,024), lands on it at 9.9 on 32/32, and falls under it on the 1,024-token prompts and the short-answer shapes, down to 1.5 on 1,024/32 where nearly all of the work is prefill. The L4 is under the floor on every shape at 256, between 0.6 and 4.5. Nothing was tested between 32 and 256, so the point where each line crosses the dotted floor sits somewhere between those two concurrencies; what the chart shows is that the L4’s crossing comes well before 256 on every shape, and the 4500’s comes at or near 256 on the chat shapes.
Both cards ran at or near their board power limits, 72W for the L4 and 165W for the 4500, on most shapes from 16 concurrent up. The short-output shapes ran lower, such as 125W on the 4500 for Qwen3.5-4B 32/32 at 32 concurrent, so tokens per GPU watt tracks throughput closely. At 256 concurrent, the 4500 produced a median 1.33x the L4’s output tokens per watt across the nine shapes, from 1.07x on the prefill-bound 1,024/32 case to 1.39x on 256/1,024. Multiply that by the 2.3x power difference, and you get the 3x throughput lead, with the 1.33x coming from Blackwell doing more per watt and the 2.3x from the 4500 spending more watts.
Gemma 4 E4B
Gemma 4 E4B is the small member of Google’s Gemma 4 family, the E standing for an effective 4B parameters, and it behaves like Qwen3.5-4B on both cards with slightly better scaling at the top of the grid. The median throughput ratio at 8 to 32 concurrent is 2.67x (range 2.31x to 3.07x), and at 256 it’s 3.20x, with the 1,024/1,024 shape reaching 3.82x because the L4 is so deep into its KV cache limit there.
On the 256/256 shape, the L4 produced 183 output tokens per second at 8 concurrent, 577 at 32, and 1,269 at 256; the 4500 produced 479, 1,678, and 4,058. On 1,024/1,024, the L4 went 180, 524, 826, and the 4500 went 480, 1,454, 3,154. The 4500’s 256-concurrent figure on the long shape is 3.8x the L4’s, and the reason is visible in the L4’s curve: it added only 302 tokens per second going from 32 to 256 concurrent, because it had no memory left to add sessions with.
Because the L4’s Gemma 4 job included 64 and 128 concurrent, this model shows exactly where the L4 gives out. Its median time to first token on the 256/256 shape was 184ms at 32 concurrent, just over a second at 64, and 9.1 seconds at 256. On 1,024/1,024, it was 339ms at 32, over a second at 64, 60 seconds at 128, and 191 seconds at 256. The 4500 stayed at 64ms at 32 concurrent and 264ms at 256 on the shorter shape, and on the longer shape it was 142ms at 32 and crossed one second only at 256, at 1,036ms. So for Gemma 4 E4B at chat-length prompts, the L4 in this server is a 32-to-64-session card, and the 4500 is a 256-session card with latency to spare.
The per-user view, with the L4’s 64 and 128 points filled in, shows where it crosses the floor. On 256/256, the L4 delivers 18.0 tokens per second per user at 32 concurrent, 14.5 at 64, 10.3 at 128, and 5.0 at 256, so it crosses the 10 tokens per second floor between 128 and 256 on that shape. On 256/1,024 it’s at 8.9 by 128, and on the short-answer and 1,024-input shapes it’s under the floor by 64 (7.7 on 256/32, 9.3 on 1,024/256). The 4500 sits at 52 to 57 per user on the chat shapes at 32 concurrent and is still above the floor at 256 on six of the nine shapes (15.9 to 21.6 on the 32- and 256-input shapes other than 256/32, and 12.3 on 1,024/1,024), dropping under it only on the three short-answer and long-prompt shapes (8.1 on 256/32, 7.4 on 1,024/256, and 1.6 on 1,024/32).
Per GPU watt, Gemma 4 E4B gives the 4500 a median 1.39x the L4 at 256 concurrent, with the widest margin, 1.68x, on 256/1,024 and 1.6x or better on the other long-answer shapes where the L4 is starved for KV cache, and the narrowest, 1.07x, on 32/32 where the workload is too short for either card to stretch out. On 256/256 it’s 24.6 output tokens per second per watt against 17.7.
Per-token latency for a single user matches Qwen: 14.1ms per token on the 4500 against 39.1ms on the L4 at 256/256, and 14.6ms against 39.8ms at 1,024/1,024, so a 1,024-token answer takes about 15 seconds on the 4500 and 41 seconds on the L4.
Power at the Wall
Edge sites often have little power to spare, so we tracked whole-server draw through the DL145’s iLO on every run. Under light load, on single-user runs with the GPU drawing 43 to 65W, the server sat at about 160W with either card installed. Serving at 32 concurrent on the 256/256 shape, it drew a median of 196W with the L4 and 301W with the RTX PRO 4500, and those medians held across both BF16 models; the maximum we logged was 209W with the L4 and 308W with the 4500. So the upgrade costs about 105W at the wall under load, which is close to the 93W difference in the cards’ board power ratings, plus a little conversion loss and extra fan overhead.
At 32 concurrent on 256/256, the server produced 2.7 output tokens per second per system watt on Qwen3.5-4B with the L4 and 5.1 with the 4500; on Gemma 4 E4B, 2.9 and 5.6. The system-level efficiency gain is larger than the card-level one because much of the DL145’s draw is a fixed cost: the CPU, memory, fans, and storage draw about the same whether the GPU is producing 532 tokens per second or 1,529. A site that’s paying for the box already gets 1.9x the tokens per watt at the wall from the swap, against 1.3x at the card. For a fleet sized on a per-site power budget, that’s the number to plan around.
What 32GB and NVFP4 Add
Everything above is work the L4 can do, just slower. The 4500’s other argument is a model class the L4 can’t serve natively: Gemma 4 26B-A4B in NVFP4, a 26B mixture-of-experts model with about 4B parameters active per token, served with an FP8 KV cache. At BF16, that model’s weights alone exceed 24GB; in NVFP4, they fit in the 4500’s 32GB with room for a working KV cache, and the Blackwell FP4 Tensor Cores run the format natively. vLLM can load NVFP4 weights on the L4’s Ada architecture through a weight-only kernel, and the 4-bit weights, at about 14GB, would fit in 24GB, but the L4 has no native FP4 compute, so we didn’t run it there. For scale, we ran the same job on an RTX PRO 6000 Blackwell Workstation Edition, a 96GB, 600W card, in a Dell Pro Max Tower T2. That’s a different host and a different class of card, and we include it only so the 4500’s numbers have a reference point.
On 256/256, the 4500 produced 647 output tokens per second at 8 concurrent, 1,613 at 32, and 4,607 at 256; the RTX PRO 6000 produced 1,198, 3,252, and 11,682. From 8 to 32 concurrent, the 4500 sits at 48% to 55% of the 6000 on every shape except the prefill-bound 1,024/32 case, where it falls to 29% at 32 concurrent, and it does that at 165W against 266W to 471W measured on the 6000 for the shapes charted. At 256 concurrent, it drops to 23% to 42% as its 32GB fills.
First-token latency tells the same story as the BF16 models: at 32 concurrent, the 4500’s median was 67ms on 256/256 and 103ms on 1,024/1,024, within a few tens of milliseconds of the 6000 either way, and it stayed at 204ms at 256 concurrent on the short shape. On 1,024/1,024 at 256 concurrent, the 4500 queued to 62 seconds, while the 6000, with three times the memory, held at 560ms. For a 26B MoE, the 4500 is a 32-session card at long context and a 256-session card at chat length, and it runs the model on FP4 hardware the L4 doesn’t have.
The DL145 drew a maximum of 306W serving the 26B model on the 4500, against a 308W peak on the 4B models, because the card is at 165W either way. A 26B-class model in an edge box at 300W at the wall is the capability upgrade here, and it doesn’t cost any more power than the small models do.
Conclusion
NVIDIA rates the RTX PRO 4500 Blackwell Server Edition at over 5x the L4, a figure that relies on FP4 and NIM. In a like-for-like BF16 comparison in the same DL145 Gen11, the 4500 delivered a median 2.7x the L4’s output throughput on two small models at working concurrency and 3.0x to 3.2x at saturation, generated tokens 2.8x faster for a single user, and held first-token latency under a second at loads where the L4 had been queuing for tens of seconds. The L4 keeps every user above 10 output tokens per second out to between 32 and 256 sessions on most shapes, and only to 8 on the prefill-heavy 1,024/32 shape, while the 4500 does it at 256 on chat-length prompts. The 4500 also served Gemma 4 26B-A4B in NVFP4, a model the L4 can’t run natively.
The cost is 105W at the wall, from about 196W to about 301W for the whole server under load, and 2.3x the board power at the card. Measured at the GPU, the 4500 is 1.3 to 1.4x as efficient as the L4 per output token; measured at the wall, where the rest of the system’s draw is consumed regardless, it’s about 1.9x. Sites that size on a per-box power budget will find the 4500 the better use of the watts they’re already spending, and sites that size on slots will find one 4500 matching the box throughput of three L4s at 32 sessions each, with 2.8x faster generation per user and a single 32GB pool.
The scope caveat from the top of the review stands, because this is one card against its predecessor in one edge server on three small models, and the RTX PRO 4500 Server Edition’s place among the L40S, the RTX PRO 6000 Server Edition, and the rest of the current inference lineup is a separate question we’re working through in a broader piece. For the question a DL145 owner with an L4 is asking, the answer is that the 4500 is the upgrade the platform is built to take, provided the chassis has the full-length riser and 16-pin feed to take it, and that HPE’s pairing of this card with this server for Nemotron at the edge is well matched to what the hardware does.
The DL145 Gen11 itself appears in the edge section of our Best Servers leaderboard, where this result now anchors the edge inference story.




Amazon