Advancing AI 2026 was AMD’s biggest launch event yet, big enough that we split our coverage in two. Our first article covered the Instinct MI455X and the 72-GPU Helios rack; this sister piece covers what could not comfortably fit alongside them: the 6th Gen EPYC server CPUs, codenamed Venice, and the Verano host processor. The CPU deserves its own headline. Venice brings up to 256 cores and 512 threads per socket, 1.6TB/s of memory bandwidth, the first PCIe Gen 6 in a server CPU, and 18× the throughput of the first-generation EPYC from 2017.
Venice is also not a chip. It is a portfolio, a point AMD repeated in every session: one CPU profile does not fit all. The same Zen 6 generation fans out into a density part, an enterprise part, a stacked-cache HPC part, and a low-power LPDDR host. So before the speeds and feeds, let’s look at the lineup.
The Venice and Verano lineup
AMD sorts the modern data center into three server classes, and the portfolio is built to populate all of them. General-purpose CPU servers run the web gateways, caches, application tiers, databases, and storage that everything else leans on. GPU servers need a host CPU that keeps the accelerators fed, a job where single-threaded speed and I/O bandwidth beat core count. The third class is dense CPU servers for the orchestration and tool-execution code that has grown up around AI services; that code is branch-heavy and stall-prone, and it wants threads above all.
The lineup consists of four products, and platforms arrive in waves: Venice SP7 in Q4 2026, Venice SP8 in the first half of 2027, and Venice-X and Verano in the second half. AMD splits the same silicon six ways, and the finer cut maps straight onto the three tiers. General purpose gets Venice SP7, SP8, and Venice-X. The host bucket pairs a high-frequency Venice bin with Verano and its LPDDR memory. Dense compute gets what the deck labels Venice 256c: high core count at low power, built to pack as many threads into a rack as the power budget allows. Verano moonlights too: AMD says select customers will deploy it as a general-purpose processor wherever performance per system watt is the binding constraint.
We get a sneak peek at the SKUs thanks to the footnotes. The 256-core EPYC 9996 leads the stack, the 9956 carries the same cores at 400W, and the 96-core 9686F is the high-frequency host part. Helios has its own version of that host silicon, the EPYC 9G76, in every compute tray.
The no-compromises tier
Before the spec tables, a word about where the big socket lands. Venice SP7 is reserved for AMD’s highest-end, no-compromise performance tier: datacenter, hyperscaler, and HPC compute deployments, where throughput per rack is the deciding factor, with no expense spared. The server vendors are already re-tiering around it. Dell has split its long-established PowerEdge naming schema to make room, and this class of compute now sits in a new 9000 series led by the PowerEdge R9825 and Rack Scale M9825, the flagships with 2 Venice chips in a 3U chassis.
The Headline Specs
The whole family builds on one platform, and nearly every corner of it is new since Turin.
Start with the cores. Zen 6 comes in two flavors: the compact Zen 6c packs 256 cores and 512 threads into a socket at up to 600W, while the standard core stops at 96 but clocks to 5GHz. AMD claims upwards of 20% more per-core performance than competing parts at matched core counts, and density does not thin the cache: even the 256-core die keeps about 4MB of L3 per core, a full gigabyte on the flagship. AVX-512 handles the tokenizers and the 3-to-8-billion-parameter models that increasingly run on the CPU itself.
Feeding those cores requires a much larger memory system: 16 channels of DDR5 at 8,000MT/s, or MRDIMM at 12,800, good for 1.6TB/s per socket, whereas Turin managed 614GB/s from 12 channels. I/O makes the same jump. Venice is the first server CPU with PCIe Gen 6, 128 lanes of it at 64 GT/s, double the per-lane bandwidth of anything else attaching to accelerators today, and the foundation of the host-node claim we unpack later. CXL 3.1 rides on those lanes for memory expansion, and select two-socket AI host platforms trade inter-socket xGMI width for I/O to reach 160 usable lanes. Two quieter additions move data without burning cores: SDXI offloads memory copies and 30 to 50 cores’ worth of crypto, and Smart Data Cache Injection drops network packets straight into cache instead of routing them through DRAM.
Power management gets three new knobs, and they matter because every rack comparison AMD makes later happens inside a power budget. UPP merges the SoC and DIMM budgets into one, so a memory-bound workload shifts watts to the DIMMs and a compute-bound one pulls them back, which either raises performance inside a fixed budget or frees rack power for more nodes. UBPS caps power at low utilization, trading the usual light-load performance bump for a flat, predictable load line; operators who prefer the bump can switch it off. FAST splits one SoC into two personalities, keeping priority cores at high frequency while background cores run slower, so that latency-critical work sharing a socket with batch jobs doesn’t suffer.
For security, Venice extends a confidential-computing lineage that runs from SEV on EPYC 7002 through encrypted state, secure nested paging, and the Trusted I/O that arrived with Turin. New this generation: an enhanced root of trust with RSA-4K and post-quantum algorithms, physical and side-channel attack mitigations, FIPS 140-3 Level 1 certification, and Device Provenance, an attestable manufacturing history that Venice will be among the first AMD products to carry.
Inside the Portfolio
The four parts differ more than the shared name suggests.
| Specification | 9006 SP7 | 9006 SP8 | 9006X SP7 | 9006 LP “Verano” |
|---|---|---|---|---|
| Cores | Up to 256 (512 threads) | 8 to 128 | Up to 96 | Up to 72 |
| Peak frequency | 5.0GHz with 96 cores (HF) | HF options offered | ~5.15GHz | Up to 5.0GHz |
| Memory | 16ch, up to 1.6TB/s | 8ch with 2 DIMMs per channel | 16ch MRDIMM at 12,800 MT/s | 24ch LPDDR5X, SOCAMM2 |
| L3 cache | Up to 1024MB | Core-count dependent | 1152MB (3D V-Cache) | Core-count dependent |
| I/O | PCIe Gen 6, 128 lanes 1P | 128 PCIe lanes, 1P and 2P | PCIe Gen 6 | Enhanced xGMI at 112 GT/s |
| Aimed at | Hyperscale density, AI host | Enterprise, edge, NEBS-friendly | HPC, AI pre-processing | Rack-scale AI host |
EPYC 9006 SP7
SP7 is the flagship socket and the one in production now; partner platforms follow in Q4. Its high-frequency bins succeed the Turin HF parts AMD says were among its fastest-ramping SKUs, and their biggest deployment is Helios, where the host CPU in every compute tray joins the rack’s coherent memory domain over Infinity Fabric with 1TB of DDR5 behind it. Because the socket is standard, any SP7 SKU up to the 256-core flagship drops in, a point we covered in the sister article.
EPYC 9006 SP8
SP8 trades the giant socket for system economics. It spans 8 to 128 cores, runs 8 memory channels with 2 DIMMs per channel, maintains 128 PCIe lanes, and comes in 1P and 2P with high-frequency and NEBS-friendly options for telco and edge deployments. AMD’s pitch is right-sized performance for the enterprise, where the metric that wins deals is performance per system dollar rather than per rack.
EPYC 9006X “Venice-X”
Venice-X stacks 3D V-Cache on the 96-core high-frequency configuration and takes L3 to 1152MB, roughly 3× the cache per core of the standard SP7 parts. Clocks reach about 5.15GHz, and the full 16-channel, 12,800 MT/s memory system carries over. The targets are simulation and modeling, large-scale analytics, in-memory databases, and AI pre-processing that feeds training pipelines—workloads where a working set held in cache matters more than extra cores.
EPYC 9006 LP “Verano”
Verano is the most focused part in the family: an AI host node first, with up to 72 cores at 5GHz, 24 channels of LPDDR5X, and an enhanced 112 GT/s xGMI link for CPU-to-GPU traffic. The LPDDR sits on SOCAMM2 modules that can be replaced in the field, which matters in a fleet where a soldered-down memory failure would otherwise scrap a whole board. It is also AMD’s direct answer to NVIDIA’s Vera, a matchup we will come back to, because AMD certainly did.
How Venice stacks up
AMD made its case in two passes, leading with Turin to show the lead it already holds, then Venice to show how far it is pulling away.
Turin’s lead today
AMD aimed most of its comparisons at NVIDIA’s Vera, the Arm CPU in the Vera Rubin platform, and it opened with the generation it already ships. According to AMD’s rack-level modeling, a rack of Turin delivers 2.4× Vera’s general-purpose throughput and twice the agents per watt. The model fixes both racks at a 100kW power budget, matches a 2P EPYC 9965 with 384 cores against a 2P Vera board with 176, and averages the results across six workloads: estimated SPECrate 2017, server-side Java, NGINX, Redis, Memcached, and a TPC-C derivative.
Two caveats apply here and to the rest of this article. Vera has not shipped, so every Vera figure is an AMD estimate of the 88-core, 450W part NVIDIA has described in its public materials. And because no standard benchmark exists for agentic workloads, AMD counts hardware threads as a proxy for agents in its agents-per-watt math.
The Intel comparison rests on firmer ground, since both sides are shipping parts that AMD could benchmark directly. Turin’s host CPUs reach 5.0GHz, while the comparable Xeon tops out at 3.9GHz, a 28% frequency advantage that matters when single threads feed GPUs. Socket-for-socket against the 128-core Xeon 6980P, AMD measures Turin at up to 1.4X on SPEC CPU, 1.4X on enterprise Java, 1.8X on HPC, and 1.7X on CPU-based AI.
Venice widens every gap
Venice extends each of those leads. Under the same 100kW rack model, now with 512 Venice cores against Vera’s 176, AMD claims 3.3X Vera’s general-purpose performance. Venice achieves 2.8× the agents per watt of the 136-core Arm AGI, and 1.8× the tokens per second on frontier models when Venice hosts GPUs. That 1.8× figure measures a narrower scenario than the headline suggests, and we return to it in the host-node section below.
Head-to-Head with Vera
On SPECrate 2026, AMD estimates a 2P Venice 9996 at 2070 against 925 for Vera, a 2.2X throughput advantage, and puts the 96-core high-frequency part about 1.2× ahead of Vera per core. Both results were compiled with GCC 15.2, the same toolchain NVIDIA used for its published Vera numbers.
The per-core claim actually grew during launch week. Ravi Kuppuswamy, who runs AMD’s CPU engineering, told the press that AMD had originally claimed a 10% per-core advantage. After NVIDIA published its own Vera numbers earlier in the week, its team reran the comparison using the same compiler and settings and measured a 20% lead, with tuning still unfinished. The 2.2× throughput gap, he said, matched AMD’s internal projections all along. A separate endnote recalculates the per-core comparison on SPECrate 2017 using AMD’s own AOCC compiler and arrives at 1.7×, making 20% the more conservative of the two calculations.
Intel and Arm
AMD also showed a broader SPECrate 2017 ranking. Venice 9996 leads at 4900 (256 cores, 600W, $14,904 at 1Ku), followed by Turin 9965 at 3240, Intel’s Xeon 6980P at 2510 (128 cores, 500W, $13,955), Arm AGI at 1944 (136 cores, 300W), and Vera at 1459. The 9996, AGI, and Vera figures are AMD estimates; the Turin and Intel scores are published results. Against Intel, that works out to roughly twice the throughput at a comparable list price, or double the performance per dollar. On a per-core basis, a 128-core, 500W Venice configuration scores 1.3× the 6980P, while Arm AGI lands at 0.7×. Unlike the Vera comparison above, each vendor’s score here was produced with its own best toolchain: AOCC for AMD, OneAPI for Intel, and GCC 13 for Arm.
Cloud native and HPC
AMD also broke the comparison down by workload, with every result indexed to Intel’s 6980P at 1.0. On the cloud-native side, the 256-core Venice 9996 posts 3.5× on MongoDB, 2.9X on Redis, 3.7X on NGINX, and 2.6X on MySQL. Turin lands between 1.6X and 2.4X on the same tests, and Graviton5 between 1.2X and 1.5X.
In HPC, Venice scores 3.1X on GROMACS and NAMD, 2.9X on WRF, and 1.8X on Quantum Espresso. Memory configuration moves these numbers substantially. With standard 8,000 MT/s RDIMMs, Venice scores 2.81X the Xeon baseline on GROMACS; 12,800 MT/s MRDIMMs raise that to 3.13X, and the same upgrade takes WRF from 2.25X to 2.90X. AMD summarizes its HPC lead as anywhere from 1.8X to 3.5X depending on workload and memory. The Intel baseline ran 8,800 MT/s MRDIMMs, so the gap does not come from pairing upgraded AMD memory against a slow Intel configuration.
The host node and the 1.8× claim
This brings us back to the 1.8X tokens-per-second claim. The hardware improvements behind it are easy to list: the 5GHz host part now carries 96 cores instead of 64, host memory bandwidth rises from 614GB/s to 1.6TB/s, and PCIe Gen 6 doubles the bandwidth of the CPU-to-GPU link.
The first thing to check is the baseline. AMD indexed this chart to Turin at 1.0 rather than to Intel, and placed the 6th Gen Xeon at 0.9, which is why the same claim approaches 2× when restated against the Xeon 6960P. The projected gain comes almost entirely from I/O bandwidth. AMD ran a CPU-offload test that streamed Qwen3-30B’s BF16 weights from host memory to the GPU and found the PCIe Gen 5 x16 receive path running at 89% to 95% of its theoretical ceiling, while host DRAM reads stayed below 10% of their theoretical ceiling. Decode speed was limited by the transfer, so doubling the link with PCIe Gen 6 yields a 1.8X to 1.9X improvement in this test. Venice is currently the only server CPU that offers PCIe Gen 6 to accelerators, so the advantage is real and, for now, exclusive. Its scope is narrow, though: the 1.8X describes one bandwidth-bound offload pattern, and it should not be read as a promise that GPUs serve every model 1.8X faster on Venice hosts. In hardware shipping today, a Turin host beats a Xeon 6960P by a geometric mean of 13% in time-to-first-token across vLLM and NIM workloads on the same 8X B200 system, with a best-case improvement of 36%.
Rack density
The last claim is density: 49,152 Venice cores in a rack, against 36,864 for Turin and 22,528 for Vera, with each part carrying twice as many threads as cores. Of all the claims in the deck, this one changes the most once the endnotes are applied.
The headline is 2.2X Vera’s cores, but the Venice rack reaches it while drawing 269kW, compared to Vera’s 185kW, so the two racks are not compared at matched power. AMD’s endnotes supply normalized versions. With both racks held to a 100kW envelope, the 9996 delivers 2.08X Vera’s cores, 1.86X Turin, and 1.24X Intel’s 6980P. Substituting the 400W EPYC 9956 for the flagship, AMD estimates 1.91X Vera’s cores at 26% less power, resulting in 2.57X cores per rack watt. Depending on which power budget is held fixed, the lead lands between 1.9X and 2.2X, and, as with every Vera figure in this article, the NVIDIA rack is modeled from published specifications rather than measured hardware. The density advantage survives normalization; it is simply smaller than the 2.2X headline.
The big bully
Set the individual claims aside: at the top end of the server CPU market, AMD currently has no direct rival. Intel’s best part trails Turin, the generation AMD is already replacing. The Arm challengers sit below Turin on every chart AMD showed. Vera has not shipped, and AMD’s estimate puts its throughput at less than half that of a Turin 9965, which has been on the market for a year and a half, with Venice roughly doubling Turin’s throughput again. That position has carried AMD to 46% of server CPU revenue, per Mercury Research. The market evidence points in the same way. Demand for top-end EPYC parts currently outruns supply, and lead times have stretched as a result.
We can add our own data point here. Our 314-trillion-digit Pi world record ran on a Dell PowerEdge R7725 with two 192-core EPYC 9965s and 1.5TB of DDR5, the same 384-core 2P configuration AMD models its Turin racks on. The computation kept every core loaded for 110 days and finished without a second of downtime; a single memory error or crash at any point would have scrapped the attempt. Power draw averaged about 1,600W, and the complete run consumed 4,305 kWh, or 13.7 kWh per trillion digits, against an estimated 33,600 kWh over roughly 225 days for the previous 300-trillion-digit record.
Conclusion
Venice is the strongest server CPU launch AMD has ever staged, and its real subject is breadth. One Zen 6 generation spans from a 256-core-density part through an enterprise socket and an 1152MB stacked-cache HPC chip to an LPDDR host node, so a customer can build every tier of a data center on the same architecture and software base. The numbers that anchor the launch hold up once the endnotes are read: roughly double Intel’s throughput at a comparable list price, a 20% per-core lead over Vera on a matched compiler, and rack density between 1.9X and 2.2X of NVIDIA’s host CPU, depending on the power envelope.
The commercial half of the story needs no forecasting. SP7 is in production now, with partner platforms due in Q4; SP8 follows in the first half of 2027, with Venice-X and Verano behind it. The company selling them already books 46% of server CPU revenue, and demand for its current parts far outstrips supply, so buyers wait in line rather than defect. Venice could slip a year, and the comparisons in this article would still read as leadership, because Turin already outscores everything else on the chart, shipped or unshipped. AMD is so far ahead that its nearest competitor, for the moment, is its own last generation.




Amazon