NVIDIA and Microsoft opened pre-orders for RTX Spark Windows laptops today, with availability set for October 16, compact desktops to follow in November, and a preview of DGX Station for Windows that NVIDIA will release in the fourth quarter. The announcements came at Microsoft’s Windows AI and Surface event in San Francisco, where Jensen Huang joined Satya Nadella on stage, and Microsoft declared its agent containment layer, Microsoft Execution Containers, generally available on Windows 11.
We covered the first of these systems earlier in the day in our Dell XPS 16 Creator Edition write-up, and the platform itself at Computex in June. This piece is about what the two companies published around the event that wasn’t known before: power envelopes for the N1X superchip, how Windows divides the unified memory pool, measured throughput for a 137-billion-parameter coding model on a laptop, the mechanics of the sandbox agents will run in, and a full specification table for the Windows DGX Station.
RTX Spark N1X: Two Chips, Three Power Envelopes
NVIDIA’s RTX Spark specification table now separates laptops from desktops, and the split is more informative than the headline numbers. Laptops get two N1X variants: a 6,144-core Blackwell RTX GPU with a 20-core Grace CPU and up to 128GB of unified LPDDR5X, and a 5,120-core GPU with an 18-core CPU that tops out at 64GB; both carry a 45 to 80W TDP. Desktops get only the 6,144-core, 20-core part with up to 128GB, and NVIDIA lists it at 140W.
That 140W figure is the same TDP NVIDIA publishes for the GB10 in DGX Spark, covering CPU and GPU together, which fits what NVIDIA said in June about the two chips being the same system tuned for different platforms. It also means a laptop N1X runs at 32% to 57% of the power the desktop part is rated for, while every system in the family is sold with the same “up to 1 petaflop” of FP4. Microsoft’s Surface footnotes describe that number as theoretical FP4 performance using sparsity.
| Specificazione | RTX Spark N1X (Laptop) | RTX Spark N1X (Laptop) | RTX Spark N1X (Desktop) | GB10 (DGX Spark) |
|---|---|---|---|---|
| GPU | 6,144-core Blackwell RTX | 5,120-core Blackwell RTX | 6,144-core Blackwell RTX | Blackwell, 5th-gen Tensor Cores, 4th-gen RT Cores |
| CPU | 20-core Grace | 18-core Grace | 20-core Grace | 20 Arm cores (10 Cortex-X925, 10 Cortex-A725) |
| TDP | Da 45 a 80 W. | Da 45 a 80 W. | 140W | 140W (240W power supply) |
| Memoria unificata | LPDDR128X fino a 5 GB | LPDDR64X fino a 5 GB | LPDDR128X fino a 5 GB | 64GB or 128GB LPDDR5X, 273GB/s |
| Media engines | 1x 9th-gen NVENC, 1x 6th-gen NVDEC | 1x 9th-gen NVENC, 1x 6th-gen NVDEC | 1x 9th-gen NVENC, 1x 6th-gen NVDEC | 1x NVENC, 1x NVDEC |
| Display and I/O | PCIe Gen5, HDMI 2.1b, DisplayPort 2.1b | PCIe Gen5, HDMI 2.1b, DisplayPort 2.1b | PCIe Gen5, HDMI 2.1b, DisplayPort 2.1b | HDMI 2.1a, ConnectX-7 200Gb/s, 10GbE |
| Sistema operativo | Windows 11 | Windows 11 | Windows 11 | Sistema operativo NVIDIA DGX |
NVIDIA doesn’t publish a dense FP4 figure for N1X; on the GB300, where it publishes both, the dense number is 75% of the sparse one. We expect OEM thermal design to separate these laptops in a way it never separated DGX Spark units, and a 35W spread inside the laptop TDP range, before any vendor tuning, is the first hard evidence for it.
Microsoft adds a wrinkle on the desktop side. Its Surface RTX Spark Dev Box page quotes a 100W thermal envelope inside an aluminum chassis with 1,000 air vents, which is 40W under the TDP NVIDIA lists for desktop N1X. The two companies may be measuring different things, a chip rating in one case and a chassis design point in the other, but a buyer comparing compact desktops in November should ask each vendor what sustained power its box holds.
Memory speed is the other platform difference from GB10. Dell’s XPS spec sheet lists LPDDR5X at 9,400 MT/s, which on a 256-bit interface like the GB10’s works out to 300.8GB/s, matching the 300GB/s NVIDIA quoted in June and about 10% above DGX Spark’s 273GB/s. The CPU and GPU halves connect over NVLink-C2C at 600GB/s. The rest of the table is common to all three parts: fifth-generation Tensor Cores, fourth-generation RT cores, DLSS 5, Reflex 2 with Frame Warp, and HDMI 2.1b at up to 4K 480Hz or 8K 120Hz with DSC.
NVIDIA’s marketplace lists 17 laptop configurations at pre-order, and they show how much of the range sits below the headline. Of the 15 we read through, seven use the 5,120-core chip, and memory runs 24GB, 32GB, 48GB, and 128GB, with storage split between PCIe Gen4 and Gen5 drives.
How Windows Divides 128GB of Unified Memory
The memory is dynamically shared between CPU and GPU, and the maximum the GPU can address depends on configuration. A Task Manager screenshot in Microsoft and GitHub’s engineering post shows what that looks like on a Surface Laptop Ultra with the 6,144-core N1X. Windows reports 74.9GB of dedicated GPU memory, 35.5GB of shared GPU memory, 110GB of total GPU memory, and 51.5GB of system memory.
The dedicated and system figures sum to 126.4GB, so Windows is presenting the unified pool the way it presents a discrete card: a fixed dedicated allocation, with the GPU able to borrow from system memory up to a cap. On that unit, the GPU’s ceiling is 110GB, or 86% of the installed 128GB, and weights, key-value cache, and runtime overhead all have to fit under it. The same screenshot shows driver 32.0.16.1630, dated July 31, and an NPU listed as a separate device from the GPU.
The post is explicit that model weights are only part of the memory budget; the operating system, applications, the inference runtime, and the KV cache that holds attention state for an agent’s growing context all draw from the same pool. For the 24GB and 32GB configurations that make up much of the pre-order list, that leaves room for 27B-class models at 4-bit and not much above. The large-model claims in the launch material all assume the 128GB SKU.
What Fits: Parameter Counts and Bits per Weight
The parameter ceilings quoted for the same chip appear a little misaligned. NVIDIA’s product page says models up to 120 billion parameters with a 1 million token context, and its blog cites Qwen 3.8 Flash Next as a 125B example. Microsoft says models exceeding 120 billion, then names MAI Code 1.1 Flash at 137 billion and DeepSeek V4 Flash at 284 billion as models coming to RTX Spark, and Dell quotes 200 billion. The spread comes down to how many bits each weight gets.
| Modello | Scheda Sintetica | Precisione | Orma | Fonte |
|---|---|---|---|---|
| MAI Code 1.1 Flash | 137B total, 6.8B active (mixture of experts) | Mixed precision, about 3.3 bits per weight | 53GB weights; 75.5GB peak at 256K context | Microsoft, GitHub |
| NVIDIA Nemotron (upcoming) | Oltre 70 miliardi | 2-bit | Just over 20GB | Microsoft |
| Flash DeepSeek V4 | 284B | Non specificato | Non specificato | Microsoft |
| Qwen 3.8 Flash Next | 125B | Non specificato | Non specificato | NVIDIA |
| Qwen3.5 27B | 27B | Q4_K_M | Non specificato | Microsoft benchmark footnote |
Microsoft’s numbers for MAI Code 1.1 Flash are the most complete. The cloud model in BF16 would be 274GB; the on-device build is 53GB, which Microsoft calls an 80% reduction, at roughly 3.3 bits per weight with mixed-precision quantization. The Nemotron figure checks out the same way, since 2 bits across 70 billion weights is 17.5GB before overhead. By our math, DeepSeek V4 Flash at 284 billion parameters needs about 3.1 to 3.3 bits per weight or fewer to fit under a 110GB GPU ceiling, before any context is loaded.
The quality data is what makes the 3-bit claim worth attention. Microsoft reports the quantized on-device model at 70.8% on SWE-Bench, verified against 72.6% for the BF16 cloud version, and 66.29% on Terminal-Bench 2.1 against 62.9%. The second result has the quantized model ahead, which, on an 89-task set, is a three-task difference and reads to us as noise. Unsloth’s GPT-OSS-120B GGUF, the comparison point, scored 32.0% and 23.6% on the same two tests.
Throughput comes from a Surface Laptop Ultra on a Windows Arm64 build of llama.cpp with the CUDA backend and what Microsoft calls DFlash2 sliding-window speculative decoding, tested with a synthetic code-generation workload. Microsoft’s chart carries no data labels, so these are our estimates: about 63 tokens per second at a 2K prompt, 61 at 8K, 57 at 32K, 52 at 64K, 55 at 120K, and 39 at 256K. Peak memory at 256K context is 75.5GB, which lands right at the dedicated GPU figure in the Task Manager capture.
Prompt processing is the number to lock in on for agent work. Microsoft gives 923.5 tokens per second at 64K context and 769.8 at 128K. A cold 64K prompt takes about 71 seconds to process, and a cold 128K prompt takes about 170 seconds before the first output token appears. The blog post makes the related point that keeping a model loaded doesn’t guarantee constant response time, and that Copilot weighs cache state when it routes work, which is why prefill is the figure we’ll measure first on review hardware.
The MacBook Comparison and Its Footnotes
Microsoft’s comparison puts RTX Spark PCs against a 16-inch MacBook Pro with M5 Pro and 64GB: up to 2.1x faster time to first token, 4.3x faster image generation, and 6.2x faster video generation. The first comes from Microsoft-commissioned testing in llama.cpp with Qwen3.5 27B at Q4_K_M and a fixed 8,192-token prompt. The other two are NVIDIA’s tests in ComfyUI, with FLUX.2 Klein 4B at four steps and 1024 x 1024, and LTX 2.3 22B at 121 frames, eight steps, and 1280 x 720.
The footnotes carry several qualifiers, starting with the hardware: all of the tests were run in September on preproduction RTX Spark systems with 64GB, and the Apple system is the M5 Pro, with no M5 Max in the comparison. Both ComfyUI tests used the models at NVFP4 precision, and the footnotes don’t describe the Mac’s software configuration. The language-model result is a prefill measurement at one prompt length; neither company published a decode tokens-per-second comparison, which is the number most local LLM users quote.
When it comes to gaming, NVIDIA says RTX Spark runs AAA titles at 1440p above 100 frames per second with DLSS 5, Reflex, and G-SYNC, without naming titles or settings. Microsoft’s contribution is Gears of War: E-Day with DirectX ray tracing, variable rate shading, DirectStorage, and Advanced Shader Delivery, which Microsoft says cuts first-launch shader compilation from minutes to seconds, plus a commitment that Call of Duty comes to RTX Spark in 2027. The Prism and anti-cheat details are in our Dell coverage.
Microsoft Execution Containers: How the Agent Sandbox Works
MXC is the piece of the announcement with the longest reach, since it applies to any Windows 11 PC and ships as an open-source library and multi-language SDK with backends for macOS and Linux. A developer declares what a workload needs in a JSON policy: the files it can modify, the files it can only read, inbound and outbound network access, and whether it can touch the desktop. MXC maps that policy onto an isolation backend and enforces it at runtime, and the policy sits outside the workload, so generated code can’t grant itself more access.
| MXC Backend | Piattaforme | Meccanismo | Destinazione d'uso |
|---|---|---|---|
| Contenitore di processo | Windows 11, macOS, Linux | AppContainer on Windows, Seatbelt on macOS, Bubblewrap on Linux | Low-latency containment for model-generated code and tool execution |
| Session container | Solo Windows 11 | Separate Windows account and session with its own desktop, clipboard, UI, and input | Long-running agents and automation that need a desktop |
| WSL container (WSLc) | Solo Windows 11 | Linux execution environment through WSL | Linux-first agent toolchains |
| MicroVM | Windows 11 and Linux (experimental) | Hardware-enforced virtualized boundary | Higher-risk workloads and untrusted code |
Three operating modes cover the policy authoring cycle: Enforcement blocks anything outside the policy, and Learning blocks it and records it to a JSON activity report. Permissive allows the operation and records what would have been denied, so a team can run an agent through real work and build a least-privilege policy from the evidence. The activity report is available only for process containers on Windows. Organizations can add their own constraints over a developer’s policy through management tools, and Microsoft says Intune management of MXC process containers is coming soon.
It’s worth being clear about what is generally available, and containment is what shipped on October 7. Agent identity, where Microsoft Entra distinguishes an agent’s actions from the signed-in user’s, is described as coming soon, as is extending Agent 365 controls to local agents. Until those arrive, an MXC container limits what an agent can reach, but security tooling still sees its activity under the user’s identity, and the fleet-level governance Microsoft describes isn’t in place yet; its own footnote adds that governance at scale may require additional services.
GitHub’s description of its own integration shows where the boundary sits. Copilot uses the base tier of the process container on Windows, and shell commands, along with local Model Context Protocol servers and language servers by default, run inside it. Copilot’s built-in file tools run in the agent harness, where requests are checked against policy but aren’t isolated by the operating system, and remote MCP servers are outside the local sandbox entirely. Local inference doesn’t change any of this, since a shell command inherits the same access whichever model asked for it.
Agents that support MXC today include GitHub Copilot, OpenAI Codex, OpenClaw, Replit, LM Studio, and Unsloth AI, with Claude Code, Perplexity, Manus, and Hermes Agent among those listed as coming. NVIDIA has integrated OpenShell into MXC, which Microsoft says adds policy controls for agent access to files and inference services, advanced network controls, credential management, and OCSF auditing for enterprises.
Huang put the weight of the announcement on the container layer: “Just as Windows and DirectX revolutionized how applications were built, MXC is going to revolutionize how agents are built and deployed.” Our view is narrower: the policy model is sound, and the cross-platform SDK lowers the cost of adopting it, but the one integration described in detail so far uses the lightest backend, and the isolation an agent gets is only as strong as the backend its developer chose.
Hybrid Routing: HydraFusion, Windows ML, and llama.cpp
The local models reach developers through GitHub Copilot. HydraFusion, which GitHub launched earlier this year to route each task to the right cloud model, is being extended to models on the device, with an experimental preview in the Copilot app, Copilot CLI, and Visual Studio Code later in October. Developers can leave placement on Auto or select a local model explicitly, either MAI Code 1.1 Flash through the Windows ML provider or any model behind an OpenAI-compatible local endpoint.
Microsoft also added llama.cpp support to Windows ML, which is the runtime the MAI Code numbers above were produced on, and it gives the open-model ecosystem a supported path onto the platform; our I migliori strumenti LLM locali page tracks the runtimes and front ends built on it. Microsoft’s own sizing of the claim is modest, though. Its Dev Box page says intelligent routing can offload up to 20% of Copilot workloads to local models, which frames local inference as a way to stretch a token budget, with the cloud still doing most of the work.
Systems, Prices, and Dates
Six OEMs have laptops on pre-order for October 16: the ASUS ProArt P16 and P14, Dell XPS 16 Creator Edition, HP OmniBook Ultra 16, Lenovo Yoga 9n 2-in-1, MSI Prestige N16 Flip AI+, and Microsoft’s Surface Laptop Ultra. NVIDIA’s marketplace also lists an HP OmniBook X 14. MSI’s machine has a 16-inch UHD+ tandem OLED and a 99.9Wh battery. NVIDIA additionally names Acer and Gigabyte among the vendors shipping RTX Spark systems.
Surface Laptop Ultra starts at $2,599, the lowest RTX Spark price we’ve seen, in a chassis Microsoft puts under 18mm and 4.5 pounds. Microsoft claims 2.5x the thermal capacity of its current Surface Laptops, a 15-inch 120Hz touchscreen at 3270 x 2180 with 2,000 nits of peak HDR brightness measured over a 10% window, user-removable storage, and a magnetic USB-C charging port that still carries video and data. A 140W power supply ships with select configurations. The lowest configuration on NVIDIA’s marketplace pairs the 5,120-core chip with 24GB and a 512GB Gen4 SSD.
The Surface RTX Spark Dev Box is the first compact desktop with a price, $5,999 with 128GB, sold only through Microsoft.com in the US and shipping in November, with Visual Studio Code, Git, GitHub Copilot, WSL, Python, and Node preinstalled. That is $1,000 above the $4,999 64GB DGX Spark, with twice the memory. Microsoft lists two USB-C ports, USB-A, HDMI, Ethernet, and a headphone jack, and makes no mention of the ConnectX-7 ports that let DGX Sparks cluster, which is one answer to the desktop networking question we raised in June.
DGX Station for Windows: Fourth Quarter, With the Full Specification Table
NVIDIA’s page for DGX Station for Windows now reads “Coming in Q4,” and Microsoft says the systems arrive later this year, naming the Dell Pro Precision with GB300 and the HP ZGX Fury AI Station first. NVIDIA’s partner list adds ASUS, Gigabyte, MSI, and Supermicro. The hardware is the GB300 Grace Blackwell Ultra Desktop Superchip we’ve tested on Ubuntu in the MSI XpertStation WS300, il ASUS ExpertCenter Pro ET900N G3, e un two-tower cluster.
| Specificazione | NVIDIA DGX Station for Windows |
|---|---|
| GPU | 1x NVIDIA Blackwell Ultra, 252GB HBM3e at 7.1TB/s |
| CPU | 1x Grace, 72 Neoverse V2 cores, 496GB LPDDR5X at 396GB/s |
| Coherent memory | Up to 748GB; NVLink-C2C at 900GB/s |
| Nucleo tensoriale FP4 | 20 PFLOPS with sparsity; 15 PFLOPS without |
| FP8 and FP6 Tensor Core | 10 PFLOP |
| FP16 and BF16 Tensor Core | 5 PFLOP |
| Nucleo Tensoriale TF32 | 2.5 PFLOP |
| FP32 / FP64 | 80 TFLOPS / 1.3 TFLOPS |
| Networking | ConnectX-8 SuperNIC, up to 800Gb/s: 2x QSFP112 at 400Gb/s per port 1x 10GbE RJ-45, 1x 1GbE RJ-45 (BMC) |
| Archiviazione | 4x M.2 PCIe Gen5 slots |
| Slot PCIe | 1x PCIe Gen5 x16; 2x PCIe Gen5 x16 (x8 electrical) |
| Supported RTX PRO GPUs | RTX PRO 6000 Workstation Edition, RTX PRO 6000 Blackwell Max-Q Workstation Edition, RTX PRO 4000 Blackwell SFF Edition, RTX PRO 2000 Blackwell (support may vary by system) |
| decoder | 7 NVDEC, 7 nvJPEG |
| Sistema di alimentazione | 1,600W; 20A circuit required |
| Sistema operativo | Microsoft Windows; Linux toolchains through WSL |
Microsoft describes two uses for the Windows DGX Station. One is a personal AI supercomputer for engineering and research work, running models it lists as Llama 4 Maverick, Kimi K2.6, and DeepSeek V4 Pro, on hardware its footnote says supports models up to 1 trillion parameters. The other is what it calls a Windows token factory, a shared system supporting 32 or more simultaneous agents for a team to reduce cloud token consumption. Microsoft says IT can manage these systems with the same tools it uses across its Windows fleet.
The optional RTX PRO card matters more on Windows than it did on Linux. As we noted at Computex, the Blackwell Ultra GPU in the GB300 has no RT cores, so ray-traced visualization in Windows design and engineering applications runs on the add-in card, with the superchip handling AI. NVIDIA lists four supported RTX PRO Blackwell cards, from the RTX PRO 2000 to the 6000, and suggests that support may vary by system.
The open question from June is still open. Our GB300 work has all been on Ubuntu with vLLM and the NVIDIA container stack, and neither company has said which serving stacks run natively on Windows on this hardware and which run under WSL, or what WSL costs in throughput on a 72-core Grace host with two 400Gb/s ports. Those are the first things we’ll test when a Windows unit reaches the lab.
A cosa ammonta
The launch material answers several of the questions we left open at Computex. Laptops run the chip at 45 to 80W and desktops at 140W; the 128GB option is tied to the larger of two chips; Windows caps GPU-addressable memory below the installed total; and Microsoft’s first compact desktop lists no high-speed networking. The line from Pavan Davuluri, Microsoft’s executive vice president of Windows and Devices, that “you can run models on this laptop that simply don’t fit on a traditional machine” holds for the 128GB configurations, and Microsoft’s 137B coding model at about 60 tokens per second is a demonstration of it.
What’s still missing is everything that depends on a specific chassis: sustained power under a long agent session, decode throughput against Apple’s M5 Max and AMD’s Strix Halo systems, battery life under inference, and pricing for most of the 128GB laptops. We also don’t know when Windows reaches the GB10-based DGX Spark, which picked up a Windows boot certificate in firmware last month. Review units will go through the same suite behind our I migliori laptop per l'IA locale and I migliori computer desktop per l'intelligenza artificiale locale classifiche.




Amazon