AI/News
Microsoft's Surface Laptop Ultra runs models "exceeding 120B parameters" on the laptop for $2,599, as Windows makes its agent sandbox generally available
Microsoft's two pages: NVIDIA RTX Spark with up to 128 GB of unified memory and "up to 1 petaflop" (FP4 with sparsity, says the footnote), pre-orders today and shipping 16 October; the $5,999 Dev Box ships in November; Microsoft Execution Containers "becomes generally available on Windows 11" with Codex, GitHub Copilot and Replit already on it and Claude Code "releasing support"; MAI Code 1.1 Flash runs locally at 3-bit.
By Daily Aletheia · Checked against the primary source · 8 October 2026 · 3 min read

Microsoft used a 7 October event to put a name on its direction for the PC - "hybrid intelligence", in the words of Pavan Davuluri, who runs Windows and devices: "a platform where agents can run locally when it makes sense, reach the cloud when they need to". Two things on the pages are concrete today.
The laptop. Surface Laptop Ultra "unites an NVIDIA Blackwell RTX GPU with up to 6,144 cores, an NVIDIA Grace CPU with up to 20 cores, and up to 128 GB of unified memory", and, the page says, can "Run AI models exceeding 120B parameters locally, with up to 1 petaflop of AI performance" - a figure the footnote defines as "theoretical FP4 performance of 1 petaflop using the sparsity feature". Pre-orders open today "starting at $2,599 (MSRP)", "with availability beginning October 16". The desk version, Surface RTX Spark Dev Box, is "$5,999 (MSRP), exclusively on Microsoft.com in the U.S." and "will begin shipping to customers this November". ASUS, Dell, HP, Lenovo and MSI laptops on the same chip "begin shipping October 16". Microsoft's own comparison is to Apple: up to "2.1x faster time to first token" than a MacBook Pro 16-inch with M5 Pro, and "up to $1,000 cash back when you trade in an eligible MacBook Pro".
The sandbox. "Today, Microsoft Execution Containers (MXC), becomes generally available on Windows 11, enabling organizations to define which files and networks agents can access, with those policies enforced at runtime." Already on it: "Codex from OpenAI, GitHub Copilot, OpenClaw, Replit, LM Studio, OpenShell from NVIDIA, and Unsloth AI". Coming: "Anthropic Claude Code, Box, Egnyte, Heidi Health, Hermes Agent by Nous Research, Manus, Perplexity, Raycast, and Simular amongst others", plus Meta's "Muse for Windows" as a native app.
The models that run on the machine. MAI Code 1.1 Flash, "137 billion total and 6.8 billion active parameters", comes to the device "using 3-bit precision to reduce the model size by nearly 80%" with "a 256K context window locally"; an "upcoming Nemotron model from NVIDIA" at "over 70 billion parameter" in 2-bit "to use just over 20GB of memory"; and "DeepSeek V4 Flash, a 284B parameter model". GitHub's HydraFusion routing reaches local models "in experimental preview later in October". Above all of it, "DGX Station for Windows" on the GB300 chip, "later this year", for models "at more than a trillion parameters".
The reason Microsoft gives. "customers' needs are outpacing what their cloud budgets can support. They want to make every AI token count without giving up frontier capabilities". And its scale claim: "over 2 trillion inferences are run locally per month" on Copilot+ PCs, with "over 40% of laptops being built for business" now in that class.
What the pages do not say. No battery figure for the laptop, no tokens-per-second for any model on it, no price for DGX Station for Windows, and no date for the Copilot features beyond "the coming months". The Apple comparisons are Microsoft's, on its own tests.