Shop for Local AI Hardware by Memory, Not by the Price Tag
On 7 October 2026, Microsoft opened pre-orders for the first Windows PCs built on Nvidia's RTX Spark chip. The 15-inch Surface Laptop Ultra starts at $2,599, and its top configuration costs $5,899.99. The Surface RTX Spark Dev Box costs $5,999. The same day, HP's US store briefly showed prices for its own RTX Spark laptops before taking them down.
That is the news, and plenty of sites have covered it. If you self-host, the launch itself matters less than one number on the spec sheet. That number is memory, and it decides what you can actually run.
What it costs to run AI locally now, and why price is the wrong number
For a long time, the honest answer to running AI locally was a gaming GPU in a tower, or a lot of patience. Now there is a mainstream price range for machines sold as local AI computers. Microsoft's two machines run from $2,599 to $5,999.
HP's prices are less certain. A leaked listing on HP's US store, later pulled, showed the 14-inch OmniBook X from $2,999.99 and the 16-inch OmniBook Ultra from $3,199.99. The same listing showed 64 GB configurations at $4,499.99 for the X 14 and $4,999.99 for the Ultra 16. HP announced these laptops in September 2026 without pricing and has not confirmed these numbers, so treat them as a leaked listing and nothing more.
Here is the problem with shopping by those prices. The cheapest Surface Laptop Ultra and the most expensive one share a name and a chip family, but they are very different machines for AI work. The $2,599 model has 24 GB of RAM and the 18-core RTX Spark. The $5,899.99 model has 128 GB of RAM and the 20-core part. For running models, that memory gap matters far more than the extra cores.
The same caution applies to the speed claims. Nvidia says RTX Spark delivers up to 2.1 times faster time-to-first-token, 4.3 times faster image generation and 6.2 times faster video generation than a MacBook Pro with an M5 Pro. Those are Nvidia's own figures from preliminary company testing. They are not independent benchmarks, and they tell you nothing about whether a specific model will fit on a specific configuration.
The spec that actually decides: unified memory
RTX Spark is an Arm-based chip that pairs an Nvidia Grace CPU with a Blackwell RTX GPU. The CPU and GPU share one pool of memory, up to 128 GB. That is what unified memory means. There is no separate graphics card with its own smaller pool of video memory.
This matters because a model has to fit in memory the GPU can reach before it can run at any useful speed. On a normal desktop, that limit is the video memory on the graphics card. On a unified memory machine, it is a large share of the whole pool.
Microsoft showed this on a 128 GB machine. Windows reported about 110 GB of GPU-addressable memory, roughly 79.9 GB dedicated plus 30 GB shared. Notice that it is not the full 128 GB. The operating system and your other apps need their share too.
So how much does a model need? A useful rule of thumb: at 4-bit quantization, which is how most people run models at home, the weights take roughly half a gigabyte per billion parameters. You then need extra room for the context, which grows with longer conversations and documents.
- Small models (around 7 to 8 billion parameters): roughly 5 GB at 4-bit. These run on a wide range of hardware.
- Mid-size models (around 30 billion parameters): roughly 15 to 20 GB at 4-bit, before context. A 24 GB machine gets tight once the operating system takes its share.
- Large models (around 70 billion parameters): roughly 35 to 40 GB at 4-bit, plus context. This is where 64 GB and 128 GB machines start to make sense.
Run at 8-bit for better quality and you need roughly double these numbers. These are estimates, not guarantees. Check the actual file size of the model you want and leave headroom.
Memory is also where the money goes. If you have watched RAM prices push your homelab budget around, the same pressure explains why memory prices move your whole build. Here, the jump from 24 GB to 128 GB is most of the gap between the cheapest and most expensive Surface Laptop Ultra.
Windows on Arm or Linux: the operating-system question
The RTX Spark machines ship with Windows on Arm. Nvidia also sells the DGX Spark, a compact local AI computer that runs Linux (DGX OS) and has been available since late 2025. Same class of machine, different operating system, and that changes how it fits into a setup you already run.
Windows on Arm suits someone who wants one machine for daily work and for running models. You get a normal desktop, your usual apps, and a laptop you can carry. The question is software. Before you buy, check that the tools you rely on, from your model runner to your container setup, have native Arm builds for Windows, or at least run acceptably under emulation.
Linux suits someone who thinks in servers. A DGX Spark can sit on a shelf, run headless, and serve models to every device on your network over SSH and an API. If your stack is already Docker on Linux, your habits and scripts carry over with less friction.
Neither option is free of lock-in. Both tie you to Nvidia's hardware. DGX OS is Nvidia's own Linux distribution, and Windows on Arm ties you to Microsoft's platform and its pace of Arm support. Pick the one that matches how you already work, not the one with the louder launch.
Do you need to buy anything at all?
Probably not to get started. Small models run on hardware many readers already own: a desktop with a mid-range graphics card, a recent laptop, or a mini PC with enough RAM. You can learn the tools, test your workflows and find out which models you actually use before spending anything.
This launch raises the ceiling. It makes it easier to run large models on one quiet machine. It does not raise the floor. The small model that handles your notes, code snippets or home automation prompts today works the same as it did last week.
There is also the option of not owning the hardware. Renting a model in the cloud costs nothing up front and gives you access to models far larger than 128 GB can hold. Owning gives you privacy, no per-token bills and no dependency on someone else's service. I have written up how renting an AI model compares to owning one if you want to work through the trade-off before you spend $2,599 or more.
Three questions to ask before you spend
- Which model do you want to run? Name it. Not a family of models, a specific one at a specific size.
- How much memory does it need? Check its size at the quantization you plan to use, add room for context, and add room for the operating system. Then compare that to the memory the GPU can actually reach, not the number on the box.
- Laptop or box at home? If you need models on the road, a laptop makes sense. If you mostly work at a desk or want to serve models to several devices, a small box on your network is often the better fit.
If the answers point to a small model, you may already own what you need. If they point to a 70-billion-parameter model on the move, a 64 GB or 128 GB configuration starts to justify its price.
Getting the sizing right
The new RTX Spark machines are a real option for running models at home, and pre-orders opened on 7 October 2026. Just buy the memory your models need, not the machine with the best launch event.
If you would rather have someone size the hardware for the models you actually want to run, take a look at my services page. You can also reach me directly. Fiverr: hiteshsaini459 · Upwork: hiteshsaini25