NVIDIA has announced a version of DGX Spark with 64 GB of unified memory instead of 128 GB. According to the NVIDIA blog, it will be sold only through partners: Acer, ASUS, Dell, Gigabyte, HP and MSI, starting October 23, at $4,999.
What stays the same
The rest of the platform is unchanged: the GB10 Grace Blackwell chip, DGX OS, the CUDA software stack and the ConnectX-7 network card. The Register adds that memory bandwidth also stays at 273 GB/s, and that storage is halved along with memory.
NVIDIA says the 64 GB model handles models of up to 100 billion parameters. Two units can be connected directly with a cable through ConnectX-7, and NVIDIA Sync Cluster Assistant sets up the network itself. A pair gives 128 GB of memory and, per NVIDIA, support for models up to 200 billion parameters. In NVIDIA's own test with Qwen 3.8 27B, two units were up to 1.7 times faster than one.
Why $4,999 is not cheap
The 64 GB model only looks affordable next to the new 128 GB price. On the same day NVIDIA raised it to $6,950, The Register reports. A year ago the 128 GB DGX Spark started at $3,999, so the new 64 GB model costs 25% more than the full version did at launch. The reason is memory prices.

Per gigabyte of memory the 64 GB model comes out more expensive: about $78 against $54 for the 128 GB one. Two 64 GB units in a cluster cost $9,998, which is more than one 128 GB unit, though you also get twice the compute and twice the bandwidth. Availability and prices in Ukraine have not been announced.
Who it is for
The use case NVIDIA has in mind is local inference: an agent or a model in the 27–35B range running around the clock, privately, without the cloud. 64 GB is enough for that, with room for long context. 70B models fit in 4-bit quantization, but the 273 GB/s bandwidth limits the speed: our rough estimate is no more than 7 tokens per second when generating. Fine-tuning is a weak spot of the 64 GB model, The Register notes.
Before buying, compare it with the alternatives on your own model: a Mac Studio with the same amount of memory, a mini PC on AMD's new Ryzen AI Max+ chips, or a regular PC with a discrete NVIDIA GPU, if your model fits in its video memory. The deciding factors are tokens per second at the context length you actually use, and whether you need CUDA.


No comments yet