NVIDIA announced a 64GB DGX Spark on 2 October 2026, starting at $4,999 from Acer, ASUS, Dell, Gigabyte, HP and MSI on Friday, 23 October. It keeps the same GB10 Grace Blackwell chip and 273 GB/s of memory bandwidth as the 128GB model, with half the memory. The same day, Hardware Busters reported the 128GB model now costs $6,950. A year ago it launched at $3,999.

So the new "entry" Spark costs $1,000 more than the original did, for half the memory. NVIDIA says it runs models of up to 100 billion parameters. We did not have a unit to test, so we checked the claim the way a buyer can today: we pulled the real file sizes of 16 popular local setups from Hugging Face, from the Qwen3.8 27B model NVIDIA uses in its own demo to FLUX.2, LTX-2.5 and Wan 2.2 video pipelines, and compared them with what 64GB actually leaves you. Eight of the 16 do not fit.

What NVIDIA announced on October 2

The 64GB model is sold only through partners; there is no NVIDIA Founders Edition this time. NVIDIA's blog lists the same software as the 128GB box: DGX OS, the NVIDIA Agent Toolkit, Nemotron models, and Ollama, vLLM and PyTorch with CUDA. The new piece is the NVIDIA Sync Cluster Assistant, which links two units over a QSFP cable and configures the ConnectX-7 network for you. A Model Launcher follows at the end of October and will set up Qwen3.8 27B and OpenCode on one Spark or a pair.

DGX Spark 64GBDGX Spark 128GBTwo 64GB, clustered
PriceFrom $4,999$6,950 (reported)$9,998 plus a QSFP cable
Unified memory64GB128GB128GB pooled
Memory bandwidth273 GB/s273 GB/s546 GB/s combined
NVIDIA's model ceilingUp to 100B parametersUp to 200B parametersUp to 200B parameters
AvailabilityPartners, from 23 OctoberOn sale nowPartners, from 23 October

NVIDIA has not published the 64GB model's SSD size, and Back2Gaming notes that storage and final prices will vary by partner. Treat $4,999 as a floor.

The price per gigabyte went the wrong way

When The Register reviewed the first DGX Spark in October 2025, the 128GB Founders Edition listed at $3,999. It rose to $4,699 in February 2026 and is now reported at $6,950, a 74% increase in a year and 48% since February. Hardware Busters ties it to the memory shortage, quoting Micron's warning that supply stays tight through 2028.

ConfigurationWhenPriceMemoryPrice per GB
DGX Spark 128GBOctober 2025 launch$3,999128GB$31
DGX Spark 128GBFebruary 2026$4,699128GB$37
DGX Spark 128GBOctober 2026$6,950128GB$54
DGX Spark 64GBFrom 23 October 2026$4,99964GB$78
Two DGX Spark 64GBFrom 23 October 2026$9,998128GB pooled$78

Per gigabyte, the 64GB model is the most expensive Spark ever sold, at $78 against $54 for today's 128GB unit. It is the cheaper box, not the better deal. That matters because memory, not compute, decides what you can run on this machine.

DGX Spark price per GB of memory: $31 at launch, $37 in February, $54 for 128GB now, $78 for the new 64GB model
Price per GB of memory: 128GB at launch, February, today, and the new 64GB model.

What fits in 64GB: 16 setups checked

DGX Spark has no separate graphics memory. The CPU, the GPU, the operating system and your model share one pool. On the 128GB model, Linux reports 121 GiB in total, according to Frank Denneman's free -h output, so about 7 GiB is never visible to the system. NVIDIA has not published the figure for the 64GB model. Our working assumption: the same 7 GiB goes, leaving 57 GiB, and DGX OS, the desktop and the runtime take another 4 GiB, leaving about 53 GiB for weights plus context. That is our budget, not an NVIDIA number.

For each setup we summed the actual weight files on Hugging Face (fetched 3 October 2026), including the text encoders and VAEs a ComfyUI workflow loads alongside the main model. Sizes below are in GiB, the unit Linux reports.

SetupWeights (GiB)DGX Spark 64GB128GB or two 64GB
Qwen3.8 27B, Q4_K_M (NVIDIA's demo model)15.3Fits, room for contextFits
Nemotron 3.5 Lightning 30B-A3B, NVFP420.1Fits, room for contextFits
Qwen3.8 27B, official FP828.8Fits, room for contextFits
Gemma 4 31B, NVFP4 (NVIDIA)30.4Fits, room for contextFits
Qwen-Image-2.1, bf16 with bf16 text encoder30.2Fits, room to workFits
LTX-2.5 22B distilled, nvfp4 with int8 text encoder33.5Fits, room to workFits
Wan 2.2 T2V 14B, fp8 high and low noise34.2Fits, room to workFits
FLUX.2 dev, fp8mixed with fp8 text encoder50.1Fits, little roomFits
Nemotron 3 Super 120B, UD-Q3_K_XL58.3Does not fitFits
gpt-oss-120b, MXFP459.0Does not fitFits
Wan 2.2 T2V 14B, fp16 high and low noise65.1Does not fitFits
LTX-2.5 22B distilled, bf16 with bf16 text encoder65.3Does not fitFits
FLUX.2 dev, fp8mixed with bf16 text encoder66.5Does not fitFits
Qwen3.8 Flash Next, UD-Q2_K_XL73.5Does not fitFits
Nemotron 3 Super 120B, NVFP4 (NVIDIA)74.8Does not fitFits
DeepSeek V4.1 Flash, ds4 Q2340.6Does not fitDoes not fit

Eight of 16 fit; eight do not. Everything at 34 GiB or under leaves real headroom for long context or large video latents. Only one setup lands in the 45 to 53 GiB band where it loads but leaves little room: FLUX.2 dev with an fp8 text encoder. The sizes come straight from the repositories, for example ggml-org's gpt-oss-120b GGUF (63.39 GB) and NVIDIA's own Nemotron 3 Super NVFP4 (80.32 GB in decimal units, 74.8 GiB).

Model sizes in GiB against the 64GB DGX Spark: 15.3 and 34.2 fit, 59.0 (gpt-oss-120b) and 74.8 (Nemotron 3 Super) do not
Weights in GiB: Qwen3.8 27B and Wan 2.2 fit; gpt-oss-120b and Nemotron 3 Super do not.

The 120B models sit just past the line

NVIDIA's "up to 100 billion parameters" is honest: a 100B model at 4-bit is roughly 50 GiB of weights, which fits our budget. The trouble is that the two open models most people mean by "a big local model" are slightly bigger. gpt-oss-120b in its native MXFP4 format is 59.0 GiB, more than the 57 GiB we expect Linux to see. NVIDIA's own Nemotron 3 Super 120B is 74.8 GiB in NVFP4, the 4-bit format GB10 is built to accelerate.

For a sense of scale, Denneman's 128GB Spark showed 94 GiB in use and 27 GiB available with Nemotron 3 Super loaded. On a 64GB machine you would be forced down to a 3-bit or 2-bit community quant, such as Unsloth's UD-Q2_K_XL at 50.9 GiB, and accept the quality loss. If a 120B-class model is the reason you want a Spark, the 64GB model is the wrong one.

The very largest local models are out of reach for both. Salvatore Sanfilippo's DwarfStar 4 (ds4) engine runs DeepSeek V4.1 Flash at 2-bit, but even that file is 340.6 GiB, and its own page lists a 512GB Mac Studio for V4.1 at Q4. A pair of Sparks does not get you there.

For image and video work, precision decides

On a PC, ComfyUI can park a text encoder in system RAM while the diffusion model runs on the GPU. On DGX Spark there is no second pool: the 64GB is the system RAM. Whatever your workflow keeps resident must fit in the same budget as the operating system.

The result is a precision rule rather than a model rule. FLUX.2 dev fits at 50.1 GiB with the fp8 Mistral text encoder and does not fit at 66.5 GiB with the bf16 one. LTX-2.5 distilled is 33.5 GiB with the nvfp4 transformer and int8 Gemma encoder, and 65.3 GiB in bf16. Wan 2.2 text-to-video needs both its high-noise and low-noise 14B experts: 34.2 GiB in fp8, 65.1 GiB in fp16.

Weights are not the whole story for video. Latents, attention buffers and the VAE decode grow with resolution and clip length, so a 34 GiB pipeline is comfortable and a 50 GiB one will hit the wall on longer clips. NVIDIA also names Blender as an early creator app, with a prebuilt installer "coming soon".

Speed: the 273 GB/s ceiling

Memory size decides what loads; memory bandwidth decides how fast a dense model writes. To produce each token, a dense language model reads every weight once, so the bandwidth divided by the weight size is a hard upper bound on decode speed. Real throughput is lower, but the ceiling tells you which models will feel slow before you buy.

Dense modelWeights (GB)Ceiling, one Spark (tokens/s)Ceiling at 546 GB/s (tokens/s)
Qwen3.8 27B, Q4_K_M16.4616.633.2
Qwen3.8 27B, FP830.878.817.7
Gemma 4 31B, NVFP432.638.416.7

That is why NVIDIA's speed claim for Qwen3.8 27B uses a cluster: NVIDIA says two clustered 64GB units ran it at up to 1.7x a single unit, below the 2x the doubled bandwidth would allow. Mixture-of-experts models such as Nemotron 3.5 Lightning 30B-A3B read only their active experts per token and run far faster than their file size suggests, which makes them the better fit for a single 64GB box. For a measured comparison of the platform against AMD's alternative, The Register tested Strix Halo against DGX Spark last December.

Decode speed ceilings on one DGX Spark: 16.6 tokens per second for Qwen3.8 27B Q4, 8.8 for FP8, 8.4 for Gemma 4 31B NVFP4
Ceiling in tokens per second at 273 GB/s: Qwen3.8 27B Q4, Qwen3.8 27B FP8, Gemma 4 31B.

Two 64GB units or one 128GB?

NVIDIA pitches the 64GB model as a place to start, with a second unit later. Priced today, the path costs more than buying big. Two 64GB units are $9,998 before the cable, $3,048 more than one 128GB unit with the same memory.

The pair does buy something real: two GB10 chips, double the combined bandwidth and NVIDIA's 1.7x on Qwen3.8 27B. If you serve one dense model to several agents at once, that throughput is worth paying for. If you want to load one big model, such as gpt-oss-120b or Nemotron 3 Super, a single 128GB box holds it without splitting it over a network link, and costs less.

The 64GB model makes sense in one case: everything you plan to run is under about 50 GiB, and you will never need a 120B-class model. Then you save $1,951 against the 128GB unit and lose nothing you use.

DGX Spark prices: $4,999 for one 64GB unit, $6,950 for one 128GB unit, $9,998 for two clustered 64GB units
One 64GB unit, one 128GB unit, or two 64GB units clustered to 128GB.

How to decide before October 23

  1. List every model you would keep loaded at the same time. For ComfyUI, that is the diffusion model plus its text encoder and VAE; for agents, the language model plus any embedding or speech model running beside it.
  2. Open each model's "Files and versions" tab on Hugging Face and note the exact size of the precision you would run, not the bf16 original.
  3. Convert to GiB by dividing gigabytes by 1.074, and add them up.
  4. Compare the total with 50 GiB. Under it, the 64GB Spark will do the job with room for context. Between 50 and 57 GiB, it loads but leaves little room. Over 57 GiB, it does not fit.
  5. Check the dense-model ceiling: 273 divided by the weight size in GB. If the answer is under 10 tokens per second, plan on a smaller quant or a mixture-of-experts model.
  6. Price the upgrade path honestly. If you can see yourself buying a second unit within a year, buy the 128GB model now and keep the $3,048 difference.
  7. Wait for partner listings on 23 October before ordering, to see SSD sizes and whether any partner sells below $4,999.

NVIDIA's DGX Spark playbooks show which runtimes are supported; vLLM and OpenClaw playbooks are marked "coming soon" for the 64GB model.

Frequently asked questions

How much does the DGX Spark 64GB cost?

It starts at $4,999 from Acer, ASUS, Dell, Gigabyte, HP and MSI. NVIDIA does not sell a Founders Edition of it, and partners set their own final prices and SSD sizes. The 128GB model is reported at $6,950.

When can I buy the DGX Spark 64GB?

Friday, 23 October 2026, from the six partners. NVIDIA announced it on 2 October 2026. The NVIDIA Sync Model Launcher, which sets up Qwen3.8 27B with one click, arrives at the end of October.

Can the DGX Spark 64GB run gpt-oss-120b?

Not in its native MXFP4 format by our count. The model file is 59.0 GiB, and if the 64GB unit hides the same 7 GiB as the 128GB model, Linux sees about 57 GiB in total. You would need the 128GB model, two clustered 64GB units, or a smaller community quant.

Can it run FLUX.2 dev or video models in ComfyUI?

Yes, at reduced precision. FLUX.2 dev fits at 50.1 GiB with an fp8 text encoder, LTX-2.5 distilled at 33.5 GiB in nvfp4, and Wan 2.2 at 34.2 GiB in fp8. The bf16 and fp16 versions of all three exceed 64GB.

Are two 64GB units the same as one 128GB unit?

They pool to the same 128GB and double the combined bandwidth, and NVIDIA measured up to 1.7x the speed of one unit on Qwen3.8 27B. They cost $3,048 more than one 128GB unit and need a QSFP cable.

How fast will Qwen3.8 27B run on one DGX Spark 64GB?

No faster than about 16.6 tokens per second at Q4_K_M and 8.8 at FP8, because each token reads every weight through the 273 GB/s memory bus. Real speed is lower. NVIDIA's 1.7x speedup figure comes from a two-unit cluster.