NVIDIA DGX Spark for local LLMs: what it runs and what it really costs
A personal machine with 128 GB of memory that runs models of up to 200 billion parameters on your desk. We look at the price, the published measurements and the catches.
The NVIDIA DGX Spark is a desktop computer for running large language models locally. Roughly the size of a thick book, it handles models of up to 200 billion parameters without sending your data anywhere. It has 128 GB of unified memory, and since the week of 23 February 2026 its list price has been $4,699, up from $3,999. The biggest limit, however, is not how much memory it has but how fast that memory can be read: a bandwidth of 273 GB/s decides how quickly a large model answers.
In short
- List price $4,699 since the week of 23 February 2026, up from $3,999. NVIDIA put the $700 rise down to memory shortages; the hardware did not change.
- 128 GB of LPDDR5x unified memory: models of up to 200 billion parameters on one machine, and up to 700 billion with four machines linked together.
- A bandwidth of 273 GB/s is the real ceiling: in LMSYS measurements Llama 3.1 70B generated 2.7 tokens per second.
- Smaller models run briskly: GPT-OSS 20B reached 49.7 tokens per second, and Llama 3.1 8B reached 368 tokens per second in total across 32 parallel requests.
- A 240 W power supply, 4 TB of NVMe storage and the GB10 Grace Blackwell chip, rated at up to 1 PFLOP at FP4 precision.
What you actually get
The DGX Spark is a ready-made workstation built around the GB10 Grace Blackwell chip, with a 20-core Arm processor and NVIDIA DGX OS. The design decision that matters most is the memory: the processor and the graphics unit share a single pool, so the model does not have to be shuttled back and forth between two separate ones. That architecture is what lets the machine hold models that simply do not fit on a consumer graphics card.
- 128 GB of unified memory, with NVIDIA quoting support for inference on models of up to 200 billion parameters.
- 273 GB/s of memory bandwidth, according to the manufacturer's specification.
- Up to 1 PFLOP at FP4 precision and the full CUDA ecosystem, which means compatibility with the tools your team probably already uses.
- 4 TB of encrypted NVMe storage and a ConnectX-7 200 Gbps network card for linking machines together.
- Up to four machines in a cluster, which NVIDIA ties to models of up to 700 billion parameters.
Where the $700 price rise came from
At launch the DGX Spark cost $3,999. In an announcement published in the week of 23 February 2026, NVIDIA raised the list price to $4,699 and said plainly that the increase was a consequence of industry-wide constraints on memory supply. It made two points worth remembering: the hardware has not changed, and orders placed before the change are fulfilled at the old price. The rise applied to every market.
This is not an isolated case but a symptom of something broader: demand from AI has been pushing up the price of memory across the market, which we covered in our piece on the AI-driven memory supercycle (in Polish). If you are planning a purchase, treat the list price as a snapshot of today rather than a constant.
Memory capacity tells you how large a model will fit on the machine. Memory bandwidth tells you how fast that model will answer. They are two entirely different promises.
Capacity is not speed
When a model generates an answer, it has to read its weights from memory for every new token. Output speed therefore depends not on how much memory you have but on how quickly it can be read. At this bandwidth a large model answers slowly, because every single token means reading tens of gigabytes.
The measurements published by the LMSYS team, which tested the DGX Spark on several models, show exactly that. Llama 3.1 70B in FP8 processed 803 tokens per second while reading in the prompt, but managed only 2.7 tokens per second while generating the answer. Smaller models tell a completely different story: GPT-OSS 20B running in Ollama delivered 49.7 tokens per second, and Llama 3.1 8B in SGLang produced 20.5 tokens per second for a single request and 368 tokens per second in total across 32 simultaneous requests. Same machine, different models and frameworks, and the results sit an order of magnitude apart.
For comparison, the authors also measured graphics cards. On GPT-OSS 20B the consumer RTX 5090 turned out to be about four times faster. The LMSYS conclusion is consistent with what NVIDIA itself promises: this is a machine for prototyping, experimentation and local inference, not a replacement for full-size cards in a server room.
Why run a model locally at all
There are three solid reasons. Privacy: data never leaves the machine, which matters for client documents, HR records and anything covered by data protection law such as the GDPR. Predictable cost: you pay once for the hardware and then for electricity, instead of paying for every token. Control: nobody can withdraw the model, change the limits or rewrite the terms of service halfway through your rollout.
Set that against the other side. A local machine is your responsibility: updates, backups, remote access and someone who actually looks after it. A local model also does not look anything up on its own. Unless you connect it to a search tool, it answers from what it learned in training and what you hand it, whereas the public assistants ground their answers in live web results.
If the plan is an assistant that answers from your own files, the model is only one part of the system. In a typical set-up a retrieval layer first cuts your documents into chunks and passes only the best matches to the model, the same retrieval-augmented generation that public assistants run after they split a question into sub-queries. How those chunks are cut decides what the model can find at all, which we explain in how chunking works.
What the list price does not include
NVIDIA's list price in dollars is not necessarily what you will pay. Depending on where you buy, local taxes and the reseller's margin come on top, so ask for a written quote in your own currency before you treat the purchase as settled. The second running cost is power: the power supply is rated at 240 W, so a machine working flat out draws roughly what a powerful desktop computer does, not what a server rack does.
The most honest way to judge whether it pays off is also the dullest: pull the invoices for the last three months and add up what you really spend on AI APIs and subscriptions. If the total is small change next to the price of the machine, the hardware will never pay for itself. If the bills grow every month, first check whether the problem lies in how you use the models rather than in where they run. Buying hardware fixes the cost only when usage is genuinely heavy and steady.
The caveats
- The price has already changed once within a year, with no change to the hardware. Nobody can guarantee it will stay at its current level.
- Street prices vary by country and reseller. We quote NVIDIA's list price because it is the only confirmed figure.
- The performance results come from one team and one configuration. A different framework, quantisation or context length produces different numbers, as the spread within that single set of tests already shows.
- 200 billion parameters is a promise about capacity, not comfort. A model that size will fit in memory, but at this bandwidth it will not answer the way you are used to from a chat window in the browser.
- The competition is not standing still. We have covered the open RISC-V approach and Qualcomm's talks to acquire Tenstorrent (in Polish), as well as the Jalapeño chip OpenAI is building with Broadcom (in Polish). Either path could change the economics of local LLMs within a year.
How to tell whether you need one
Before you spend the money, work through these four steps.
- Add up your spending from invoices, not from memory. The last three months, all AI tools together.
- Check how large a model you really need. Tasks such as summarising, classifying and extracting data from documents are often handled well by a model of a few billion parameters, and small models increasingly match far larger ones (in Polish) in narrow uses.
- Work out how many tokens per second you need. An assistant writing in the background can afford to be slow. A chatbot talking to a customer cannot.
- Establish whether privacy is the real reason. If it is, the maths looks different, because some data simply cannot go to someone else's cloud, and comparing costs alone then makes no sense.
The DGX Spark makes sense for a team that works with models every day, needs a lot of memory and wants to keep its data in-house. For everyone else it is a very expensive way of doing something that can be rented by the hour.
Common questions
How much does the NVIDIA DGX Spark cost?
The list price has been $4,699 since the week of 23 February 2026. Before that the machine cost $3,999. NVIDIA raised the price by $700, citing industry-wide constraints on memory supply, and did not change the hardware.
How large a model can I run on the DGX Spark?
NVIDIA quotes models of up to 200 billion parameters on a single machine, thanks to 128 GB of unified memory, and up to 700 billion with four units linked together. Capacity does not mean speed, though: in LMSYS measurements Llama 3.1 70B generated 2.7 tokens per second.
Is the DGX Spark faster than a graphics card?
Not at generating text. In LMSYS tests on GPT-OSS 20B, the consumer RTX 5090 was about four times faster. The advantage of the DGX Spark is its 128 GB of unified memory, which holds models that simply do not fit on a graphics card.
Who is the DGX Spark not worth it for?
Companies that use models only occasionally and have no requirement to keep data on their own premises. For occasional use, pay-as-you-go cloud usually works out cheaper, and there is no machine to administer.
Sources
Read next
Find out whether AI recommends your company.
Start with the free SEO and GEO audit, delivered in 5 working days. We check how the models describe your brand and hand back a prioritised list of changes.