Skip to main content

vs RTX 5090

Four times faster. Four times less room.

The 5090 is a monster at what it can load — and it cannot load a single model in our library, at any quantisation anyone publishes. 32GB is the wall, and the interesting models live on the other side of it.

Buy the 5090 if your workload tops out at 35B

We mean it — and saying so is why you can trust the rest of this page.

Qwen3.5-35B at 165 tok/s against our 43. Gaming on the side. Image generation twice as fast. If 30–35B models cover your work, a 5090 workstation is the better machine and we will tell you so in the order flow. The Spark exists for the other case: DeepSeek V4 Flash at 104GB, Laguna at 107GB, MiniMax at 101GB — models a 5090 cannot open, at any quant, from any publisher.

Measured head-to-head — including where we lose

Same models, both machines, community-measured. The Spark takes the vision-language pair; the 5090 takes most of the rest of what fits it.

WorkloadRTX 5090DGX SparkWinner
gpt-oss-20B1,338 tok/s1,094 tok/s5090
Qwen3-4B class1,446 tok/s1,105 tok/s5090
OCR 3B1,577 tok/s696 tok/s5090
Qwen3-VL-4B1,005 tok/s1,237 tok/sSpark
Qwen3-VL-8B868 tok/s972 tok/sSpark
Qwen-Image (one image)46s98s5090
ASR realtime factor0.324×0.342×≈ tie

The rig maths

A single-5090 workstation lands around $7,500 — ~60% more than a Spark, ~4× the power draw, still walled at 32GB. Two 5090s with tensor split reach 64GB (enough for gpt-oss-120b) at roughly $12–15k, 1,150W of GPU draw, and tensor-parallel serving complexity. The 100GB class stays out of reach at any card count you would put under a desk.

The fit wall, concretely

Smallest published quants: DeepSeek V4 Flash 82.5GB · Laguna 54GB · MiniMax 93GB · gpt-oss-120b 59GB · Qwen3.5-122B 77GB. Against 32GB of VRAM, every one is a no — not slow, impossible. Memory is the axis where 2026’s interesting models moved, and cards did not move with it.

Questions, answered straight

What about offloading to system RAM?
CPU offload runs — at CPU speeds for the offloaded layers. Community measurements of big MoE models split across a 5090 and system RAM land in single-digit tok/s, below what the Spark does natively with the whole model in unified memory.
Isn't the 5090 better value per token?
For models that fit it, often yes. Value per token of a model that cannot load is zero — that is the whole comparison.
Configure for the 100GB class