Skip to main content

The library

Every model that fits, byte-verified

Sizes are summed from the actual files on Hugging Face, speeds are measured on real hardware, and the list of what does NOT fit is published right below the list of what does. That is the whole method. Last verified: 5 August 2026.

Text & agents

The working set for reasoning, coding and agent loops — largest fitting quant shown with its exact size.

DeepSeek V4 Flash 0731

DeepSeek · 304B MoE · 13B active

The frontier open flagship — deep reasoning and agentic work, uncensored.

UD-IQ3_XXS · 104.2GB · DS4 Q2 87GB

MIT

Hands-on on one Spark: 103GB resident, 18GB free, 131K context, 15.5–16.6 tok/s decode, 411 tok/s prefill. The DS4 engine (antirez's DwarfStar, MIT) fits it in 87GB by quantising only the routed experts — its Blackwell fork reports 776 tok/s sustained prefill and 59 tok/s aggregate across 12 concurrent agents on one box (project-reported, Aug 2026). 1M context architecturally.

Weights on Hugging Face

Laguna S 2.1

Poolside · US · 118B MoE · 8B active

Agentic coding at frontier-adjacent quality — the US-made permissive pick.

NVFP4 71GB · Q6_K_XL 107.1GB

OpenMDW-1.1

Poolside's own launch: “small enough to run on a single NVIDIA DGX Spark.” Hands-on: 31 tok/s with speculative decoding, 121.9 tok/s aggregate at 8 streams, 262K context. FP8 (131.3GB) does not fit — NVFP4 only.

Weights on Hugging Face

Qwen3.5 122B A10B

Alibaba · 125B MoE · 10B active

The long-context workhorse — 262K measured on one box with ~52GB free at NVFP4.

UD-Q6_K_XL 112.4GB · NVFP4 75.6GB

Apache-2.0

Weights on Hugging Face

Inkling-Small

Thinking Machines · US · 276B

A second US frontier lab in the library — reasoning-heavy work.

UD-IQ3_S 107.5GB

Apache-2.0

Weights on Hugging Face

MiniMax-M2.5

MiniMax · 230B MoE · 10B active

High-throughput generalist — ~26 tok/s measured, 128K context.

UD-Q3_K_XL 101.3GB

Community

Requires --no-mmap on unified memory.

Weights on Hugging Face

gpt-oss-120b

OpenAI · US · 120B MoE · 5.1B active

The concurrency king: 286 tok/s sustained aggregate at 256 streams.

Q8_0 63.4GB — half the box free

Apache-2.0

Natively MXFP4 — every quant lands 62.6–64.5GB; quantising does not shrink it.

Weights on Hugging Face

Ornith-1.0 35B

DeepReinforce · 35B MoE

Full-precision inference with room to spare — zero quantisation loss.

BF16 70.2GB — no quantisation at all

MIT

Weights on Hugging Face

Ling-3.0-flash

inclusionAI (Ant) · 127B MoE

Ant's brand-new flash-tier MoE — frontier-adjacent, MIT, huge headroom on one box.

IQ4_XS 66.4GB · NVFP4 81.4GB

MIT

Released 2026-08-02; added to this library 2026-08-05. Sizes are summed file bytes from community GGUF/NVFP4 builds. The OFFICIAL FP8 build is 128.4GB — it does not fit one box, but fits a Y Tower 192 or a two-Spark cluster whole. Our own on-box measurement pending.

Weights on Hugging Face

Gemma 4 26B-A4B

Google · US · 26B MoE · 3.8B active

The fast daily driver — a purpose-built DGX Spark GGUF, 3,125 downloads/month.

Q4_K_M 16.8GB — Spark-specific build

Gemma licence

Weights on Hugging Face

LFM2.5-2.6B

Liquid AI · US · 2.6B dense hybrid · conv + GQA

The tool-calling sidekick — plans, calls tools and runs multi-step loops at near-zero cost while the flagship thinks.

~3GB class — dozens fit alongside the flagship

LFM open-weight — read before commercial use

Trained with agentic RL, 128K context, ~34T tokens. Vendor-reported: ToolSandbox 77.83, Multi-IF 80.07, IFStruct 85.49 — ahead of Gemma-4-E4B on all three (Liquid's own harness; labelled as such). Built for phones — on this box it is effectively free to keep resident next to the flagship for cheap tool-call turns.

Weights on Hugging Face

What does not fit — and we will not pretend otherwise

MoE 'active parameters' is a compute figure, not a memory figure: all experts must be resident. Anyone selling you these on a desktop box is confusing the two.

  • Kimi K3 (2.8T)594GB even at 1-bit — Not on one box, not on four.
  • GLM-5.2 (744B)~405GB Int4-Int8Mix — Measured on 4× Spark at 22–42 tok/s — we build that cluster.
  • Hunyuan Hy3 (295B)IQ1_M 181.2GB — Two boxes minimum.
  • LongCat 2.0 (1.6T)no GGUF exists — Datacentre class.
  • Qwen3.8 / Qwen-Max (≈2.4T)— — Not open weights at all — API only.

Image generation — ten models, all measured on a real Spark

Generation time for a standard image via the official DGX Spark ComfyUI playbook. Every one fits.

ModelFamilyGenerationMemory
SD 3.5 mediumStable Diffusion34s22GB
SD 3.5 largeStable Diffusion82s29GB
FLUX.2-klein 9BFLUX.295s66GB
FLUX.1-schnellFLUX.1110s72GB
FLUX.1-devFLUX.1111s95GB
FLUX.1-Kontext-devFLUX.1111s66GB
Z-Image-TurboZ-Image178s22GB
Z-ImageZ-Image179s22GB
Qwen-Image-2512Qwen212s63GB
FLUX.2-dev (4-bit)FLUX.2397s34GB

Video, voice & music

The rest of the media stack — licences read before anything went on this page.

LTX-2.3

Lightricks · 22B DiT

Video generation with synced audio in one pass, on your desk. No geographic restriction.

~33GB working set (FP8 DiT + NVFP4 encoder)

LTX-2 Community — free < $10M ARR

Weights on Hugging Face

Wan 2.2

Alibaba · MoE video

The cleanest video licence in the field — Alibaba claims no rights over outputs.

Fits comfortably

Apache-2.0

Weights on Hugging Face

Whisper large-v3

OpenAI · 1.55B

Transcribes 99 languages, fully offline.

~3GB

MIT

Weights on Hugging Face

Parakeet TDT 0.6B v3

NVIDIA · 600M

Beats Whisper on European languages (6.32% vs 7.44% WER) at 5–10× the speed.

~600MB

CC-BY-4.0

Weights on Hugging Face

Chatterbox

Resemble AI · 0.5B

Clones a voice from ~10 seconds of audio — MIT licensed, commercially safe.

Small

MIT

65.3% blind-test preference over ElevenLabs is vendor-reported.

Weights on Hugging Face

Kokoro-82M

Community · 82M

Fast, commercially-safe narration. No cloning — sometimes that is the feature.

~300MB

Apache-2.0

Weights on Hugging Face

ACE-Step v1.5 XL Turbo

ACE Studio · DiT 9.97GB

Full music tracks, locally — pre-bundled in the DGX Spark ComfyUI distro.

~15GB with encoders

Open weights

Weights on Hugging Face

Unlimited-OCR

Baidu · 3.3B

Document OCR at scale — contracts, scans, handwriting — entirely on the box.

6.7GB BF16 — runs alongside anything

MIT

2.7M monthly downloads on Hugging Face; the workhorse for the document pipelines our regulated-industry buyers run. Added 2026-08-05.

Weights on Hugging Face

Notably absent: MiniMax-H3, this summer’s hot video release — its licence excludes the EU, UK and US by name, which is why it is in our research file and not on this page. Licences get read here.

What they make — same brief, different models

Every image and clip below was generated from the identical prompt (“A compact matte-black AI mini computer on a walnut desk, golden-hour window light…”) using each model's published open weights — the differences you see are the models.

Output generated by FLUX.1-dev
FLUX.1-dev95GB on the box · 111s per image measured on a Spark
Output generated by Qwen-Image-2512
Qwen-Image-251263GB · 212s per image measured
Output generated by SD 3.5 Medium
SD 3.5 Medium22GB · 34s per image — the fast option
LTX videoThe video headline — ~33GB working set, free commercial under $10M ARR
Wan 2.2Apache-2.0 — the cleanest video licence in the field
Generated voice
ChatterboxMIT voice — clones from ~10s of audio
Generated music
ACE-Step v1.5A 30-second instrumental, generated locally-runnable weights

Demo batch generated from each model’s published weights via hosted inference — identical weights to what we install on the box. On-box generation times shown are separate hands-on measurements from the table above. We re-shoot this gallery on our own hardware the day the bench unit lands.

Questions, answered straight

How much of the 128GB can models actually use?
Operators consistently report 119–121GB allocatable after the OS and driver reserve. Our fit calls use ~115GB as the safe working pool, which is why a 104GB quant with KV cache headroom counts as fitting and a 118GB one does not.
Why do you quote 3-bit quants — doesn't quality collapse?
DeepSeek V4 Flash is quantisation-aware-trained with its routed experts natively in MXFP4, so repacking is unusually gentle. Unsloth's published figures show ~96% top-1 token agreement at Q4-class sizes. Below Q4 the published record stops — we are measuring that curve ourselves and will publish it.
What is the single best model to start with?
For agentic and reasoning work, DeepSeek V4 Flash 0731 at UD-IQ3_XXS — the heaviest thing that runs on one box, verified hands-on at 15.5–16.6 tok/s with 131K context. For coding with a Western licence chain, Laguna S 2.1 at NVFP4. We install both.
Can I run video and image models alongside a text model?
Not simultaneously at the large end — the text flagship occupies ~103GB. The practical pattern is switching workloads (models load in tens of seconds) or dedicating a second box to media. The install sets both patterns up.
Get it pre-installed