Does it fit your GPU?
31.6B total parameters across 128 experts (6 active, ~3B active per token), 52 layers of which only 6 are attention layers — the rest are Mamba2 state-space or MoE blocks. Native 262K context. Released under NVIDIA's own open model licence rather than Apache or MIT.
| VRAM | Fit |
|---|---|
| 16 GB + 32 GB RAM | UD-IQ3_XXS (19.8 GB) with the idle experts in system RAM — only ~3B params are active per token, so it stays quick · 262K context |
| 24 GB | UD-Q3_K_XL (21.2 GB) or MXFP4_MOE (23.2 GB) fully in VRAM · 262K context |
| 32 GB | UD-Q4_K_M (25.3 GB) or UD-Q5_K_M (30.2 GB) fully in VRAM · 262K context |
| 48 GB+ | Q8_0 (35.0 GB) fully in VRAM · 262K context |
unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUFHow to run it
Install TurboLLM
npx turbollm— no install step, works on Windows, macOS, and Linux. It detects your GPU and auto-provisions a matchingllama-serverbuild (no CUDA toolkit, no Python, no compiler).Get the model
Paste
unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUFinto TurboLLM's in-app Hugging Face search — it lists every quant with a VRAM-fit verdict against your real free VRAM, and downloads are resumable and SHA-256 verified.Let auto-fit pick the setup
Flip the auto-fit toggle and TurboLLM chooses the GPU/CPU layer split (and, for a Mixture-of-Experts model, the expert-offload split) for your card — the same manual flag-hunting every other guide walks you through by hand.
Load it
TurboLLM benchmarks the model on your exact GPU as it loads and shows the real measured tokens/sec — not a number copied from someone else's card.
Use it
Chat in the built-in UI, or point any OpenAI- or Anthropic-compatible tool (including Claude Code) at TurboLLM's local API — see the API docs.
This needs a llama.cpp build with nemotron_h / Mamba2 support — older engines refuse the file. TurboLLM provisions a current build on first run, so this only matters if you have pinned an older engine on the Engines screen. Note also that the quant ladder is unusually flat at the bottom: UD-IQ2_XXS and UD-IQ3_XXS are within 0.4 GB of each other, so there is no reason to take the 2-bit one.
FAQ
Will NVIDIA Nemotron 3.5 Lightning 30B-A3B run on my GPU?
Check the VRAM table above for the tier closest to your card. If you're not sure which tier you're in, or want picks tailored to your exact hardware, use What can I run? — enter your VRAM (or Apple unified memory) and it suggests the model, quant, and offload split that fits, the same logic TurboLLM's auto-fit runs when you load a model.
Why does NVIDIA Nemotron 3.5 Lightning 30B-A3B suddenly get slow or stutter?
The usual cause is silent VRAM spill: once a model's weights plus its KV cache stop fitting in VRAM, the driver quietly offloads the overflow to system RAM over PCIe, and that path runs 5–10× slower with no error message — just a sudden slowdown. TurboLLM's auto-fit sizes the quant and GPU/CPU split against your card's real free VRAM (including KV-cache growth at your chosen context) specifically to avoid this, and shows a VRAM-fit verdict before you load rather than after it's already crawling. Full mechanics in VRAM spill explained.
Which quant should I use for NVIDIA Nemotron 3.5 Lightning 30B-A3B?
Pick the highest quant that fits your VRAM with room left over for the KV cache and compute buffer — not just the model weights. As a starting point, Q4_K_M is usually the sweet spot (small quality loss for a large size drop); step up to Q5_K_M/Q6_K/Q8_0 if you have headroom, or down to Q3_K_M/an IQ quant if you don't. TurboLLM shows a VRAM-fit verdict for every quant against your card's real free VRAM before you download — see Quantization explained for what the letters and numbers actually mean, or llama.cpp flags that actually matter for the other settings worth knowing.
Run it now
One command detects your GPU, provisions the right engine, and opens the UI. New here? Start with Install & first run and Quantization explained. On different hardware, see RTX 3090, RTX 4090 or RTX 5090. Also see how to run it or how to run it.