DeepSeek
DeepSeek V4 Pro (1600B parameters) requires approximately 865.4 GB of VRAM with NVFP4 quantization. As a Mixture of Experts model with 49B active parameters, it uses less memory than its total parameter count suggests. For the best balance of quality and speed, we recommend hardware with at least 996 GB of VRAM.
Get started
— copy & paste to run locallyCopy-paste commands to run DeepSeek V4 Pro on your machine.
Run
docker run --rm -it ghcr.io/ggerganov/llama.cpp:full \
--hf-repo "deepseek-ai/DeepSeek-V4-Pro" \
--hf-file "DeepSeek-V4-Pro-NVFP4.gguf" \
-c 4096 -ngl 99Quick specs
About this model
Related models
Inference speed
Estimated decode speed (tokens/sec) for DeepSeek V4 Pro at NVFP4 across popular GPUs and Apple Silicon, including multi-GPU rigs, using the fastest local runtime per device. Fastest is RTX 5090 32GB at ~2 tok/s. Speed is memory-bandwidth bound, so cards that fit the whole model in VRAM run far faster than ones that offload to system RAM.
| GPU / Mac | Memory | Quant | Speed (tok/s) | Fits? |
|---|---|---|---|---|
| 32 GB | NVFP4 | 2.0 | Too big | |
| 24 GB | NVFP4 | 2.0 | Too big | |
| 16 GB | NVFP4 | 2.0 | Too big | |
| 24 GB | NVFP4 | 2.0 | Too big | |
| 12 GB | NVFP4 | 2.0 | Too big | |
| 12 GB | NVFP4 | 2.0 | Too big | |
| 8 GB | NVFP4 | 2.0 | Too big | |
RX 7900 XTX 24GB | 24 GB | NVFP4 | 2.0 | Too big |
MacBook Pro M4 Max 128GB | 128 GB | NVFP4 | 2.0 | Too big |
Mac Studio M3 Ultra 256GB | 256 GB | NVFP4 | 2.0 | Too big |
Mac Studio M2 Ultra 128GB | 128 GB | NVFP4 | 2.0 | Too big |
Mac Studio M1 Ultra 128GB | 128 GB | NVFP4 | 2.0 | Too big |
MacBook Pro M4 Max 64GB | 64 GB | NVFP4 | 2.0 | Too big |
MacBook Pro M3 Max 64GB | 64 GB | NVFP4 | 2.0 | Too big |
MacBook Pro M1 Max 64GB | 64 GB | NVFP4 | 2.0 | Too big |
MacBook Pro M4 Pro 48GB | 48 GB | NVFP4 | 2.0 | Too big |
| 48 GB | NVFP4 | 2.0 | Too big | |
| 48 GB | NVFP4 | 2.0 | Too big | |
2× RX 7900 XTX 24GB | 48 GB | NVFP4 | 2.0 | Too big |
| 48 GB | NVFP4 | 2.0 | Too big |
Estimates for single-stream decoding at NVFP4; real tokens/sec varies with prompt length, context, batch size, and runtime build. Prompt processing (prefill) is faster than the decode figures shown here.
Quantization
How much VRAM DeepSeek V4 Pro (1600B) needs at each GGUF quant, and whether it fits a 24 GB card (RTX 4090 / 3090). The recommended NVFP4 uses ~896 GB — about 48% less VRAM than Q8_0, at a small quality cost.
| Quant | Bits | VRAM (weights) | Quality | Fits 24 GB? |
|---|---|---|---|---|
| Q2_K | 2 | 624 GB | Low | Too big |
| Q3_K_S | 3 | 784 GB | Low | Too big |
| NVFP4recommended | 4 | 896 GB | Medium | Too big |
| Q4_K_M | 4 | 976 GB | Medium | Too big |
| Q5_K_M | 5 | 1152 GB | High | Too big |
| Q6_K | 6 | 1312 GB | High | Too big |
| Q8_0 | 8 | 1712 GB | Very High | Too big |
| F16 | 16 | 3280 GB | Maximum | Too big |
VRAM shown is quantized weights only; add ~1–3 GB runtime overhead plus KV cache for your context length. Lower quants trade quality for memory — Q4_K_M is the usual sweet spot; Q2/Q3 only when you must fit a bigger model.
Quality benchmarks
Coding
Reasoning
Source: vendor-reported · 2026-04-24
Hardware compatibility
Computing compatibility...
Memory breakdown
Frequently asked questions
DeepSeek V4 Pro (1600B parameters) requires approximately 865.4 GB of VRAM with NVFP4 quantization. Lower quantizations like Q4_K_M use less memory but may reduce quality.
The recommended quantization for DeepSeek V4 Pro is NVFP4, which offers the best balance between model quality and memory efficiency. Higher quantizations preserve more quality but require more VRAM.
Yes, DeepSeek V4 Pro is well-suited for reasoning as well as agentic, coding, long-context. It was designed with these use cases in mind.
See also