Cohere
Command R+ 104B (104B parameters) requires approximately 68.4 GB of VRAM with Q4_K_M quantization. For the best balance of quality and speed, we recommend hardware with at least 79 GB of VRAM.
Get started
— copy & paste to run locallyCopy-paste commands to run Command R+ 104B on your machine.
Run
ollama run command-r-plusQuick specs
About this model
Related models
Inference speed
Estimated decode speed (tokens/sec) for Command R+ 104B at Q4_K_M across popular GPUs and Apple Silicon, including multi-GPU rigs, using the fastest local runtime per device. Fastest is MacBook Pro M4 Max 128GB at ~10 tok/s. Speed is memory-bandwidth bound, so cards that fit the whole model in VRAM run far faster than ones that offload to system RAM.
| GPU / Mac | Memory | Quant | Speed (tok/s) | Fits? |
|---|---|---|---|---|
MacBook Pro M4 Max 128GB | 128 GB | Q4_K_M | 10.3 | Tight |
Quick picks
Best hardware
Run this model
Quantization
How much VRAM Command R+ 104B (104B) needs at each GGUF quant, and whether it fits a 24 GB card (RTX 4090 / 3090). The recommended Q4_K_M uses ~63.4 GB — about 43% less VRAM than Q8_0, at a small quality cost.
| Quant | Bits | VRAM (weights) | Quality | Fits 24 GB? |
|---|---|---|---|---|
| Q2_K | 2 | 40.6 GB | Low | Too big |
| Q3_K_S | 3 | 51 GB | Low | Too big |
| NVFP4 | 4 | 58.2 GB | Medium | Too big |
| Q4_K_Mrecommended | 4 | 63.4 GB | Medium | Too big |
Quality benchmarks
Coding
Reasoning
General
Hardware compatibility
Computing compatibility...
Memory breakdown
Frequently asked questions
Command R+ 104B (104B parameters) requires approximately 68.4 GB of VRAM with Q4_K_M quantization. Lower quantizations like Q4_K_M use less memory but may reduce quality.
Yes, MacBook Pro M3 Max 128GB can run Command R+ 104B with a compatibility score of 62/100. It provides 128 GB of memory and achieves approximately 4.1 tokens per second.
The recommended quantization for Command R+ 104B is Q4_K_M, which offers the best balance between model quality and memory efficiency. Higher quantizations preserve more quality but require more VRAM.
The top recommended hardware for Command R+ 104B: NVIDIA GH200 96GB (score: 73/100), NVIDIA H20 96GB (score: 73/100), AMD Instinct MI300A 128GB (score: 71/100). These provide the best combination of memory, bandwidth, and compute for running this model locally.
Yes, Command R+ 104B is well-suited for chat as well as rag, reasoning. It was designed with these use cases in mind.
See also
| 256 GB |
| Q4_K_M |
| 9.5 |
| Fits |
Mac Studio M2 Ultra 128GB | 128 GB | Q4_K_M | 8.0 | Tight |
Mac Studio M1 Ultra 128GB | 128 GB | Q4_K_M | 7.5 | Tight |
2× RX 7900 XTX 24GB | 48 GB | Q4_K_M | 6.3 | Too big |
MacBook Pro M4 Max 64GB | 64 GB | Q4_K_M | 5.5 | Too big |
| 48 GB | Q4_K_M | 4.4 | Too big |
| 48 GB | Q4_K_M | 3.8 | Too big |
| 48 GB | Q4_K_M | 3.3 | Too big |
MacBook Pro M4 Pro 48GB | 48 GB | Q4_K_M | 2.9 | Too big |
MacBook Pro M3 Max 64GB | 64 GB | Q4_K_M | 2.2 | Too big |
| 32 GB | Q4_K_M | 2.0 | Too big |
| 24 GB | Q4_K_M | 2.0 | Too big |
| 16 GB | Q4_K_M | 2.0 | Too big |
| 24 GB | Q4_K_M | 2.0 | Too big |
| 12 GB | Q4_K_M | 2.0 | Too big |
| 12 GB | Q4_K_M | 2.0 | Too big |
| 8 GB | Q4_K_M | 2.0 | Too big |
RX 7900 XTX 24GB | 24 GB | Q4_K_M | 2.0 | Too big |
MacBook Pro M1 Max 64GB | 64 GB | Q4_K_M | 2.0 | Too big |
Estimates for single-stream decoding at Q4_K_M; real tokens/sec varies with prompt length, context, batch size, and runtime build. Prompt processing (prefill) is faster than the decode figures shown here.
| Q5_K_M |
| 5 |
| 74.9 GB |
| High |
| Too big |
| Q6_K | 6 | 85.3 GB | High | Too big |
| Q8_0 | 8 | 111.3 GB | Very High | Too big |
| F16 | 16 | 213.2 GB | Maximum | Too big |
VRAM shown is quantized weights only; add ~1–3 GB runtime overhead plus KV cache for your context length. Lower quants trade quality for memory — Q4_K_M is the usual sweet spot; Q2/Q3 only when you must fit a bigger model.
Source: official · 2024-04-04