Determining the Required VRAM for Running LLMs Locally
The central consideration when deploying a local LLM is straightforward: will it reside within your GPU's memory? The answer hinges on the model's size, the quantization method applied, and the length of the context. This guide provides a practical foundation for selecting the appropriate amount of VRAM.
The impact of quantization on VRAM
Quantization lowers the precision used to store model weights. Reducing the bit count shrinks the model size and decreases VRAM usage, though this comes at the cost of some quality degradation.
| Quant | Bits per weight | Typical use |
|---|---|---|
| Q8_0 | 8 | Very high quality |
| Q6_K | ~6.6 | Very good quality |
| Q5_K_M | ~5.5 | Good quality and size |
| Q4_K_M | ~4.5 | Good balance of size and quality |
| Q3_K_M | ~3.5 | Lower VRAM, more quality loss |
Q4_K_M is a standard choice when VRAM is constrained. If you have additional VRAM available, opting for Q5 or Q6 allows you to run the same model with reduced quantization, thereby preserving more quality.
Approximate VRAM requirements by model size
The following figures are rough estimates for the model weights alone. The actual VRAM demand is higher because the runtime environment, KV cache, and context data also consume memory.
| Model size | Q8_0 | Q6_K | Q4_K_M | Q3_K_M |
|---|---|---|---|---|
| 4B | ~5 GB | ~4 GB | ~3 GB | ~2.5 GB |
| 8B | ~9 GB | ~7 GB | ~5.5 GB | ~4.5 GB |
| 12B | ~13 GB | ~10 GB | ~8 GB | ~6.5 GB |
| 14B | ~16 GB | ~12 GB | ~9 GB | ~7.5 GB |
| 27B | ~30 GB | ~22 GB | ~17 GB | ~13 GB |
| 32B | ~36 GB | ~27 GB | ~20 GB | ~16 GB |
| 70B | ~80 GB | ~60 GB | ~42 GB | ~34 GB |
These are estimates rather than strict limits. Variations in model architecture and quantization formats can influence the actual memory footprint.
Capabilities based on VRAM capacity
| VRAM | Practical range | Current examples |
|---|---|---|
| 8 GB | Small models around 4B to 9B | Gemma 4 E4B, Qwen3.5 9B |
| 12 GB | Small to mid-sized models around 9B to 14B | Gemma 4 12B, Qwen3.5 9B |
| 16 GB | 12B to 27B with lower quantization | Gemma 4 26B-A4B, Qwen3.6 27B at Q4 |
| 24 GB | 27B to 35B at Q4 to Q6 | Qwen3.8 27B, Gemma 4 31B |
| 32 GB | 27B to 35B at higher quantization | Qwen3.8 27B, Gemma 4 31B |
| 48 GB | Large dense models at lower quantization | 70B-class models at Q3 to Q4 |
| 80 GB | Large dense models at higher quantization | 70B-class models at Q4 to Q6 |
These ranges apply to models whose weights can fully reside on the GPU. Large MoE models differ significantly: while only a subset of parameters is active for each token, the model must still store its entire set of weights. Consequently, a model with 100B or more total parameters will not fit within a 100B-sized VRAM budget simply due to having fewer active parameters.
MoE models
Mixture-of-Experts (MoE) models consist of multiple parameter groups known as experts. Since only specific experts are activated per token, inference can be more efficient compared to a dense model with an equivalent total parameter count.
However, the inactive experts remain part of the model structure. Thus, large MoE models often demand significantly more memory than their active parameter count implies. Very large models may necessitate multiple GPUs or the offloading of system RAM.
Context length and VRAM usage
Model weights represent only a portion of the total memory requirements. The KV cache expands as context length increases, meaning that running the same model with a 64K context can demand substantially more VRAM than a 4K context.
- Longer context windows require more VRAM.
- KV-cache precision impacts memory consumption.
- Batch size and concurrent user counts also elevate memory usage.
- Reserve sufficient VRAM for the runtime instead of allocating all GPU memory to model weights.
Practical recommendations
- Verify the actual size of the quantized model you intend to run.
- Avoid assuming the model file size equals the exact VRAM requirement. Allow space for the KV cache and runtime processes.
- If a model cannot fit entirely in VRAM, portions can be offloaded to system RAM, but inference speed will typically decrease.
- For long-context or agentic workloads, budget for more VRAM than the model weights alone demand.
- Multiple GPUs can be used to distribute a model if a single GPU lacks sufficient VRAM.
Run it on DaDesktop
There is no need to purchase a GPU to run a local LLM. DaDesktop provides a cloud desktop equipped with the necessary VRAM, allowing you to run models directly without owning the underlying hardware.
Select the VRAM tier that matches your model, load it, and begin using it immediately. This eliminates the need for setup, hardware purchases, or driver troubleshooting. Visit available GPUs to view the options.