GPU VRAM Calculator for LLMs
Pick a model or enter its size, choose a precision, context length and batch size, and see how much GPU memory you'll need for inference, LoRA/QLoRA or full fine-tuning — broken down into weights, KV cache, activations and overhead.
Model architecture (from the model's config.json)
Estimated VRAM needed
18.0 GB
- Model weights
- 14.9 GB
- KV cache
- 1.00 GB
- Framework & CUDA overhead
- 2.09 GB
8.00B parameters × 2 bytes (BF16)
2 × 32 layers × 8 KV heads × 128 × 8192 tokens × 1 batch × 2 B
≈10% + 0.5 GB for the CUDA context, buffers and fragmentation
Which GPUs fit?
| GPU | VRAM | Result |
|---|---|---|
| GeForce RTX 5080 | 16 GB | 2 GPUs |
| GeForce RTX 3090 | 24 GB | ✓ Fits |
| GeForce RTX 4090 | 24 GB | ✓ Fits |
| GeForce RTX 5090 | 32 GB | ✓ Fits |
| NVIDIA L40S | 48 GB | ✓ Fits |
| RTX PRO 6000 Blackwell | 96 GB | ✓ Fits |
| NVIDIA A100 80GB | 80 GB | ✓ Fits |
| NVIDIA H100 80GB | 80 GB | ✓ Fits |
| NVIDIA H200 | 141 GB | ✓ Fits |
| NVIDIA B200 | 180 GB | ✓ Fits |
| AMD Instinct MI300X | 192 GB | ✓ Fits |
| AMD Instinct MI325X | 256 GB | ✓ Fits |
“Fits” leaves about 5% of VRAM free. Splitting a model across several GPUs (tensor or pipeline parallelism) adds communication overhead, so multi-GPU counts are a lower bound.
Data last updated:
Results are estimates for planning only. Prices, limits and real-world usage vary — always confirm against the provider's official documentation or your own measurements before making decisions.
How much GPU memory does an LLM need?
GPU memory (VRAM) is usually the limiting factor when you run or fine-tune a large language model yourself. If the model, its working memory and the framework don’t all fit, it either won’t load or will crawl as it spills into system memory. This calculator estimates VRAM for inference, LoRA or QLoRA fine-tuning and full fine-tuning, shows where the memory goes, and checks which common GPUs have enough.
How to use the calculator
- Choose the workload: inference, LoRA/QLoRA or full fine-tuning.
- Pick a model preset or enter the parameter count and architecture from the model’s config.json.
- Choose the precision — BF16/FP16 for full quality, INT8 or INT4 for quantized models.
- Set the context (sequence) length and batch size you plan to use.
- Read the total, the breakdown, and which GPUs fit on their own or need several cards.
What uses the memory
Model weights. Parameters × bytes per parameter: 4 bytes for FP32, 2 for FP16 or BF16, 1 for INT8 and about 0.5 for INT4. An 8-billion-parameter model is roughly 16 GB in BF16 but only about 4 GB at 4-bit.
KV cache (inference). For every token in the context, each layer stores a key and a value vector for each KV head: 2 × layers × KV heads × head dimension × context length × batch size × bytes per value. Models with grouped-query attention, such as Llama 3.1 and Qwen2.5, have far fewer KV heads than attention heads, which keeps this manageable — but at long contexts or large batches, the KV cache can still exceed the size of the weights.
Training memory. Full fine-tuning with mixed-precision AdamW needs about 16 bytes per parameter for the weights, gradients and optimizer states, plus activations that grow with sequence length and batch size. Gradient checkpointing trades extra compute for much lower activation memory. LoRA freezes the base model and trains small adapters, usually under 2% of the parameters, and QLoRA also loads the frozen model in 4-bit.
Overhead. The CUDA context, framework buffers and memory fragmentation add more. The calculator adds 10% plus 0.5 GB, and counts a GPU as a fit only if about 5% of its memory is left free.
Why this is an estimate
Real memory use depends on the inference engine (vLLM, TensorRT-LLM, llama.cpp, Hugging Face Transformers), the quantization format, attention kernels and settings such as how much memory an engine pre-allocates for the KV cache. Use these numbers to shortlist hardware, then confirm with a short test run. GPU memory sizes were last checked on 2026-09-27.
Comparing self-hosting with an API? Estimate API spend with the LLM API Cost Calculator, count prompt lengths with the LLM Token Counter, and tidy up config.json files with the JSON Formatter.
Frequently asked questions
How much VRAM do I need to run a 7B or 8B model?
At FP16/BF16, the weights alone take about 2 bytes per parameter — roughly 15–16 GB for an 8B model — plus the KV cache and overhead. Quantized to INT4, the same model needs only around 5–6 GB, which fits on many consumer GPUs.
What is the KV cache?
While generating text, a model stores the keys and values from its attention layers for every token so far. This KV cache grows with context length and batch size, and for long contexts it can use more memory than the weights.
Why does full fine-tuning need so much more memory?
Training stores gradients and optimizer states as well as the weights. With mixed-precision AdamW, that's about 16 bytes per parameter — around 128 GB for an 8B model before activations.
What's the difference between LoRA and QLoRA?
LoRA freezes the base model and trains small adapter matrices, so only a tiny fraction of parameters needs gradients and optimizer states. QLoRA also loads the frozen base model in 4-bit, cutting memory further. Choose INT4 as the base precision to estimate QLoRA.
How accurate is this estimate?
It uses standard formulas and is usually within a few gigabytes, but real usage depends on the framework (vLLM, llama.cpp, PyTorch), kernels, quantization format and settings. Leave some headroom and confirm with a test run.