BigTool

GPU VRAM Calculator for LLMs

Pick a model or enter its size, choose a precision, context length and batch size, and see how much GPU memory you'll need for inference, LoRA/QLoRA or full fine-tuning — broken down into weights, KV cache, activations and overhead.

Model architecture (from the model's config.json)

Estimated VRAM needed

18.0 GB

Model weights
14.9 GB

8.00B parameters × 2 bytes (BF16)

KV cache
1.00 GB

2 × 32 layers × 8 KV heads × 128 × 8192 tokens × 1 batch × 2 B

Framework & CUDA overhead
2.09 GB

≈10% + 0.5 GB for the CUDA context, buffers and fragmentation

Which GPUs fit?

GPUVRAMResult
GeForce RTX 508016 GB2 GPUs
GeForce RTX 309024 GB✓ Fits
GeForce RTX 409024 GB✓ Fits
GeForce RTX 509032 GB✓ Fits
NVIDIA L40S48 GB✓ Fits
RTX PRO 6000 Blackwell96 GB✓ Fits
NVIDIA A100 80GB80 GB✓ Fits
NVIDIA H100 80GB80 GB✓ Fits
NVIDIA H200141 GB✓ Fits
NVIDIA B200180 GB✓ Fits
AMD Instinct MI300X192 GB✓ Fits
AMD Instinct MI325X256 GB✓ Fits

“Fits” leaves about 5% of VRAM free. Splitting a model across several GPUs (tensor or pipeline parallelism) adds communication overhead, so multi-GPU counts are a lower bound.

Data last updated:

Results are estimates for planning only. Prices, limits and real-world usage vary — always confirm against the provider's official documentation or your own measurements before making decisions.

How much GPU memory does an LLM need?

GPU memory (VRAM) is usually the limiting factor when you run or fine-tune a large language model yourself. If the model, its working memory and the framework don’t all fit, it either won’t load or will crawl as it spills into system memory. This calculator estimates VRAM for inference, LoRA or QLoRA fine-tuning and full fine-tuning, shows where the memory goes, and checks which common GPUs have enough.

How to use the calculator

  1. Choose the workload: inference, LoRA/QLoRA or full fine-tuning.
  2. Pick a model preset or enter the parameter count and architecture from the model’s config.json.
  3. Choose the precision — BF16/FP16 for full quality, INT8 or INT4 for quantized models.
  4. Set the context (sequence) length and batch size you plan to use.
  5. Read the total, the breakdown, and which GPUs fit on their own or need several cards.

What uses the memory

Model weights. Parameters × bytes per parameter: 4 bytes for FP32, 2 for FP16 or BF16, 1 for INT8 and about 0.5 for INT4. An 8-billion-parameter model is roughly 16 GB in BF16 but only about 4 GB at 4-bit.

KV cache (inference). For every token in the context, each layer stores a key and a value vector for each KV head: 2 × layers × KV heads × head dimension × context length × batch size × bytes per value. Models with grouped-query attention, such as Llama 3.1 and Qwen2.5, have far fewer KV heads than attention heads, which keeps this manageable — but at long contexts or large batches, the KV cache can still exceed the size of the weights.

Training memory. Full fine-tuning with mixed-precision AdamW needs about 16 bytes per parameter for the weights, gradients and optimizer states, plus activations that grow with sequence length and batch size. Gradient checkpointing trades extra compute for much lower activation memory. LoRA freezes the base model and trains small adapters, usually under 2% of the parameters, and QLoRA also loads the frozen model in 4-bit.

Overhead. The CUDA context, framework buffers and memory fragmentation add more. The calculator adds 10% plus 0.5 GB, and counts a GPU as a fit only if about 5% of its memory is left free.

Why this is an estimate

Real memory use depends on the inference engine (vLLM, TensorRT-LLM, llama.cpp, Hugging Face Transformers), the quantization format, attention kernels and settings such as how much memory an engine pre-allocates for the KV cache. Use these numbers to shortlist hardware, then confirm with a short test run. GPU memory sizes were last checked on 2026-09-27.

Comparing self-hosting with an API? Estimate API spend with the LLM API Cost Calculator, count prompt lengths with the LLM Token Counter, and tidy up config.json files with the JSON Formatter.

Frequently asked questions

How much VRAM do I need to run a 7B or 8B model?

At FP16/BF16, the weights alone take about 2 bytes per parameter — roughly 15–16 GB for an 8B model — plus the KV cache and overhead. Quantized to INT4, the same model needs only around 5–6 GB, which fits on many consumer GPUs.

What is the KV cache?

While generating text, a model stores the keys and values from its attention layers for every token so far. This KV cache grows with context length and batch size, and for long contexts it can use more memory than the weights.

Why does full fine-tuning need so much more memory?

Training stores gradients and optimizer states as well as the weights. With mixed-precision AdamW, that's about 16 bytes per parameter — around 128 GB for an 8B model before activations.

What's the difference between LoRA and QLoRA?

LoRA freezes the base model and trains small adapter matrices, so only a tiny fraction of parameters needs gradients and optimizer states. QLoRA also loads the frozen base model in 4-bit, cutting memory further. Choose INT4 as the base precision to estimate QLoRA.

How accurate is this estimate?

It uses standard formulas and is usually within a few gigabytes, but real usage depends on the framework (vLLM, llama.cpp, PyTorch), kernels, quantization format and settings. Leave some headroom and confirm with a test run.