A model’s disk size is a useful starting point, but it does not describe all memory used during operation. The server must hold weights, the state of current requests and temporary computation data. This guide covers inference: obtaining answers from an already trained model. Training and fine-tuning require a separate calculation.
1. Split the budget into components
Weights are the model’s stored parameters. Start with an exact checkpoint revision, its format and the size of its weight files. Do not treat an entire download directory as weights; it may contain supporting material. Conversely, stored tensor size need not match GPU allocation after loading and conversion.
Add KV cache, intermediate activations, computational workspaces and runtime overhead. Headroom must cover a specific scenario. “Free after loading” does not tell you what happens with several long requests. Fix the units too: decimal GB and binary GiB are different.
2. Treat precision as a scenario to validate
For a rough lower bound, multiply the parameter count by the size of one value. The table describes nominal storage only: quantization scales, metadata and layers kept at another precision add overhead. A mixed-format checkpoint must be assessed from its actual contents.
Reducing numerical precision is called quantization. It does not guarantee higher speed or preserved quality on your task. Check whether the accelerator and runtime support the specific format, whether suitable weights exist, and how they perform on evaluation samples. Set weight and KV-cache precision separately; they are different settings.
| Format | Nominal storage per parameter | Check separately |
|---|---|---|
| BF16 | 2 bytes | Actual weight contents and working memory |
| FP8 | 1 byte | Format support, scales and layer precision |
| INT4 | 0.5 bytes | Quantization method, auxiliary data and quality |
3. Budget for long context
The KV cache stores intermediate attention data for text already processed. Its footprint depends on architecture, context length, storage format and the number of concurrent sequences. An implementation may pre-allocate memory; some architectures limit cache growth. A single formula without model details is therefore insufficient.
Specify input length, maximum output and concurrency together for the test. Compare ordinary operation with peak demand, not just one short request. Moving cache to system RAM or quantizing it may reduce GPU memory use, but changes the operating conditions. Decide using measured latency, throughput and quality.
4. Count all resident MoE weights
A mixture-of-experts model, or MoE, selects some model blocks to process a token. Its active parameter count describes the portion used in computation, not total storage. When all weights reside in GPU memory, the budget includes every expert and shared layer. Do not substitute the active count for the full resident weight count.
If some weights move to RAM or span nodes, that is a separate execution arrangement with its own data-transfer requirements. Ask where each part resides, what stays on each GPU and which measurements support the arrangement. A low active parameter count does not itself establish a small memory footprint.
5. Validate distribution and remaining capacity
Multiple GPUs do not automatically form one large memory device. The runtime partitions the model, for example by splitting computation within layers or placing different layers on different devices. Check support for that arrangement, interconnects, communication overhead, replicated data and peak usage on the busiest GPU.
Build a planning budget first, then run the chosen weights in the agreed environment. Measure each GPU’s memory, latency and errors under the target demand. Test startup and restart separately. Remaining capacity should match the agreed growth scenario, rather than an arbitrary percentage. If the test fails, change context, concurrency, format or configuration and measure again.
A positive estimate identifies a scenario worth testing. Readiness requires a repeatable test of the actual weights, software and workload; total VRAM in a specification cannot establish it.
Six checks before choosing memory capacity
Sources and documentation
- Hugging Face Transformers: KV cache strategies ↗
- Hugging Face: mixture of experts explained ↗
- vLLM: parallelism and scaling ↗
These guides help you prepare requirements. The exact configuration and terms are set out in the quotation.
