Context and performance

One million tokens: budgeting for long context

Calculate KV memory, distinguish MHA, GQA and MLA, and account for several long requests running together.

7 min readUpdated 6 October 2026
Editorial illustration

A million-token context changes server requirements even when model weights stay the same. The system retains the state of the text already read and must process the input before producing its first answer. This guide isolates the KV budget and workload test. One million means 1,000,000 tokens here: an illustrative scenario, not a supported limit for every model.

KV / MEMORY LAB

How much memory does context use?

An educational estimate for conventional full-attention KV caches: MHA, GQA or MQA. The initial values are an example, not a profile of a catalogue model.

Input + generated tokens. Every request has this length; no shared prefix is assumed.

KV cache only, aggregate327.68 GB≈ 305.18 GiB

2 × layers × KV heads × dimension × bytes × tokens × requests

GB = 10⁹ bytes · GiB = 2³⁰ bytes. FP8/INT8 uses 1 byte per value before quantization metadata; engine support needs separate verification.

Excludes weights, activations, CUDA graphs, buffers and overhead. Does not calculate MLA, compressed, sliding-window or hybrid caches. A model and engine must separately support a one-million-token context.

1. Establish the context the model actually supports

Find the context limit in the exact model card and deployment recipe. Record the revision, tokenizer, positional encoding and runtime requirements. Raising a length setting does not establish that the model reliably uses an entire document. Test facts near the beginning, middle and end, relationships between distant passages, and the quality of the final answer.

The budget includes system instructions, conversation history, retrieved documents, control tokens and the future answer. File size does not map to a fixed token count: language, format and tokenizer affect it. Tokenize representative inputs first, then define the usual length, maximum length and generation allowance.

2. Calculate KV separately from weights

For conventional full attention with equal key and value dimensions, the baseline in bytes is: 2 × layers × KV heads × head dimension × bytes per value × the sum of active sequence lengths. The factor of two covers keys and values. This is tensor payload before weights, alignment, metadata and working buffers.

For example, use 80 layers, 8 KV heads, a head dimension of 128 and two-byte BF16 values. One request retaining a million tokens needs approximately 327.68 GB, or 305.18 GiB, for KV alone. The table changes KV-head count with everything else held constant. It does not identify a checkpoint or establish million-token support.

Illustrative architectureKV headsKV for 1 million tokens
MHA642621.44 GB
GQA8327.68 GB
MQA140.96 GB

3. Use the calculation for your architecture

MHA keeps separate keys and values for attention heads. GQA shares KV across groups of query heads; MQA uses a single shared KV head. Consequently, query-head count cannot automatically substitute for KV-head count. Read the relevant dimensions from the model configuration instead of inferring them from its parameter total.

MLA, used in DeepSeek-V3, retains a compressed latent representation and needs a different calculation that matches the cache implementation. Sliding-window attention bounds history only in the affected layers; hybrid architectures require a layer-by-layer estimate. Quantized weights do not imply a quantized cache. Check cache-format support, overhead and quality separately.

4. Count active sequences rather than user accounts

Four independent million-token requests in the GQA example require 1310.72 GB of KV before weights. A hundred registered users do not necessarily create a hundred simultaneous requests: admission control and the running sequence count matter. For different lengths, add the actual retained lengths, including tokens already generated.

Scheduler batch size and maximum context length are different limits. Chunked prefill processes a long input in pieces so it can share execution with ongoing answers; it does not remove the entire retained KV state. A compatible engine can reuse a repeated prefix, but this saving should not be assumed for independent documents.

5. Measure the first answer and mixed traffic

Run separate tests with an empty cache and with a repeated prefix. Record time to first token, subsequent generation speed, peak memory on every GPU, and any preemption or recomputation. Submit short requests alongside the long document: acceptable performance for one document does not establish a responsive service.

If the budget fails, compare shorter inputs, retrieval of relevant passages, lower concurrency and supported cache compression. Moving state into CPU RAM changes latency and interconnect traffic. Choose using task quality and completion time, and retain the test configuration so the result can be reproduced.

Context support, sufficient memory and useful speed are separate checks. KV arithmetic identifies a starting configuration to test.

Checklist

Before a million-token test

Sources and documentation

These guides help you prepare requirements. The exact configuration and terms are set out in the quotation.