Current example
The requested sequences fit this memory budget
This uniform decoder-only attention model budgets weights, reserve and both key/value tensors at the full selected context. It models memory, not serving throughput or latency. GiB means 2^30 bytes. Tensor parallelism, prefix sharing, sliding windows, paging fragmentation and allocator behavior are omitted. A zero cache budget is an exhausted state; negative headroom shows the shortfall.
- Memory-only capacity (sequences)
- 56 sequences
Example: 80 GiB total, 16 GiB weights, 8 GiB non-KV reserve; 32 layers, 8 KV heads, 128 dimensions/head, 2 bytes/element, 8,192 retained tokens per sequence and 8 sequences. These are model inputs, not a named accelerator or deployed model.
Memory envelope
- Total device memory (GiB)
- 80
Sequence demand
- Retained tokens per sequence
- 8192
- Concurrent sequences
- 8
Cache representation
- KV storage width
- 2 bytes per element
02 / Follow the flow
Four boundaries. One connected explanation.
Each card explains a node in the diagram. Highlights show which boundaries contribute to the current result.
Resident weights are the first allocation, not the last
Parameter count multiplied by storage bytes is a starting estimate. Quantization scales, padding, tied-weight handling and runtime representation change the actual footprint. Kubernetes GPU resource limits select an advertised device allocation; they do not turn the remaining device memory into a per-request reservation. This reader takes the weights allocation as an explicit input instead of pretending to infer it from a model name.
Reserve
Leave room for the execution engine
Prefill activations, kernel workspaces, graph capture, communication buffers and runtime allocations compete with the cache. Their peaks depend on model, engine version, batch shapes and parallelism. The reserve here is one visible budget, not a universal percentage. Size it from representative execution and failure tests. If weights and reserve exceed the device, zero cache capacity is the correct result even before the first request arrives.
Retained tokens multiply through every attention layer
For uniform full attention, each cached token stores keys and values for every layer, KV head and head dimension. The factor of two counts K and V. Grouped-query attention changes KV-head count, not necessarily query-head count. Longer retained sequences consume more cache even when only one new token is emitted per decode step. Prefix sharing, sliding-window attention and tensor-parallel layouts are deliberately excluded rather than folded into a misleading generic formula.
Memory fit is necessary; meeting the SLO is separate
The floor of available cache divided by per-sequence cache gives an arithmetic maximum for these identical sequences. It is not a recommended production concurrency. Queueing, prefill bursts, allocation granularity and latency requirements may demand a lower cap. Continuous batching improves how a serving engine schedules active work but does not create memory. Pair admission limits with queue age, time to first token, inter-token latency and completed useful work.
03 / Keep the model honest
Model assumptions
KV bytes/token = 2 × layers × KV heads × head dimension × bytes/element
KV GiB/sequence = bytes/token × retained tokens / 2³⁰
Cache budget GiB = max(0, total − weights − reserve)
Arithmetic sequence limit = floor(cache budget / KV GiB/sequence)
Required GiB = weights + reserve + concurrency × KV GiB/sequence
Headroom GiB = total − required- A single replica on one device, 32 uniform decoder attention layers, head dimension 128 and no tensor/pipeline parallel sharding. The values are modeling inputs, not engine configuration advice.
- Binary units throughout: 1 GiB = 2³⁰ bytes. Device capacity, cache tokens and units must use the same boundary.
- No cross-request prefix sharing, sliding window, speculative branches, beam search, heterogeneous attention, paging fragmentation or allocator rounding.
- The reserve must already cover non-KV peaks and desired margin. A negative headroom result is shown explicitly; the reader does not hide oversubscription by clamping it to zero.
- The storage-width choice models byte count only. Quantized KV support, calibration and quality need validation on the selected backend and model.
- This is not a scheduler or capacity recommendation. No cluster or GPU is accessed. The initial calculation and explanations work without JavaScript.
04 / Think it through
Questions behind the example.
Primary documentation

