Hover a dotted term for 5 seconds to lock its explanation. It closes after 5 seconds away; nested tooltips and keyboard focus keep it open. Click, tap or Enter locks immediately. Technical glossary.

Kubernetes & inference / Interactive field guide

How many sequences fit before the KV cache runs out?

A scheduled GPU is not an unlimited token budget. Allocate the memory explicitly, then distinguish an arithmetic capacity limit from a safe serving policy.

4 linked boundariesLocal teaching modelNo account required

01 / Follow the explanation

The system, step by step.

A brief visual sequence plays automatically. The example and its assumptions are already here—nothing to configure.

An interactive teaching model drawn from real delivery work. No agent, cloud account, GPU or cluster is accessed.

Memory budget stages: weights allocation, non-KV reserve, remaining KV cache, sequence admission.A memory-accounting diagram, not physical buffer addresses. Follow each allocation before choosing a concurrency cap.01Weights02Reserve03KV cache04Admission
A memory-accounting diagram, not physical buffer addresses. Follow each allocation before choosing a concurrency cap.

Current example

The requested sequences fit this memory budget

This uniform decoder-only attention model budgets weights, reserve and both key/value tensors at the full selected context. It models memory, not serving throughput or latency. GiB means 2^30 bytes. Tensor parallelism, prefix sharing, sliding windows, paging fragmentation and allocator behavior are omitted. A zero cache budget is an exhausted state; negative headroom shows the shortfall.

Memory-only capacity (sequences)
56 sequences

Example: 80 GiB total, 16 GiB weights, 8 GiB non-KV reserve; 32 layers, 8 KV heads, 128 dimensions/head, 2 bytes/element, 8,192 retained tokens per sequence and 8 sequences. These are model inputs, not a named accelerator or deployed model.

Memory envelope

Total device memory (GiB)
80
Weights allocation (GiB)
16
Non-KV reserve (GiB)
8
Available KV budget (GiB)
56.000 GiB

Sequence demand

Retained tokens per sequence
8192
Concurrent sequences
8
KV storage per sequence (GiB)
1.000 GiB
Total memory required (GiB)
32.000 GiB
Remaining headroom (GiB)
48.000 GiB

Cache representation

KV heads per layer
8 KV heads
KV storage width
2 bytes per element
KV storage per token (bytes)
131072 bytes

02 / Follow the flow

Four boundaries. One connected explanation.

Each card explains a node in the diagram. Highlights show which boundaries contribute to the current result.

Weights

Resident weights are the first allocation, not the last

Parameter count multiplied by storage bytes is a starting estimate. Quantization scales, padding, tied-weight handling and runtime representation change the actual footprint. Kubernetes GPU resource limits select an advertised device allocation; they do not turn the remaining device memory into a per-request reservation. This reader takes the weights allocation as an explicit input instead of pretending to infer it from a model name.

Reserve

Leave room for the execution engine

Prefill activations, kernel workspaces, graph capture, communication buffers and runtime allocations compete with the cache. Their peaks depend on model, engine version, batch shapes and parallelism. The reserve here is one visible budget, not a universal percentage. Size it from representative execution and failure tests. If weights and reserve exceed the device, zero cache capacity is the correct result even before the first request arrives.

KV cache

Retained tokens multiply through every attention layer

For uniform full attention, each cached token stores keys and values for every layer, KV head and head dimension. The factor of two counts K and V. Grouped-query attention changes KV-head count, not necessarily query-head count. Longer retained sequences consume more cache even when only one new token is emitted per decode step. Prefix sharing, sliding-window attention and tensor-parallel layouts are deliberately excluded rather than folded into a misleading generic formula.

Admission

Memory fit is necessary; meeting the SLO is separate

The floor of available cache divided by per-sequence cache gives an arithmetic maximum for these identical sequences. It is not a recommended production concurrency. Queueing, prefill bursts, allocation granularity and latency requirements may demand a lower cap. Continuous batching improves how a serving engine schedules active work but does not create memory. Pair admission limits with queue age, time to first token, inter-token latency and completed useful work.

03 / Keep the model honest

Model assumptions

Transparent reasoning / Units and boundaries
KV bytes/token = 2 × layers × KV heads × head dimension × bytes/element
KV GiB/sequence = bytes/token × retained tokens / 2³⁰
Cache budget GiB = max(0, total − weights − reserve)
Arithmetic sequence limit = floor(cache budget / KV GiB/sequence)
Required GiB = weights + reserve + concurrency × KV GiB/sequence
Headroom GiB = total − required
  • A single replica on one device, 32 uniform decoder attention layers, head dimension 128 and no tensor/pipeline parallel sharding. The values are modeling inputs, not engine configuration advice.
  • Binary units throughout: 1 GiB = 2³⁰ bytes. Device capacity, cache tokens and units must use the same boundary.
  • No cross-request prefix sharing, sliding window, speculative branches, beam search, heterogeneous attention, paging fragmentation or allocator rounding.
  • The reserve must already cover non-KV peaks and desired margin. A negative headroom result is shown explicitly; the reader does not hide oversubscription by clamping it to zero.
  • The storage-width choice models byte count only. Quantized KV support, calibration and quality need validation on the selected backend and model.
  • This is not a scheduler or capacity recommendation. No cluster or GPU is accessed. The initial calculation and explanations work without JavaScript.

04 / Think it through

Questions behind the example.

  1. How would doubling retained context affect per-sequence cache and the arithmetic sequence limit?

  2. Why would using the query-head count in place of eight KV heads overestimate this GQA model’s cache?

  3. What happens to admission headroom when weights and reserve already exceed total memory?

Primary documentation

References & further reading

Engineering notes

Read the project behind the model.

Real client engagements and the engineering behind them.

A useful next conversation

What needs to work better?

A system, a delivery bottleneck, or an engineering opportunity. Tell me what you are building and where you want to go.

Let’s talk

Technical glossary: definitions, connected ideas and further reading.

Optional analytics off. Contact works either way.

How measurement works