Current example
Fits the budget; throughput remains an optimistic bound
Dense 70B warm decode: 80 full-attention layers, 64 query heads, 8 KV heads and head dimension 128; total modeling budget of 180 decimal GB (180e9 bytes), not a conversion of nominal B200 memory. Weights, reserve and cache all count against this budget; actual device inventory and allocatable memory must be checked separately. Entered bandwidth and compute are scenario inputs for the model. The larger traffic/compute time is a lower bound, not achieved latency or throughput. No tensor parallelism, MoE, prefix sharing, padding, launch cost or collectives.
- Warm decode step lower bound
- 30.148 ms/step
Initial model: one B200 with a total modeling budget of 180 decimal GB (180e9 bytes), a dense 70B decoder, batch 8, 4,096 retained tokens, 140 GB resident weights, 18 GB reserve, two-byte KV storage, 5,000 GB/s bandwidth and 800 TFLOP/s effective math rate.
Decode workload
Memory budget
- KV storage width
- 2 bytes per element
- Budget headroom
- 11.263 GB
Effective engine rates
- Effective math rate (TFLOP/s)
- 800
- Traffic time lower bound
- 30.148 ms/step
- Compute time lower bound
- 1.507 ms/step
02 / Follow the flow
Four boundaries. One connected explanation.
Each card explains a node in the diagram. Highlights show which boundaries contribute to the current result.
Measure the representation the engine actually holds
Packed low-precision weights may require scales, padding, transforms and tensors left at a higher precision. The model therefore accepts resident weight bytes directly rather than pretending that BF16, FP8 or NVFP4 determines the whole allocation. Inspect a versioned engine and checkpoint under the intended kernel backend. Reducing weight storage does not automatically reduce every activation, cache or communication allocation.
Math
Keep effective math rates tied to a numerical contract
The decoder approximation includes dense parameter work and a full-attention term with fixed query-head geometry. It excludes MoE routing, tensor parallelism and engine overhead. Enter a rate for the relevant operation and precision; do not compare sparse low-precision peaks with dense higher-precision work. Numerical quality and backend support must be checked independently before treating a lower-precision path as a valid optimization.
Batching trades reuse against retained state
More sequences can amortize a weight read across useful output tokens, but every retained sequence also consumes cache and contributes attention traffic. The reference shape has 80 attention layers, eight KV heads and 128 dimensions per head. The model uses uniform sequence lengths and no prefix sharing or sliding window. Paged allocation can manage a pool more efficiently without making its finite capacity disappear.
An optimistic ceiling is not useful serving throughput
The larger of compute and memory time is an idealized per-step lower bound. Dividing active sequences by that bound yields a conditional ceiling only when required memory fits the total modeling budget. It is not achieved tokens per second, time to first token or an end-to-end latency prediction. Real decisions need workload distributions, offered load, queueing, response quality, tail latency, failed requests and completed useful work at the same service-level target.
# One warm decode step; decimal GB and user-entered effective rates
kv_bytes_per_token = 2 * 80 * 8 * 128 * cacheBytes
cache_GB = batch * context * kv_bytes_per_token / 1e9
required_GB = weightGB + reserveGB + cache_GB
# 180 decimal GB is the total modeling budget, not physical capacity.
headroom_GB = 180 - required_GB
traffic_GB = weightGB + cache_GB + batch * kv_bytes_per_token / 1e9
memory_ms = traffic_GB / bandwidthGBs * 1000
flops = batch * (2 * 70e9 + 4 * 80 * 64 * 128 * context)
compute_ms = flops / (computeTF * 1e12) * 1000
step_ms = max(memory_ms, compute_ms)
tokens_ceiling = batch / step_ms * 1000 if headroom_GB >= 0 else None03 / Keep the model honest
Model assumptions
KV bytes/token = 2 × 80 layers × 8 KV heads × 128 dimensions × cache bytes
Cache GB = batch × retained tokens × KV bytes/token / 10⁹
Required GB = resident weights + reserve + cache GB
Step traffic GB = resident weights + cache GB + new-token KV writes
Memory ms = step traffic GB / effective GB/s × 10³
FLOPs = batch × (2 × 70×10⁹ + 4 × 80 × 64 × 128 × retained tokens)
Compute ms = FLOPs / (effective TFLOP/s × 10¹²) × 10³
Optimistic step lower bound = max(memory ms, compute ms)- Named reference: one NVIDIA B200, not an eight-GPU HGX total. The 180 decimal GB (180e9 bytes) are the total modeling budget, not a conversion of NVIDIA’s nominal GB specification. Device-reported nvidia-smi inventory in MiB and usable engine allocations are distinct; check the actual device and runtime. The decoder shape and effective rates are model inputs, not specifications of a deployed workload.
- A dense 70B decoder with 80 uniform full-attention layers, 64 query heads, eight KV heads and head dimension 128. No tensor/pipeline parallelism, MoE or distributed collectives.
- One warm decode iteration. No prefill, launch overhead, speculative decoding, prefix sharing, allocator granularity, cache misses beyond the stated traffic approximation or variable-length batch simulation.
- Resident weights must include actual representation overhead. Weights, cache and reserve all count against the total modeling budget. The reserve must cover other allocations and the desired margin; a zero reserve is an optimistic mathematical case, not deployment advice.
- When required memory exceeds the budget, the reader withholds the numerical throughput ceiling and preserves negative budget headroom. This is not proof of physical single-GPU infeasibility, and fitting the budget does not guarantee an engine allocation.
- All timings are modeled lower bounds. This page performs no GPU benchmark, security boundary, numerical-quality evaluation or engine validation.
04 / Think it through
Questions behind the example.
Why can a larger batch exceed the memory budget while its arithmetic throughput ceiling still looks attractive?
Why does more math throughput stop helping a memory-bound decode step?
Which quality, kernel, graph and memory measurements would be needed to validate a smaller weight footprint?
Primary documentation

