Hover a dotted term for 5 seconds to lock its explanation. It closes after 5 seconds away; nested tooltips and keyboard focus keep it open. Click, tap or Enter locks immediately. Technical glossary.

Accelerator engineering

B200 Inference Optimization: vLLM, TensorRT-LLM and Cost per Token

A real client engagement. The engineering and the results are described below.

A legal-document search product we worked with has a Blackwell budget and a long queue of customers waiting on answers. Its team compares serving configurations rather than promising that a B200 or B300 label will automatically make the product economical.

Request a time through the inquiry form. A meeting is confirmed separately by email.

The business problem behind the technology

Which serving configuration makes the same answer cheaper?

A higher tokens-per-second headline cannot justify an expensive fleet if long prompts, queueing or memory pressure make useful answers miss the service objective.

Read the client engagement ↓

Client engagement / Delivered results

The GPU upgrade needed a workload contract

A document-retrieval product we worked with plans 24 million accepted responses per month. It has to serve that mix with a fixed model-quality threshold and explicit latency limits.

The constraint

A higher tokens-per-second headline cannot justify an expensive fleet if long prompts, queueing or memory pressure make useful answers miss the service objective.

The engineering decision

The team compares cache policy, batching and numerical formats under one controlled workload. The accepted configuration needed 18,000 rather than 24,000 GPU-hours a month at the same $6 hourly rate.

The delivered outcome

The monthly GPU compute charge moved from $144,000 to $108,000: $0.006 versus $0.0045 per accepted response for the fixed workload.

Measured inputs and delivered differences
Measure / unitBeforeAfterDifference
GPU compute charge for the fixed response workload
USD/month
144,000108,00036,000

GPU compute charge for the fixed response workload. 24,000 × $6 versus 18,000 × $6. Actual output-token counts and all other costs would be needed for a cost-per-token claim.

The conditions behind the results

  • The 24 million responses retain the same prompt/output distribution, quality and latency constraints.
  • The 25% GPU-hour reduction came from the engagement’s controlled workload comparison.
  • The $6 rate applies to the same exact device and commercial boundary.
  • The ledger is separate from the 70B warm-decode teaching model; networking, storage, staff and migration costs are excluded.

What this does not prove. The lower GPU charges are scoped to this workload and commercial boundary; networking, storage, staff and migration costs were excluded.

Evidence to collect for your own decision

Key decisions

Blackwell changes the available implementation choices, not the need to identify a bottleneck. Start with an explicit dense-model memory and decode bound, then test scheduling, precision, kernels, and topology against a reproducible workload and latency contract.

Follow the decision

Which serving configuration makes the same answer cheaper?

Select a step to follow its reasoning, then continue into the technical chapters.

Problem → boundary → decision → evidence

Name the configuration

Separate device, node and allocatable-memory claims.

Read every component and connection
Name the configuration · Problem
Separate device, node and allocatable-memory claims.
Follow the request · Boundary
Distinguish queueing, prefill, decode and useful completion.
Test the implementation · Decision
Inspect cache policy, batching, precision and selected kernels.
Compare honestly · Evidence
Keep the workload and quality contract fixed across engines.
  • Name the configuration → Follow the request: identify the constraint
  • Follow the request → Test the implementation: choose a bounded change
  • Test the implementation → Compare honestly: check the outcome

A conceptual decision map for this article, not a measured timeline, physical topology or a depiction of a specific client system.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Name the B200 device before borrowing a system specification

The client’s search product names its exact Blackwell configuration before treating any memory figure as available capacity.

The sourced cloud-node diagrams below add real production offerings: AWS P6-B200 and P6-B300. They retain the provider product-table memory and network figures rather than substitute a DGX chassis specification. The numerical companion reader remains a single-B200 model; it is not a B300 analysis, live tenant inspection or a guarantee of currently available rental capacity.

NVIDIA’s HGX AI Factory component table gives nominal memory labels of 180 GB HBM3e per B200 SXM GPU and 1.44 TB per node; it does not define those labels in bytes. Those are different accounting boundaries. This reader uses one GPU, not the sum of eight GPUs in an HGX system, and does not treat their memory as a transparently shared allocation. It also does not substitute a GB200 Grace Blackwell system, a B300, or a consumer Blackwell GPU for B200.

Throughout the numerical model, GB means 10^9 bytes and bandwidth GB/s uses the same decimal unit. The fixed 180e9 bytes are the total modeling budget we planned against, not an inferred physical B200 capacity or a guarantee of allocatable memory. Keep nominal vendor GB labels, device-reported inventory, and usable engine allocations separate. Vultr’s verified cookbook instance reports 183,359 MiB per B200 through nvidia-smi; MiB means 2^20 bytes. That is an inventory observation for that instance, not a universal capacity or allocation guarantee. Check the actual device and runtime: weights, cache, activations, workspaces, and captured graphs all consume memory.

The reader defaults to a 5,000 GB/s memory rate and 800 TFLOP/s arithmetic rate. Neither is a vendor peak or a sustained measurement. Keep compute precision, dense versus structured-sparse arithmetic, power configuration, device count, and interconnect boundary attached to any specification used in a real experiment. A sparse Tensor Core headline is not a dense-model rate, and node aggregate bandwidth is not one GPU’s HBM bandwidth.

Published offering · p6-b200.48xlarge

Inside a real P6-B200 cloud node

Eight B200 GPUs · 1,432 GB reported GPU memory · 192 vCPUs · 2,048 GiB host RAM

Host and staging

Eight separate GPU memory domains

Scale up within the node

Scale out, storage and isolation

Host CPUs

Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.

Read every component and connection
Host CPUs · 192 vCPUs
Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.
Host DRAM · 2,048 GiB
CPU-addressable memory stages data and hosts application state. It is separate from GPU HBM; copies, mappings and supported direct-I/O paths have distinct costs.
GPU 0 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 1 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 2 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 3 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 4 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 5 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 6 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 7 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
NVLink fabric · GPU ↔ GPU
The scale-up fabric connects compatible GPU peers for transfers and collectives. The runtime must partition work and enable the relevant paths; summed HBM is not one automatic allocation.
EFA networking · 3.2 Tbps aggregate
EFA provides the scale-out transport between instances. NCCL and the supported transport stack coordinate collectives; inter-node networking is not the same link as NVLink.
Local NVMe · 8 × 3.84 TB
Local instance storage can cache datasets and model artifacts. The listed raw drive capacity is not durable object storage or a promise of the same usable RAID capacity.
EBS path · 100 Gbps listed
The product table lists EBS bandwidth separately from local NVMe and EFA. Volume provisioning and workload behavior still constrain achieved storage throughput.
AWS Nitro · Virtualization / I/O
Nitro combines specialized hardware and firmware for the EC2 isolation and I/O boundary. Guest Linux, NVIDIA drivers and application permissions remain distinct responsibilities.

AWS product-table figures, checked September 8, 2026. This is a conceptual map of a published production offering, not access to a tenant, a live health feed or a guarantee of spare capacity. Ordinals 0–7 are diagram labels. Links show logical responsibilities, not proprietary board traces.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Published offering · p6-b300.48xlarge

Inside a real P6-B300 cloud node

Eight B300 GPUs · 2,144 GB reported GPU memory · 192 vCPUs · 4,096 GiB host RAM

Host and staging

Eight separate GPU memory domains

Scale up within the node

Scale out, storage and isolation

Host CPUs

Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.

Read every component and connection
Host CPUs · 192 vCPUs
Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.
Host DRAM · 4,096 GiB
CPU-addressable memory stages data and hosts application state. It is separate from GPU HBM; copies, mappings and supported direct-I/O paths have distinct costs.
GPU 0 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 1 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 2 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 3 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 4 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 5 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 6 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 7 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
NVLink fabric · GPU ↔ GPU
The scale-up fabric connects compatible GPU peers for transfers and collectives. The runtime must partition work and enable the relevant paths; summed HBM is not one automatic allocation.
EFA networking · 6.4 Tbps aggregate
EFA provides the scale-out transport between instances. NCCL and the supported transport stack coordinate collectives; inter-node networking is not the same link as NVLink.
Local NVMe · 8 × 3.84 TB
Local instance storage can cache datasets and model artifacts. The listed raw drive capacity is not durable object storage or a promise of the same usable RAID capacity.
EBS path · 100 Gbps listed
The product table lists EBS bandwidth separately from local NVMe and EFA. Volume provisioning and workload behavior still constrain achieved storage throughput.
AWS Nitro · Virtualization / I/O
Nitro combines specialized hardware and firmware for the EC2 isolation and I/O boundary. Guest Linux, NVIDIA drivers and application permissions remain distinct responsibilities.

AWS product-table figures, checked September 8, 2026. This is a conceptual map of a published production offering, not access to a tenant, a live health feed or a guarantee of spare capacity. Ordinals 0–7 are diagram labels. Links show logical responsibilities, not proprietary board traces.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Optimize the workload mix, not the phase name

The trace separates two accelerator workloads inside one response. Their names alone do not identify the dominant cost.

Prefill processes prompt positions and builds KV state. Large prompts can expose substantial matrix parallelism, while attention work and temporary memory grow with prompt shape. Decode usually processes one new position per active sequence, reuses resident weights across that iteration’s batch, reads prior KV, and writes the new cache entry. Low-batch decode often pressures memory traffic, but neither every prefill nor every decode has a universal bottleneck.

Define at least a short-prompt interactive class, a long-document class, and a long-output class when those exist in the service. Keep their prompt/output length distributions, arrival schedule, shared-prefix distribution, and cancellation behavior. A fixed batch of equally long requests is useful for isolating kernels but erases the changing shapes that an inference engine must schedule.

Separate pure prefill, warm decode, mixed serving, cold initialization, and burst recovery experiments. A faster prefill implementation can worsen active streams if admission permits larger uninterrupted chunks. A better decode kernel can leave user latency unchanged when tokenization, admission waiting, serialization, or the network dominates. Measure the boundary where the intended benefit should appear.

Make TTFT, inter-token latency, and goodput explicit

You turn a vague complaint about slowness into explicit user-visible measures. Goodput prevents rejected work from improving the apparent result.

Client time to first token, TTFT, starts at the declared request-send timestamp and ends when the first output token is received. It includes transport, server admission, prompt processing, and first-token delivery. Server TTFT starts elsewhere and must be labeled separately. Inter-token latency is an individual gap between successive output-token timestamps; a slow scheduling iteration can create a visible stall even when the mean gap looks acceptable.

For N output tokens with N greater than one, define per-request time per output token as TPOT = (last-token time - first-token time) / (N - 1). It is an average, not the worst inter-token gap. A one-token output has no post-first-token TPOT. If the endpoint streams chunks containing several tokens, chunk timestamps do not reveal the individual token times; report chunk gaps or document the instrumentation rather than manufacturing token-level samples.

Define request goodput as successfully completed responses meeting the declared latency limits divided by the observation interval. Token goodput counts output tokens only in those passing responses over the same interval. Report both so a verbosity change cannot masquerade as more useful answers. Preserve offered requests, failures, rejections, cancellations, and pending work in the report, and evaluate task quality separately. Never add stage p99 values and label the sum end-to-end p99.

Derive cache fit using GQA’s KV heads

The model needs room for more than weights. The KV-cache calculation follows the actual attention structure rather than the largest head-count label.

For this fixed full-attention decoder, raw KV bytes per cached token are K = 2 × 80 × 8 × 128 × cacheBytes. The two accounts for keys and values. Grouped-query attention shares 8 KV heads among 64 query heads, so storage uses 8, not 64. With two-byte cache elements, K is 327,680 bytes, or 320 KiB. Changing a cache format is an engine choice; changing a trained model’s KV-head count is not a harmless serving flag.

Let B be active batch size and T the cached context per sequence before this iteration. The cache storage estimate is C = B × T × K / 10^9 GB. Required memory is W + R + C, where W is actual resident weight GB and R is the declared reserve. Headroom within the total modeling budget is 180 - (W + R + C) GB; reserve is subtracted inside that budget, not added to it. T includes prompt and generated history, not only the original prompt. The model uses equal lengths; a real batch requires the sum of its sequence lengths.

At B = 8, T = 4,096, W = 140 GB, R = 18 GB, and two-byte cache elements, C is 10.73741824 GB, required memory is 168.73741824 GB, and headroom within the budget is 11.26258176 GB. This is tensor-accounting arithmetic, not evidence that an engine allocation succeeds. The reserve must cover peak prefill activations, graph pools, cache metadata and padding, and other runtime allocations. A zero-headroom mathematical fit offers no allowance for omitted overhead or the next cache growth.

Read the warm-decode roofline as a conditional bound

A warm-decode bound gives the investigation a reference point. The reader’s assumptions stay beside its result so they cannot become a benchmark by implication.

The reader assumes a dense weight sweep once per batch iteration, a read of the existing raw KV tensor, and an append of one KV entry per sequence. Thus stepTrafficGB = W + C + B × K / 10^9. Its arithmetic estimate is F = B × (2 × 70 × 10^9 + 4 × 80 × 64 × 128 × T) FLOPs: a simplified dense parameter term plus query-key and attention-value work. The attention arithmetic uses query heads; KV storage uses KV heads.

Memory time in milliseconds is stepTrafficGB / bandwidthGBs × 1,000. Compute time is F / (computeTF × 10^12) × 1,000. Taking their maximum gives an optimistic lower bound under the stated resource rates and work counts, with ideal overlap rather than adding both times. B / stepMs × 1,000 is an optimistic aggregate output-token ceiling only when required memory fits the total modeling budget. It is not a throughput measurement, TTFT, per-user inter-token latency, or end-to-end prediction.

The default inputs give approximately 30.148 ms for memory, 1.507 ms for compute, and a conditional ceiling of 265.36 output tokens/s across the batch. These are the calculated planning figures. The simplified work counts omit softmax, normalization, sampling, conversions, extra reads, and scheduling; they are not a rigorous lower bound for every alternative implementation. The model excludes tensor parallelism, MoE, speculative decoding, prefix sharing, allocator padding, launch overhead, and collectives. If required memory exceeds the total modeling budget of 180 decimal GB, this model withholds the token ceiling; that is not proof of physical single-GPU infeasibility.

Doubling the batch can amortize the weight sweep across more outputs, but also increases cache storage and attention work. Lowering W can reduce memory traffic until another term dominates. The compute rate must match the actual arithmetic path; do not double it just because a dropdown says FP8. Use the reader to formulate a bottleneck hypothesis, then test that hypothesis with the engine.

01 / Follow the explanation

What limits one warm decode step on Blackwell?

A brief visual sequence plays automatically. The example and its assumptions are already here—nothing to configure.

An interactive teaching model drawn from real delivery work. No agent, cloud account, GPU or cluster is accessed.

Read the full explanation and assumptions

Warm decode constraints: resident weights, arithmetic, retained KV cache and token completion.A one-step accounting model, not a hardware diagram or profiler trace. No GPU code executes in this browser.01Weights02Math03KV cache04Tokens
A one-step accounting model, not a hardware diagram or profiler trace. No GPU code executes in this browser.

Current example

Fits the budget; throughput remains an optimistic bound

Dense 70B warm decode: 80 full-attention layers, 64 query heads, 8 KV heads and head dimension 128; total modeling budget of 180 decimal GB (180e9 bytes), not a conversion of nominal B200 memory. Weights, reserve and cache all count against this budget; actual device inventory and allocatable memory must be checked separately. Entered bandwidth and compute are scenario inputs for the model. The larger traffic/compute time is a lower bound, not achieved latency or throughput. No tensor parallelism, MoE, prefix sharing, padding, launch cost or collectives.

Warm decode step lower bound
30.148 ms/step
Optimistic aggregate token ceiling
265.357 tokens/s (ceiling)

Initial model: one B200 with a total modeling budget of 180 decimal GB (180e9 bytes), a dense 70B decoder, batch 8, 4,096 retained tokens, 140 GB resident weights, 18 GB reserve, two-byte KV storage, 5,000 GB/s bandwidth and 800 TFLOP/s effective math rate.

Decode workload

Active decode sequences
8
Retained tokens per sequence
4096
KV cache
10.737 GB

Memory budget

Resident weights (decimal GB)
140
Non-KV reserve (decimal GB)
18
KV storage width
2 bytes per element
Required modeled memory
168.737 GB
Budget headroom
11.263 GB

Effective engine rates

Effective bandwidth (GB/s)
5000
Effective math rate (TFLOP/s)
800
Traffic time lower bound
30.148 ms/step
Compute time lower bound
1.507 ms/step
Instructional excerpt, not executed here
# One warm decode step; decimal GB and user-entered effective rates
kv_bytes_per_token = 2 * 80 * 8 * 128 * cacheBytes
cache_GB = batch * context * kv_bytes_per_token / 1e9
required_GB = weightGB + reserveGB + cache_GB
# 180 decimal GB is the total modeling budget, not physical capacity.
headroom_GB = 180 - required_GB
traffic_GB = weightGB + cache_GB + batch * kv_bytes_per_token / 1e9
memory_ms = traffic_GB / bandwidthGBs * 1000
flops = batch * (2 * 70e9 + 4 * 80 * 64 * 128 * context)
compute_ms = flops / (computeTF * 1e12) * 1000
step_ms = max(memory_ms, compute_ms)
tokens_ceiling = batch / step_ms * 1000 if headroom_GB >= 0 else None
python / example
"""Standalone arithmetic from the capacity plan: no engine, GPU access, or benchmark."""
batch, context, cache_bytes = 8, 4096, 2
weight_gb, reserve_gb = 140.0, 18.0
bandwidth_gbs, compute_tf = 5000.0, 800.0  # planning inputs, not peaks
kv_bytes = 2 * 80 * 8 * 128 * cache_bytes
cache_gb = batch * context * kv_bytes / 1e9
required_gb = weight_gb + reserve_gb + cache_gb
# 180 decimal GB is the total modeling budget, not physical capacity.
headroom_gb = 180.0 - required_gb
traffic_gb = weight_gb + cache_gb + batch * kv_bytes / 1e9
flops = batch * (2 * 70e9 + 4 * 80 * 64 * 128 * context)
memory_ms = traffic_gb / bandwidth_gbs * 1000
compute_ms = flops / (compute_tf * 1e12) * 1000
step_ms = max(memory_ms, compute_ms)
print(f"Required: {required_gb:.6f} GB; budget headroom: {headroom_gb:.6f} GB")
print(f"Resource bounds: {memory_ms:.6f} ms memory, {compute_ms:.6f} ms compute")
if headroom_gb >= 0:
    print(f"Conditional optimistic ceiling: {batch / step_ms * 1000:.2f} tokens/s")
else:
    print("Required memory exceeds the budget: model token ceiling is not applicable")

Paged cache and prefix cache solve different waste

Cache optimization offers several mechanisms. You distinguish wasted allocation from reusable prefixes before deciding what to change.

TensorRT-LLM documents a paged cache as blocks allocated, tracked, and recycled by a cache manager rather than one maximum-sized contiguous tensor per request. Paging improves allocation flexibility; it does not shrink the underlying full-attention KV values. Block tails can be partly unused, metadata consumes space, and allocation policy determines when a request is preempted, recomputed, or rejected.

Prefix caching reuses already computed KV for a matching prompt prefix. vLLM’s documented hash keys include prior-prefix identity, block tokens, and relevant extra identities such as adapters, multimodal inputs, and cache salts. Its design caches full blocks. Similar natural-language meaning is not a cache hit: token IDs and the relevant execution identity must match. Dynamic system messages, differing chat templates, and request-specific early tokens can destroy reuse.

Prefix hits can avoid repeated prefill work and share retained blocks, but they do not eliminate autoregressive generation or the need to attend to prior context. Report warm-prefix and cold-prefix traffic separately, including hit definitions, eviction policy, and cache occupancy. Decide whether tenant isolation requires separate cache identities; a cache key is not a blanket confidentiality guarantee. This reader intentionally counts independent full caches with no prefix-sharing credit.

Sliding-window layers, hybrid attention, latent/compressed KV, KV offload, and quantized-cache metadata need different accounting. The reader’s one-byte option halves raw tensor storage but is not an implementation of every FP8 cache layout. In particular it does not model NVFP4 KV storage. Weight precision and KV precision remain independent settings subject to the selected engine’s support matrix.

Continuous batching couples scheduler policy to kernel shape

Its cost-per-response model now depends on serving the same workload rather than making the benchmark easier.

Continuous batching, also called in-flight or iteration-level batching, admits and removes sequences between engine iterations. The active set can change without waiting for every request in a static batch to finish. TensorRT-LLM 1.1.0 documents packed input for in-flight batching and distinguishes maximum request count, maximum sequence length, and the per-iteration token budget. Those limits are not interchangeable.

A decode request may need one token of new work while a new prompt needs thousands. Chunked prefill spreads that prompt across iterations so it can coexist with decode. Smaller chunks may protect token-gap latency but increase scheduling overhead or delay first-token completion; larger chunks can improve prefill efficiency while stalling ongoing streams. Tune the request and token limits together on the mixed workload, not just on pure decode.

vLLM and TensorRT-LLM expose similar scheduling concepts through different versioned configurations. Record resolved values and the engine backend rather than assuming the same flag name has identical semantics. Bound admission queue length, maximum prompt plus output length, request deadline, and cancellation cleanup. Stop a concurrency sweep when cache churn or waiting time rises without a goodput benefit; accepting more requests is not equivalent to finishing more useful work.

Treat BF16, FP8, and NVFP4 as numerical contracts

A smaller precision format offers another possibility. It must pass the numerical and quality contract before its speed matters.

BF16 uses two bytes per stored element and is a useful reference precision, not a guarantee of exact outputs or universal kernel support. FP8 names a family of formats and scaling recipes, not a single drop-in datatype. Specify the value format, per-tensor/per-channel/per-block scale layout, static calibration versus dynamic activation scaling, accumulator/output types, and which tensors remain at higher precision. A checkpoint’s storage dtype does not prove its GEMMs execute on the intended low-precision path.

NVIDIA describes NVFP4 as E2M1 four-bit values with E4M3 FP8 scales for 16-value micro-blocks and a second-level FP32 tensor scale. This differs from generic INT4 and from MXFP4. Four data bits plus eight scale bits per 16 values already means 4.5 bits per value before global scales, padding, and higher-precision tensors. Dividing a BF16 file size by four therefore does not determine resident NVFP4 memory.

Record quantizer and export versions, calibration corpus and preprocessing, scale tensors, exclusions from quantization, and checkpoint digest. Verify that the exact vLLM or TensorRT-LLM release, CUDA/runtime stack, B200 architecture target, model implementation, and chosen kernel backend support that checkpoint recipe. TensorRT-LLM publishes separate model and hardware support matrices; a supported GPU row alone does not establish support for every model or KV format. Documentation under stable or latest can move, so save the version used in the experiment.

Compare application accuracy, long-context retrieval, numerical stability, structured-output validity, and representative difficult prompts against the BF16 reference using fixed evaluation rules. Quality differences need not be obvious in a short fluent answer. Recheck output-length and stopping distributions: a precision change that changes generated work can invalidate a throughput comparison. Lower storage and faster supported kernels are opportunities, never a guaranteed speedup or guaranteed lossless quantization.

Compile and capture before measuring the warm path

The warm path is not ready to time until compilation and capture have been handled deliberately. Otherwise, different experiments can measure different work.

Compilation and CUDA Graph capture address different costs. Compilation can specialize, fuse, or choose an implementation; a CUDA Graph records an eligible execution sequence for replay with less repeated launch work. vLLM documents separate full, piecewise, and decode-specific graph modes, with dispatch depending on batch shape and attention-backend capability. Requesting a mode does not prove that every mixed batch uses it: unsupported cases can take a different path.

Warm the actual prefill, mixed, and decode shape buckets required by the experiment. Record cold model loading, compilation, graph capture, and warm-up duration separately from steady state. Graph pools and shape specialization also consume memory; reducing launch overhead while crowding out cache can lower serving goodput. A warmed batch size of one is not evidence that a later large mixed batch avoids compilation or eager fallback.

Use startup logs and traces to confirm the resolved compiler and capture behavior. Keep an eager or non-graph baseline for isolating the benefit, holding numerical settings and workload fixed. Do not report a cached second launch as a cold-start result. A server becoming responsive before model initialization, allocation, and bounded warm-up finish is not evidence of a usable warm replica.

Profile the implementation actually selected on B200

The engine’s label is not an execution trace. You inspect which B200 kernels vLLM or TensorRT-LLM actually selects for the workload.

Attention kernels differ across prefill and decode in tiling, reduction strategy, supported head dimensions, cache layout, mask handling, GQA mapping, and precision. A FlashAttention, FlashInfer, Triton, or TensorRT-LLM label is not itself an execution trace. Confirm the selected implementation and fallback paths for the exact B200, engine build, model, and input shape. An implementation optimized for a different architecture can be a poor baseline even when it runs.

Start with a timeline of CPU scheduling, CUDA launches, transfers, kernels, and stream dependencies. Look for idle gaps, synchronizations, repeated layout conversions, allocation, or compilation during the supposedly warm interval. Only then inspect expensive kernels for memory traffic, arithmetic throughput, occupancy limits, register pressure, and shape inefficiency. Occupancy and utilization are clues rather than optimization objectives; fusion can reduce HBM traffic while increasing register pressure.

For an isolated device interval, use timing events on the stream containing the work and wait for completion before consuming elapsed time. Multi-stream or distributed work needs an explicit completion dependency covering every relevant stream or rank. Host end-to-end timing instead includes dispatch and synchronization at its declared boundary. A host timer around an asynchronous launch alone mostly measures enqueueing. Profile a bounded representative slice separately from uninstrumented performance runs because profiling itself can perturb execution.

Conceptual Blackwell execution and memory map

From HBM bytes to a result inside a GPU

Off-chip DRAM → memory controllers → caches → registers / execution → stores. Instruction issue and asynchronous copies coordinate different paths.

Large off-chip storage

On-chip reuse and staging

Inside one representative streaming multiprocessor

Device-wide work and peers

HBM3e DRAM

Stacked DRAM stores large device-resident arrays. DRAM cells need refresh; capacity and sustained bandwidth are different limits. HBM is not the register file.

Read every component and connection
HBM3e DRAM · Weights / KV / arrays
Stacked DRAM stores large device-resident arrays. DRAM cells need refresh; capacity and sustained bandwidth are different limits. HBM is not the register file.
Memory controllers · Channels and requests
Controllers organize reads and writes to memory channels. Access patterns, contention and the memory technology influence service time; bandwidth is not zero latency.
L2 cache · Device-wide reuse
L2 can satisfy repeated requests without another HBM access. Its capacity and residency behavior affect traffic; a cache hit is not a new DRAM transfer.
L1 / shared memory · Caching / explicit tiles
B200 combines L1, texture and shared-memory resources. Shared memory is software-managed block storage with synchronization rules; it is not an automatic replacement for registers.
Register file · Thread operands
The B200 tuning guide specifies 64K 32-bit registers per SM. Threads use registers for live values; spills can create device-memory traffic. Registers are not off-chip DRAM.
ALU pipelines · Integer / floating point
Execution pipelines perform supported arithmetic and logic on operands. CMOS gates underlie those circuits. Floating-point operations include more work than the integer full adder shown.
Tensor cores · Matrix operations
Specialized matrix instructions use supported operand formats and accumulation paths. Tensor throughput is not scalar ALU throughput, and not every kernel can use tensor cores.
Load / store units · Addresses and movement
Load/store machinery forms and services memory operations. Coalescing groups useful lane accesses; dependencies prevent a consumer from using a value before it is ready.
Warp schedulers · Ready instruction issue
Schedulers issue eligible warp instructions subject to dependencies and resource availability. Other ready warps can hide a wait; occupancy alone does not prove throughput.
Instruction path · Fetch / decode / issue
Compiled machine instructions reach the SM instruction machinery. PTX is a virtual ISA; a compatible cubin or driver compilation supplies hardware-executable code.
Async copy / TMA · Tile movement
Supported asynchronous transfer paths can stage tiles while computation proceeds. Barriers and producer/consumer ordering still apply; overlap is not permission to read unfinished data.
GPU front end · Submitted work
Device work submission and scheduling machinery distribute kernel work. Block resource requirements influence residency. Kubernetes does not choose a warp or allocate an SM register.
NVLink interface · Peer devices
Peer access and collectives move data between compatible GPUs. The application/runtime manages distributed work; aggregate device memory is not one automatically shared allocation.

This is a functional map, not a floorplan or cycle-accurate simulator. One representative SM is expanded; it is not the GPU’s SM count. Cache bypass, asynchronous copies, distributed shared memory and specialized tensor accumulator paths mean not every operation follows every arrow.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Dense, MoE, and tensor parallelism need different models

The model architecture changes the resource story. Dense, MoE and parallel configurations need their own explicit assumptions.

The companion bound assumes dense weights with the declared 70B parameter term. A mixture-of-experts model activates only selected experts for each token, but its resident weight requirement can include many inactive experts. Routing distribution, expert imbalance, grouped GEMM shapes, dispatch/combine work, and expert-parallel all-to-all traffic affect execution. Neither total parameter count nor active parameter count alone can replace the dense formula.

Tensor parallelism shards operations across GPUs and introduces communication on the inference path. Some layers require collectives between partial results; a small low-batch matrix operation may complete quickly enough that collective latency dominates. KV-head placement can shard or replicate cache depending on head counts and implementation, so dividing every memory term by the tensor-parallel degree is not valid.

A model that fits on one B200 may benefit from independent replicas for aggregate traffic without paying per-step model-parallel communication, but that does not make replication the right answer for every latency or capacity target. If testing an HGX group, record GPU count, placement, NVLink/NVSwitch connectivity, host NUMA alignment, and any cross-host NIC path. Benchmark collectives at the message sizes and concurrency of the model rather than substituting an aggregate fabric headline.

Weight or KV offload and prefill/decode disaggregation add transfer and lifetime boundaries. Include host/device or inter-node bytes, overlap constraints, and failure handling in their own experiment. None of those alternatives is represented by the single-device reader, and a configuration that exceeds its total modeling budget cannot be rescued by interpreting its arithmetic token ceiling as an offloaded result.

A bounded timing example makes the acceptance rule executable

The acceptance rule becomes executable. The bounded timing example keeps synchronization and the timing boundary in view.

The complete standard-library Python example below reduces a small recorded trace; it does not contact an engine or measure a GPU. The declared observation window is one second. Each successful row contains exactly one client receipt timestamp per output token on the same monotonic clock as the send timestamp. A real chunked API does not automatically satisfy that assumption. The latency limits are policy inputs, not recommended B200 targets.

A response passes only if it completes inside the window, meets a 250 ms TTFT limit, and has no individual token gap greater than 50 ms. A one-token output passes the gap condition vacuously but has no TPOT value. The failed row remains in the offered count. The second response in the trace has a 60 ms gap and fails even though its average gap is 40 ms. The resulting goodput is therefore two requests/s and four tokens/s, not all emitted tokens divided by the window.

For a real run, record the engine/model/tokenizer digests, resolved scheduling and precision configuration, exact measurement boundaries, and request-to-output accounting alongside the trace. Define how responses crossing the cutoff are drained or classified; this small example rejects them from goodput. It demonstrates a timing contract completely without pretending to be a runnable deployment.

python / example
"""Recorded client token timestamps; no engine or hardware is measured."""
config = {"start_s": 0.0, "end_s": 1.0, "ttft_ms": 250.0, "max_gap_ms": 50.0}
rows = [
    {"sent": 0.00, "tokens": [0.10, 0.13, 0.16], "done": 0.17, "ok": True},
    {"sent": 0.10, "tokens": [0.20, 0.22, 0.28], "done": 0.29, "ok": True},
    {"sent": 0.20, "tokens": [0.30], "done": 0.31, "ok": True},
    {"sent": 0.40, "tokens": [], "done": 0.45, "ok": False},
]
window_s = config["end_s"] - config["start_s"]
assert window_s > 0
passed_requests = passed_tokens = 0
for index, row in enumerate(rows):
    times = row["tokens"]
    assert config["start_s"] <= row["sent"] < config["end_s"]
    assert row["done"] >= row["sent"]
    assert times == sorted(times)
    assert all(config["start_s"] <= t < config["end_s"] for t in times)
    complete = row["ok"] and row["done"] < config["end_s"]
    ttft_ok = bool(times) and (times[0] - row["sent"]) * 1000 <= config["ttft_ms"]
    gaps = [round((b - a) * 1000) for a, b in zip([row["sent"], *times], [*times, row["done"]])][1:]
    gap_ok = all(g <= config["max_gap_ms"] for g in gaps)
    if complete and ttft_ok and gap_ok:
        passed_requests += 1
        passed_tokens += len(times)
print(f"Goodput: {passed_requests / window_s:.2f} requests/s, {passed_tokens / window_s:.2f} tokens/s")

Compare engines with one experiment contract and explicit stops

The final comparison must support the fixed response volume and quality contract before the compute difference became a real proposal.

Compare vLLM and TensorRT-LLM using the same checkpoint semantics, tokenizer, prompt corpus, requested output limits, sampling rules, prefix-cache state, arrival schedule, B200 allocation, and quality acceptance rule. Save engine/backend versions, container and model digests, CUDA and driver versions, quantization recipe, graph settings, scheduling limits, topology, and device power/sharing state. Do not assume a TensorRT-LLM PyTorch backend and a built TensorRT engine have identical setup or feature behavior.

First establish a correct supported baseline, then change one factor: batch/token budget, prefill chunking, graph mode, attention backend, or precision. Use both isolated phase experiments and mixed-service confirmation. An open-loop offered-load schedule exposes queue accumulation; a closed-loop fixed-concurrency test backs off automatically when responses slow. Report which was used, verify the load generator is not limiting the run, and retain the same workload distribution across candidates.

Repeat bounded warm runs and report variation, sample counts, offered/completed/pending totals, client TTFT and token-gap distributions, request/token goodput, cache occupancy, preemptions, memory peaks, and output-quality results. Keep startup and warm-up costs separate. A brief run cannot support a stable extreme percentile simply because a reporting tool prints one. Compare performance at shared quality and service objectives rather than ranking maximum unconstrained token counts.

Choose stop conditions before the sweep: invalid outputs or quality regression, allocation failure, wrong precision/backend fallback, persistent preemption, unbounded queue age, rising timeout/rejection rates, or violated latency objectives. Stop increasing a setting when repeated runs show no meaningful goodput improvement relative to variation. Restore the last recorded valid configuration rather than stacking additional flags on an unexplained regression. This is the proposed experiment procedure; no engine comparison, failure injection, or hardware timing is claimed here.

Questions behind the decision

Is vLLM or TensorRT-LLM faster on B200?

The answer depends on the model, format, workload mix, cache policy, selected kernels and software versions. Compare accepted work under the same quality and latency limits instead of transferring another benchmark’s ranking.

Can prefix caching lower time to first token?

Reusable prefixes can avoid some repeated prefill work when the cache is valid and available. They do not eliminate queueing, uncached input processing or decode costs, and low reuse can change the economic result.

References & further reading

Engineering notes

Keep following the thread.

Real client engagements and the engineering behind them.

Technical glossary: definitions, connected ideas and further reading.

Optional analytics off. Contact works either way.

How measurement works