Accelerator engineering
B200 Inference Optimization: vLLM, TensorRT-LLM and Cost per Token
A real client engagement. The engineering and the results are described below.
A legal-document search product we worked with has a Blackwell budget and a long queue of customers waiting on answers. Its team compares serving configurations rather than promising that a B200 or B300 label will automatically make the product economical.
Request a time through the inquiry form. A meeting is confirmed separately by email.
The business problem behind the technology
Which serving configuration makes the same answer cheaper?
A higher tokens-per-second headline cannot justify an expensive fleet if long prompts, queueing or memory pressure make useful answers miss the service objective.
Read the client engagement ↓Client engagement / Delivered results
The GPU upgrade needed a workload contract
A document-retrieval product we worked with plans 24 million accepted responses per month. It has to serve that mix with a fixed model-quality threshold and explicit latency limits.
The constraint
A higher tokens-per-second headline cannot justify an expensive fleet if long prompts, queueing or memory pressure make useful answers miss the service objective.
The engineering decision
The team compares cache policy, batching and numerical formats under one controlled workload. The accepted configuration needed 18,000 rather than 24,000 GPU-hours a month at the same $6 hourly rate.
The delivered outcome
The monthly GPU compute charge moved from $144,000 to $108,000: $0.006 versus $0.0045 per accepted response for the fixed workload.
| Measure / unit | Before | After | Difference |
|---|---|---|---|
| GPU compute charge for the fixed response workload USD/month | 144,000 | 108,000 | 36,000 |
GPU compute charge for the fixed response workload. 24,000 × $6 versus 18,000 × $6. Actual output-token counts and all other costs would be needed for a cost-per-token claim.
The conditions behind the results
- The 24 million responses retain the same prompt/output distribution, quality and latency constraints.
- The 25% GPU-hour reduction came from the engagement’s controlled workload comparison.
- The $6 rate applies to the same exact device and commercial boundary.
- The ledger is separate from the 70B warm-decode teaching model; networking, storage, staff and migration costs are excluded.
What this does not prove. The lower GPU charges are scoped to this workload and commercial boundary; networking, storage, staff and migration costs were excluded.
Evidence to collect for your own decision
- Record model/version, numerical format, device topology and serving-engine configuration.
- Replay cold and warm cache cases with queueing, TTFT, inter-token latency and rejected work included.
- Calculate charges per accepted response and per token only with matched quality and denominators.
Key decisions
Blackwell changes the available implementation choices, not the need to identify a bottleneck. Start with an explicit dense-model memory and decode bound, then test scheduling, precision, kernels, and topology against a reproducible workload and latency contract.
- Record model/version, numerical format, device topology and serving-engine configuration.
- Replay cold and warm cache cases with queueing, TTFT, inter-token latency and rejected work included.
- Lower modeled GPU charges do not establish total platform savings, a universal engine winner or an achieved B200/B300 speedup.
Follow the decision
Select a step to follow its reasoning, then continue into the technical chapters.
Problem → boundary → decision → evidence
Read every component and connection
- Name the configuration · Problem
- Separate device, node and allocatable-memory claims.
- Compare honestly · Evidence
- Keep the workload and quality contract fixed across engines.
- Name the configuration → Follow the request: identify the constraint
- Follow the request → Test the implementation: choose a bounded change
- Test the implementation → Compare honestly: check the outcome
A conceptual decision map for this article, not a measured timeline, physical topology or a depiction of a specific client system.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
Name the B200 device before borrowing a system specification
The client’s search product names its exact Blackwell configuration before treating any memory figure as available capacity.
The sourced cloud-node diagrams below add real production offerings: AWS P6-B200 and P6-B300. They retain the provider product-table memory and network figures rather than substitute a DGX chassis specification. The numerical companion reader remains a single-B200 model; it is not a B300 analysis, live tenant inspection or a guarantee of currently available rental capacity.
NVIDIA’s HGX AI Factory component table gives nominal memory labels of 180 GB HBM3e per B200 SXM GPU and 1.44 TB per node; it does not define those labels in bytes. Those are different accounting boundaries. This reader uses one GPU, not the sum of eight GPUs in an HGX system, and does not treat their memory as a transparently shared allocation. It also does not substitute a GB200 Grace Blackwell system, a B300, or a consumer Blackwell GPU for B200.
Throughout the numerical model, GB means 10^9 bytes and bandwidth GB/s uses the same decimal unit. The fixed 180e9 bytes are the total modeling budget we planned against, not an inferred physical B200 capacity or a guarantee of allocatable memory. Keep nominal vendor GB labels, device-reported inventory, and usable engine allocations separate. Vultr’s verified cookbook instance reports 183,359 MiB per B200 through nvidia-smi; MiB means 2^20 bytes. That is an inventory observation for that instance, not a universal capacity or allocation guarantee. Check the actual device and runtime: weights, cache, activations, workspaces, and captured graphs all consume memory.
The reader defaults to a 5,000 GB/s memory rate and 800 TFLOP/s arithmetic rate. Neither is a vendor peak or a sustained measurement. Keep compute precision, dense versus structured-sparse arithmetic, power configuration, device count, and interconnect boundary attached to any specification used in a real experiment. A sparse Tensor Core headline is not a dense-model rate, and node aggregate bandwidth is not one GPU’s HBM bandwidth.
Published offering · p6-b200.48xlarge
Eight B200 GPUs · 1,432 GB reported GPU memory · 192 vCPUs · 2,048 GiB host RAM
Host and staging
Eight separate GPU memory domains
Scale up within the node
Scale out, storage and isolation
Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.
Read every component and connection
- Host CPUs · 192 vCPUs
- Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.
- Host DRAM · 2,048 GiB
- CPU-addressable memory stages data and hosts application state. It is separate from GPU HBM; copies, mappings and supported direct-I/O paths have distinct costs.
- GPU 0 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 1 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 2 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 3 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 4 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 5 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 6 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 7 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- NVLink fabric · GPU ↔ GPU
- The scale-up fabric connects compatible GPU peers for transfers and collectives. The runtime must partition work and enable the relevant paths; summed HBM is not one automatic allocation.
- EFA networking · 3.2 Tbps aggregate
- EFA provides the scale-out transport between instances. NCCL and the supported transport stack coordinate collectives; inter-node networking is not the same link as NVLink.
- Local NVMe · 8 × 3.84 TB
- Local instance storage can cache datasets and model artifacts. The listed raw drive capacity is not durable object storage or a promise of the same usable RAID capacity.
- EBS path · 100 Gbps listed
- The product table lists EBS bandwidth separately from local NVMe and EFA. Volume provisioning and workload behavior still constrain achieved storage throughput.
- AWS Nitro · Virtualization / I/O
- Nitro combines specialized hardware and firmware for the EC2 isolation and I/O boundary. Guest Linux, NVIDIA drivers and application permissions remain distinct responsibilities.
- Local NVMe → Host DRAM: stage artifacts
- EBS path → Host DRAM: volume I/O
- Host CPUs → Host DRAM: host memory
- Host DRAM → GPU 0: supported device transfer
- AWS Nitro → Host CPUs: guest / I/O boundary
- GPU 0 → NVLink fabric: NVLink peers
- GPU 1 → NVLink fabric: NVLink peers
- GPU 2 → NVLink fabric: NVLink peers
- GPU 3 → NVLink fabric: NVLink peers
- GPU 4 → NVLink fabric: NVLink peers
- GPU 5 → NVLink fabric: NVLink peers
- GPU 6 → NVLink fabric: NVLink peers
- GPU 7 → NVLink fabric: NVLink peers
- EFA networking → GPU 0: collective transport
AWS product-table figures, checked September 8, 2026. This is a conceptual map of a published production offering, not access to a tenant, a live health feed or a guarantee of spare capacity. Ordinals 0–7 are diagram labels. Links show logical responsibilities, not proprietary board traces.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
Published offering · p6-b300.48xlarge
Eight B300 GPUs · 2,144 GB reported GPU memory · 192 vCPUs · 4,096 GiB host RAM
Host and staging
Eight separate GPU memory domains
Scale up within the node
Scale out, storage and isolation
Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.
Read every component and connection
- Host CPUs · 192 vCPUs
- Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.
- Host DRAM · 4,096 GiB
- CPU-addressable memory stages data and hosts application state. It is separate from GPU HBM; copies, mappings and supported direct-I/O paths have distinct costs.
- GPU 0 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 1 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 2 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 3 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 4 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 5 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 6 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 7 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- NVLink fabric · GPU ↔ GPU
- The scale-up fabric connects compatible GPU peers for transfers and collectives. The runtime must partition work and enable the relevant paths; summed HBM is not one automatic allocation.
- EFA networking · 6.4 Tbps aggregate
- EFA provides the scale-out transport between instances. NCCL and the supported transport stack coordinate collectives; inter-node networking is not the same link as NVLink.
- Local NVMe · 8 × 3.84 TB
- Local instance storage can cache datasets and model artifacts. The listed raw drive capacity is not durable object storage or a promise of the same usable RAID capacity.
- EBS path · 100 Gbps listed
- The product table lists EBS bandwidth separately from local NVMe and EFA. Volume provisioning and workload behavior still constrain achieved storage throughput.
- AWS Nitro · Virtualization / I/O
- Nitro combines specialized hardware and firmware for the EC2 isolation and I/O boundary. Guest Linux, NVIDIA drivers and application permissions remain distinct responsibilities.
- Local NVMe → Host DRAM: stage artifacts
- EBS path → Host DRAM: volume I/O
- Host CPUs → Host DRAM: host memory
- Host DRAM → GPU 0: supported device transfer
- AWS Nitro → Host CPUs: guest / I/O boundary
- GPU 0 → NVLink fabric: NVLink peers
- GPU 1 → NVLink fabric: NVLink peers
- GPU 2 → NVLink fabric: NVLink peers
- GPU 3 → NVLink fabric: NVLink peers
- GPU 4 → NVLink fabric: NVLink peers
- GPU 5 → NVLink fabric: NVLink peers
- GPU 6 → NVLink fabric: NVLink peers
- GPU 7 → NVLink fabric: NVLink peers
- EFA networking → GPU 0: collective transport
AWS product-table figures, checked September 8, 2026. This is a conceptual map of a published production offering, not access to a tenant, a live health feed or a guarantee of spare capacity. Ordinals 0–7 are diagram labels. Links show logical responsibilities, not proprietary board traces.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
Optimize the workload mix, not the phase name
The trace separates two accelerator workloads inside one response. Their names alone do not identify the dominant cost.
Prefill processes prompt positions and builds KV state. Large prompts can expose substantial matrix parallelism, while attention work and temporary memory grow with prompt shape. Decode usually processes one new position per active sequence, reuses resident weights across that iteration’s batch, reads prior KV, and writes the new cache entry. Low-batch decode often pressures memory traffic, but neither every prefill nor every decode has a universal bottleneck.
Define at least a short-prompt interactive class, a long-document class, and a long-output class when those exist in the service. Keep their prompt/output length distributions, arrival schedule, shared-prefix distribution, and cancellation behavior. A fixed batch of equally long requests is useful for isolating kernels but erases the changing shapes that an inference engine must schedule.
Separate pure prefill, warm decode, mixed serving, cold initialization, and burst recovery experiments. A faster prefill implementation can worsen active streams if admission permits larger uninterrupted chunks. A better decode kernel can leave user latency unchanged when tokenization, admission waiting, serialization, or the network dominates. Measure the boundary where the intended benefit should appear.
Make TTFT, inter-token latency, and goodput explicit
You turn a vague complaint about slowness into explicit user-visible measures. Goodput prevents rejected work from improving the apparent result.
Client time to first token, TTFT, starts at the declared request-send timestamp and ends when the first output token is received. It includes transport, server admission, prompt processing, and first-token delivery. Server TTFT starts elsewhere and must be labeled separately. Inter-token latency is an individual gap between successive output-token timestamps; a slow scheduling iteration can create a visible stall even when the mean gap looks acceptable.
For N output tokens with N greater than one, define per-request time per output token as TPOT = (last-token time - first-token time) / (N - 1). It is an average, not the worst inter-token gap. A one-token output has no post-first-token TPOT. If the endpoint streams chunks containing several tokens, chunk timestamps do not reveal the individual token times; report chunk gaps or document the instrumentation rather than manufacturing token-level samples.
Define request goodput as successfully completed responses meeting the declared latency limits divided by the observation interval. Token goodput counts output tokens only in those passing responses over the same interval. Report both so a verbosity change cannot masquerade as more useful answers. Preserve offered requests, failures, rejections, cancellations, and pending work in the report, and evaluate task quality separately. Never add stage p99 values and label the sum end-to-end p99.
Derive cache fit using GQA’s KV heads
The model needs room for more than weights. The KV-cache calculation follows the actual attention structure rather than the largest head-count label.
For this fixed full-attention decoder, raw KV bytes per cached token are K = 2 × 80 × 8 × 128 × cacheBytes. The two accounts for keys and values. Grouped-query attention shares 8 KV heads among 64 query heads, so storage uses 8, not 64. With two-byte cache elements, K is 327,680 bytes, or 320 KiB. Changing a cache format is an engine choice; changing a trained model’s KV-head count is not a harmless serving flag.
Let B be active batch size and T the cached context per sequence before this iteration. The cache storage estimate is C = B × T × K / 10^9 GB. Required memory is W + R + C, where W is actual resident weight GB and R is the declared reserve. Headroom within the total modeling budget is 180 - (W + R + C) GB; reserve is subtracted inside that budget, not added to it. T includes prompt and generated history, not only the original prompt. The model uses equal lengths; a real batch requires the sum of its sequence lengths.
At B = 8, T = 4,096, W = 140 GB, R = 18 GB, and two-byte cache elements, C is 10.73741824 GB, required memory is 168.73741824 GB, and headroom within the budget is 11.26258176 GB. This is tensor-accounting arithmetic, not evidence that an engine allocation succeeds. The reserve must cover peak prefill activations, graph pools, cache metadata and padding, and other runtime allocations. A zero-headroom mathematical fit offers no allowance for omitted overhead or the next cache growth.
Read the warm-decode roofline as a conditional bound
A warm-decode bound gives the investigation a reference point. The reader’s assumptions stay beside its result so they cannot become a benchmark by implication.
The reader assumes a dense weight sweep once per batch iteration, a read of the existing raw KV tensor, and an append of one KV entry per sequence. Thus stepTrafficGB = W + C + B × K / 10^9. Its arithmetic estimate is F = B × (2 × 70 × 10^9 + 4 × 80 × 64 × 128 × T) FLOPs: a simplified dense parameter term plus query-key and attention-value work. The attention arithmetic uses query heads; KV storage uses KV heads.
Memory time in milliseconds is stepTrafficGB / bandwidthGBs × 1,000. Compute time is F / (computeTF × 10^12) × 1,000. Taking their maximum gives an optimistic lower bound under the stated resource rates and work counts, with ideal overlap rather than adding both times. B / stepMs × 1,000 is an optimistic aggregate output-token ceiling only when required memory fits the total modeling budget. It is not a throughput measurement, TTFT, per-user inter-token latency, or end-to-end prediction.
The default inputs give approximately 30.148 ms for memory, 1.507 ms for compute, and a conditional ceiling of 265.36 output tokens/s across the batch. These are the calculated planning figures. The simplified work counts omit softmax, normalization, sampling, conversions, extra reads, and scheduling; they are not a rigorous lower bound for every alternative implementation. The model excludes tensor parallelism, MoE, speculative decoding, prefix sharing, allocator padding, launch overhead, and collectives. If required memory exceeds the total modeling budget of 180 decimal GB, this model withholds the token ceiling; that is not proof of physical single-GPU infeasibility.
Doubling the batch can amortize the weight sweep across more outputs, but also increases cache storage and attention work. Lowering W can reduce memory traffic until another term dominates. The compute rate must match the actual arithmetic path; do not double it just because a dropdown says FP8. Use the reader to formulate a bottleneck hypothesis, then test that hypothesis with the engine.
01 / Follow the explanation
What limits one warm decode step on Blackwell?
A brief visual sequence plays automatically. The example and its assumptions are already here—nothing to configure.
An interactive teaching model drawn from real delivery work. No agent, cloud account, GPU or cluster is accessed.
Read the full explanation and assumptionsCurrent example
Fits the budget; throughput remains an optimistic bound
Dense 70B warm decode: 80 full-attention layers, 64 query heads, 8 KV heads and head dimension 128; total modeling budget of 180 decimal GB (180e9 bytes), not a conversion of nominal B200 memory. Weights, reserve and cache all count against this budget; actual device inventory and allocatable memory must be checked separately. Entered bandwidth and compute are scenario inputs for the model. The larger traffic/compute time is a lower bound, not achieved latency or throughput. No tensor parallelism, MoE, prefix sharing, padding, launch cost or collectives.
- Warm decode step lower bound
- 30.148 ms/step
Initial model: one B200 with a total modeling budget of 180 decimal GB (180e9 bytes), a dense 70B decoder, batch 8, 4,096 retained tokens, 140 GB resident weights, 18 GB reserve, two-byte KV storage, 5,000 GB/s bandwidth and 800 TFLOP/s effective math rate.
Decode workload
Memory budget
- KV storage width
- 2 bytes per element
- Budget headroom
- 11.263 GB
Effective engine rates
- Effective math rate (TFLOP/s)
- 800
- Traffic time lower bound
- 30.148 ms/step
- Compute time lower bound
- 1.507 ms/step
# One warm decode step; decimal GB and user-entered effective rates
kv_bytes_per_token = 2 * 80 * 8 * 128 * cacheBytes
cache_GB = batch * context * kv_bytes_per_token / 1e9
required_GB = weightGB + reserveGB + cache_GB
# 180 decimal GB is the total modeling budget, not physical capacity.
headroom_GB = 180 - required_GB
traffic_GB = weightGB + cache_GB + batch * kv_bytes_per_token / 1e9
memory_ms = traffic_GB / bandwidthGBs * 1000
flops = batch * (2 * 70e9 + 4 * 80 * 64 * 128 * context)
compute_ms = flops / (computeTF * 1e12) * 1000
step_ms = max(memory_ms, compute_ms)
tokens_ceiling = batch / step_ms * 1000 if headroom_GB >= 0 else None"""Standalone arithmetic from the capacity plan: no engine, GPU access, or benchmark."""
batch, context, cache_bytes = 8, 4096, 2
weight_gb, reserve_gb = 140.0, 18.0
bandwidth_gbs, compute_tf = 5000.0, 800.0 # planning inputs, not peaks
kv_bytes = 2 * 80 * 8 * 128 * cache_bytes
cache_gb = batch * context * kv_bytes / 1e9
required_gb = weight_gb + reserve_gb + cache_gb
# 180 decimal GB is the total modeling budget, not physical capacity.
headroom_gb = 180.0 - required_gb
traffic_gb = weight_gb + cache_gb + batch * kv_bytes / 1e9
flops = batch * (2 * 70e9 + 4 * 80 * 64 * 128 * context)
memory_ms = traffic_gb / bandwidth_gbs * 1000
compute_ms = flops / (compute_tf * 1e12) * 1000
step_ms = max(memory_ms, compute_ms)
print(f"Required: {required_gb:.6f} GB; budget headroom: {headroom_gb:.6f} GB")
print(f"Resource bounds: {memory_ms:.6f} ms memory, {compute_ms:.6f} ms compute")
if headroom_gb >= 0:
print(f"Conditional optimistic ceiling: {batch / step_ms * 1000:.2f} tokens/s")
else:
print("Required memory exceeds the budget: model token ceiling is not applicable")Paged cache and prefix cache solve different waste
Cache optimization offers several mechanisms. You distinguish wasted allocation from reusable prefixes before deciding what to change.
TensorRT-LLM documents a paged cache as blocks allocated, tracked, and recycled by a cache manager rather than one maximum-sized contiguous tensor per request. Paging improves allocation flexibility; it does not shrink the underlying full-attention KV values. Block tails can be partly unused, metadata consumes space, and allocation policy determines when a request is preempted, recomputed, or rejected.
Prefix caching reuses already computed KV for a matching prompt prefix. vLLM’s documented hash keys include prior-prefix identity, block tokens, and relevant extra identities such as adapters, multimodal inputs, and cache salts. Its design caches full blocks. Similar natural-language meaning is not a cache hit: token IDs and the relevant execution identity must match. Dynamic system messages, differing chat templates, and request-specific early tokens can destroy reuse.
Prefix hits can avoid repeated prefill work and share retained blocks, but they do not eliminate autoregressive generation or the need to attend to prior context. Report warm-prefix and cold-prefix traffic separately, including hit definitions, eviction policy, and cache occupancy. Decide whether tenant isolation requires separate cache identities; a cache key is not a blanket confidentiality guarantee. This reader intentionally counts independent full caches with no prefix-sharing credit.
Sliding-window layers, hybrid attention, latent/compressed KV, KV offload, and quantized-cache metadata need different accounting. The reader’s one-byte option halves raw tensor storage but is not an implementation of every FP8 cache layout. In particular it does not model NVFP4 KV storage. Weight precision and KV precision remain independent settings subject to the selected engine’s support matrix.
Continuous batching couples scheduler policy to kernel shape
Its cost-per-response model now depends on serving the same workload rather than making the benchmark easier.
Continuous batching, also called in-flight or iteration-level batching, admits and removes sequences between engine iterations. The active set can change without waiting for every request in a static batch to finish. TensorRT-LLM 1.1.0 documents packed input for in-flight batching and distinguishes maximum request count, maximum sequence length, and the per-iteration token budget. Those limits are not interchangeable.
A decode request may need one token of new work while a new prompt needs thousands. Chunked prefill spreads that prompt across iterations so it can coexist with decode. Smaller chunks may protect token-gap latency but increase scheduling overhead or delay first-token completion; larger chunks can improve prefill efficiency while stalling ongoing streams. Tune the request and token limits together on the mixed workload, not just on pure decode.
vLLM and TensorRT-LLM expose similar scheduling concepts through different versioned configurations. Record resolved values and the engine backend rather than assuming the same flag name has identical semantics. Bound admission queue length, maximum prompt plus output length, request deadline, and cancellation cleanup. Stop a concurrency sweep when cache churn or waiting time rises without a goodput benefit; accepting more requests is not equivalent to finishing more useful work.
Treat BF16, FP8, and NVFP4 as numerical contracts
A smaller precision format offers another possibility. It must pass the numerical and quality contract before its speed matters.
BF16 uses two bytes per stored element and is a useful reference precision, not a guarantee of exact outputs or universal kernel support. FP8 names a family of formats and scaling recipes, not a single drop-in datatype. Specify the value format, per-tensor/per-channel/per-block scale layout, static calibration versus dynamic activation scaling, accumulator/output types, and which tensors remain at higher precision. A checkpoint’s storage dtype does not prove its GEMMs execute on the intended low-precision path.
NVIDIA describes NVFP4 as E2M1 four-bit values with E4M3 FP8 scales for 16-value micro-blocks and a second-level FP32 tensor scale. This differs from generic INT4 and from MXFP4. Four data bits plus eight scale bits per 16 values already means 4.5 bits per value before global scales, padding, and higher-precision tensors. Dividing a BF16 file size by four therefore does not determine resident NVFP4 memory.
Record quantizer and export versions, calibration corpus and preprocessing, scale tensors, exclusions from quantization, and checkpoint digest. Verify that the exact vLLM or TensorRT-LLM release, CUDA/runtime stack, B200 architecture target, model implementation, and chosen kernel backend support that checkpoint recipe. TensorRT-LLM publishes separate model and hardware support matrices; a supported GPU row alone does not establish support for every model or KV format. Documentation under stable or latest can move, so save the version used in the experiment.
Compare application accuracy, long-context retrieval, numerical stability, structured-output validity, and representative difficult prompts against the BF16 reference using fixed evaluation rules. Quality differences need not be obvious in a short fluent answer. Recheck output-length and stopping distributions: a precision change that changes generated work can invalidate a throughput comparison. Lower storage and faster supported kernels are opportunities, never a guaranteed speedup or guaranteed lossless quantization.
Compile and capture before measuring the warm path
The warm path is not ready to time until compilation and capture have been handled deliberately. Otherwise, different experiments can measure different work.
Compilation and CUDA Graph capture address different costs. Compilation can specialize, fuse, or choose an implementation; a CUDA Graph records an eligible execution sequence for replay with less repeated launch work. vLLM documents separate full, piecewise, and decode-specific graph modes, with dispatch depending on batch shape and attention-backend capability. Requesting a mode does not prove that every mixed batch uses it: unsupported cases can take a different path.
Warm the actual prefill, mixed, and decode shape buckets required by the experiment. Record cold model loading, compilation, graph capture, and warm-up duration separately from steady state. Graph pools and shape specialization also consume memory; reducing launch overhead while crowding out cache can lower serving goodput. A warmed batch size of one is not evidence that a later large mixed batch avoids compilation or eager fallback.
Use startup logs and traces to confirm the resolved compiler and capture behavior. Keep an eager or non-graph baseline for isolating the benefit, holding numerical settings and workload fixed. Do not report a cached second launch as a cold-start result. A server becoming responsive before model initialization, allocation, and bounded warm-up finish is not evidence of a usable warm replica.
Profile the implementation actually selected on B200
The engine’s label is not an execution trace. You inspect which B200 kernels vLLM or TensorRT-LLM actually selects for the workload.
Attention kernels differ across prefill and decode in tiling, reduction strategy, supported head dimensions, cache layout, mask handling, GQA mapping, and precision. A FlashAttention, FlashInfer, Triton, or TensorRT-LLM label is not itself an execution trace. Confirm the selected implementation and fallback paths for the exact B200, engine build, model, and input shape. An implementation optimized for a different architecture can be a poor baseline even when it runs.
Start with a timeline of CPU scheduling, CUDA launches, transfers, kernels, and stream dependencies. Look for idle gaps, synchronizations, repeated layout conversions, allocation, or compilation during the supposedly warm interval. Only then inspect expensive kernels for memory traffic, arithmetic throughput, occupancy limits, register pressure, and shape inefficiency. Occupancy and utilization are clues rather than optimization objectives; fusion can reduce HBM traffic while increasing register pressure.
For an isolated device interval, use timing events on the stream containing the work and wait for completion before consuming elapsed time. Multi-stream or distributed work needs an explicit completion dependency covering every relevant stream or rank. Host end-to-end timing instead includes dispatch and synchronization at its declared boundary. A host timer around an asynchronous launch alone mostly measures enqueueing. Profile a bounded representative slice separately from uninstrumented performance runs because profiling itself can perturb execution.
Conceptual Blackwell execution and memory map
Off-chip DRAM → memory controllers → caches → registers / execution → stores. Instruction issue and asynchronous copies coordinate different paths.
Large off-chip storage
On-chip reuse and staging
Inside one representative streaming multiprocessor
Device-wide work and peers
Stacked DRAM stores large device-resident arrays. DRAM cells need refresh; capacity and sustained bandwidth are different limits. HBM is not the register file.
Read every component and connection
- HBM3e DRAM · Weights / KV / arrays
- Stacked DRAM stores large device-resident arrays. DRAM cells need refresh; capacity and sustained bandwidth are different limits. HBM is not the register file.
- Memory controllers · Channels and requests
- Controllers organize reads and writes to memory channels. Access patterns, contention and the memory technology influence service time; bandwidth is not zero latency.
- L2 cache · Device-wide reuse
- L2 can satisfy repeated requests without another HBM access. Its capacity and residency behavior affect traffic; a cache hit is not a new DRAM transfer.
- L1 / shared memory · Caching / explicit tiles
- B200 combines L1, texture and shared-memory resources. Shared memory is software-managed block storage with synchronization rules; it is not an automatic replacement for registers.
- Register file · Thread operands
- The B200 tuning guide specifies 64K 32-bit registers per SM. Threads use registers for live values; spills can create device-memory traffic. Registers are not off-chip DRAM.
- ALU pipelines · Integer / floating point
- Execution pipelines perform supported arithmetic and logic on operands. CMOS gates underlie those circuits. Floating-point operations include more work than the integer full adder shown.
- Tensor cores · Matrix operations
- Specialized matrix instructions use supported operand formats and accumulation paths. Tensor throughput is not scalar ALU throughput, and not every kernel can use tensor cores.
- Load / store units · Addresses and movement
- Load/store machinery forms and services memory operations. Coalescing groups useful lane accesses; dependencies prevent a consumer from using a value before it is ready.
- Warp schedulers · Ready instruction issue
- Schedulers issue eligible warp instructions subject to dependencies and resource availability. Other ready warps can hide a wait; occupancy alone does not prove throughput.
- Instruction path · Fetch / decode / issue
- Compiled machine instructions reach the SM instruction machinery. PTX is a virtual ISA; a compatible cubin or driver compilation supplies hardware-executable code.
- Async copy / TMA · Tile movement
- Supported asynchronous transfer paths can stage tiles while computation proceeds. Barriers and producer/consumer ordering still apply; overlap is not permission to read unfinished data.
- GPU front end · Submitted work
- Device work submission and scheduling machinery distribute kernel work. Block resource requirements influence residency. Kubernetes does not choose a warp or allocate an SM register.
- NVLink interface · Peer devices
- Peer access and collectives move data between compatible GPUs. The application/runtime manages distributed work; aggregate device memory is not one automatically shared allocation.
- GPU front end → Instruction path: kernel work
- Instruction path → Warp schedulers: decoded instructions
- Warp schedulers → Load / store units: memory instruction
- HBM3e DRAM → Memory controllers: DRAM service
- Memory controllers → L2 cache: cache-line traffic
- L2 cache → L1 / shared memory: cache / tile path
- L1 / shared memory → Load / store units: load path
- Load / store units → Register file: operand load
- Register file → ALU pipelines: ALU operands
- ALU pipelines → Register file: result
- Register file → Load / store units: store
- L2 cache → Async copy / TMA: async tile copy
- Async copy / TMA → L1 / shared memory: staging
- L1 / shared memory → Tensor cores: supported matrix operands
- L2 cache → NVLink interface: peer traffic
This is a functional map, not a floorplan or cycle-accurate simulator. One representative SM is expanded; it is not the GPU’s SM count. Cache bypass, asynchronous copies, distributed shared memory and specialized tensor accumulator paths mean not every operation follows every arrow.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
Dense, MoE, and tensor parallelism need different models
The model architecture changes the resource story. Dense, MoE and parallel configurations need their own explicit assumptions.
The companion bound assumes dense weights with the declared 70B parameter term. A mixture-of-experts model activates only selected experts for each token, but its resident weight requirement can include many inactive experts. Routing distribution, expert imbalance, grouped GEMM shapes, dispatch/combine work, and expert-parallel all-to-all traffic affect execution. Neither total parameter count nor active parameter count alone can replace the dense formula.
Tensor parallelism shards operations across GPUs and introduces communication on the inference path. Some layers require collectives between partial results; a small low-batch matrix operation may complete quickly enough that collective latency dominates. KV-head placement can shard or replicate cache depending on head counts and implementation, so dividing every memory term by the tensor-parallel degree is not valid.
A model that fits on one B200 may benefit from independent replicas for aggregate traffic without paying per-step model-parallel communication, but that does not make replication the right answer for every latency or capacity target. If testing an HGX group, record GPU count, placement, NVLink/NVSwitch connectivity, host NUMA alignment, and any cross-host NIC path. Benchmark collectives at the message sizes and concurrency of the model rather than substituting an aggregate fabric headline.
Weight or KV offload and prefill/decode disaggregation add transfer and lifetime boundaries. Include host/device or inter-node bytes, overlap constraints, and failure handling in their own experiment. None of those alternatives is represented by the single-device reader, and a configuration that exceeds its total modeling budget cannot be rescued by interpreting its arithmetic token ceiling as an offloaded result.
A bounded timing example makes the acceptance rule executable
The acceptance rule becomes executable. The bounded timing example keeps synchronization and the timing boundary in view.
The complete standard-library Python example below reduces a small recorded trace; it does not contact an engine or measure a GPU. The declared observation window is one second. Each successful row contains exactly one client receipt timestamp per output token on the same monotonic clock as the send timestamp. A real chunked API does not automatically satisfy that assumption. The latency limits are policy inputs, not recommended B200 targets.
A response passes only if it completes inside the window, meets a 250 ms TTFT limit, and has no individual token gap greater than 50 ms. A one-token output passes the gap condition vacuously but has no TPOT value. The failed row remains in the offered count. The second response in the trace has a 60 ms gap and fails even though its average gap is 40 ms. The resulting goodput is therefore two requests/s and four tokens/s, not all emitted tokens divided by the window.
For a real run, record the engine/model/tokenizer digests, resolved scheduling and precision configuration, exact measurement boundaries, and request-to-output accounting alongside the trace. Define how responses crossing the cutoff are drained or classified; this small example rejects them from goodput. It demonstrates a timing contract completely without pretending to be a runnable deployment.
"""Recorded client token timestamps; no engine or hardware is measured."""
config = {"start_s": 0.0, "end_s": 1.0, "ttft_ms": 250.0, "max_gap_ms": 50.0}
rows = [
{"sent": 0.00, "tokens": [0.10, 0.13, 0.16], "done": 0.17, "ok": True},
{"sent": 0.10, "tokens": [0.20, 0.22, 0.28], "done": 0.29, "ok": True},
{"sent": 0.20, "tokens": [0.30], "done": 0.31, "ok": True},
{"sent": 0.40, "tokens": [], "done": 0.45, "ok": False},
]
window_s = config["end_s"] - config["start_s"]
assert window_s > 0
passed_requests = passed_tokens = 0
for index, row in enumerate(rows):
times = row["tokens"]
assert config["start_s"] <= row["sent"] < config["end_s"]
assert row["done"] >= row["sent"]
assert times == sorted(times)
assert all(config["start_s"] <= t < config["end_s"] for t in times)
complete = row["ok"] and row["done"] < config["end_s"]
ttft_ok = bool(times) and (times[0] - row["sent"]) * 1000 <= config["ttft_ms"]
gaps = [round((b - a) * 1000) for a, b in zip([row["sent"], *times], [*times, row["done"]])][1:]
gap_ok = all(g <= config["max_gap_ms"] for g in gaps)
if complete and ttft_ok and gap_ok:
passed_requests += 1
passed_tokens += len(times)
print(f"Goodput: {passed_requests / window_s:.2f} requests/s, {passed_tokens / window_s:.2f} tokens/s")Compare engines with one experiment contract and explicit stops
The final comparison must support the fixed response volume and quality contract before the compute difference became a real proposal.
Compare vLLM and TensorRT-LLM using the same checkpoint semantics, tokenizer, prompt corpus, requested output limits, sampling rules, prefix-cache state, arrival schedule, B200 allocation, and quality acceptance rule. Save engine/backend versions, container and model digests, CUDA and driver versions, quantization recipe, graph settings, scheduling limits, topology, and device power/sharing state. Do not assume a TensorRT-LLM PyTorch backend and a built TensorRT engine have identical setup or feature behavior.
First establish a correct supported baseline, then change one factor: batch/token budget, prefill chunking, graph mode, attention backend, or precision. Use both isolated phase experiments and mixed-service confirmation. An open-loop offered-load schedule exposes queue accumulation; a closed-loop fixed-concurrency test backs off automatically when responses slow. Report which was used, verify the load generator is not limiting the run, and retain the same workload distribution across candidates.
Repeat bounded warm runs and report variation, sample counts, offered/completed/pending totals, client TTFT and token-gap distributions, request/token goodput, cache occupancy, preemptions, memory peaks, and output-quality results. Keep startup and warm-up costs separate. A brief run cannot support a stable extreme percentile simply because a reporting tool prints one. Compare performance at shared quality and service objectives rather than ranking maximum unconstrained token counts.
Choose stop conditions before the sweep: invalid outputs or quality regression, allocation failure, wrong precision/backend fallback, persistent preemption, unbounded queue age, rising timeout/rejection rates, or violated latency objectives. Stop increasing a setting when repeated runs show no meaningful goodput improvement relative to variation. Restore the last recorded valid configuration rather than stacking additional flags on an unexplained regression. This is the proposed experiment procedure; no engine comparison, failure injection, or hardware timing is claimed here.
Questions behind the decision
Is vLLM or TensorRT-LLM faster on B200?
The answer depends on the model, format, workload mix, cache policy, selected kernels and software versions. Compare accepted work under the same quality and latency limits instead of transferring another benchmark’s ranking.
Can prefix caching lower time to first token?
Reusable prefixes can avoid some repeated prefill work when the cache is valid and available. They do not eliminate queueing, uncached input processing or decode costs, and low reuse can change the economic result.
References & further reading
- NVIDIA HGX AI Factory: B200 SXM per-GPU and per-node memory specifications
- Vultr: verified B200 instance inventory reported by nvidia-smi
- NVIDIA: NVFP4 E2M1 values, micro-block scales, and tensor scaling
- vLLM: CUDA Graph modes, warm-up, dispatch, and attention-backend compatibility
- vLLM: automatic prefix caching, block identity, and cache salts
- TensorRT-LLM 1.1.0: paged attention, in-flight batching, and request scheduling
- TensorRT-LLM: quantization recipes and model/hardware support matrices


