LLM Inference Benchmarking: TTFT, Goodput and Cost per Response
A real client engagement. The engineering and the results are described below.
An AI research product we worked with gets a spectacular demo from a hosting candidate just as its annual capacity decision approaches. The engineering team replays its real request distribution before trusting the headline.
Request a time through the inquiry form. A meeting is confirmed separately by email.
The business problem behind the technology
What does the fast demo leave out of the business case?
The product must include long prompts, realistic output lengths, queueing and failures. A candidate that rejects difficult requests cannot earn a lower cost denominator.
Read the client engagement ↓Client engagement / Delivered results
The capacity contract depended on the slow requests
An AI research product we worked with plans twelve million accepted responses each month. Its procurement shortlist looks attractive when measured with short prompts and a warm prefix cache.
The constraint
The product must include long prompts, realistic output lengths, queueing and failures. A candidate that rejects difficult requests cannot earn a lower cost denominator.
The engineering decision
The team fixes the workload and acceptance criteria, records cache state and replays controlled offered load. The engagement compared $96,000 and $78,000 monthly compute charges for the same twelve million accepted responses.
The delivered outcome
Compute cost moved from $0.008 to $0.0065 per accepted response, a monthly difference of $18,000 for this client’s request mix.
| Measure / unit | Before | After | Difference |
|---|---|---|---|
| Compute charges at fixed accepted volume USD/month | 96,000 | 78,000 | 18,000 |
Compute charges at fixed accepted volume. Divide either monthly charge by twelve million accepted responses. Output-token counts are additionally required to report cost per token.
The conditions behind the results
- Both cases deliver the same twelve million accepted responses at the same quality and latency thresholds.
- Monthly charges are specific to this model and request mix and do not imply the same performance on a different one.
- Cold starts, cache state, rejected work and offered load must be disclosed in a real comparison.
What this does not prove. This cost per accepted response is specific to the client’s request mix and hosting contract.
Evidence to collect for your own decision
- Capture prompt/output length distributions, arrival rates and quality acceptance.
- Report TTFT, inter-token and complete-response latency with failures and goodput.
- Reconcile cache conditions, exact configuration and a reproducible cost denominator.
Key decisions
Compare LLM serving with a defined workload, quality threshold and latency boundary. Keep queueing, cache state, rejected work and useful throughput visible before choosing a hosting configuration.
Follow the decision
Select a step to follow its reasoning, then continue into the technical chapters.
Problem → boundary → decision → evidence
Read every component and connection
- Ask with a model · Boundary
- Use physical bounds to guide the experiment, not replace it.
- Define the result · Decision
- Keep latency, throughput and useful completion unambiguous.
- Replay and report · Evidence
- Control cache state and preserve a reproducible evidence record.
- Bound memory → Ask with a model: identify the constraint
- Ask with a model → Define the result: choose a bounded change
- Define the result → Replay and report: check the outcome
A conceptual decision map for this article, not a measured timeline, physical topology or a depiction of a specific client system.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
Estimate weights and KV cache separately
The client’s AI product accounts for the full memory obligation before comparing capacity contracts.
Start with a memory inventory: model weights, key/value cache, activations, runtime workspaces, communication buffers, and allocator overhead. Weight storage is approximately parameter count times stored bytes per parameter, with quantization metadata and implementation details added. Weight quantization does not automatically quantize the KV cache. Record both formats rather than labeling a server simply four-bit.
For a conventional decoder with full attention, each retained token stores a key and a value for every layer. If L is layer count, Hkv is the number of KV heads, D is head dimension, b is bytes per element, and Ti is the retained token count of sequence i, logical KV bytes equal 2 × L × Hkv × D × b × sum(Ti). The factor two counts keys and values. Grouped-query attention shares KV heads across query groups, so substituting the larger query-head count overestimates the cache.
The configuration below uses 32 layers, 8 KV heads, head dimension 128, two-byte elements, and one 8,192-token sequence. It needs 1 GiB of logical KV storage. This excludes block rounding, metadata, and other memory. Sliding-window, hybrid, compressed-cache, and latent-attention architectures require their own calculation. Tensor parallelism may shard or replicate cache components; do not blindly divide total memory by device count.
# Full-attention configuration from the capacity plan.
layers, kv_heads, head_dim = 32, 8, 128
bytes_per_element, retained_tokens = 2, 8192
kv_bytes = (2 * layers * kv_heads * head_dim
* bytes_per_element * retained_tokens)
print(kv_bytes / 2**30) # 1.0 GiB of logical KV storageUse theoretical bounds to ask better questions
A theoretical bound helps identify an implausible claim. It remains a question for measurement, not the final performance result.
Prefill processes prompt tokens and often exposes substantial parallel compute. Decode generates successive tokens and can be dominated by memory traffic at small batch sizes. These are tendencies, not universal labels: prompt length, batching, attention kernels, model shape, and interconnect communication can change the bottleneck.
A simple lower bound on execution time is the larger of required operations divided by peak compute and required memory traffic divided by peak bandwidth. It assumes ideal utilization and overlap and excludes scheduling, networking, synchronization, and software overhead. It is a reasoning tool, not a predicted latency. Be explicit about operation-count conventions and which memory boundary the traffic crosses.
If measured behavior is far from that bound, profile before buying more capacity. Small kernels, CPU tokenization, a saturated network link, or queueing can dominate the end-to-end experience. Increasing batch size may improve aggregate efficiency while slowing individual requests. A serving configuration is a tradeoff across memory, concurrency, latency, and output quality, not one maximum tokens-per-second number.
Define latency and throughput without ambiguity
The team chooses accepted-work metrics so its proposed cost difference cannot be improved by rejecting difficult requests.
Time to first token, or TTFT, should state its clock boundaries. Client-observed TTFT includes transport, queueing, prompt processing, and the first streamed token. Server-side prefill time does not include all those costs. Record both when available. For more than one output token, average time per output token is TPOT = (last-token time - first-token time) / (N - 1), where N is the output token count.
That TPOT average can hide uneven streaming. Measure inter-token gaps and tail percentiles as well as total response latency. A one-token response has no post-first-token interval and should not contribute a fabricated zero TPOT. Token counts must use the model tokenizer, and stream chunks do not necessarily correspond one-to-one with tokens.
Distinguish completed requests per second, per-request generation speed, aggregate output tokens per second, and combined input/output token throughput. Define the observation window and treatment of incomplete requests. Report goodput as work meeting the stated latency and quality criteria, with failures and cancellations visible. A server that accepts an ever-growing queue has not demonstrated sustainable throughput.
Control prefix caching and warmup explicitly
A warm cache can change the apparent story. You make prefix caching and warmup part of the experiment rather than an invisible advantage.
Prefix caching reuses compatible KV state for shared prompt prefixes. It reduces repeated prefill computation; it does not inherently accelerate generation of new tokens during decode. Replaying an identical prompt can therefore produce an unusually favorable TTFT distribution that does not represent independent document questions.
Run clearly labeled cache-disabled, cold-cache, and representative warm-cache workloads. Record the shared-prefix distribution and cache-hit evidence rather than assuming every repeated document hits. Prefix identity depends on actual tokenized content and serving configuration, including relevant model or adapter state. Cache isolation also matters across tenants because reuse behavior can create a security boundary.
Separate process startup, model loading, kernel compilation or graph capture, and request-cache warmup. Warm a configuration consistently before steady-state measurement, then measure cold start separately when it matters to autoscaling. Preserve those settings in the benchmark record. A restart that leaves an external cache intact may not produce the cold condition the experiment intended.
Replay realistic load and publish reproducible evidence
A reproducible replay, not the demo, must support the model’s cost per accepted response.
Build a workload matrix spanning prompt and output lengths, concurrency, arrival rate, cache reuse, and task type. A closed-loop client waits before sending its next request and can hide overload by reducing offered traffic as latency rises. An open-loop arrival schedule better exposes queue growth, provided the load generator can maintain the schedule. Report offered and achieved rates in either case.
Run repeated trials long enough to reveal steady-state behavior and burst recovery. Preserve latency distributions, errors, queue depth, memory pressure, and actual output lengths. Keep model revision, tokenizer, runtime version, precision, parallelism, sampling settings, and hardware inventory fixed or explicitly varied. Quantization and speculative decoding also require quality checks; faster but unacceptable answers are not a serving improvement.
Choose the configuration that sustains the required workload within latency, quality, and cost limits, with headroom for realistic variation. Report uncertainty and bottleneck evidence beside every result. The memory calculation can rule out an impossible configuration; only an executed, controlled benchmark can establish performance. This article gives the method and the inputs we applied.
Questions behind the decision
What is goodput in an LLM inference benchmark?
Goodput counts work that meets the declared usefulness criteria, such as quality and latency, rather than every generated token. State the acceptance rule and keep rejected or late work visible in the report.
Why report time to first token separately from total latency?
TTFT describes how long the user waits for generation to begin, including the chosen queueing and prefill boundary. Long outputs can still take substantial time after the first token, so complete-response and inter-token behavior answer different questions.
