Hover a dotted term for 5 seconds to lock its explanation. It closes after 5 seconds away; nested tooltips and keyboard focus keep it open. Click, tap or Enter locks immediately. Technical glossary.

AI infrastructure

LLM Inference on Kubernetes: GPU Scheduling, Autoscaling and Cost

A real client engagement. The engineering and the results are described below.

A support-software company we worked with wants to offer an AI assistant without letting idle model replicas consume the launch budget. Its team traces the serving queue through Kubernetes placement and model readiness before changing capacity.

Request a time through the inquiry form. A meeting is confirmed separately by email.

The business problem behind the technology

Can the service keep its promises with fewer paid GPU-hours?

The team cannot count a Pending Pod or a started process as ready serving capacity. It needs peak headroom and recovery without paying for the same warm fleet at every hour.

Read the client engagement ↓

Client engagement / Delivered results

The idle replica was cheap insurance—or expensive habit

A support platform we worked with plans eight million accepted assistant responses each month. Demand is uneven, but customers still expect a predictable first token during business hours.

The constraint

The team cannot count a Pending Pod or a started process as ready serving capacity. It needs peak headroom and recovery without paying for the same warm fleet at every hour.

The engineering decision

It measures offered load, model warmup and queue age, keeps a justified warm floor, and separates autoscaling recommendations from usable device placement. Scheduled capacity fell from 20,000 to 14,000 GPU-hours at the client’s $5 rate.

The delivered outcome

Monthly GPU compute charges moved from $100,000 to $70,000 for the fixed workload while preserving the service objective during peaks and failures.

Measured inputs and delivered differences
Measure / unitBeforeAfterDifference
GPU compute charges
USD/month
100,00070,00030,000

GPU compute charges. 20,000 × $5 versus 14,000 × $5 at the client’s rate for this workload.

The conditions behind the results

  • The same eight million accepted responses, prompt/output mix and quality threshold apply.
  • Peak and recovery reserve remain sufficient; the design does not scale all production replicas to zero.
  • The separate KV-cache reader explains a boundary and does not derive the 6,000-hour reduction.
  • Storage, networking, staff, commitments and migration costs are excluded.

What this does not prove. Fewer desired replicas are not invoice savings; savings require billable capacity to retire without losing useful service.

Evidence to collect for your own decision

Key decisions

An allocated GPU is not a serving-capacity guarantee. Follow a request from queue to prefill to decode, calculate an explicit KV-cache budget, and measure completed work against latency objectives before changing the deployment.

  • Replay arrival bursts with queueing, TTFT, output latency and rejection recorded.
  • Measure scheduling, image/artifact loading and model-ready time on exact GPU groups.
  • Fewer desired replicas are not invoice savings; savings require billable capacity to retire without losing useful service.

Follow the decision

Can the service keep its promises with fewer paid GPU-hours?

Select a step to follow its reasoning, then continue into the technical chapters.

Problem → boundary → decision → evidence

Follow demand

Separate queueing, prefill and decode at the user boundary.

Read every component and connection
Follow demand · Problem
Separate queueing, prefill and decode at the user boundary.
Bound admission · Boundary
Account for model and KV-cache capacity before accepting work.
Reach readiness · Decision
Allocation, placement and a loaded healthy model are different states.
Prove the service · Evidence
Test useful throughput, failure and reversible delivery.

A conceptual decision map for this article, not a measured timeline, physical topology or a depiction of a specific client system.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

A request has two different accelerator workloads

The client’s support platform defines a complete model replica before assigning cost to capacity.

Prefill processes the prompt, constructs the attention key/value cache, and prepares the first generated token. Many prompt positions can be processed together, so substantial matrix work can expose useful parallelism. Decode then advances each active sequence through successive generation steps, reading prior KV state and appending new state. A short prompt with a long answer and a long prompt with a short answer can consume very different resources even when their total token counts match.

For a conventional autoregressive decoder, each new generation step depends on prior output. Low-batch decode often places pressure on weight and cache memory traffic; large prefills may expose more arithmetic work. These are workload tendencies, not laws that identify the bottleneck from the phase name alone. Batch shape, attention implementation, precision, context length, speculative decoding, and communication all change the balance. Inspect a representative execution trace rather than declaring every prefill compute-bound and every decode bandwidth-bound.

Define latency at the boundary the user experiences

The user does not experience an isolated kernel. You define latency at the boundary where the answer is actually received.

Define client time to first token, or TTFT, as elapsed time from the chosen request-send boundary to receipt of the first output token. It includes transport, frontend processing, admission waiting, prefill, and first-token delivery. Server-side TTFT has a different start point and must be labeled separately. A growing admission queue can ruin client TTFT while an individual prefill kernel remains fast.

For a response with N output tokens and N greater than one, a useful per-request time-per-output-token definition is TPOT = (time of last token - time of first token) / (N - 1), in seconds per output token. Inter-token latency instead describes individual gaps. A per-request average can hide a single long stall, and a streamed chunk may contain multiple tokens; record tokenization and chunk-timestamp conventions. A one-token response has TTFT but no post-first-token TPOT under this definition.

Track client TTFT, per-request TPOT, individual gap distributions, end-to-end completion latency, failures, and queue residence by workload class. The vLLM metrics documentation describes server-side TTFT, inter-token latency, prefill/decode duration, queue duration, waiting requests, and cache occupancy. Metric names and observation boundaries depend on the pinned engine version. Do not add unrelated p99 stage values and call the sum end-to-end p99; retain per-request observations or measure the end-to-end distribution directly.

Useful tokens beat a utilization-only objective

The GPU utilization chart is only supporting evidence. Rejected or late work cannot improve the useful-token objective.

A busy GPU can be recomputing evicted state, serving outputs after their deadlines, or processing work that the client already canceled. Device utilization is a diagnostic signal about sampled activity, not a count of useful answers. High memory allocation may also reflect an engine-reserved cache pool rather than active token occupancy. Separate physical allocation from occupied KV blocks and from requests that actually finish.

For this exercise, define useful output-token throughput as output tokens in successfully completed responses that satisfy the chosen TTFT and TPOT limits, divided by the full observation interval. Also report successful requests per second and application-quality checks: verbosity should not be rewarded as efficiency. Publish all failures, timeouts, cancellations, rejected requests, and requests still pending at the cutoff. None should quietly disappear from the offered-load denominator.

Compare cost per useful output token only within comparable model quality, tokenizer, prompt/output distributions, and service objectives. Include host CPUs, accelerator reservation time, networking, storage, idle warm replicas, and failed work in the cost boundary. A lower allocation is not necessarily a lower bill, and a cheaper token that fails the application task is not an improvement.

Continuous batching is an admission policy, not free capacity

The serving queue becomes an admission decision. Continuous batching must fit both resource capacity and the latency contract.

With continuous batching, the engine can update the active request set between scheduling iterations rather than waiting for the longest sequence in a fixed batch to finish. Completed sequences release their slots and new work can join. This reduces some padding and idle-slot costs, but it does not remove the device-memory budget or make every batch shape equally fast. vLLM documents continuous batching and chunked prefill as distinct serving capabilities.

Prefill competes with active decode work. Large new prompts admitted without a scheduling budget can lengthen token gaps for existing streams. Chunking prefill spreads prompt work across iterations, trading scheduling overhead and possible first-token delay against decode responsiveness. Tune token budgets and sequence limits together with the traffic mix; measure the effect instead of copying a maximum batch size from another model.

Bound both queued work and active work. Define maximum prompt length, maximum generated length, request deadlines, and cancellation behavior before raising concurrency. If admission repeatedly exceeds cache capacity, preemption, recomputation, or offload can turn extra accepted requests into worse goodput. A queue-based scaling signal is useful only if new replicas become ready quickly enough and upstream admission still prevents an unbounded backlog.

Derive the KV budget using KV heads, not query heads

The cache calculation follows KV heads and the declared model shape. A larger query-head count does not automatically mean the same cache layout.

For the stated unsharded full-attention model, KV bytes per cached token = 2 × L × Hkv × D × b. The factor two stores keys and values; L is the number of attention layers; Hkv is the number of key/value heads; D is the per-head dimension; and b is bytes per cached element. Grouped-query attention, or GQA, shares KV heads among more query heads. Use Hkv rather than the query-head count. The formula describes tensor storage, not an allocator footprint.

With L = 32, Hkv = 8, D = 128, and b = 2, the result is 131,072 bytes per token, or 128 KiB. A sequence holding 8,192 total cached tokens uses 1 GiB in this idealized calculation. Here KiB means 2^10 bytes and GiB means 2^30 bytes. Total cached tokens include the prompt and generated history, not just the answer. Holding the other inputs fixed, 32 KV heads would require four times as much cache; this is an architectural comparison, not permission to change a trained model by editing a serving flag.

Subtract weights and the declared reserve before dividing: available KV GiB = max(0, device GiB - weights GiB - reserve GiB). Under the reader defaults, 80 - 16 - 8 = 56 GiB, so the idealized memory-only ceiling is floor(56 / 1) = 56 simultaneous 8,192-token sequences. Eight such sequences consume 8 GiB of KV and 32 GiB including the weights and reserve. Neither number establishes a latency-safe batch size or a throughput result.

text / example
kv_bytes_per_token = 2 * layers * kv_heads * head_dim * bytes_per_element
kv_gib_per_sequence = kv_bytes_per_token * cached_tokens / 2^30
available_kv_gib = max(0, device_gib - weights_gib - reserve_gib)
memory_only_sequence_ceiling = floor(available_kv_gib / kv_gib_per_sequence)

Treat the capacity reader as a bound with exclusions

You inspect the capacity reader with its exclusions visible. Its result is a conditional bound, not an allocation promise.

The companion reader deliberately holds layer count and head dimension fixed so the effect of KV-head count, cache precision, context length, and concurrency stays visible. Doubling cached length doubles KV storage in this model. Changing the element size from two bytes to one halves raw tensor storage, but real one-byte cache formats need engine and hardware support, scaling metadata, and quality validation. Weight quantization and KV-cache quantization are separate decisions.

Reserve memory for peak prefill activations, temporary workspaces, captured execution graphs, communication buffers, allocator rounding, and runtime state. Paged-cache blocks can have partially used tails; prefix caching can share some blocks but depends on actual reuse. Sliding-window, hybrid-attention, compressed-cache, and latent-attention models need different accounting. The reader excludes those effects and makes no estimate of fragmentation, prefix hits, or throughput.

Do not divide this result blindly by a tensor-parallel degree. Some implementations shard KV heads while others replicate them for particular head counts or parallel layouts. Inspect the selected model and engine placement. If weights plus reserve exceed device memory, the clamped available-cache value is zero; even an empty KV cache does not make that configuration deployable. A positive cache budget is only one necessary condition for serving.

01 / Follow the explanation

How many sequences fit before the KV cache runs out?

A brief visual sequence plays automatically. The example and its assumptions are already here—nothing to configure.

An interactive teaching model drawn from real delivery work. No agent, cloud account, GPU or cluster is accessed.

Read the full explanation and assumptions

Memory budget stages: weights allocation, non-KV reserve, remaining KV cache, sequence admission.A memory-accounting diagram, not physical buffer addresses. Follow each allocation before choosing a concurrency cap.01Weights02Reserve03KV cache04Admission
A memory-accounting diagram, not physical buffer addresses. Follow each allocation before choosing a concurrency cap.

Current example

The requested sequences fit this memory budget

This uniform decoder-only attention model budgets weights, reserve and both key/value tensors at the full selected context. It models memory, not serving throughput or latency. GiB means 2^30 bytes. Tensor parallelism, prefix sharing, sliding windows, paging fragmentation and allocator behavior are omitted. A zero cache budget is an exhausted state; negative headroom shows the shortfall.

Memory-only capacity (sequences)
56 sequences

Example: 80 GiB total, 16 GiB weights, 8 GiB non-KV reserve; 32 layers, 8 KV heads, 128 dimensions/head, 2 bytes/element, 8,192 retained tokens per sequence and 8 sequences. These are model inputs, not a named accelerator or deployed model.

Memory envelope

Total device memory (GiB)
80
Weights allocation (GiB)
16
Non-KV reserve (GiB)
8
Available KV budget (GiB)
56.000 GiB

Sequence demand

Retained tokens per sequence
8192
Concurrent sequences
8
KV storage per sequence (GiB)
1.000 GiB
Total memory required (GiB)
32.000 GiB
Remaining headroom (GiB)
48.000 GiB

Cache representation

KV heads per layer
8 KV heads
KV storage width
2 bytes per element
KV storage per token (bytes)
131072 bytes

GPU resources follow an extended-resource contract

Readiness becomes the bridge between paid devices and accepted assistant responses.

In the traditional device-plugin path covered here, node drivers and the vendor device plugin must already expose a resource such as nvidia.com/gpu. Kubernetes schedules that declared resource; it does not infer free model capacity from an utilization chart. The Kubernetes GPU scheduling documentation permits a GPU limit without an explicit request, in which case the request defaults to the limit. If both are present they must be equal; a GPU request without a limit is not the supported pattern.

These device quantities are integer allocation units, not fractional CPU-style shares. The meaning of one unit still depends on the advertised device-plugin configuration: it may represent exclusive device access, a MIG profile, or shared time-sliced access. Check resource names, discovered accelerator labels, and node-sharing policy. A container memory limit governs host memory, not an independently enforced HBM quota. CPU and host-memory sizing still matter for tokenization, model loading, and the serving frontend.

Software ownership, from cluster to silicon

A Kubernetes request does not schedule an ALU

Kubernetes places the workload. The container stack exposes a device. CUDA and the driver submit work. GPU scheduling and execution happen below that boundary.

Cluster lifecycle and placement

Application and user-space libraries

Kernel and device boundary

Observe the real service

Kubernetes scheduler

The scheduler chooses a suitable node using declared resources and placement policy. It does not place individual CUDA blocks, schedule warps or make an unready model useful.

Read every component and connection
Kubernetes scheduler · Node / resource placement
The scheduler chooses a suitable node using declared resources and placement policy. It does not place individual CUDA blocks, schedule warps or make an unready model useful.
Device plugin / kubelet · GPU allocation
A device plugin advertises supported GPU resources and participates in allocation. Kubelet starts the pod with the chosen devices. GPU sharing changes the resource contract.
GPU Operator · Lifecycle controller
The operator can manage drivers, Container Toolkit, device plugin, feature discovery and monitoring. This is a control-plane responsibility, not a mandatory hop for each tensor.
OCI runtime / Toolkit · Container device access
The configured runtime and NVIDIA Container Toolkit expose the permitted devices and libraries. Containers share their host kernel; a VM or bare-metal host sets another boundary.
Serving framework · Batching / model / NCCL
A serving engine schedules requests and calls framework/library kernels. NCCL coordinates supported multi-GPU collectives. Application queueing is distinct from GPU instruction scheduling.
CUDA runtime / driver · libcudart / libcuda
CUDA APIs manage contexts, streams, memory and launches. The user-mode driver interfaces with kernel support. Compatible machine code or supported PTX compilation is required.
Linux NVIDIA modules · Devices / memory / I/O
Kernel modules cooperate with the OS for device access, memory mapping and control. Linux permissions and isolation still matter; a container image does not replace the host kernel driver.
Queues / GPU firmware · Device work submission
Supported driver and firmware paths submit and manage device work. Firmware-mediated responsibilities vary with platform and driver; this is not a proprietary command-protocol schematic.
SMs / memory / ALUs · Execute and move data
GPU execution resources run the selected instructions and move operands through the appropriate memory paths. The gate, register and memory diagrams expand this layer.
DCGM + application telemetry · Health ≠ goodput
DCGM-based monitoring supplies device signals. Application metrics and traces must still prove latency, errors, queue age and useful throughput; a busy GPU is not the service objective.

A typical Linux/device-plugin deployment, not a universal platform configuration. GPU Operator manages components; it is not on every inference request’s data path. Driver, CUDA, framework and hardware compatibility must be validated together.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

MIG and time-slicing solve different isolation problems

Sharing the device introduces another decision. MIG and time slicing do not provide the same isolation or memory behavior.

NVIDIA documents MIG as hardware partitioning into predefined GPU instances with memory and fault isolation at that layer. A chosen profile exposes only its portion of memory and compute, and available profiles depend on the device and configuration. Capacity must fit inside the assigned instance. This is not immunity to a host failure, device-wide maintenance, or every possible driver and platform fault.

Time-slicing instead oversubscribes a device and interleaves work from sharing processes. NVIDIA explicitly distinguishes it from MIG: time-sliced replicas do not provide memory or fault isolation. Requesting two advertised replicas does not guarantee twice the compute share, and an integer resource limit is not a proportional-performance contract. A noisy neighbor can alter latency even when each pod appears correctly scheduled.

Use the isolation and predictability required by the workload to choose a deployment policy, then test the allowed colocations under equivalent load. Document shared versus exclusive access in benchmark metadata. Neither packing more pods onto a GPU nor selecting a smaller MIG slice is evidence of savings until useful throughput, tail latency, and recoverability remain acceptable.

A started process is not a model-ready replica

A process starts, but the model still has work to do before serving. Readiness needs to represent that real state.

Startup, readiness, and liveness probes have different jobs. A startup probe delays readiness and liveness checks until startup succeeds, allowing bounded initialization time. Readiness determines whether normal Service routing should send new traffic; liveness addresses failures for which restarting the process is appropriate. A listening socket or a model-name listing does not universally prove that weights, tokenizer, execution graphs, and a usable worker are ready.

The fragment below belongs inside an already defined container; it is not a deployable manifest. The application implements /readyz on port 8000 as a cheap check of engine state that becomes successful only after verified model loading and a bounded warm-up, and /livez as a process-health check independent of queue saturation. These are declared application endpoint contracts, not promises about vLLM endpoints. Verify equivalent semantics for the exact engine version before adapting the fragment.

The example permits roughly 600 seconds of unsuccessful startup checks, polls readiness every five seconds, and uses the host resources from this deployment. Choose the startup allowance from artifact fetch and cold initialization observations, not from warm inference latency. Do not execute an expensive generation on every probe or use transient overload as a liveness failure: restart storms discard warm state and reduce capacity when it is most needed. Readiness also needs an explicit draining state for shutdown; it does not by itself finish existing streams.

yaml / example
# Container fields only; application endpoint contracts described above.
resources:
  requests:
    cpu: "4"
    memory: "24Gi"
  limits:
    memory: "32Gi"
    nvidia.com/gpu: 1
startupProbe:
  httpGet:
    path: /readyz
    port: 8000
  periodSeconds: 10
  timeoutSeconds: 2
  failureThreshold: 60
readinessProbe:
  httpGet:
    path: /readyz
    port: 8000
  periodSeconds: 5
  timeoutSeconds: 2
  failureThreshold: 2
livenessProbe:
  httpGet:
    path: /livez
    port: 8000
  periodSeconds: 10
  timeoutSeconds: 2
  failureThreshold: 3

Place the parallel group, not just the pod count

Parallel replicas need compatible device groups and communication paths. Counting Pods alone cannot prove the placement.

A model that fits comfortably on one device can be replicated to serve independent requests. Tensor parallelism instead splits operations and introduces communication among devices during execution; pipeline parallelism splits layers and introduces stage transfers and scheduling bubbles. Both may enable a model that otherwise does not fit, but neither guarantees better latency. Replicas and model shards are different capacity units.

Trace the actual communication path: same-device memory, intra-host GPU links, host NUMA placement, NIC attachment, and cross-host fabric. A count of available GPUs does not describe those links. Place a collective-heavy group on the intended topology and ensure all its workers can become ready together. Verify shared-memory and communication-library requirements for the pinned engine without casually widening pod privileges.

Disaggregating prefill and decode can separate their scheduling pressures, but moving KV state becomes a real transfer with bandwidth, lifetime, and failure costs. Record the transfer boundary and its effect on TTFT before assuming phase separation helps. Across any parallel layout, a lost worker may invalidate an entire model group, so recovery headroom must be counted in groups rather than individual pod replicas.

Published offering · p6-b200.48xlarge

Inside a real P6-B200 cloud node

Eight B200 GPUs · 1,432 GB reported GPU memory · 192 vCPUs · 2,048 GiB host RAM

Host and staging

Eight separate GPU memory domains

Scale up within the node

Scale out, storage and isolation

Host CPUs

Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.

Read every component and connection
Host CPUs · 192 vCPUs
Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.
Host DRAM · 2,048 GiB
CPU-addressable memory stages data and hosts application state. It is separate from GPU HBM; copies, mappings and supported direct-I/O paths have distinct costs.
GPU 0 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 1 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 2 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 3 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 4 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 5 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 6 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 7 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
NVLink fabric · GPU ↔ GPU
The scale-up fabric connects compatible GPU peers for transfers and collectives. The runtime must partition work and enable the relevant paths; summed HBM is not one automatic allocation.
EFA networking · 3.2 Tbps aggregate
EFA provides the scale-out transport between instances. NCCL and the supported transport stack coordinate collectives; inter-node networking is not the same link as NVLink.
Local NVMe · 8 × 3.84 TB
Local instance storage can cache datasets and model artifacts. The listed raw drive capacity is not durable object storage or a promise of the same usable RAID capacity.
EBS path · 100 Gbps listed
The product table lists EBS bandwidth separately from local NVMe and EFA. Volume provisioning and workload behavior still constrain achieved storage throughput.
AWS Nitro · Virtualization / I/O
Nitro combines specialized hardware and firmware for the EC2 isolation and I/O boundary. Guest Linux, NVIDIA drivers and application permissions remain distinct responsibilities.

AWS product-table figures, checked September 8, 2026. This is a conceptual map of a published production offering, not access to a tenant, a live health feed or a guarantee of spare capacity. Ordinals 0–7 are diagram labels. Links show logical responsibilities, not proprietary board traces.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Published offering · p6-b300.48xlarge

Inside a real P6-B300 cloud node

Eight B300 GPUs · 2,144 GB reported GPU memory · 192 vCPUs · 4,096 GiB host RAM

Host and staging

Eight separate GPU memory domains

Scale up within the node

Scale out, storage and isolation

Host CPUs

Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.

Read every component and connection
Host CPUs · 192 vCPUs
Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.
Host DRAM · 4,096 GiB
CPU-addressable memory stages data and hosts application state. It is separate from GPU HBM; copies, mappings and supported direct-I/O paths have distinct costs.
GPU 0 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 1 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 2 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 3 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 4 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 5 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 6 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 7 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
NVLink fabric · GPU ↔ GPU
The scale-up fabric connects compatible GPU peers for transfers and collectives. The runtime must partition work and enable the relevant paths; summed HBM is not one automatic allocation.
EFA networking · 6.4 Tbps aggregate
EFA provides the scale-out transport between instances. NCCL and the supported transport stack coordinate collectives; inter-node networking is not the same link as NVLink.
Local NVMe · 8 × 3.84 TB
Local instance storage can cache datasets and model artifacts. The listed raw drive capacity is not durable object storage or a promise of the same usable RAID capacity.
EBS path · 100 Gbps listed
The product table lists EBS bandwidth separately from local NVMe and EFA. Volume provisioning and workload behavior still constrain achieved storage throughput.
AWS Nitro · Virtualization / I/O
Nitro combines specialized hardware and firmware for the EC2 isolation and I/O boundary. Guest Linux, NVIDIA drivers and application permissions remain distinct responsibilities.

AWS product-table figures, checked September 8, 2026. This is a conceptual map of a published production offering, not access to a tenant, a live health feed or a guarantee of spare capacity. Ordinals 0–7 are diagram labels. Links show logical responsibilities, not proprietary board traces.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Benchmark the service at controlled offered load

The service meets controlled offered load. You retain queueing, rejected work and latency rather than optimize a utilization headline.

Write the experiment contract first: model and tokenizer digests, engine/runtime/driver versions, device model and sharing mode, parallel layout, weight and KV precision, sampling settings, prompt-length and output-length distributions, prefix reuse, and quality acceptance criteria. Keep cold artifact loading, cold compilation or graph capture, warm steady-state serving, and burst recovery in separate reported phases. Do not put warm-up requests into a steady-state result without saying so.

Use an open-loop arrival schedule to study a specified request rate and its queue buildup; ensure the load generator can sustain that schedule without becoming the bottleneck. A closed-loop fixed-concurrency test is useful too, but it reduces offered traffic when responses slow and can conceal overload. Report which design was used. Sweep offered load through and beyond the intended operating range while holding the workload distribution fixed, and include representative bursts rather than only a smooth average.

For each point, retain the duration, offered and completed counts, errors and rejections, client TTFT/TPOT distributions, useful tokens per second, cache occupancy, queue depth and age, preemptions, and memory peaks. Include the drain policy for requests crossing the measurement cutoff, repeated-run variability, and sample counts supporting tail percentiles. Check a representative output sample for correctness or quality. Only compare candidates at a shared service objective; a higher total-token rate with worse deadline compliance is a different operating point, not a free improvement.

Rehearse failure and make the release reversible in Git

The proposed 6,000-hour monthly reduction has to survive peak demand and loss of a failure domain.

A useful validation plan exercises an unavailable model artifact, a checksum mismatch, initialization that exceeds its allowance, cache exhaustion under long contexts, a lost worker, and a traffic step during replacement. Observe readiness, queue growth, cancellation cleanup, partial-stream behavior, and recovery time. Distinguish host-memory OOM kills from engine-reported device-memory allocation failures. An interrupted stream cannot generally resume by blindly replaying a partially delivered answer; define what the client sees and bound retry amplification.

Version the serving image digest, model and tokenizer artifacts, engine settings, GPU resource/profile, parallel topology, probe contract, and routing change as one reviewed release. Start with a controlled candidate allocation and explicit stop conditions based on errors, TTFT/TPOT compliance, queue age, and readiness. Allow old and new replicas to coexist with enough device capacity for the replacement strategy; a surge replica that remains pending cannot provide rollout safety.

A reversible GitOps rollout restores the previous compatible desired state, not merely a mutable model alias. Retain its artifacts and device capacity, and keep tokenizer/API changes compatible while both versions serve. Argo CD documents that direct application rollback is unavailable with automated sync enabled; make a reviewed Git revision that restores the known-good configuration, or follow the controller-specific approved recovery procedure. An imperative live patch may be overwritten by reconciliation. Plan draining and in-flight request outcomes separately from configuration rollback, which cannot undo already delivered tokens. These are proposed exercises only; no release or failure injection is executed here.

Questions behind the decision

Which metric should drive LLM inference autoscaling?

Use a signal connected to unmet useful demand, such as queue age or backlog under a latency objective, and validate it against serving behavior. CPU or GPU utilization alone can miss memory pressure, blocked dependencies and long model startup.

Should an inference service scale to zero?

It depends on arrival patterns and the accepted cold-start delay. Interactive workloads often need warm capacity; a lower idle bill is not a win if startup makes requests miss their objective.

References & further reading

Engineering notes

Keep following the thread.

Real client engagements and the engineering behind them.

Technical glossary: definitions, connected ideas and further reading.

Optional analytics off. Contact works either way.

How measurement works