Hover a dotted term for 5 seconds to lock its explanation. It closes after 5 seconds away; nested tooltips and keyboard focus keep it open. Click, tap or Enter locks immediately. Technical glossary.

Production AI infrastructure

GPU Servers for AI: Hosting, Isolation and Production Readiness

A real client engagement. The engineering and the results are described below.

An enterprise AI vendor we worked with wins permission to pilot a new document service, but each rollout leaves expensive GPUs waiting for model artifacts. The business deadline turns a server specification into a question about storage, isolation and readiness.

Request a time through the inquiry form. A meeting is confirmed separately by email.

The business problem behind the technology

When does a paid GPU server become useful capacity?

The next release needs predictable preparation and recovery, without pretending a faster cold start alone establishes high availability.

Read the client engagement ↓

Client engagement / Delivered results

The launch was waiting for the platform, not the GPU

An enterprise document platform we worked with runs 240 model-instance preparations a month across evaluation and release environments. Operators discover that an allocated device is often idle while software and artifacts become ready.

The constraint

The next release needs predictable preparation and recovery, without pretending a faster cold start alone establishes high availability.

The engineering decision

The team measures artifact transfer, CPU unpacking, driver compatibility and model loading separately. It chooses local model staging and a documented image/runtime contract, while preserving tenant isolation and a tested restore path.

The delivered outcome

Preparation fell from 45 to 15 minutes per instance. Across 240 monthly preparations, aggregate preparation time falls from 180 to 60 hours.

Measured inputs and delivered differences
Measure / unitBeforeAfterDifference
Aggregate instance preparation time
hours/month
18060120

Aggregate instance preparation time. 240 × 45 / 60 versus 240 × 15 / 60. Parallel preparations mean this is not 120 hours off a calendar.

The conditions behind the results

  • The 240 preparations and 45/15-minute times are the engagement’s operating values.
  • Each instance loads the same artifacts and reaches the same real readiness condition.
  • This ledger is separate from the eight-GPU placement example and does not claim a B200/B300 benchmark.

What this does not prove. Aggregate preparation time tracked the platform work delivered; overlapping instances were not counted as sequential delay.

Evidence to collect for your own decision

  • Time every stage from allocation through successful model-ready requests.
  • Test identity, storage permissions and isolation on the exact offered node configuration.
  • Remove a failure domain and prove recovery with the required model available.

Key decisions

A production model depends on much more than a GPU. Follow the path from rack power and server hardware through the hypervisor, guest operating system and serving runtime, then test the storage, isolation and recovery assumptions that make the whole service dependable.

  • Time every stage from allocation through successful model-ready requests.
  • Test identity, storage permissions and isolation on the exact offered node configuration.
  • Aggregate preparation time is not staff time, cash savings or an availability guarantee; overlapping instances must not be counted as sequential delay.

Follow the decision

When does a paid GPU server become useful capacity?

Select a step to follow its reasoning, then continue into the technical chapters.

Problem → boundary → decision → evidence

Inspect the node

Keep physical resources and provider-reported capacities in their proper scope.

Read every component and connection
Inspect the node · Problem
Keep physical resources and provider-reported capacities in their proper scope.
Choose isolation · Boundary
A VM, container and GPU partition protect different boundaries.
Make the model ready · Decision
Storage, drivers and warmup must work before capacity can serve.
Rehearse recovery · Evidence
Prove the failure domain and service objective you intend to support.
  • Inspect the node → Choose isolation: identify the constraint
  • Choose isolation → Make the model ready: choose a bounded change
  • Make the model ready → Rehearse recovery: check the outcome

A conceptual decision map for this article, not a measured timeline, physical topology or a depiction of a specific client system.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Start below the operating system

The client’s document platform inventories the host and facility before promising that a device is ready for its pilot.

Real cloud examples make the boundary concrete. AWS currently documents p6-b200.48xlarge and p6-b300.48xlarge offerings with eight GPUs each. Its product table reports 1,432 GB and 2,144 GB of GPU memory respectively, alongside different host-memory and network capacities. These diagrams preserve that table’s scope; they are not live tenant telemetry or a claim that I operate these instances.

Do not substitute a DGX reference-system bill of materials for the cloud instance. NVIDIA’s DGX B200 guide lists 1,440 GB across eight GPUs, two fifth-generation NVSwitch devices and ConnectX-7 networking. Its DGX B300 guide lists eight 288 GB GPUs and ConnectX-8 networking. Provider-exposed memory, model allocations and reference-system specifications can differ; inspect the actual ordered configuration rather than silently reconciling unlike numbers.

A server combines CPUs, system RAM, GPUs, storage controllers, network interfaces, firmware and power supplies. PCIe and NUMA topology influence which CPU, memory bank and devices exchange data efficiently. A rack adds power distribution, cooling and network uplinks. These dependencies remain present when a cloud API hides them.

Verify rack power density, cooling compatibility, circuit redundancy, link capacity and the ownership of firmware updates before promising usable GPU capacity. Nameplate power is a planning limit, not measured energy. The facility must support the actual equipment and failover behavior, including maintenance conditions.

Four layers, four operational responsibilities
  1. Facility

    Power, cooling, physical access and network paths keep equipment available.

  2. Host

    Firmware, devices and the host or hypervisor establish the machine boundary.

  3. Guest

    A VM owns its kernel; containers may share that guest kernel.

  4. Service

    Drivers, runtime, model and application policy deliver useful requests.

A conceptual responsibility stack, not a universal boot sequence or security certification.

Published offering · p6-b200.48xlarge

Inside a real P6-B200 cloud node

Eight B200 GPUs · 1,432 GB reported GPU memory · 192 vCPUs · 2,048 GiB host RAM

Host and staging

Eight separate GPU memory domains

Scale up within the node

Scale out, storage and isolation

Host CPUs

Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.

Read every component and connection
Host CPUs · 192 vCPUs
Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.
Host DRAM · 2,048 GiB
CPU-addressable memory stages data and hosts application state. It is separate from GPU HBM; copies, mappings and supported direct-I/O paths have distinct costs.
GPU 0 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 1 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 2 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 3 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 4 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 5 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 6 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 7 · B200
One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
NVLink fabric · GPU ↔ GPU
The scale-up fabric connects compatible GPU peers for transfers and collectives. The runtime must partition work and enable the relevant paths; summed HBM is not one automatic allocation.
EFA networking · 3.2 Tbps aggregate
EFA provides the scale-out transport between instances. NCCL and the supported transport stack coordinate collectives; inter-node networking is not the same link as NVLink.
Local NVMe · 8 × 3.84 TB
Local instance storage can cache datasets and model artifacts. The listed raw drive capacity is not durable object storage or a promise of the same usable RAID capacity.
EBS path · 100 Gbps listed
The product table lists EBS bandwidth separately from local NVMe and EFA. Volume provisioning and workload behavior still constrain achieved storage throughput.
AWS Nitro · Virtualization / I/O
Nitro combines specialized hardware and firmware for the EC2 isolation and I/O boundary. Guest Linux, NVIDIA drivers and application permissions remain distinct responsibilities.

AWS product-table figures, checked September 8, 2026. This is a conceptual map of a published production offering, not access to a tenant, a live health feed or a guarantee of spare capacity. Ordinals 0–7 are diagram labels. Links show logical responsibilities, not proprietary board traces.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Published offering · p6-b300.48xlarge

Inside a real P6-B300 cloud node

Eight B300 GPUs · 2,144 GB reported GPU memory · 192 vCPUs · 4,096 GiB host RAM

Host and staging

Eight separate GPU memory domains

Scale up within the node

Scale out, storage and isolation

Host CPUs

Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.

Read every component and connection
Host CPUs · 192 vCPUs
Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.
Host DRAM · 4,096 GiB
CPU-addressable memory stages data and hosts application state. It is separate from GPU HBM; copies, mappings and supported direct-I/O paths have distinct costs.
GPU 0 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 1 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 2 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 3 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 4 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 5 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 6 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
GPU 7 · B300
One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
NVLink fabric · GPU ↔ GPU
The scale-up fabric connects compatible GPU peers for transfers and collectives. The runtime must partition work and enable the relevant paths; summed HBM is not one automatic allocation.
EFA networking · 6.4 Tbps aggregate
EFA provides the scale-out transport between instances. NCCL and the supported transport stack coordinate collectives; inter-node networking is not the same link as NVLink.
Local NVMe · 8 × 3.84 TB
Local instance storage can cache datasets and model artifacts. The listed raw drive capacity is not durable object storage or a promise of the same usable RAID capacity.
EBS path · 100 Gbps listed
The product table lists EBS bandwidth separately from local NVMe and EFA. Volume provisioning and workload behavior still constrain achieved storage throughput.
AWS Nitro · Virtualization / I/O
Nitro combines specialized hardware and firmware for the EC2 isolation and I/O boundary. Guest Linux, NVIDIA drivers and application permissions remain distinct responsibilities.

AWS product-table figures, checked September 8, 2026. This is a conceptual map of a published production offering, not access to a tenant, a live health feed or a guarantee of spare capacity. Ordinals 0–7 are diagram labels. Links show logical responsibilities, not proprietary board traces.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Know what the hypervisor isolates

The next boundary is the machine the application believes it owns. Hypervisors and containers divide responsibilities in different ways.

A hypervisor presents virtual CPUs, memory and devices to guests and controls access to underlying resources. Different platforms implement this differently: KVM uses Linux virtualization support, while cloud providers may combine a hypervisor with specialized offload hardware. Bare-metal hosting exposes a different responsibility boundary.

A VM has its own guest kernel. Ordinary containers share the kernel of their host, which may itself be a VM. Namespaces, cgroups and container policies are useful controls but are not automatically equivalent to a hardware virtualization boundary. Choose tenancy and threat assumptions before selecting the packaging format.

Software ownership, from cluster to silicon

A Kubernetes request does not schedule an ALU

Kubernetes places the workload. The container stack exposes a device. CUDA and the driver submit work. GPU scheduling and execution happen below that boundary.

Cluster lifecycle and placement

Application and user-space libraries

Kernel and device boundary

Observe the real service

Kubernetes scheduler

The scheduler chooses a suitable node using declared resources and placement policy. It does not place individual CUDA blocks, schedule warps or make an unready model useful.

Read every component and connection
Kubernetes scheduler · Node / resource placement
The scheduler chooses a suitable node using declared resources and placement policy. It does not place individual CUDA blocks, schedule warps or make an unready model useful.
Device plugin / kubelet · GPU allocation
A device plugin advertises supported GPU resources and participates in allocation. Kubelet starts the pod with the chosen devices. GPU sharing changes the resource contract.
GPU Operator · Lifecycle controller
The operator can manage drivers, Container Toolkit, device plugin, feature discovery and monitoring. This is a control-plane responsibility, not a mandatory hop for each tensor.
OCI runtime / Toolkit · Container device access
The configured runtime and NVIDIA Container Toolkit expose the permitted devices and libraries. Containers share their host kernel; a VM or bare-metal host sets another boundary.
Serving framework · Batching / model / NCCL
A serving engine schedules requests and calls framework/library kernels. NCCL coordinates supported multi-GPU collectives. Application queueing is distinct from GPU instruction scheduling.
CUDA runtime / driver · libcudart / libcuda
CUDA APIs manage contexts, streams, memory and launches. The user-mode driver interfaces with kernel support. Compatible machine code or supported PTX compilation is required.
Linux NVIDIA modules · Devices / memory / I/O
Kernel modules cooperate with the OS for device access, memory mapping and control. Linux permissions and isolation still matter; a container image does not replace the host kernel driver.
Queues / GPU firmware · Device work submission
Supported driver and firmware paths submit and manage device work. Firmware-mediated responsibilities vary with platform and driver; this is not a proprietary command-protocol schematic.
SMs / memory / ALUs · Execute and move data
GPU execution resources run the selected instructions and move operands through the appropriate memory paths. The gate, register and memory diagrams expand this layer.
DCGM + application telemetry · Health ≠ goodput
DCGM-based monitoring supplies device signals. Application metrics and traces must still prove latency, errors, queue age and useful throughput; a busy GPU is not the service objective.

A typical Linux/device-plugin deployment, not a universal platform configuration. GPU Operator manages components; it is not on every inference request’s data path. Driver, CUDA, framework and hardware compatibility must be validated together.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

GPU passthrough, MIG and sharing are not synonyms

Sharing the accelerator sounds economical until isolation and contention enter the discussion. You now distinguish passthrough, MIG and time slicing.

With device passthrough, a guest uses an assigned device through a supported driver and virtualization path. IOMMU protection and device isolation groups matter for DMA; peer devices, firmware and platform support can affect the boundary. Do not treat an unsafe passthrough workaround as tenant isolation.

On supported NVIDIA GPUs, MIG partitions particular compute and memory resources into supported instance profiles. It does not partition the entire server or remove host administration responsibilities. Time slicing schedules access to shared hardware; advertising more logical GPU allocations does not create more VRAM or equivalent hardware isolation.

Select a sharing method by its actual boundary
MethodWhat is allocatedWhat to verify
Whole-device passthroughA supported physical device assignmentIOMMU grouping, driver support and reset behavior
MIGA supported hardware resource partitionGPU generation, profile sizes, reconfiguration and tenancy
Time slicing / process sharingAccess to shared GPU execution resourcesMemory contention, scheduling and isolation requirements
ContainersAn application environment and resource policyKernel boundary, device exposure and permissions

Separate scheduling capacity from serving capacity

The scheduler can count devices, but the service needs complete, usable replicas. The placement example shows why those are different answers.

A model replica using tensor parallelism may require several compatible GPUs together. For eight GPUs, a two-GPU reserve and two GPUs per replica, the arithmetic ceiling is floor((8 − 2) / 2) = three replicas. If the devices are in incompatible hosts or links, even that ceiling may not be achievable.

Check GPU memory per device, CPU and RAM for preprocessing, network bandwidth, storage access and placement constraints. A scheduler allocation is not a loaded model, and a loaded model is not proof of healthy latency under traffic. Readiness should reflect the state required to serve useful requests.

Count whole replicas without spending the reserve
  1. Inventory

    Count compatible physical GPUs, not oversubscribed logical shares.

  2. Reserve

    Hold back the capacity required by the explicitly stated recovery plan.

  3. Place

    Allocate complete GPU groups with the needed locality and memory.

  4. Prove ready

    Load weights and validate the service before routing traffic.

This model does not divide VRAM into arbitrary slots or simulate a Kubernetes scheduler.

Whole-GPU placement ceiling3 replicas, before topology and other limits

Available physical GPUs: 8; GPUs reserved for headroom: 2; Whole GPUs per replica: 2

floor(max(0, available − reserve) / GPUs per replica). Requires compatible GPUs in a usable topology. CPU, memory, power, model VRAM and failure-domain placement may lower this ceiling. Reserve alone does not establish high availability.

Change the example assumptions
Model inputs

The worked result is readable without JavaScript. Inputs become available when the local WASM model loads; constrained connections and devices retain the static example.

Storage and CPU can dominate cold starts

The team traces cold-start work because its preparation-time model depends on making the same artifacts ready sooner.

Model startup may include image retrieval, weight download, checksum verification, deserialization, device transfer, compilation and warmup. Separate those intervals. An ideal 40 GB transfer over a sustained 2 GB/s path already needs 20 seconds before other work; that is a lower bound with decimal units, not a startup measurement.

For an active model cache, fast local NVMe can remove repeated network retrieval when lifecycle and consistency rules allow it. It is not durable model ownership: keep versioned artifacts and recovery sources. Track cache misses, storage throughput, CPU bottlenecks and per-stage startup time instead of assuming every slow start is a GPU problem.

Price facility energy without double counting

The conversation reaches the facility bill. You keep measured energy, planning assumptions and bundled colocation charges from becoming the same number.

Power Usage Effectiveness compares total facility energy with IT equipment energy over a defined boundary and period. A planning estimate with 5 kW of average IT load and PUE 1.4 implies 7 kW including facility overhead. Over a 730-hour month at $0.15/kWh, that is $766.50. Real metering, tariffs and load shape determine the actual bill.

Do not multiply a colocation quote by PUE if its power charge already includes the same overhead. Identify whether billing covers reserved circuit capacity, metered energy, a rack bundle or another basis. PUE does not measure useful model work or total environmental impact; neither a lower PUE nor a lower GPU wattage alone proves a better service.

Design around the failure domain you can lose

A reserve looks reassuring until the whole host disappears. The recovery story must survive the actual failure domain, not merely a smaller allocation.

Two spare GPUs in the same eight-GPU host cannot recover the service if that host fails. A host, rack, switch, power circuit and availability zone are different failure domains. Place surviving capacity and data where the intended failure cannot remove them together.

Exercise a drain, an abrupt node loss, a missing model artifact, a slow storage dependency and a rollout overlap. Measure time to healthy capacity and the client-visible error or queueing interval. Backups need restore tests; redundant power supplies need genuinely independent upstream power to provide the intended protection.

Make production acceptance observable

The platform review closes only when readiness and recovery—not a server specification—support the customer pilot.

Track user-facing latency, errors, queue age and useful throughput alongside GPU memory, thermal or power throttling, device errors, CPU pressure and storage or network saturation. Correlate application traces with the particular model version, replica and host. GPU utilization is diagnostic context, not the service-level objective.

Document the owners of firmware, drivers, hypervisor, guest images, credentials, model artifacts and incident response. Gate releases on compatible versions and a reproducible rollback path. Hosting selection becomes much clearer once these responsibilities and recovery requirements are priced explicitly.

Questions behind the decision

What should an AI GPU server acceptance test include?

Test the complete path: artifacts, host resources, drivers, serving runtime, readiness and recovery. Name the exact device and node scope; a device count or nominal memory figure does not establish service capacity.

Are GPU passthrough, MIG and time slicing equivalent?

No. They divide assignment, partitioning and scheduling in different ways, with different isolation and contention properties. Choose from the workload and tenant boundary rather than treating every sharing mechanism as interchangeable capacity.

References & further reading

Engineering notes

Keep following the thread.

Real client engagements and the engineering behind them.

Technical glossary: definitions, connected ideas and further reading.

Optional analytics off. Contact works either way.

How measurement works