Cloud cost and infrastructure decisions
GPU Rental Pricing vs On-Premises: A Workload-Based TCO Guide
A real client engagement. The engineering and the results are described below.
A media-analysis company we worked with gets three GPU hosting quotes before a large customer pilot. The lowest hourly rate looks like an easy win—until its finance lead asks who pays when capacity is idle, unavailable or expensive to operate.
Request a time through the inquiry form. A meeting is confirmed separately by email.
The business problem behind the technology
Which GPU hosting quote buys the required service?
The team must compare the same workload, quality and recovery boundary. An annual purchase commitment is not justified by a few busy demonstration hours.
Read the client engagement ↓Client engagement / Delivered results
The cheapest hourly rate did not settle the contract
A media-analysis firm we worked with needs a stable monthly batch of completed video-processing jobs and a smaller interactive service. It can rent specialist GPU capacity or own an installation in colocation.
The constraint
The team must compare the same workload, quality and recovery boundary. An annual purchase commitment is not justified by a few busy demonstration hours.
The engineering decision
Procurement converts both offers into monthly obligations, includes an explicit allowance for operations and failure reserve, and runs a reversible trial. The comparable all-in planning estimates came to $140,000 rented and $112,000 owned-and-operated per month.
The delivered outcome
The estimate differed by $28,000 a month, and the recommendation stayed conditional: demand volatility, exit costs and the ability to operate the owned option could reverse it.
| Measure / unit | Before | After | Difference |
|---|---|---|---|
| Comparable monthly planning estimate USD/month | 140,000 | 112,000 | 28,000 |
Comparable monthly planning estimate. The $28,000 difference compares the two fully scoped estimates built for this decision.
The conditions behind the results
- Both estimates cover the same useful work, availability target and staffing boundary.
- The ownership estimate includes capital amortization, colocation, energy, support and operations; do not add them again.
- Migration and exit costs are excluded and must be priced separately before commitment.
- These figures are independent of the article’s four-GPU teaching worksheet and public quote-normalization example.
What this does not prove. The TCO advantage was specific to this workload and quotes; it was priced as a decision input, not a universal claim that owned hardware beats GPU rental.
Evidence to collect for your own decision
- Collect dated quotes with device, node, region, term and minimum billing scope.
- Replay the workload at expected and low utilization, including failed and rejected jobs.
- Price migration, overlap, failure reserve and an executable exit before signing.
Key decisions
The cheapest advertised GPU-hour can become an expensive production service. Compare hyperscalers, specialist GPU clouds, owned hardware and colocation using the workload, availability target and operating responsibility you actually need.
- Collect dated quotes with device, node, region, term and minimum billing scope.
- Replay the workload at expected and low utilization, including failed and rejected jobs.
- A modeled TCO advantage is neither an invoice saving nor a universal claim that owned hardware beats GPU rental.
Follow the decision
Select a step to follow its reasoning, then continue into the technical chapters.
Problem → boundary → decision → evidence
Specify quality, useful throughput, latency and the recovery target.
Read every component and connection
- Define the workload · Problem
- Specify quality, useful throughput, latency and the recovery target.
- Normalize the offer · Boundary
- Match the hardware, region, capacity arrangement and included services.
- Price the obligations · Decision
- Include idle time, staff, storage, network and ownership costs.
- Test the decision · Evidence
- Run comparable work and retain a workable exit.
- Define the workload → Normalize the offer: identify the constraint
- Normalize the offer → Price the obligations: choose a bounded change
- Price the obligations → Test the decision: check the outcome
A conceptual decision map for this article, not a measured timeline, physical topology or a depiction of a specific client system.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
Define an equivalent workload before comparing prices
The client’s media company defines the useful-work contract before comparing rental and ownership estimates.
Specify the model and version, precision, prompt and output lengths, batch policy, concurrency, quality checks and latency target. Record useful throughput at that target, not peak tokens per second from a different workload. For training, compare time to a validated result and checkpoint or restart behavior.
Then define the operating envelope: region, data residency, private connectivity, recovery time, capacity guarantee and support expectations. A single interruptible GPU and a redundant reserved service do not buy the same outcome. The cheapest option that cannot meet the requirement is not a saving.
Quotes
Normalize hardware, region, duration, capacity and included services.
Ownership
Add idle time, labor, storage, power, networking and migration.
Decision
Choose a reversible fit and measure whether assumptions survive reality.
This is a decision workflow, not a vendor ranking or an investment return forecast.
Know which responsibilities each option moves
The provider category does not settle who operates the system. You trace the responsibility boundary behind each quote.
Hyperscalers usually offer a broad ecosystem of managed services and regions. Specialist GPU clouds, often called neoclouds, concentrate on accelerator capacity and related infrastructure. Both still require verification of specific capacity, network, support and service terms; neither label guarantees performance or availability.
On-premises hosting puts physical infrastructure responsibilities with the owner. Colocation places owned or contracted equipment in another operator’s facility, often leaving hardware lifecycle and application operations with you. Managed hosting changes that split again. Obtain an explicit responsibility matrix instead of relying on the category name.
| Option | Potential fit | Costs and risks to price |
|---|---|---|
| Hyperscaler | Managed dependencies, regional reach and variable demand | Compute, commitments, storage, cross-zone traffic and egress |
| Neocloud | GPU-focused capacity and workload-specific configurations | Exact GPU topology, availability, support and data exit |
| On-premises | Sustained demand and direct infrastructure control | Capital, power, cooling, space, staff, spares and refresh |
| Colocation | Owned hardware with a contracted facility boundary | Rack or power billing, cross-connects, remote hands and hardware ownership |
Normalize an actual public price carefully
A public listing gives the conversation a concrete number. You keep its node-level scope and exclusions attached when converting it into another unit.
When accessed on September 8, 2026, CoreWeave’s North America pricing page listed an eight-GPU NVIDIA HGX B200 configuration at $68.80 per on-demand node-hour. Dividing by eight gives $8.60 per allocated GPU-hour. Holding that node for a 730-hour month would be $50,224 for that line item, before any separately charged services or applicable terms.
That arithmetic does not prove a standalone GPU is available at the divided rate, establish reserved capacity, or include every production cost. Confirm the current quote, minimums, included CPU/RAM/storage, region, taxes, network charges and support. AWS and Google Cloud pricing likewise distinguish compute from other billable resources and describe commitment or reservation rules. Dated public list prices are evidence, not a substitute for a matched procurement quote.
Compare real B200 and B300 node profiles
The hardware names look similar, but the instance profiles are not interchangeable. Real B200 and B300 configurations make the comparison more specific.
AWS lists both P6-B200 and P6-B300 production instance types. The product table reports eight GPUs in each, 192 vCPUs, 2,048 versus 4,096 GiB host memory, 3.2 versus 6.4 Tbps aggregate networking, and eight 3.84 TB local drives. AWS also published B300 availability in Seoul on August 20, 2026. Availability announcements do not guarantee that an account can reserve capacity now.
Keep the memory scope visible: the AWS table reports 1,432/2,144 GB, while its feature prose uses 1,440 GB/2.1 TB. CoreWeave’s separate B300 profile lists 270 GB GPU RAM and 30.72 TB usable local storage from eight 7.68 TB drives in RAID 10. Those are different provider configurations, not interchangeable node prices. Get a current quote and actual device inventory before deriving cost per useful request.
The diagrams below expose compute, memory, scale-up fabric, scale-out networking, storage and virtualization separately. They show the published offerings checked September 8, 2026.
Published offering · p6-b200.48xlarge
Eight B200 GPUs · 1,432 GB reported GPU memory · 192 vCPUs · 2,048 GiB host RAM
Host and staging
Eight separate GPU memory domains
Scale up within the node
Scale out, storage and isolation
Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.
Read every component and connection
- Host CPUs · 192 vCPUs
- Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.
- Host DRAM · 2,048 GiB
- CPU-addressable memory stages data and hosts application state. It is separate from GPU HBM; copies, mappings and supported direct-I/O paths have distinct costs.
- GPU 0 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 1 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 2 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 3 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 4 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 5 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 6 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 7 · B200
- One of eight B200 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- NVLink fabric · GPU ↔ GPU
- The scale-up fabric connects compatible GPU peers for transfers and collectives. The runtime must partition work and enable the relevant paths; summed HBM is not one automatic allocation.
- EFA networking · 3.2 Tbps aggregate
- EFA provides the scale-out transport between instances. NCCL and the supported transport stack coordinate collectives; inter-node networking is not the same link as NVLink.
- Local NVMe · 8 × 3.84 TB
- Local instance storage can cache datasets and model artifacts. The listed raw drive capacity is not durable object storage or a promise of the same usable RAID capacity.
- EBS path · 100 Gbps listed
- The product table lists EBS bandwidth separately from local NVMe and EFA. Volume provisioning and workload behavior still constrain achieved storage throughput.
- AWS Nitro · Virtualization / I/O
- Nitro combines specialized hardware and firmware for the EC2 isolation and I/O boundary. Guest Linux, NVIDIA drivers and application permissions remain distinct responsibilities.
- Local NVMe → Host DRAM: stage artifacts
- EBS path → Host DRAM: volume I/O
- Host CPUs → Host DRAM: host memory
- Host DRAM → GPU 0: supported device transfer
- AWS Nitro → Host CPUs: guest / I/O boundary
- GPU 0 → NVLink fabric: NVLink peers
- GPU 1 → NVLink fabric: NVLink peers
- GPU 2 → NVLink fabric: NVLink peers
- GPU 3 → NVLink fabric: NVLink peers
- GPU 4 → NVLink fabric: NVLink peers
- GPU 5 → NVLink fabric: NVLink peers
- GPU 6 → NVLink fabric: NVLink peers
- GPU 7 → NVLink fabric: NVLink peers
- EFA networking → GPU 0: collective transport
AWS product-table figures, checked September 8, 2026. This is a conceptual map of a published production offering, not access to a tenant, a live health feed or a guarantee of spare capacity. Ordinals 0–7 are diagram labels. Links show logical responsibilities, not proprietary board traces.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
Published offering · p6-b300.48xlarge
Eight B300 GPUs · 2,144 GB reported GPU memory · 192 vCPUs · 4,096 GiB host RAM
Host and staging
Eight separate GPU memory domains
Scale up within the node
Scale out, storage and isolation
Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.
Read every component and connection
- Host CPUs · 192 vCPUs
- Host threads tokenize, prepare batches, run the serving scheduler and submit CUDA work. A vCPU count is not a GPU execution-lane count.
- Host DRAM · 4,096 GiB
- CPU-addressable memory stages data and hosts application state. It is separate from GPU HBM; copies, mappings and supported direct-I/O paths have distinct costs.
- GPU 0 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 1 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 2 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 3 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 4 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 5 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 6 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- GPU 7 · B300
- One of eight B300 devices, each with its own memory and execution resources. This ordinal is a diagram label, not a hardware UUID or a claimed running allocation.
- NVLink fabric · GPU ↔ GPU
- The scale-up fabric connects compatible GPU peers for transfers and collectives. The runtime must partition work and enable the relevant paths; summed HBM is not one automatic allocation.
- EFA networking · 6.4 Tbps aggregate
- EFA provides the scale-out transport between instances. NCCL and the supported transport stack coordinate collectives; inter-node networking is not the same link as NVLink.
- Local NVMe · 8 × 3.84 TB
- Local instance storage can cache datasets and model artifacts. The listed raw drive capacity is not durable object storage or a promise of the same usable RAID capacity.
- EBS path · 100 Gbps listed
- The product table lists EBS bandwidth separately from local NVMe and EFA. Volume provisioning and workload behavior still constrain achieved storage throughput.
- AWS Nitro · Virtualization / I/O
- Nitro combines specialized hardware and firmware for the EC2 isolation and I/O boundary. Guest Linux, NVIDIA drivers and application permissions remain distinct responsibilities.
- Local NVMe → Host DRAM: stage artifacts
- EBS path → Host DRAM: volume I/O
- Host CPUs → Host DRAM: host memory
- Host DRAM → GPU 0: supported device transfer
- AWS Nitro → Host CPUs: guest / I/O boundary
- GPU 0 → NVLink fabric: NVLink peers
- GPU 1 → NVLink fabric: NVLink peers
- GPU 2 → NVLink fabric: NVLink peers
- GPU 3 → NVLink fabric: NVLink peers
- GPU 4 → NVLink fabric: NVLink peers
- GPU 5 → NVLink fabric: NVLink peers
- GPU 6 → NVLink fabric: NVLink peers
- GPU 7 → NVLink fabric: NVLink peers
- EFA networking → GPU 0: collective transport
AWS product-table figures, checked September 8, 2026. This is a conceptual map of a published production offering, not access to a tenant, a live health feed or a guarantee of spare capacity. Ordinals 0–7 are diagram labels. Links show logical responsibilities, not proprietary board traces.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
Build the ownership worksheet from real obligations
Its finance team now prices the obligations that the attractive GPU-hour headline omitted.
For the four-GPU purchase in this review, take $90,000 amortized over 36 months with no residual value: $2,500 per month. Add $750 for facility space/connectivity, $1,500 allocated operations labor, $500 support/spares and $450 power. The modeled total is $5,700 per month. These are the client’s planning inputs, not prices for B200 hardware or a colocation offer.
This simplified scenario assumes no financing, additional taxes or software-license cost. Replace those zero assumptions when they apply. Keep capital cash outlay separate from the monthly economic allocation; depreciation does not make the purchase free. Staff capacity is a cost allocation, not necessarily incremental payroll, and must not be counted again as cash savings.
Monthly boundary
Include the obligations assigned to this workload.
Capacity
Count compatible owned GPUs and the modeled calendar hours.
Useful fraction
Exclude idle or nonproductive time using a stated definition.
Unit cost
Divide cost by comparable productive GPU-hours, not activity alone.
At 60% productive utilization, four GPUs supply 1,752 productive GPU-hours in this modeled month.
All-in monthly cost (USD): 5700; Owned GPUs: 4; Productive utilization (%): 60
monthly cost / (GPUs × 730 hours × productive fraction). A 730-hour month, not a quote. Count only comparable useful work; billing hours and GPU activity are not necessarily productive hours.
Change the example assumptions
The worked result is readable without JavaScript. Inputs become available when the local WASM model loads; constrained connections and devices retain the static example.
Find the utilization threshold without hiding assumptions
The spreadsheet now meets uncertain demand. Changing productive utilization can change the answer without changing the monthly obligation.
At 60% productive utilization, this system costs about $3.25 per productive GPU-hour. At 20%, it costs about $9.76; at 90%, about $2.17. These are different uses of the same monthly obligation, not predicted demand or guaranteed throughput. Recovery reserve and workload fragmentation can prevent the higher utilization case.
For a rental alternative costing $4 per equivalent productive GPU-hour after all relevant costs, the simple ownership threshold is 5,700 / (4 GPUs × 730 hours × $4), about 48.8% utilization. This assumes the rental can genuinely be billed with useful demand and both options deliver the same quality, throughput and availability. Add minimum bookings, idle rental time or different operations costs before applying that comparison.
Keep commitment discounts separate from capacity guarantees
A discount enters the negotiation, followed by a different question: will the required capacity actually be available?
A spend commitment can lower a rate while retaining an obligation during idle periods. A capacity reservation can secure availability while billing unused reserved capacity. Provider-specific terms determine how these interact; a discount does not automatically reserve a scarce GPU in the desired zone.
Spot or interruptible capacity adds restart, checkpoint, queueing and completion-time risk. It can suit restartable work, but evaluate the cost per completed useful job, including failed attempts and lost progress. Do not assume a discount of a particular size or uninterrupted supply from a historical listing.
Include migration, data gravity and the exit
The move itself has a bill and an operational risk. Data gravity, overlap and a credible exit belong in cloud cost optimization strategies.
Price model and dataset transfer, cross-zone and inter-region traffic, storage performance, backups and retrieval. Include a migration overlap period when old and new capacity run together. Private connectivity, remote hands, physical replacement and security operations can dominate a small nominal GPU-rate difference.
A useful exit plan names artifact formats, export bandwidth, deletion or retention obligations, replacement capacity and the recovery test. Avoid amortizing an optimistic three-year saving while ignoring an expensive early exit or a demand collapse. Run low, base and high demand cases and record which assumption changes the decision.
Choose with a measured, reversible trial
The team used the trial and an explicit exit to decide whether the monthly difference deserved a commitment.
Request matching quotes and run the same representative workload with the same accuracy gate. Compare latency distributions, useful throughput, failures, startup behavior and total billed resources. Record dates, hardware, software versions and the exact capacity arrangement so another engineer can reproduce the comparison.
Start with the smallest trial that resolves the uncertain decision, not a wholesale migration. Establish stop conditions and restore the previous path before committing the wider workload. The outcome should be a priced responsibility boundary and observed workload fit, not a universal claim that cloud or ownership always wins.
Questions behind the decision
How should GPU rental prices be compared?
Normalize the exact hardware, node configuration, region, billing minimum and commitment before comparing useful work. Include idle time, failure coverage, data movement and the operations that the provider does not perform.
When does on-premises GPU hosting become cheaper?
There is no universal utilization threshold. Calculate it from the dated rental alternative and the full ownership obligation, then stress-test demand, maintenance, replacement and exit assumptions.
