Hover a dotted term for 5 seconds to lock its explanation. It closes after 5 seconds away; nested tooltips and keyboard focus keep it open. Click, tap or Enter locks immediately. Technical glossary.

GPU architecture

How a GPU Works: CUDA, Memory and Better Capacity Decisions

A real client engagement. The engineering and the results are described below.

A product-catalog company we worked with prepares an AI search launch and receives a proposal to double its GPU allocation. Before buying, the engineering team follows a request through the CPU, GPU memory and execution units to discover what the hardware is actually waiting for.

Request a time through the inquiry form. A meeting is confirmed separately by email.

The business problem behind the technology

Does the product need more GPUs—or better-fed GPUs?

The team must keep search quality and latency while meeting the same nightly enrichment deadline. Buying devices before understanding CPU preparation, transfers and memory stalls could lock in the wrong expense.

Read the client engagement ↓

Client engagement / Delivered results

The procurement question hidden inside a slow response

A product-catalog service we worked with supports a growing batch-enrichment workload alongside interactive search. Its first capacity proposal is based on device utilization, not completed work.

The constraint

The team must keep search quality and latency while meeting the same nightly enrichment deadline. Buying devices before understanding CPU preparation, transfers and memory stalls could lock in the wrong expense.

The engineering decision

It separates the host pipeline, memory traffic and execution phases, then requires a representative replay before approving extra allocation. Feeding the existing fleet more effectively avoided 4,000 additional rented GPU-hours each month.

The delivered outcome

At $4 per GPU-hour, the plan falls from 12,000 to 8,000 GPU-hours: $16,000 of monthly compute charges the client did not incur.

Measured inputs and delivered differences
Measure / unitBeforeAfterDifference
Planned monthly GPU compute charge
USD/month
48,00032,00016,000

Planned monthly GPU compute charge. 12,000 × $4 versus 8,000 × $4 at the client’s contracted rate.

The conditions behind the results

  • The same useful jobs complete at the same quality and latency objectives.
  • The avoided 4,000 GPU-hours come from the engagement’s capacity plan, not from the diagrams.
  • Only flexible compute charges are counted; storage, networking, support and engineering cost remain outside this comparison.

What this does not prove. The avoided spend required billable capacity to retire without losing useful service.

Evidence to collect for your own decision

  • Profile CPU preparation, transfers, memory stalls and completed GPU work.
  • Replay the actual interactive and batch workload mix under the same SLO.
  • Reconcile a dated rental quote and minimum commitments with the proposed allocation.

Key decisions

A GPU is not simply a faster CPU. It combines many execution lanes, specialized arithmetic and a memory hierarchy designed to keep parallel work moving. Trace one workload through those layers and see why memory, batching and data layout often matter as much as arithmetic.

  • Profile CPU preparation, transfers, memory stalls and completed GPU work.
  • Replay the actual interactive and batch workload mix under the same SLO.
  • Avoided planned spend is not an observed saving or proof that utilization improvements remove a particular number of GPUs.

Follow the decision

Does the product need more GPUs—or better-fed GPUs?

Select a step to follow its reasoning, then continue into the technical chapters.

Problem → boundary → decision → evidence

Submit work

The host, runtime and driver prepare a device operation.

Read every component and connection
Submit work · Problem
The host, runtime and driver prepare a device operation.
Find the operands · Boundary
Memory hierarchy and data layout determine which bytes must move.
Execute instructions · Decision
Schedulers, registers and arithmetic resources do different jobs.
Return useful output · Evidence
The application still has to meet its quality and latency requirements.
  • Submit work → Find the operands: identify the constraint
  • Find the operands → Execute instructions: choose a bounded change
  • Execute instructions → Return useful output: check the outcome

A conceptual decision map for this article, not a measured timeline, physical topology or a depiction of a specific client system.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Follow work from the CPU to the GPU

The client’s catalog team follows host work first, before approving another GPU allocation.

The CPU prepares buffers and submits commands through a runtime and driver. Data may move over PCIe or another supported link into device memory. A kernel launch describes many threads; GPU hardware schedules their work onto streaming multiprocessors, or SMs. The result stays on the device until a consumer needs it elsewhere.

Submission is commonly asynchronous. With the right streams, hardware and dependencies, transfers and computation can overlap. The diagram separates responsibilities, not four mandatory nonoverlapping time intervals. Small jobs can lose more time to launch and transfer overhead than they save in parallel execution.

A request crosses several different resources
  1. Host

    The CPU prepares inputs and submits work through the driver.

  2. Device memory

    HBM or other GPU memory holds tensors and intermediate data.

  3. Multiprocessors

    SMs schedule warps and issue arithmetic or memory instructions.

  4. Consumer

    A later kernel or host operation waits on the required dependency.

This is a dependency map; real transfers and kernels can overlap.

Software ownership, from cluster to silicon

A Kubernetes request does not schedule an ALU

Kubernetes places the workload. The container stack exposes a device. CUDA and the driver submit work. GPU scheduling and execution happen below that boundary.

Cluster lifecycle and placement

Application and user-space libraries

Kernel and device boundary

Observe the real service

Kubernetes scheduler

The scheduler chooses a suitable node using declared resources and placement policy. It does not place individual CUDA blocks, schedule warps or make an unready model useful.

Read every component and connection
Kubernetes scheduler · Node / resource placement
The scheduler chooses a suitable node using declared resources and placement policy. It does not place individual CUDA blocks, schedule warps or make an unready model useful.
Device plugin / kubelet · GPU allocation
A device plugin advertises supported GPU resources and participates in allocation. Kubelet starts the pod with the chosen devices. GPU sharing changes the resource contract.
GPU Operator · Lifecycle controller
The operator can manage drivers, Container Toolkit, device plugin, feature discovery and monitoring. This is a control-plane responsibility, not a mandatory hop for each tensor.
OCI runtime / Toolkit · Container device access
The configured runtime and NVIDIA Container Toolkit expose the permitted devices and libraries. Containers share their host kernel; a VM or bare-metal host sets another boundary.
Serving framework · Batching / model / NCCL
A serving engine schedules requests and calls framework/library kernels. NCCL coordinates supported multi-GPU collectives. Application queueing is distinct from GPU instruction scheduling.
CUDA runtime / driver · libcudart / libcuda
CUDA APIs manage contexts, streams, memory and launches. The user-mode driver interfaces with kernel support. Compatible machine code or supported PTX compilation is required.
Linux NVIDIA modules · Devices / memory / I/O
Kernel modules cooperate with the OS for device access, memory mapping and control. Linux permissions and isolation still matter; a container image does not replace the host kernel driver.
Queues / GPU firmware · Device work submission
Supported driver and firmware paths submit and manage device work. Firmware-mediated responsibilities vary with platform and driver; this is not a proprietary command-protocol schematic.
SMs / memory / ALUs · Execute and move data
GPU execution resources run the selected instructions and move operands through the appropriate memory paths. The gate, register and memory diagrams expand this layer.
DCGM + application telemetry · Health ≠ goodput
DCGM-based monitoring supplies device signals. Application metrics and traces must still prove latency, errors, queue age and useful throughput; a busy GPU is not the service objective.

A typical Linux/device-plugin deployment, not a universal platform configuration. GPU Operator manages components; it is not on every inference request’s data path. Driver, CUDA, framework and hardware compatibility must be validated together.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

An ALU result begins with gates and stored operands

Zoom in until a large computation becomes a few bits. The adder makes the distinction between combinational logic and stored state concrete.

CMOS transistors implement logic gates by controlling pull-up and pull-down paths. Signals take time to charge and discharge capacitance, and real devices also have leakage and other losses. A gate is combinational logic; a register stores state across clocked stages. Do not confuse a register file with off-chip DRAM or treat a drawn pulse as a measured transistor delay.

The full adder below is a working generic Boolean example: sum = A XOR B XOR carry-in, while carry-out combines generated and propagated carry. Change the three input bits and inspect the intermediate gates. Wider integer arithmetic builds on bit-level logic with implementation-specific carry networks and pipelines. Floating-point units add alignment, normalization, rounding and special-value behavior. This teaching circuit is not a claim about NVIDIA’s proprietary ALU layout.

Generic digital logic · not a NVIDIA die schematic

One addition bit, all the way down to gates

Sum = A XOR B XOR carry-in · carry-out = (A AND B) OR ((A XOR B) AND carry-in)

Input voltages interpreted as bits

First gate layer

Combine the partial results

Output bits captured by later storage

1 + 0 + 1 = 2 · sum 0, carry 1
A

A logical 1 or 0 represents a voltage range, not an infinitely precise voltage. CMOS transistors charge and discharge capacitances when a gate switches.

Read every component and connection
A · Input bit
A logical 1 or 0 represents a voltage range, not an infinitely precise voltage. CMOS transistors charge and discharge capacitances when a gate switches.
B · Input bit
The second operand bit enters both the XOR and AND paths. Gates compute Boolean functions; a clocked register stores a result for a later pipeline stage.
Carry in · Previous position
Carry-in is the carry from a less-significant bit position. Wide adders must manage the carry dependency; this does not claim NVIDIA uses a ripple-carry implementation.
XOR · A XOR B
XOR is 1 when exactly one input is 1. This partial sum feeds both the final sum gate and the carry-propagation gate.
AND · A AND B
When A and B are both 1, this path generates a carry regardless of carry-in. AND can itself be built from other CMOS gate arrangements.
XOR · Partial sum XOR carry
The second XOR produces the least-significant result bit. For 1 + 0 + 1, this bit is 0 while carry-out is 1.
AND · Partial sum AND carry
This path propagates carry-in when exactly one operand bit is 1. It is distinct from the carry generated by A AND B.
OR · Generated OR propagated
OR combines generated and propagated carry. Logically equivalent NAND networks or other carry designs can implement the same truth function.
Sum · Result bit
The sum bit and carry bit together represent A + B + carry-in. They are a two-bit result, not two independent arithmetic answers.
Carry out · Next bit position
Carry-out contributes weight two to this one-bit addition. The complete truth table must work for all eight input combinations.
  • A → XOR: A
  • B → XOR: B
  • A → AND: A
  • B → AND: B
  • XOR → XOR: partial sum
  • Carry in → XOR: carry-in
  • XOR → AND: partial sum
  • Carry in → AND: carry-in
  • AND → OR: generated
  • AND → OR: propagated
  • XOR → Sum: sum
  • OR → Carry out: carry

A locally calculated full-adder truth function. CMOS pull-up/pull-down transistor networks implement logic gates; real GPU adders use implementation-specific carry networks and pipelines. Animation speed is chosen for readability, not a propagation delay or a transistor simulation.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Threads become warps, not independent little CPUs

Zoom back out to a warp. The active lanes in one instruction tell you something useful, but not the whole story of GPU utilization.

CUDA groups threads into blocks and 32-thread warps. Threads have their own registers and logical identity; instructions execute with an active-lane mask. Divergent branches can require different subsets of lanes to execute different paths. Modern scheduling details do not make communication between threads automatically safe: use the documented synchronization primitives.

In the example, 20 active lanes use 62.5% of the lanes for that instruction. That is not a claim of 62.5% GPU utilization. Another warp may run while this one waits, and instruction throughput depends on the operation, dependencies and hardware pipelines. Occupancy instead measures resident warps relative to the supported maximum.

An instruction mask is not GPU utilization
  1. Group

    A CUDA warp has 32 logical lanes.

  2. Mask

    Only the lanes participating in this instruction are active.

  3. Issue

    The scheduler needs eligible work, not merely resident threads.

  4. Repeat

    The next instruction can have a different mask and bottleneck.

Change the active-lane count. The result describes one instruction, not a speedup.

Active-lane fraction62.5 % of one 32-lane CUDA warp

Active lanes in this instruction: 20

active / 32 × 100. This shows one instruction mask, not occupancy, GPU utilization or a throughput prediction. Other architectures can use different group sizes.

Change the example assumptions
Model inputs

The worked result is readable without JavaScript. Inputs become available when the local WASM model loads; constrained connections and devices retain the static example.

Memory is a hierarchy with different sharing rules

The operands need a place to live. Registers, caches, shared memory and HBM are different resources, not interchangeable names for GPU memory.

Registers hold thread-local values. Shared memory lets threads in a block cooperate, subject to synchronization and capacity limits. Caches reduce some repeated accesses, while off-chip device memory holds much larger arrays. Names and capacities vary by architecture; do not turn a diagram into a claim about every GPU.

Capacity determines whether a workload fits; bandwidth determines how quickly bytes can move. A model that fits in VRAM can still be bandwidth-bound. Excessive register demand can spill into memory, increasing traffic. More shared memory per block can reduce how many blocks fit on an SM at once.

Conceptual Blackwell execution and memory map

From HBM bytes to a result inside a GPU

Off-chip DRAM → memory controllers → caches → registers / execution → stores. Instruction issue and asynchronous copies coordinate different paths.

Large off-chip storage

On-chip reuse and staging

Inside one representative streaming multiprocessor

Device-wide work and peers

HBM3e DRAM

Stacked DRAM stores large device-resident arrays. DRAM cells need refresh; capacity and sustained bandwidth are different limits. HBM is not the register file.

Read every component and connection
HBM3e DRAM · Weights / KV / arrays
Stacked DRAM stores large device-resident arrays. DRAM cells need refresh; capacity and sustained bandwidth are different limits. HBM is not the register file.
Memory controllers · Channels and requests
Controllers organize reads and writes to memory channels. Access patterns, contention and the memory technology influence service time; bandwidth is not zero latency.
L2 cache · Device-wide reuse
L2 can satisfy repeated requests without another HBM access. Its capacity and residency behavior affect traffic; a cache hit is not a new DRAM transfer.
L1 / shared memory · Caching / explicit tiles
B200 combines L1, texture and shared-memory resources. Shared memory is software-managed block storage with synchronization rules; it is not an automatic replacement for registers.
Register file · Thread operands
The B200 tuning guide specifies 64K 32-bit registers per SM. Threads use registers for live values; spills can create device-memory traffic. Registers are not off-chip DRAM.
ALU pipelines · Integer / floating point
Execution pipelines perform supported arithmetic and logic on operands. CMOS gates underlie those circuits. Floating-point operations include more work than the integer full adder shown.
Tensor cores · Matrix operations
Specialized matrix instructions use supported operand formats and accumulation paths. Tensor throughput is not scalar ALU throughput, and not every kernel can use tensor cores.
Load / store units · Addresses and movement
Load/store machinery forms and services memory operations. Coalescing groups useful lane accesses; dependencies prevent a consumer from using a value before it is ready.
Warp schedulers · Ready instruction issue
Schedulers issue eligible warp instructions subject to dependencies and resource availability. Other ready warps can hide a wait; occupancy alone does not prove throughput.
Instruction path · Fetch / decode / issue
Compiled machine instructions reach the SM instruction machinery. PTX is a virtual ISA; a compatible cubin or driver compilation supplies hardware-executable code.
Async copy / TMA · Tile movement
Supported asynchronous transfer paths can stage tiles while computation proceeds. Barriers and producer/consumer ordering still apply; overlap is not permission to read unfinished data.
GPU front end · Submitted work
Device work submission and scheduling machinery distribute kernel work. Block resource requirements influence residency. Kubernetes does not choose a warp or allocate an SM register.
NVLink interface · Peer devices
Peer access and collectives move data between compatible GPUs. The application/runtime manages distributed work; aggregate device memory is not one automatically shared allocation.

This is a functional map, not a floorplan or cycle-accurate simulator. One representative SM is expanded; it is not the GPU’s SM count. Cache bypass, asynchronous copies, distributed shared memory and specialized tensor accumulator paths mean not every operation follows every arrow.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Different resources answer different questions
ResourceRoleTypical constraint
RegistersPer-thread live valuesRegister pressure and spilling
Shared memoryExplicit cooperation within a blockCapacity, bank access and synchronization
L1 / L2 cachesReuse of memory accessesWorking set and access pattern
Device memoryLarge tensors, weights and KV cacheCapacity and sustained bandwidth
PCIe / interconnectHost-device or device-device exchangeTopology, latency and transfer volume

Coalescing makes useful bytes cheaper to fetch

Operand movement is now a purchasing question: useful work, not the device label, must justify the bill.

Suppose 32 active lanes each read a neighboring four-byte value. They request 128 useful bytes in a compact aligned span. If those lanes instead read widely separated addresses, the memory system may need many more sectors to deliver the same useful data. Caching, alignment and the device generation determine the actual transactions.

Reusing a loaded tile can reduce repeated memory traffic; simply adding more threads does not. Count bytes at the relevant memory boundary and compare them with useful operations. Arithmetic intensity is operations per byte, not the size of the input file. A kernel can be limited at shared memory or an instruction pipeline even when HBM is not saturated.

Tensor cores accelerate specific matrix operations

A matrix operation takes another route through the chip. Tensor cores help only when the operation, format and software path can actually use them.

Tensor cores are specialized matrix multiply-accumulate hardware. They do not accelerate every branch, lookup or scalar instruction. Supported data types, layouts, dimensions and accumulation behavior depend on the architecture and software path. The application must actually select a compatible operation.

Precision changes require quality evaluation. FP32, TF32, BF16, FP16 and lower-precision formats are not interchangeable labels for the same arithmetic. Advertised sparse throughput also assumes a supported sparsity pattern; do not compare it with a dense workload as if both perform the same useful computation.

Inference changes character between prefill and decode

The model begins generating an answer, and the workload changes. Prefill, decode and queueing now need separate explanations.

During prefill, a model processes the input context, often exposing substantial matrix parallelism. Autoregressive decode generates subsequent tokens with repeated access to weights and a growing KV cache. Small-batch decode can be dominated by memory movement; batching may improve reuse but can add queueing and affect time to first token.

Measure the actual prompt and output lengths, batch policy, concurrency and quality target. Track time to first token, inter-token latency, completed tokens per second and peak memory together. Neither high device utilization nor maximum throughput alone proves the service meets its latency objective.

Several GPUs introduce communication and placement

One device is no longer enough for the chosen configuration. Communication and placement enter the story alongside arithmetic and memory capacity.

Tensor or pipeline parallelism partitions model work across devices. The partitions exchange activations or reductions over available links. PCIe, NVLink and network fabrics have different topologies and constraints. A fast link on a specification sheet does not establish the bandwidth available between the particular devices assigned to a job.

Adding the VRAM labels on several cards does not automatically create one transparently usable memory pool. The runtime must partition or address data appropriately, and communication can become the limiting resource. Compare the complete serving configuration, including CPU, network and failure recovery, not only the GPU count.

Choose hardware from the workload outward

The proposed fleet expansion returns to workload evidence; the avoided rental hours stayed conditional on representative replay.

Start with memory fit and the required precision, then identify the operations and data movement. Estimate physical compute and bandwidth bounds, profile a representative run, and test the actual service-level target. A serial, branch-heavy or very small computation may remain a better CPU workload.

For procurement, carry the measured batch, quality and latency conditions into every provider comparison. If the code is the bottleneck, continue to the kernel-development guide. If the model waits on storage, CPU preprocessing or network, buying more GPU arithmetic may leave the problem intact.

Questions behind the decision

Why is GPU utilization not enough for capacity planning?

A busy device can still spend time on inefficient memory access, small kernels or work that misses the user’s objective. Pair device activity with useful throughput, latency, quality and the full host-to-device path.

Do Tensor Cores speed up every GPU workload?

They accelerate supported matrix operations with particular shapes and numerical formats. A memory-bound operator, unsupported shape or CPU bottleneck can remain unchanged even when the GPU has a higher advertised Tensor Core peak.

References & further reading

Engineering notes

Keep following the thread.

Real client engagements and the engineering behind them.

Technical glossary: definitions, connected ideas and further reading.

Optional analytics off. Contact works either way.

How measurement works