GPU architecture
How a GPU Works: CUDA, Memory and Better Capacity Decisions
A real client engagement. The engineering and the results are described below.
A product-catalog company we worked with prepares an AI search launch and receives a proposal to double its GPU allocation. Before buying, the engineering team follows a request through the CPU, GPU memory and execution units to discover what the hardware is actually waiting for.
Request a time through the inquiry form. A meeting is confirmed separately by email.
The business problem behind the technology
Does the product need more GPUs—or better-fed GPUs?
The team must keep search quality and latency while meeting the same nightly enrichment deadline. Buying devices before understanding CPU preparation, transfers and memory stalls could lock in the wrong expense.
Read the client engagement ↓Client engagement / Delivered results
The procurement question hidden inside a slow response
A product-catalog service we worked with supports a growing batch-enrichment workload alongside interactive search. Its first capacity proposal is based on device utilization, not completed work.
The constraint
The team must keep search quality and latency while meeting the same nightly enrichment deadline. Buying devices before understanding CPU preparation, transfers and memory stalls could lock in the wrong expense.
The engineering decision
It separates the host pipeline, memory traffic and execution phases, then requires a representative replay before approving extra allocation. Feeding the existing fleet more effectively avoided 4,000 additional rented GPU-hours each month.
The delivered outcome
At $4 per GPU-hour, the plan falls from 12,000 to 8,000 GPU-hours: $16,000 of monthly compute charges the client did not incur.
| Measure / unit | Before | After | Difference |
|---|---|---|---|
| Planned monthly GPU compute charge USD/month | 48,000 | 32,000 | 16,000 |
Planned monthly GPU compute charge. 12,000 × $4 versus 8,000 × $4 at the client’s contracted rate.
The conditions behind the results
- The same useful jobs complete at the same quality and latency objectives.
- The avoided 4,000 GPU-hours come from the engagement’s capacity plan, not from the diagrams.
- Only flexible compute charges are counted; storage, networking, support and engineering cost remain outside this comparison.
What this does not prove. The avoided spend required billable capacity to retire without losing useful service.
Evidence to collect for your own decision
Key decisions
A GPU is not simply a faster CPU. It combines many execution lanes, specialized arithmetic and a memory hierarchy designed to keep parallel work moving. Trace one workload through those layers and see why memory, batching and data layout often matter as much as arithmetic.
Follow the decision
Select a step to follow its reasoning, then continue into the technical chapters.
Problem → boundary → decision → evidence
The host, runtime and driver prepare a device operation.
Read every component and connection
- Submit work · Problem
- The host, runtime and driver prepare a device operation.
- Find the operands · Boundary
- Memory hierarchy and data layout determine which bytes must move.
- Execute instructions · Decision
- Schedulers, registers and arithmetic resources do different jobs.
- Return useful output · Evidence
- The application still has to meet its quality and latency requirements.
- Submit work → Find the operands: identify the constraint
- Find the operands → Execute instructions: choose a bounded change
- Execute instructions → Return useful output: check the outcome
A conceptual decision map for this article, not a measured timeline, physical topology or a depiction of a specific client system.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
Follow work from the CPU to the GPU
The client’s catalog team follows host work first, before approving another GPU allocation.
The CPU prepares buffers and submits commands through a runtime and driver. Data may move over PCIe or another supported link into device memory. A kernel launch describes many threads; GPU hardware schedules their work onto streaming multiprocessors, or SMs. The result stays on the device until a consumer needs it elsewhere.
Submission is commonly asynchronous. With the right streams, hardware and dependencies, transfers and computation can overlap. The diagram separates responsibilities, not four mandatory nonoverlapping time intervals. Small jobs can lose more time to launch and transfer overhead than they save in parallel execution.
Host
The CPU prepares inputs and submits work through the driver.
Device memory
HBM or other GPU memory holds tensors and intermediate data.
Consumer
A later kernel or host operation waits on the required dependency.
This is a dependency map; real transfers and kernels can overlap.
Software ownership, from cluster to silicon
Kubernetes places the workload. The container stack exposes a device. CUDA and the driver submit work. GPU scheduling and execution happen below that boundary.
Cluster lifecycle and placement
Application and user-space libraries
Kernel and device boundary
Observe the real service
The scheduler chooses a suitable node using declared resources and placement policy. It does not place individual CUDA blocks, schedule warps or make an unready model useful.
Read every component and connection
- Kubernetes scheduler · Node / resource placement
- The scheduler chooses a suitable node using declared resources and placement policy. It does not place individual CUDA blocks, schedule warps or make an unready model useful.
- Device plugin / kubelet · GPU allocation
- A device plugin advertises supported GPU resources and participates in allocation. Kubelet starts the pod with the chosen devices. GPU sharing changes the resource contract.
- GPU Operator · Lifecycle controller
- The operator can manage drivers, Container Toolkit, device plugin, feature discovery and monitoring. This is a control-plane responsibility, not a mandatory hop for each tensor.
- OCI runtime / Toolkit · Container device access
- The configured runtime and NVIDIA Container Toolkit expose the permitted devices and libraries. Containers share their host kernel; a VM or bare-metal host sets another boundary.
- Serving framework · Batching / model / NCCL
- A serving engine schedules requests and calls framework/library kernels. NCCL coordinates supported multi-GPU collectives. Application queueing is distinct from GPU instruction scheduling.
- CUDA runtime / driver · libcudart / libcuda
- CUDA APIs manage contexts, streams, memory and launches. The user-mode driver interfaces with kernel support. Compatible machine code or supported PTX compilation is required.
- Linux NVIDIA modules · Devices / memory / I/O
- Kernel modules cooperate with the OS for device access, memory mapping and control. Linux permissions and isolation still matter; a container image does not replace the host kernel driver.
- Queues / GPU firmware · Device work submission
- Supported driver and firmware paths submit and manage device work. Firmware-mediated responsibilities vary with platform and driver; this is not a proprietary command-protocol schematic.
- Kubernetes scheduler → Device plugin / kubelet: placement / allocation
- Device plugin / kubelet → OCI runtime / Toolkit: start with devices
- GPU Operator → Device plugin / kubelet: manage plugin
- GPU Operator → OCI runtime / Toolkit: manage toolkit
- GPU Operator → Linux NVIDIA modules: manage driver
- OCI runtime / Toolkit → Serving framework: run workload
- Serving framework → CUDA runtime / driver: library calls
- CUDA runtime / driver → Linux NVIDIA modules: driver interface
- Linux NVIDIA modules → Queues / GPU firmware: submit / manage
- Queues / GPU firmware → SMs / memory / ALUs: device work
- SMs / memory / ALUs → DCGM + application telemetry: device signals
- Serving framework → DCGM + application telemetry: service signals
A typical Linux/device-plugin deployment, not a universal platform configuration. GPU Operator manages components; it is not on every inference request’s data path. Driver, CUDA, framework and hardware compatibility must be validated together.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
An ALU result begins with gates and stored operands
Zoom in until a large computation becomes a few bits. The adder makes the distinction between combinational logic and stored state concrete.
CMOS transistors implement logic gates by controlling pull-up and pull-down paths. Signals take time to charge and discharge capacitance, and real devices also have leakage and other losses. A gate is combinational logic; a register stores state across clocked stages. Do not confuse a register file with off-chip DRAM or treat a drawn pulse as a measured transistor delay.
The full adder below is a working generic Boolean example: sum = A XOR B XOR carry-in, while carry-out combines generated and propagated carry. Change the three input bits and inspect the intermediate gates. Wider integer arithmetic builds on bit-level logic with implementation-specific carry networks and pipelines. Floating-point units add alignment, normalization, rounding and special-value behavior. This teaching circuit is not a claim about NVIDIA’s proprietary ALU layout.
Generic digital logic · not a NVIDIA die schematic
Sum = A XOR B XOR carry-in · carry-out = (A AND B) OR ((A XOR B) AND carry-in)
Input voltages interpreted as bits
First gate layer
Combine the partial results
Output bits captured by later storage
A logical 1 or 0 represents a voltage range, not an infinitely precise voltage. CMOS transistors charge and discharge capacitances when a gate switches.
Read every component and connection
- A · Input bit
- A logical 1 or 0 represents a voltage range, not an infinitely precise voltage. CMOS transistors charge and discharge capacitances when a gate switches.
- B · Input bit
- The second operand bit enters both the XOR and AND paths. Gates compute Boolean functions; a clocked register stores a result for a later pipeline stage.
- Carry in · Previous position
- Carry-in is the carry from a less-significant bit position. Wide adders must manage the carry dependency; this does not claim NVIDIA uses a ripple-carry implementation.
- XOR · A XOR B
- XOR is 1 when exactly one input is 1. This partial sum feeds both the final sum gate and the carry-propagation gate.
- AND · A AND B
- When A and B are both 1, this path generates a carry regardless of carry-in. AND can itself be built from other CMOS gate arrangements.
- XOR · Partial sum XOR carry
- The second XOR produces the least-significant result bit. For 1 + 0 + 1, this bit is 0 while carry-out is 1.
- AND · Partial sum AND carry
- This path propagates carry-in when exactly one operand bit is 1. It is distinct from the carry generated by A AND B.
- OR · Generated OR propagated
- OR combines generated and propagated carry. Logically equivalent NAND networks or other carry designs can implement the same truth function.
- Sum · Result bit
- The sum bit and carry bit together represent A + B + carry-in. They are a two-bit result, not two independent arithmetic answers.
- Carry out · Next bit position
- Carry-out contributes weight two to this one-bit addition. The complete truth table must work for all eight input combinations.
A locally calculated full-adder truth function. CMOS pull-up/pull-down transistor networks implement logic gates; real GPU adders use implementation-specific carry networks and pipelines. Animation speed is chosen for readability, not a propagation delay or a transistor simulation.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
Threads become warps, not independent little CPUs
Zoom back out to a warp. The active lanes in one instruction tell you something useful, but not the whole story of GPU utilization.
CUDA groups threads into blocks and 32-thread warps. Threads have their own registers and logical identity; instructions execute with an active-lane mask. Divergent branches can require different subsets of lanes to execute different paths. Modern scheduling details do not make communication between threads automatically safe: use the documented synchronization primitives.
In the example, 20 active lanes use 62.5% of the lanes for that instruction. That is not a claim of 62.5% GPU utilization. Another warp may run while this one waits, and instruction throughput depends on the operation, dependencies and hardware pipelines. Occupancy instead measures resident warps relative to the supported maximum.
Mask
Only the lanes participating in this instruction are active.
Repeat
The next instruction can have a different mask and bottleneck.
Change the active-lane count. The result describes one instruction, not a speedup.
Active lanes in this instruction: 20
active / 32 × 100. This shows one instruction mask, not occupancy, GPU utilization or a throughput prediction. Other architectures can use different group sizes.
Change the example assumptions
The worked result is readable without JavaScript. Inputs become available when the local WASM model loads; constrained connections and devices retain the static example.
Memory is a hierarchy with different sharing rules
The operands need a place to live. Registers, caches, shared memory and HBM are different resources, not interchangeable names for GPU memory.
Registers hold thread-local values. Shared memory lets threads in a block cooperate, subject to synchronization and capacity limits. Caches reduce some repeated accesses, while off-chip device memory holds much larger arrays. Names and capacities vary by architecture; do not turn a diagram into a claim about every GPU.
Capacity determines whether a workload fits; bandwidth determines how quickly bytes can move. A model that fits in VRAM can still be bandwidth-bound. Excessive register demand can spill into memory, increasing traffic. More shared memory per block can reduce how many blocks fit on an SM at once.
Conceptual Blackwell execution and memory map
Off-chip DRAM → memory controllers → caches → registers / execution → stores. Instruction issue and asynchronous copies coordinate different paths.
Large off-chip storage
On-chip reuse and staging
Inside one representative streaming multiprocessor
Device-wide work and peers
Stacked DRAM stores large device-resident arrays. DRAM cells need refresh; capacity and sustained bandwidth are different limits. HBM is not the register file.
Read every component and connection
- HBM3e DRAM · Weights / KV / arrays
- Stacked DRAM stores large device-resident arrays. DRAM cells need refresh; capacity and sustained bandwidth are different limits. HBM is not the register file.
- Memory controllers · Channels and requests
- Controllers organize reads and writes to memory channels. Access patterns, contention and the memory technology influence service time; bandwidth is not zero latency.
- L2 cache · Device-wide reuse
- L2 can satisfy repeated requests without another HBM access. Its capacity and residency behavior affect traffic; a cache hit is not a new DRAM transfer.
- L1 / shared memory · Caching / explicit tiles
- B200 combines L1, texture and shared-memory resources. Shared memory is software-managed block storage with synchronization rules; it is not an automatic replacement for registers.
- Register file · Thread operands
- The B200 tuning guide specifies 64K 32-bit registers per SM. Threads use registers for live values; spills can create device-memory traffic. Registers are not off-chip DRAM.
- ALU pipelines · Integer / floating point
- Execution pipelines perform supported arithmetic and logic on operands. CMOS gates underlie those circuits. Floating-point operations include more work than the integer full adder shown.
- Tensor cores · Matrix operations
- Specialized matrix instructions use supported operand formats and accumulation paths. Tensor throughput is not scalar ALU throughput, and not every kernel can use tensor cores.
- Load / store units · Addresses and movement
- Load/store machinery forms and services memory operations. Coalescing groups useful lane accesses; dependencies prevent a consumer from using a value before it is ready.
- Warp schedulers · Ready instruction issue
- Schedulers issue eligible warp instructions subject to dependencies and resource availability. Other ready warps can hide a wait; occupancy alone does not prove throughput.
- Instruction path · Fetch / decode / issue
- Compiled machine instructions reach the SM instruction machinery. PTX is a virtual ISA; a compatible cubin or driver compilation supplies hardware-executable code.
- Async copy / TMA · Tile movement
- Supported asynchronous transfer paths can stage tiles while computation proceeds. Barriers and producer/consumer ordering still apply; overlap is not permission to read unfinished data.
- GPU front end · Submitted work
- Device work submission and scheduling machinery distribute kernel work. Block resource requirements influence residency. Kubernetes does not choose a warp or allocate an SM register.
- NVLink interface · Peer devices
- Peer access and collectives move data between compatible GPUs. The application/runtime manages distributed work; aggregate device memory is not one automatically shared allocation.
- GPU front end → Instruction path: kernel work
- Instruction path → Warp schedulers: decoded instructions
- Warp schedulers → Load / store units: memory instruction
- HBM3e DRAM → Memory controllers: DRAM service
- Memory controllers → L2 cache: cache-line traffic
- L2 cache → L1 / shared memory: cache / tile path
- L1 / shared memory → Load / store units: load path
- Load / store units → Register file: operand load
- Register file → ALU pipelines: ALU operands
- ALU pipelines → Register file: result
- Register file → Load / store units: store
- L2 cache → Async copy / TMA: async tile copy
- Async copy / TMA → L1 / shared memory: staging
- L1 / shared memory → Tensor cores: supported matrix operands
- L2 cache → NVLink interface: peer traffic
This is a functional map, not a floorplan or cycle-accurate simulator. One representative SM is expanded; it is not the GPU’s SM count. Cache bypass, asynchronous copies, distributed shared memory and specialized tensor accumulator paths mean not every operation follows every arrow.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
| Resource | Role | Typical constraint |
|---|---|---|
| Registers | Per-thread live values | Register pressure and spilling |
| Shared memory | Explicit cooperation within a block | Capacity, bank access and synchronization |
| L1 / L2 caches | Reuse of memory accesses | Working set and access pattern |
| Device memory | Large tensors, weights and KV cache | Capacity and sustained bandwidth |
| PCIe / interconnect | Host-device or device-device exchange | Topology, latency and transfer volume |
Coalescing makes useful bytes cheaper to fetch
Operand movement is now a purchasing question: useful work, not the device label, must justify the bill.
Suppose 32 active lanes each read a neighboring four-byte value. They request 128 useful bytes in a compact aligned span. If those lanes instead read widely separated addresses, the memory system may need many more sectors to deliver the same useful data. Caching, alignment and the device generation determine the actual transactions.
Reusing a loaded tile can reduce repeated memory traffic; simply adding more threads does not. Count bytes at the relevant memory boundary and compare them with useful operations. Arithmetic intensity is operations per byte, not the size of the input file. A kernel can be limited at shared memory or an instruction pipeline even when HBM is not saturated.
Tensor cores accelerate specific matrix operations
A matrix operation takes another route through the chip. Tensor cores help only when the operation, format and software path can actually use them.
Tensor cores are specialized matrix multiply-accumulate hardware. They do not accelerate every branch, lookup or scalar instruction. Supported data types, layouts, dimensions and accumulation behavior depend on the architecture and software path. The application must actually select a compatible operation.
Precision changes require quality evaluation. FP32, TF32, BF16, FP16 and lower-precision formats are not interchangeable labels for the same arithmetic. Advertised sparse throughput also assumes a supported sparsity pattern; do not compare it with a dense workload as if both perform the same useful computation.
Inference changes character between prefill and decode
The model begins generating an answer, and the workload changes. Prefill, decode and queueing now need separate explanations.
During prefill, a model processes the input context, often exposing substantial matrix parallelism. Autoregressive decode generates subsequent tokens with repeated access to weights and a growing KV cache. Small-batch decode can be dominated by memory movement; batching may improve reuse but can add queueing and affect time to first token.
Measure the actual prompt and output lengths, batch policy, concurrency and quality target. Track time to first token, inter-token latency, completed tokens per second and peak memory together. Neither high device utilization nor maximum throughput alone proves the service meets its latency objective.
Several GPUs introduce communication and placement
One device is no longer enough for the chosen configuration. Communication and placement enter the story alongside arithmetic and memory capacity.
Tensor or pipeline parallelism partitions model work across devices. The partitions exchange activations or reductions over available links. PCIe, NVLink and network fabrics have different topologies and constraints. A fast link on a specification sheet does not establish the bandwidth available between the particular devices assigned to a job.
Adding the VRAM labels on several cards does not automatically create one transparently usable memory pool. The runtime must partition or address data appropriately, and communication can become the limiting resource. Compare the complete serving configuration, including CPU, network and failure recovery, not only the GPU count.
Choose hardware from the workload outward
The proposed fleet expansion returns to workload evidence; the avoided rental hours stayed conditional on representative replay.
Start with memory fit and the required precision, then identify the operations and data movement. Estimate physical compute and bandwidth bounds, profile a representative run, and test the actual service-level target. A serial, branch-heavy or very small computation may remain a better CPU workload.
For procurement, carry the measured batch, quality and latency conditions into every provider comparison. If the code is the bottleneck, continue to the kernel-development guide. If the model waits on storage, CPU preprocessing or network, buying more GPU arithmetic may leave the problem intact.
Questions behind the decision
Why is GPU utilization not enough for capacity planning?
A busy device can still spend time on inefficient memory access, small kernels or work that misses the user’s objective. Pair device activity with useful throughput, latency, quality and the full host-to-device path.
Do Tensor Cores speed up every GPU workload?
They accelerate supported matrix operations with particular shapes and numerical formats. A memory-bound operator, unsupported shape or CPU bottleneck can remain unchanged even when the GPU has a higher advertised Tensor Core peak.
