Hover a dotted term for 5 seconds to lock its explanation. It closes after 5 seconds away; nested tooltips and keyboard focus keep it open. Click, tap or Enter locks immediately. Technical glossary.

Kernel development / Interactive field guide

Where does a GPU kernel spend its time?

A faster arithmetic unit cannot remove a data-movement bottleneck. Follow the data and see what tiling can—and cannot—buy.

4 linked boundariesLocal teaching modelNo account required

01 / Follow the explanation

The system, step by step.

A brief visual sequence plays automatically. The example and its assumptions are already here—nothing to configure.

An interactive teaching model drawn from real delivery work. No agent, cloud account, GPU or cluster is accessed.

Matrix multiplication path: global memory, shared-memory tile, multiply and accumulate, output store.Follow the memory-to-math path. Highlighting identifies modeled constraints, not a profiler trace. Each node links to its explanation.01DRAM02Shared tile03Math04Output
Follow the memory-to-math path. Highlighting identifies modeled constraints, not a profiler trace. Each node links to its explanation.

Current example

Memory traffic sets this lower bound

An optimistic roofline lower bound under approximate uncached tiled Float32 traffic, not predicted latency or measured speedup. The larger of compute time and memory time wins; they are not added. Cache reuse, occupancy, register pressure, instruction mix, launch and synchronization costs are omitted. GB/s and TFLOP/s use decimal units.

Combined lower bound (ms)
0.5411 ms

Example: 1024 × 1024 square matrices, float32 elements, a 16 × 16 tile, 1,000 GB/s bandwidth and 60 TFLOP/s compute. Hardware rates are model inputs, not B200 or TPU specifications.

Shape & reuse

Matrix dimension N
1024
Shared-memory tile
16 × 16
Approximate DRAM traffic (GB)
0.541 GB
Shared storage per block (KiB)
2.000 KiB

Effective ceilings

Compute (TFLOP/s)
60
Memory lower bound (ms)
0.5411 ms
Compute lower bound (ms)
0.0358 ms

02 / Follow the flow

Four boundaries. One connected explanation.

Each card explains a node in the diagram. Highlights show which boundaries contribute to the current result.

DRAM

Load contiguous data, not scattered wishes

Neighboring threads should request useful adjacent data so the memory system can combine accesses. This model counts requested A and B values, not actual memory transactions. Caches may satisfy repeated reads, while misalignment, striding and partially used transactions may increase the traffic observed by a profiler. Establish a row-major indexing contract before comparing layouts; a transposed input is a different access pattern.

Shared tile

Pay for a tile once; reuse it within a block

A T × T output tile repeatedly loads one A tile and one B tile into shared memory. Threads synchronize before consuming those values and before another iteration overwrites them. The teaching approximation reduces repeated input loads by T. Real performance also depends on shared-memory bank conflicts, barriers, register allocation and occupancy. A 32 × 32 one-output-per-thread block has 1,024 threads; legal on suitable devices does not mean optimal.

Math

Count the arithmetic actually executed

Square dense matrix multiplication performs approximately 2N³ floating-point operations when a multiply-add counts as two operations. The compute rate must describe the same precision and execution path as the kernel. A simple float32 CUDA teaching kernel does not automatically use Tensor Cores. Instruction overhead and dependency chains are excluded here, so the compute-time term is a ceiling-derived lower bound, not a measurement.

Output

Keep the result correct before keeping it fast

The model writes N² output elements once and does not read an existing C matrix. An alpha-AB-plus-beta-C contract with nonzero beta changes the traffic. Compare against a trusted implementation with absolute and relative tolerances, including zeros, mixed signs and awkward dimensions in the actual project. The slider deliberately uses divisible sizes; the companion kernel article explains masking the edges that this simplified model does not simulate.

03 / Keep the model honest

Model assumptions

Transparent reasoning / Units and boundaries
FLOPs ≈ 2 × N³
Traffic bytes ≈ (2 × N³ / tile + N²) × 4
Arithmetic intensity = FLOPs / traffic bytes
Compute ms = FLOPs / (TFLOP/s × 10¹²) × 10³
Memory ms = traffic bytes / (GB/s × 10⁹) × 10³
Optimistic lower bound = max(compute ms, memory ms)
  • Float32 A, B and C; square dimensions; one output write; no beta-C read. The untiled baseline does not allocate a shared-memory tile.
  • The traffic approximation requires ideal reuse inside each tile and no reuse between tiles. It is not a simulation of the cache hierarchy or DRAM transactions.
  • Compute and memory can overlap in this idealized roofline bound. Adding the two terms would describe a different execution model.
  • Launches, barriers, instruction overhead, occupancy, register spilling, thermal behavior and communication are excluded. The result is not predicted latency or a speedup claim.
  • No GPU code executes in this reader. The worked example is calculated during site generation; all diagrams and explanations remain readable without JavaScript.

04 / Think it through

Questions behind the example.

  1. Why does a 16 × 16 tile change the traffic term without changing the matrix multiplication FLOP count?

  2. Why does a memory-bound calculation stop benefiting from a higher theoretical compute rate?

  3. What occupancy and synchronization evidence would be needed before choosing a larger tile in a real kernel?

Primary documentation

References & further reading

Engineering notes

Read the project behind the model.

Real client engagements and the engineering behind them.

A useful next conversation

What needs to work better?

A system, a delivery bottleneck, or an engineering opportunity. Tell me what you are building and where you want to go.

Let’s talk

Technical glossary: definitions, connected ideas and further reading.

Optional analytics off. Contact works either way.

How measurement works