Kernel development / Interactive field guide
Where does a GPU kernel spend its time?
A faster arithmetic unit cannot remove a data-movement bottleneck. Follow the data and see what tiling can—and cannot—buy.
Current example
Memory traffic sets this lower bound
An optimistic roofline lower bound under approximate uncached tiled Float32 traffic, not predicted latency or measured speedup. The larger of compute time and memory time wins; they are not added. Cache reuse, occupancy, register pressure, instruction mix, launch and synchronization costs are omitted. GB/s and TFLOP/s use decimal units.
- Combined lower bound (ms)
- 0.5411 ms
Example: 1024 × 1024 square matrices, float32 elements, a 16 × 16 tile, 1,000 GB/s bandwidth and 60 TFLOP/s compute. Hardware rates are model inputs, not B200 or TPU specifications.
Shape & reuse
- Matrix dimension N
- 1024
- Shared-memory tile
- 16 × 16
- Arithmetic intensity (FLOP/B)
- 3.969 FLOP/B
- Shared storage per block (KiB)
- 2.000 KiB
Effective ceilings
- Compute (TFLOP/s)
- 60
- Memory lower bound (ms)
- 0.5411 ms
- Compute lower bound (ms)
- 0.0358 ms
02 / Follow the flow
Four boundaries. One connected explanation.
Each card explains a node in the diagram. Highlights show which boundaries contribute to the current result.
Load contiguous data, not scattered wishes
Neighboring threads should request useful adjacent data so the memory system can combine accesses. This model counts requested A and B values, not actual memory transactions. Caches may satisfy repeated reads, while misalignment, striding and partially used transactions may increase the traffic observed by a profiler. Establish a row-major indexing contract before comparing layouts; a transposed input is a different access pattern.
Shared tile
Pay for a tile once; reuse it within a block
A T × T output tile repeatedly loads one A tile and one B tile into shared memory. Threads synchronize before consuming those values and before another iteration overwrites them. The teaching approximation reduces repeated input loads by T. Real performance also depends on shared-memory bank conflicts, barriers, register allocation and occupancy. A 32 × 32 one-output-per-thread block has 1,024 threads; legal on suitable devices does not mean optimal.
Math
Count the arithmetic actually executed
Square dense matrix multiplication performs approximately 2N³ floating-point operations when a multiply-add counts as two operations. The compute rate must describe the same precision and execution path as the kernel. A simple float32 CUDA teaching kernel does not automatically use Tensor Cores. Instruction overhead and dependency chains are excluded here, so the compute-time term is a ceiling-derived lower bound, not a measurement.
Output
Keep the result correct before keeping it fast
The model writes N² output elements once and does not read an existing C matrix. An alpha-AB-plus-beta-C contract with nonzero beta changes the traffic. Compare against a trusted implementation with absolute and relative tolerances, including zeros, mixed signs and awkward dimensions in the actual project. The slider deliberately uses divisible sizes; the companion kernel article explains masking the edges that this simplified model does not simulate.
03 / Keep the model honest
Model assumptions
FLOPs ≈ 2 × N³
Traffic bytes ≈ (2 × N³ / tile + N²) × 4
Arithmetic intensity = FLOPs / traffic bytes
Compute ms = FLOPs / (TFLOP/s × 10¹²) × 10³
Memory ms = traffic bytes / (GB/s × 10⁹) × 10³
Optimistic lower bound = max(compute ms, memory ms)- Float32 A, B and C; square dimensions; one output write; no beta-C read. The untiled baseline does not allocate a shared-memory tile.
- The traffic approximation requires ideal reuse inside each tile and no reuse between tiles. It is not a simulation of the cache hierarchy or DRAM transactions.
- Compute and memory can overlap in this idealized roofline bound. Adding the two terms would describe a different execution model.
- Launches, barriers, instruction overhead, occupancy, register spilling, thermal behavior and communication are excluded. The result is not predicted latency or a speedup claim.
- No GPU code executes in this reader. The worked example is calculated during site generation; all diagrams and explanations remain readable without JavaScript.
04 / Think it through
Questions behind the example.
Why does a 16 × 16 tile change the traffic term without changing the matrix multiplication FLOP count?
Why does a memory-bound calculation stop benefiting from a higher theoretical compute rate?
What occupancy and synchronization evidence would be needed before choosing a larger tile in a real kernel?
Primary documentation

