Hover a dotted term for 5 seconds to lock its explanation. It closes after 5 seconds away; nested tooltips and keyboard focus keep it open. Click, tap or Enter locks immediately. Technical glossary.

Kubernetes workloads / Interactive field guide

Will the workload stay available while it changes?

A controller wants replicas. A scheduler needs capacity. A Service needs ready endpoints. A disruption budget answers a different question again.

4 linked boundariesLocal WASM modelNo account required

01 / Follow the explanation

The system, step by step.

A brief visual sequence plays automatically. The example and its assumptions are already here—nothing to configure.

An interactive teaching model drawn from real delivery work. No agent, cloud account, GPU or cluster is accessed.

Workload responsibilities: controller intent, scheduler capacity, readiness and disruption handling.Separate budgets, not one health score. Follow the boundary that prevents useful capacity from becoming available.01Controller02Scheduler03Readiness04Disruption
Separate budgets, not one health score. Follow the boundary that prevents useful capacity from becoming available.

Current example

The requested surge fits this snapshot

One GPU slot per Pod; no existing surge or in-flight disruptions. Rollout headroom is not a controller deletion count. The separate PDB limits voluntary eviction, not controller updates or unplanned failure. Allocation does not establish model readiness; CPU, RAM, affinity and quota are outside this snapshot.

Initial snapshot: desired 3, ready 3, max surge 1, max unavailable 1, one free GPU slot, PDB minimum available 2, and one requested voluntary eviction.

Replica health

Desired replicas
3
Currently ready replicas
3
Rollout minimum ready
2 Pods
Rollout removal headroom
1 Pods

Rollout capacity

Maximum surge replicas
1
Maximum unavailable replicas
1
Free one-GPU pod slots
1
Surge that fits
1 Pods
Surge without a slot
0 Pods

Voluntary disruption

PDB minimum available
2
Requested voluntary evictions
1
Voluntary eviction headroom
1 Pods
Requested voluntary evictions
Within this PDB budget

02 / Follow the flow

Four boundaries. One connected explanation.

Each card explains a node in the diagram. Highlights show which boundaries contribute to the current result.

Controller

Choose a controller from the recovery contract

Deployments manage replaceable replicas and rollout strategy. StatefulSets add stable identities and ordered behavior subject to their configuration. Jobs pursue completion rather than a permanent serving replica count. Those responsibilities are not interchangeable. The arithmetic here is a simplified Deployment snapshot, not a simulation of all three controllers or an instruction to delete a particular pod.

Scheduler

A free accelerator is only one placement constraint

Requests, extended GPU resources, taints, affinity, topology and quotas determine whether a pod can be placed. A resource allocation does not reserve future token capacity or prove that a model loaded. The surge fit calculation deliberately counts abstract one-GPU slots; investigate actual scheduling events and all requested resources before treating an apparently free device as usable capacity.

Readiness

A running process is not necessarily a serving endpoint

Startup, readiness and liveness checks answer different questions. Startup gates can protect long initialization from premature liveness checks; readiness controls endpoint eligibility. A process-level health response alone is not proof that the required model, dependency or data snapshot is usable. Graceful termination also needs admission to stop and accepted work to drain within the declared deadline.

Disruption

Do not substitute a PDB for rollout or failure planning

A PodDisruptionBudget constrains supported voluntary evictions. It does not stop a controller from following its rollout policy, create spare nodes or prevent an unplanned machine failure. The reader separates readiness-based rollout headroom from readiness-based eviction headroom. In production, consult current controller and PDB status, pending disruptions, readiness transitions and recovery capacity rather than acting on these simplified formulas alone.

Instructional excerpt, not executed here
# Snapshot equations; PDB and rollout are independent
rollout_min = max(0, desired - maxUnavailable)
rollout_room = max(0, ready - rollout_min)
surge_fit = min(maxSurge, freeSlots)
surge_pending = max(0, maxSurge - freeSlots)
pdb_room = max(0, ready - pdbMinAvailable)
evictions_allowed = evictions <= pdb_room

03 / Keep the model honest

Model assumptions

Transparent reasoning / Units and boundaries
Rollout minimum ready = max(0, desired − maxUnavailable)
Modeled removal headroom = max(0, ready − rollout minimum)
Surge fit = min(maxSurge, free slots)
Pending surge = max(0, maxSurge − free slots)
PDB headroom = max(0, ready − PDB minimum)
Requested evictions fit when request ≤ PDB headroom
  • Absolute integer counts, not percentages. The model assumes no existing surge pods, in-flight disruptions, availability delay or terminating endpoints.
  • The rollout result is numerical headroom, not a prediction of ReplicaSet selection, ordering or the number of pods a real controller will delete.
  • A free slot represents one pod requesting one GPU. Other scheduling constraints and engine-level memory admission remain separate.
  • Both maxSurge and maxUnavailable set to zero is not a valid Deployment RollingUpdate configuration. The equations can still be explored, but do not constitute configuration validation.
  • A PDB does not protect against every outage or controller update. A modeled zero-eviction request is a no-op, not evidence that a workload is healthy.
  • No Kubernetes API is contacted. The snapshot is self-contained and remains readable without JavaScript.

04 / Think it through

Questions behind the example.

  1. Why can a rollout stall with no free GPU slot even when all existing replicas are ready?

  2. How does voluntary-eviction headroom differ from rollout removal headroom when a PDB minimum exceeds the Ready count?

  3. Which telemetry distinguishes slow startup, a bad readiness probe, an unhealthy dependency and insufficient capacity?

Primary documentation

References & further reading

Engineering notes

Read the project behind the model.

Real client engagements and the engineering behind them.

A useful next conversation

What needs to work better?

A system, a delivery bottleneck, or an engineering opportunity. Tell me what you are building and where you want to go.

Let’s talk

Technical glossary: definitions, connected ideas and further reading.

Optional analytics off. Contact works either way.

How measurement works