Kubernetes workloads / Interactive field guide
Will the workload stay available while it changes?
A controller wants replicas. A scheduler needs capacity. A Service needs ready endpoints. A disruption budget answers a different question again.
Current example
The requested surge fits this snapshot
One GPU slot per Pod; no existing surge or in-flight disruptions. Rollout headroom is not a controller deletion count. The separate PDB limits voluntary eviction, not controller updates or unplanned failure. Allocation does not establish model readiness; CPU, RAM, affinity and quota are outside this snapshot.
Initial snapshot: desired 3, ready 3, max surge 1, max unavailable 1, one free GPU slot, PDB minimum available 2, and one requested voluntary eviction.
02 / Follow the flow
Four boundaries. One connected explanation.
Each card explains a node in the diagram. Highlights show which boundaries contribute to the current result.
Choose a controller from the recovery contract
Deployments manage replaceable replicas and rollout strategy. StatefulSets add stable identities and ordered behavior subject to their configuration. Jobs pursue completion rather than a permanent serving replica count. Those responsibilities are not interchangeable. The arithmetic here is a simplified Deployment snapshot, not a simulation of all three controllers or an instruction to delete a particular pod.
A free accelerator is only one placement constraint
Requests, extended GPU resources, taints, affinity, topology and quotas determine whether a pod can be placed. A resource allocation does not reserve future token capacity or prove that a model loaded. The surge fit calculation deliberately counts abstract one-GPU slots; investigate actual scheduling events and all requested resources before treating an apparently free device as usable capacity.
A running process is not necessarily a serving endpoint
Startup, readiness and liveness checks answer different questions. Startup gates can protect long initialization from premature liveness checks; readiness controls endpoint eligibility. A process-level health response alone is not proof that the required model, dependency or data snapshot is usable. Graceful termination also needs admission to stop and accepted work to drain within the declared deadline.
Disruption
Do not substitute a PDB for rollout or failure planning
A PodDisruptionBudget constrains supported voluntary evictions. It does not stop a controller from following its rollout policy, create spare nodes or prevent an unplanned machine failure. The reader separates readiness-based rollout headroom from readiness-based eviction headroom. In production, consult current controller and PDB status, pending disruptions, readiness transitions and recovery capacity rather than acting on these simplified formulas alone.
# Snapshot equations; PDB and rollout are independent
rollout_min = max(0, desired - maxUnavailable)
rollout_room = max(0, ready - rollout_min)
surge_fit = min(maxSurge, freeSlots)
surge_pending = max(0, maxSurge - freeSlots)
pdb_room = max(0, ready - pdbMinAvailable)
evictions_allowed = evictions <= pdb_room03 / Keep the model honest
Model assumptions
Rollout minimum ready = max(0, desired − maxUnavailable)
Modeled removal headroom = max(0, ready − rollout minimum)
Surge fit = min(maxSurge, free slots)
Pending surge = max(0, maxSurge − free slots)
PDB headroom = max(0, ready − PDB minimum)
Requested evictions fit when request ≤ PDB headroom- Absolute integer counts, not percentages. The model assumes no existing surge pods, in-flight disruptions, availability delay or terminating endpoints.
- The rollout result is numerical headroom, not a prediction of ReplicaSet selection, ordering or the number of pods a real controller will delete.
- A free slot represents one pod requesting one GPU. Other scheduling constraints and engine-level memory admission remain separate.
- Both maxSurge and maxUnavailable set to zero is not a valid Deployment RollingUpdate configuration. The equations can still be explored, but do not constitute configuration validation.
- A PDB does not protect against every outage or controller update. A modeled zero-eviction request is a no-op, not evidence that a workload is healthy.
- No Kubernetes API is contacted. The snapshot is self-contained and remains readable without JavaScript.
04 / Think it through
Questions behind the example.
Why can a rollout stall with no free GPU slot even when all existing replicas are ready?
How does voluntary-eviction headroom differ from rollout removal headroom when a PDB minimum exceeds the Ready count?
Which telemetry distinguishes slow startup, a bad readiness probe, an unhealthy dependency and insufficient capacity?
Primary documentation

