Hover a dotted term for 5 seconds to lock its explanation. It closes after 5 seconds away; nested tooltips and keyboard focus keep it open. Click, tap or Enter locks immediately. Technical glossary.

Platform engineering

Kubernetes Deployments: Readiness, Rollouts and Production Risk

A real client engagement. The engineering and the results are described below.

An ordering platform we worked with has a release dashboard full of green checks while customers wait for capacity that is not ready. The engineering team follows the gap between desired replicas and useful service before the next commercial launch.

Request a time through the inquiry form. A meeting is confirmed separately by email.

The business problem behind the technology

What does a green deployment actually prove to the business?

The next product launch needs a repeatable path to confirm serving capacity and recover incompatible changes. A faster dashboard update is not a business outcome.

Read the client engagement ↓

Client engagement / Delivered results

The release meeting was waiting on the wrong green light

An ordering platform we worked with spent 90 minutes investigating and recovering a problematic release. Teams repeatedly confuse scheduling, startup, readiness and supported eviction with the same condition.

The constraint

The next product launch needs a repeatable path to confirm serving capacity and recover incompatible changes. A faster dashboard update is not a business outcome.

The engineering decision

The team narrows probe contracts, budgets rollout overlap separately from voluntary eviction, rehearses draining and verifies a compatible rollback. Those checks reduced the next comparable release investigation-and-recovery exercise to 30 minutes.

The delivered outcome

The exercise saved 60 elapsed minutes per problematic release.

Measured inputs and delivered differences
Measure / unitBeforeAfterDifference
Comparable release investigation and recovery
minutes/release
903060

Comparable release investigation and recovery. These are the team’s exercise times under the same injected fault.

The conditions behind the results

  • The before/after exercise has the same injected fault and recovery criteria.
  • The 90/30-minute times come from the engagement; they are not calculated by the seven-input rollout widget.
  • No lost orders, revenue, headcount or incident-probability reduction is inferred.

What this does not prove. The shorter recovery exercise is the team’s rehearsal result under the same injected fault.

Evidence to collect for your own decision

  • Check successful requests at the application boundary after rollout and failure.
  • Rehearse startup delay, loss of capacity and supported voluntary eviction separately.
  • Time recovery with compatible application, configuration and data state.

Key decisions

A replica count expresses intent, not completed work. Follow the contracts between workload controllers, scheduling, probes, admission, and recovery, then use a small snapshot model to distinguish rollout headroom from a PodDisruptionBudget.

  • Check successful requests at the application boundary after rollout and failure.
  • Rehearse startup delay, loss of capacity and supported voluntary eviction separately.
  • A shorter modeled recovery exercise is not an uptime guarantee, prevented revenue loss or measured production improvement.

Follow the decision

What does a green deployment actually prove to the business?

Select a step to follow its reasoning, then continue into the technical chapters.

Problem → boundary → decision → evidence

Declare responsibility

Choose the workload controller and resource contract deliberately.

Read every component and connection
Declare responsibility · Problem
Choose the workload controller and resource contract deliberately.
Reach readiness · Boundary
Separate placement, startup and application readiness.
Change safely · Decision
Treat rollout overlap and voluntary eviction as different budgets.
Verify recovery · Evidence
Retain a reversible release and proof at each boundary.
  • Declare responsibility → Reach readiness: identify the constraint
  • Reach readiness → Change safely: choose a bounded change
  • Change safely → Verify recovery: check the outcome

A conceptual decision map for this article, not a measured timeline, physical topology or a depiction of a specific client system.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Choose the controller by the responsibility it owns

The client’s ordering team first identifies who owns each step between desired replicas and useful service.

A Deployment maintains interchangeable replicas through ReplicaSets and manages replacement when its Pod template changes. It suits a service whose requests can move among replicas without relying on a particular Pod identity. Scaling changes the desired replica count; changing the template starts a rollout. The controller does not know whether a response is correct or a business transaction committed. Those are application contracts, exposed through appropriate readiness and operational evidence.

A StatefulSet adds stable ordinal identity and a relationship between replacement Pods and their persistent storage. Its default OrderedReady policy also orders creation and scaling; policy and update settings matter. Choose it when identity, storage association, or ordering is required, not merely because a process keeps data in memory. A StatefulSet is not a database replication protocol, a backup, or proof of quorum. Storage provisioning, fencing, data replication, recovery, and compatibility remain separate responsibilities.

A Job runs work toward completion, with explicit completion, parallelism, and failure policies. It does not maintain a continuously ready serving pool. A successful container exit contributes to Job completion, but only the application can make that exit mean the required output was durably committed. Retry limits bound attempts; they do not provide exactly-once side effects. Use a Deployment for long-lived queue consumers when that is the intended service lifecycle, or Jobs for bounded tasks with an explicit completion contract.

Pending, startup, and Ready answer different questions

The status column becomes the first clue. Pending, startup and Ready each describe a different point in the journey.

Pending is a Pod phase, not a diagnosis of insufficient GPUs. It includes waiting for scheduling and time spent preparing containers, including pulling images. First distinguish an unbound Pod with PodScheduled false from a bound Pod whose container is Waiting. Scheduler events may identify insufficient resources or incompatible placement; container waiting reasons can instead identify an image, mount, or configuration problem. Adding nodes cannot repair a bad image reference.

Startup is an application initialization interval, not a Kubernetes Pod phase. A process may already be Running while it loads weights, reconstructs indexes, warms caches, or connects to dependencies. A startup probe gives that initialization its own bounded failure policy. Ready is a Pod condition derived from container readiness and any configured readiness gates. A scheduled GPU, a running process, and a successful TCP connection do not establish that the application can accept useful work.

Deployment availability can be stricter than instantaneous Ready: minReadySeconds requires a new Pod to remain ready without a container crash for a specified period. The reader deliberately treats its ready count as available capacity by assuming minReadySeconds is zero and excluding terminating Pods. In an actual rollout, record readyReplicas and availableReplicas separately; do not feed the first into every availability decision without that qualification.

Requests place work; limits constrain its execution

The scheduler sees requests while the process experiences contention and limits. Those views must describe a compatible workload.

The scheduler uses declared resource requests and node allocatable capacity, not a graph of momentary CPU or memory utilization. A low utilization sample does not create schedulable capacity already reserved by other Pods. Understating requests may pack Pods densely while leaving them competing during initialization or peak demand. Account for sidecars, init-container requirements, and Pod overhead where applicable, rather than measuring only the main process.

On Linux, CPU limits constrain execution through throttling; memory limits can lead to reactive OOM termination. A CPU request is not a CPU limit, and the host memory request is not a reservation of accelerator HBM. Specify units deliberately: 1 CPU equals 1,000 millicpu, while Gi denotes binary gibibytes. Select host resource values using cold-start and representative load observations. The values in the instructional fragment below are assumptions, not validated sizing recommendations.

In the traditional GPU device-plugin path, drivers, a compatible container runtime, and the NVIDIA plugin must already make a resource such as nvidia.com/gpu available. A GPU limit alone also supplies the request; if both are present, they must match. A GPU request without a limit is not the supported pattern. These extended-resource quantities are integer allocation units, unlike fractional CPU requests. Allocation does not certify application compatibility, usable model memory, or throughput.

A free GPU slot still needs eligible placement

A GPU appears available, yet placement still has constraints. You follow the whole resource contract rather than one free-device count.

A toleration permits a Pod to pass the matching taint restriction; it neither forces the Pod onto that node nor supplies capacity. Required node affinity filters eligible nodes, while preferred affinity expresses a preference. GPU product labels, driver and runtime compatibility, storage attachment, and network locality may make two nominally free devices non-interchangeable. Keep a capacity inventory tied to the actual resource name and required node properties.

Topology spread constraints and Pod anti-affinity can keep replicas from sharing a failure domain, but they also reduce the placements available during a rollout or node drain. DoNotSchedule is a hard spread constraint; ScheduleAnyway expresses a placement preference. Check the actual node labels and eligible domains. A globally free GPU in the wrong zone cannot satisfy a hard requirement, and three replicas on one node do not tolerate that node failing.

The meaning of an NVIDIA allocation unit depends on plugin configuration. MIG strategies can advertise profile-specific resources; time-slicing can advertise multiple shared accesses to a device. NVIDIA documents that time-sliced clients do not gain separate memory or fault isolation. Do not count those shared accesses as independent exclusive GPUs or promise proportional throughput. The reader assumes one exclusive, interchangeable GPU slot per Pod and does not model MIG profiles, sharing, interconnects, multi-Pod groups, or gang scheduling.

Give startup, readiness, and liveness narrow contracts

The application starts answering probes. Each probe now needs a narrow meaning that will remain useful during a failure.

A startup probe suppresses readiness and liveness checks until it succeeds. Use it to allow a bounded cold initialization interval, not to conceal initialization that never completes. A readiness probe answers whether the replica should receive new ordinary Service traffic. Failure makes it unready without itself restarting the container. A liveness probe detects a failure for which restarting is a useful remedy, such as an unrecoverable local deadlock.

Do not use a dependency outage or a full request queue as an automatic liveness failure. Restarting every warm replica during overload discards useful state and can amplify the outage. Dependency readiness should reflect whether this replica can actually serve its contract, not recursively fail because any optional downstream system is degraded. Keep checks cheap and bounded; an expensive inference executed on every probe adds its own load.

The following is only a field fragment for an existing container, not a deployable manifest. The custom application implements /readyz on port 8000, returning success only after initialization and a bounded warm-up have completed and while it is not draining. /livez on the same port reports local process progress independently of queue saturation. These are declared endpoint semantics, not built-in Kubernetes or inference-engine endpoints. The startup failure allowance is roughly 60 checks at ten-second intervals, not an exact deadline.

yaml / example
# Fields inside an existing application container; custom endpoints per the contract above.
resources:
  requests:
    cpu: "2"
    memory: "8Gi"
  limits:
    cpu: "4"
    memory: "12Gi"
    nvidia.com/gpu: 1
startupProbe:
  httpGet:
    path: /readyz
    port: 8000
  periodSeconds: 10
  timeoutSeconds: 2
  failureThreshold: 60
readinessProbe:
  httpGet:
    path: /readyz
    port: 8000
  periodSeconds: 5
  timeoutSeconds: 2
  failureThreshold: 2
livenessProbe:
  httpGet:
    path: /livez
    port: 8000
  periodSeconds: 10
  timeoutSeconds: 2
  failureThreshold: 3

A rollout spends replacement capacity, not a PDB

The Kubernetes deployment begins replacing replicas. Temporary overlap consumes capacity before the new version can carry traffic.

For a RollingUpdate Deployment, maxSurge allows extra replicas above the desired count and maxUnavailable bounds how many desired replicas may be unavailable during replacement. The two cannot both be zero in a valid rolling-update strategy. Kubernetes also accepts percentages, rounding surge up and unavailable down; the reader intentionally accepts only absolute integer counts. This strategy does not apply unchanged to a StatefulSet or Job.

The fragment below belongs directly under an existing Deployment spec. With three desired replicas, maxSurge one, and maxUnavailable one, the availability floor is two and the extra-replica allowance is one. A replacement still needs eligible CPU, memory, storage, and GPU capacity and must pass readiness. If there is no free GPU and maxUnavailable is zero, a surge-based replacement can stall because no old capacity may be taken away to free a slot. A failed or slow startup creates a different stall after placement.

Surge is not a strict cap on all resource consumption during shutdown: terminating Pods can retain resources while replacement Pods are created. Account for the termination interval as well as cold-start overlap. A Deployment progress deadline reports stalled progress; it does not automatically restore the last good application. Availability policy, error handling, and an approved recovery action must be designed together.

yaml / example
# Fields directly under an existing Deployment spec; not a manifest.
replicas: 3
minReadySeconds: 0
strategy:
  type: RollingUpdate
  rollingUpdate:
    maxSurge: 1
    maxUnavailable: 1

PDBs constrain voluntary eviction through a separate boundary

A node needs maintenance, and a different budget applies. The eviction path cannot be inferred from the rollout settings.

A PodDisruptionBudget describes tolerated disruption for the Pods selected by its label selector. Maintenance tooling that uses the Eviction API, such as a normal node drain, can be refused until enough selected Pods are healthy. A minAvailable of two with three healthy selected replicas illustrates one healthy voluntary eviction of headroom. Verify the selector and workload ownership: an arithmetic minimum attached to the wrong set of Pods protects the wrong contract.

A PDB does not constrain a Deployment or StatefulSet controller performing its own rolling update. Those controllers use their workload-specific update policies. Direct Pod or controller deletion also bypasses the PDB. Unplanned node failures and node-pressure evictions cannot be prevented by a PDB, although unavailable Pods consume the healthy capacity used to calculate the remaining budget. A green PDB therefore does not promise uninterrupted service.

Real eviction decisions use current PDB status, including observed generation, healthy counts, disrupted Pods, and the unhealthyPodEvictionPolicy. Multiple overlapping budgets or disruptions already in flight require additional care. A pending replacement can hold up a drain even when a rollout setting looked permissive. The reader isolates a single budget, assumes proposed evictions target healthy selected Pods, and deliberately does not emulate these API decisions.

Read all seven widget inputs as one declared snapshot

The release exercise now separates rollout overlap from eviction protection so its recovery claim has a clear boundary.

The Kubernetes workloads reader is an arithmetic teaching model implemented with WebAssembly, not a Kubernetes scheduling simulator. Its fixed example inputs are declared explicitly, not discovered from a cluster. The model has no timeline, existing surge Pods, in-flight disruptions or competing workloads. Its seven parameters are integer counts; no percentage rounding is performed.

The inputs remain independent. It is useful to explore a surplus ready count or a PDB minimum that exceeds current readiness rather than silently repair those inputs. The allowed arithmetic domain also includes maxSurge and maxUnavailable both zero; that combination is useful for showing zero change allowance but is not a valid RollingUpdate strategy to copy into Kubernetes. A zero-eviction proposal consumes no budget even when availability is already below the PDB minimum.

01 / Follow the explanation

Will the workload stay available while it changes?

A brief visual sequence plays automatically. The example and its assumptions are already here—nothing to configure.

An interactive teaching model drawn from real delivery work. No agent, cloud account, GPU or cluster is accessed.

Read the full explanation and assumptions

Workload responsibilities: controller intent, scheduler capacity, readiness and disruption handling.Separate budgets, not one health score. Follow the boundary that prevents useful capacity from becoming available.01Controller02Scheduler03Readiness04Disruption
Separate budgets, not one health score. Follow the boundary that prevents useful capacity from becoming available.

Current example

The requested surge fits this snapshot

One GPU slot per Pod; no existing surge or in-flight disruptions. Rollout headroom is not a controller deletion count. The separate PDB limits voluntary eviction, not controller updates or unplanned failure. Allocation does not establish model readiness; CPU, RAM, affinity and quota are outside this snapshot.

Initial snapshot: desired 3, ready 3, max surge 1, max unavailable 1, one free GPU slot, PDB minimum available 2, and one requested voluntary eviction.

Replica health

Desired replicas
3
Currently ready replicas
3
Rollout minimum ready
2 Pods
Rollout removal headroom
1 Pods

Rollout capacity

Maximum surge replicas
1
Maximum unavailable replicas
1
Free one-GPU pod slots
1
Surge that fits
1 Pods
Surge without a slot
0 Pods

Voluntary disruption

PDB minimum available
2
Requested voluntary evictions
1
Voluntary eviction headroom
1 Pods
Requested voluntary evictions
Within this PDB budget
Instructional excerpt, not executed here
# Snapshot equations; PDB and rollout are independent
rollout_min = max(0, desired - maxUnavailable)
rollout_room = max(0, ready - rollout_min)
surge_fit = min(maxSurge, freeSlots)
surge_pending = max(0, maxSurge - freeSlots)
pdb_room = max(0, ready - pdbMinAvailable)
evictions_allowed = evictions <= pdb_room
  • desired: 1–32 intended serving replicas, corresponding to the Deployment replica target in this engagement; default 3.
  • ready: 0–64 observed ready, nonterminating replicas selected for this exercise; default 3. Treat them as available only under the stated minReadySeconds = 0 assumption.
  • maxSurge: 0–32 extra replicas permitted for an attempted replacement wave; default 1. It is an allowance, not the number a controller will necessarily create.
  • maxUnavailable: 0–32 desired replicas permitted to be unavailable in the arithmetic rollout floor; default 1. Values above desired yield a floor of zero.
  • freeSlots: 0–32 currently free, interchangeable exclusive GPU allocation units; default 1. One new Pod consumes one slot, with all other placement constraints satisfied.
  • pdbMinAvailable: 0–32 healthy replicas required by the separate PDB; default 2. This does not change the rollout floor.
  • evictions: 0–32 proposed healthy voluntary evictions against that PDB snapshot; default 1. This is a proposal count, not a command or an observed drain result.

Derive the six outputs without turning them into promises

The calculated outputs are ready to interpret. You keep their assumptions visible instead of turning headroom into an availability promise.

The outputs are counts except for the final yes/no comparison. They describe independent arithmetic boundaries: changing the PDB minimum cannot alter rollout headroom, and adding a free GPU slot cannot turn an existing unready Pod into a ready one. The formulas below are the full model. No measured latency, queue length, failure probability, or utilization enters them.

With the defaults, minReady is 2, removalHeadroom is 1, surgeFit is 1, pendingSurge is 0, pdbHeadroom is 1, and the one-eviction proposal passes the comparison. Set freeSlots to zero and only the immediate surge fit changes to zero while pendingSurge becomes one. Set pdbMinAvailable to three and PDB headroom becomes zero while the rollout arithmetic still permits one removal. That disagreement is the lesson, not an inconsistency to hide.

These calculations exclude CPU and RAM fit, quota and admission failures, affinity, taints, storage, topology, multiple GPU resource types, startup duration, readiness changes, minReadySeconds delays, termination overlap, actual ReplicaSet state, and concurrent disruptions. The removal headroom is not an actual controller deletion count. Pending surge is only an immediate slot shortfall for the stated wave, not a prediction that a particular Pod will enter Pending or how long it would remain there. Real evidence must come from the relevant controllers and application.

  • Rollout minimum, minReady = max(0, desired - maxUnavailable): the available-replica floor, not a guarantee that failures cannot push capacity below it.
  • Rollout removal headroom, removalHeadroom = max(0, ready - minReady): currently ready replicas above that floor, not a deletion plan or a PDB allowance.
  • Surge fit, surgeFit = min(maxSurge, freeSlots): how many of the allowed extra one-slot Pods fit in the immediate GPU-only wave.
  • Pending surge, pendingSurge = max(0, maxSurge - freeSlots): the allowed extra Pods without an immediate GPU slot under these assumptions.
  • PDB headroom, pdbHeadroom = max(0, ready - pdbMinAvailable): healthy voluntary-eviction headroom for the independent budget.
  • Eviction comparison, evictionsAllowed = 1 if evictions <= pdbHeadroom, otherwise 0: whether the entire proposal fits the arithmetic budget, displayed as yes or no rather than an API authorization.
text / example
minReady = max(0, desired - maxUnavailable)
removalHeadroom = max(0, ready - minReady)
surgeFit = min(maxSurge, freeSlots)
pendingSurge = max(0, maxSurge - freeSlots)
pdbHeadroom = max(0, ready - pdbMinAvailable)
evictionsAllowed = evictions <= pdbHeadroom ? 1 : 0

Admission, HPA, and node capacity run different control loops

Other controllers enter the story. Admission, HPA and node provisioning must not be mistaken for the same control loop.

Application admission decides whether new work is accepted, queued within a bound, or rejected with a documented retry policy. Queue age, deadlines, cancellation, and per-tenant concurrency should remain bounded while scaling reacts. This is not the same as Kubernetes API admission, which can reject or modify a Pod specification before scheduling. Neither layer is replaced by declaring more replicas.

An HPA adjusts the scale of a supported workload from observed metrics; it does not reserve nodes or guarantee new replicas become ready before a deadline. CPU utilization targets are relative to CPU requests, not limits. Queue metrics can describe demand more directly for some consumers, but need an appropriate metrics adapter, meaningful units, and a tested target. Observe the time from demand through replica intent, scheduling, initialization, and readiness before expecting a scale-out response to rescue a short traffic burst.

Node capacity is another prerequisite. More desired Pods can remain unschedulable until an appropriately configured node provisioning system can supply eligible capacity, subject to provider availability, quota, labels, and startup time. Bound HPA growth and retain intentional warm headroom where the response-time requirement demands it. A deeper queue, a larger replica target, and more nodes are three different interventions, each with its own cost and delay.

Drain work before the grace period drains away

The old process receives a termination signal. Draining useful work now matters more than a fast exit on a chart.

A graceful application shutdown should stop admitting new work, report a draining readiness state, finish or safely hand off existing work, flush required durable state, and exit within its termination budget. The main process must receive and handle the configured stop signal; on ordinary Linux container runtimes this is commonly SIGTERM unless the image specifies another signal. A shell wrapper that does not forward signals can defeat a correct handler in the child process.

Kubernetes starts the termination grace-period clock before running a preStop hook. Hook time therefore consumes the same budget as application shutdown; it is not a bonus interval before SIGTERM. A hook should coordinate real draining when needed, not assume an arbitrary delay proves every load balancer has stopped sending traffic. Once the budget expires, remaining processes can be forcibly killed. Choose the budget from bounded work and shutdown observations, with explicit behavior for requests that cannot finish.

EndpointSlice termination state and local shutdown progress are asynchronous. Terminating endpoints are marked not ready for ordinary routing, but existing connections, external load balancers, caches, and clients may behave differently. Readiness removal does not complete a stream or make an interrupted write safe to retry. Exercise the actual routing path, connection lifetime, cancellation handling, and drain behavior rather than claiming zero dropped requests from a YAML field alone.

Durable work must survive a new Pod identity

A replacement Pod has a new identity. Durable state and work ownership need to survive that ordinary lifecycle event.

A replacement Pod has a new UID, even when a StatefulSet gives it the same name. A checkpoint stored only in a container filesystem or emptyDir is not durable across Pod replacement. Persist enough state to a storage system with the required durability and access semantics, and verify that a new process can read and validate it. A mounted persistent volume is one component of this design, not evidence that application state was flushed, is consistent, or has a recoverable backup.

For resumable computation, define checkpoint contents, version, input identity, integrity checks, and an atomic publication rule. Write a complete candidate before making it the latest checkpoint. Bound the interval of work that can be lost, and test recovery from a partial write and an incompatible checkpoint. Shutdown checkpointing alone is insufficient because an abrupt node failure may give the process no shutdown opportunity.

Kubernetes explicitly warns that a Job program can sometimes be started twice even with parallelism one, completions one, and restartPolicy Never. Make externally visible effects idempotent using a durable work identifier and a transactional commit or conditional write appropriate to the destination. For queue work, acknowledge only after the durable outcome is established and handle lease expiry and redelivery. Define how duplicates, incomplete output, stale locks, and exhausted retry budgets become visible; a successful retry must not silently duplicate a charge or an output artifact.

A reversible GitOps change restores a compatible contract

The release needs a way back. Reverting a tag alone is not enough when the surrounding contract has changed.

Review the immutable image and artifact references, resource requests, placement, probes, replica ownership, rollout policy, and routing together. Retain the prior artifacts and enough capacity to run them. If an HPA owns replica count, configure the GitOps field-ownership policy so continuous reconciliation does not fight scaling. Reversibility is not merely keeping an old YAML file: the old executable must still understand the current API, storage schema, queue messages, and checkpoint format.

For an automatically reconciled application, a reviewed Git revision restoring the previous compatible desired state is usually clearer than an untracked live patch. Argo CD documents that its direct application rollback is unavailable while automated sync is enabled. Follow the selected controller’s recovery procedure rather than assuming an imperative Deployment rollback will remain in place. A Git revert cannot un-send a response, undo external side effects, or reverse an incompatible data migration.

Before a rollout, state stop conditions and who can act on them: rising errors, queue-age breaches, unavailable replicas, initialization failures, and stalled progress need bounded responses. A PDB is not that decision-maker. Separate restoring serving configuration from repairing durable work or reconciling partial side effects, and retain evidence linking the observed state to the reviewed source revision.

Require observable proof at each boundary

The team times a comparable recovery exercise and checks real requests, not a green object alone.

For a real deployment, record controller generation and observed generation, old and new ReplicaSet counts, ready and available replicas, and rollout conditions. Pair scheduler events and node allocatable resources with PodScheduled, container waiting reasons, startup failures, restart counts, and readiness transitions. Inspect PDB status separately from Deployment status. An unready application on a successfully allocated GPU is evidence about initialization or serving, not proof that the scheduler failed.

Application evidence should include accepted, rejected, canceled, and completed work, queue depth and oldest age, latency by request class, duplicate or failed task outcomes, and shutdown completion. Correlate that evidence with the release revision and a bounded observation interval. A successful rollout condition without successful application work is incomplete proof; an aggregate utilization chart cannot establish durability or drain safety.

An authorized rehearsal should distinguish no eligible slot, a slow startup, readiness loss without restart, a genuinely stuck process, a voluntary drain blocked by a PDB, and an abrupt node loss that bypasses graceful cleanup. Verify checkpoint recovery and duplicate-work handling separately. These are proposed validation scenarios, not claims that they have been executed here. The companion widget demonstrates only the six declared equations and cannot substitute for any of this cluster or application evidence.

Questions behind the decision

What is the difference between readiness and liveness probes?

Readiness decides whether a Pod should receive traffic; liveness can trigger a restart when its declared condition fails. Conflating transient dependency trouble with process failure can amplify an incident instead of restoring useful capacity.

Does a PodDisruptionBudget prevent HPA scale-down?

No. A PDB constrains supported voluntary eviction paths; it is not a minimum-replica setting for ordinary HPA or Deployment changes. Configure and test replica floors, rollout strategy and eviction policy at their separate boundaries.

References & further reading

Engineering notes

Keep following the thread.

Real client engagements and the engineering behind them.

Technical glossary: definitions, connected ideas and further reading.

Optional analytics off. Contact works either way.

How measurement works