Hover a dotted term for 5 seconds to lock its explanation. It closes after 5 seconds away; nested tooltips and keyboard focus keep it open. Click, tap or Enter locks immediately. Technical glossary.

Kubernetes autoscaling / Interactive field guide

From a busy microservice to ready Kubernetes capacity.

A traffic spike does not create a node directly. Follow one microservice through the control loops, then inspect the requests, policies and readiness conditions between a recommendation and usable capacity.

4 linked boundariesLocal WASM modelNo account required

01 / Follow the explanation

The system, step by step.

A brief visual sequence plays automatically. The example and its assumptions are already here—nothing to configure.

An interactive teaching model drawn from real delivery work. No agent, cloud account, GPU or cluster is accessed.

API: 3 → 6 Pods; 0 pending.A CPU-only target layout, not live infrastructure. Bars show requested / usable CPU. Flow pulses depict control messages, not measured requests. Proposed nodes add serving capacity only after provisioning, scheduling and readiness.MICROSERVICES / ONE SELECTED HPAGateway2 → 2 Pods250m request / PodAPI3 → 6 Pods500m request / PodWorkers2 → 2 Pods1000m request / PodHPA / policy before scale6 raw6 desired2 existing + 1 proposed3 nodes required by CPUNode 1Existing2/2 CPU req.Node 2Existing2/2 CPU req.Node 3Proposed1.5/2 CPU req.Node 4Existing0/2 CPU req.Node 5Existing0/2 CPU req.Node 6Existing0/2 CPU req.Node 7Existing0/2 CPU req.Node 8Existing0/2 CPU req.Node 9Existing0/2 CPU req.Node 10Existing0/2 CPU req.Node 11Existing0/2 CPU req.Node 12Existing0/2 CPU req.Node 13Existing0/2 CPU req.Node 14Existing0/2 CPU req.Node 15Existing0/2 CPU req.Node 16Existing0/2 CPU req.Node 17Existing0/2 CPU req.Node 18Existing0/2 CPU req.Node 19Existing0/2 CPU req.Node 20Existing0/2 CPU req.Ready is a separate gateAPI: 3 → 6 Pods; 0 pending.A CPU-only target layout, not live infrastructure. Bars show requested / usable CPU. Flow pulses depict control messages, not measured requests. Proposed nodes add serving capacity only after provisioning, scheduling and readiness.MICROSERVICES / ONE SELECTED HPAGateway2 → 2 Pods250m request / PodAPI3 → 6 Pods500m request / PodWorkers2 → 2 Pods1000m request / PodHPA / policy before scale6 raw6 desired2 existing + 1 proposedNode 1Existing2/2 CPU req.Node 2Existing2/2 CPU req.Node 3Proposed1.5/2 CPU req.Node 4Existing0/2 CPU req.Node 5Existing0/2 CPU req.Node 6Existing0/2 CPU req.Node 7Existing0/2 CPU req.Node 8Existing0/2 CPU req.Node 9Existing0/2 CPU req.Node 10Existing0/2 CPU req.Node 11Existing0/2 CPU req.Node 12Existing0/2 CPU req.Node 13Existing0/2 CPU req.Node 14Existing0/2 CPU req.Node 15Existing0/2 CPU req.Node 16Existing0/2 CPU req.Node 17Existing0/2 CPU req.Node 18Existing0/2 CPU req.Node 19Existing0/2 CPU req.Node 20Existing0/2 CPU req.Ready is a separate gate
  1. 01 / Observe

    Observe the selected microservice. Other services have independent scaling decisions.

  2. 02 / Recommend

    HPA computes a replica recommendation, then applies tolerance, bounds and behavior.

  3. 03 / Create Pods

    The scale target changes. Deployment and ReplicaSet controllers manage the Pods.

  4. 04 / Schedule

    The scheduler needs compatible capacity. Excess target Pods remain pending.

  5. 05 / Provision

    A node autoscaler may provision nodes; provider capacity and initialization still matter.

  6. 06 / Ready gate

    Only Ready Pods add serving capacity. Consolidation needs a separate safe drain decision.

A CPU-only target layout, not live infrastructure. Bars show requested / usable CPU. Flow pulses depict control messages, not measured requests. Proposed nodes add serving capacity only after provisioning, scheduling and readiness.

Current example

API: 3 → 6 desired Pods. Scale-up requested.

Desired replicas are not ready capacity. 0 target Pods still lack CPU capacity; 0 nodes are candidates for a separate safe drain decision.

Capacity at $2/node-hour
$4 → $6/hour

Initial example: three API Pods use 600m CPU each against a 500m request and a 60% target. The raw recommendation is six Pods. Two existing nodes have 2,000m usable CPU each; one additional node is required. The animation compresses conceptual stages, not real provisioning time.

Observe one microservice

Deployment controlled by this HPA
API service
Current Ready Pods
3
Observed CPU per Pod (millicores)
600
CPU request per Pod
500m CPU
CPU usage / request
120%
Raw recommendation
6

Apply HPA behavior

Target CPU utilization (%)
60
Maximum replicas
12
Highest prior recommendation in 5 minutes
3
Desired Pods
6
Policy decision
Scale-up requested

Find schedulable capacity

Usable CPU per node
2,000m after system reservations
Additional nodes permitted
2
CPU-only node count
2 existing / 3 required
Additional nodes
1
Unschedulable target Pods
0
Consolidation candidates
0

02 / Follow the flow

Four boundaries. One connected explanation.

Each card explains a node in the diagram. Highlights show which boundaries contribute to the current result.

CPU signal

Measure the right denominator

HPA reads a metric for its scale target. For CPU utilization, usage is compared with requested CPU across the relevant containers; CPU limits are not the denominator. This engagement used identical requests and complete samples from Ready Pods. Actual HPA conservatively handles missing samples and initializing Pods, discards failed or terminating Pods for resource metrics, and can use custom or external metrics instead. Queue deadlines often need a backlog signal rather than CPU.

HPA target

Normalize a recommendation before writing scale

For this complete metric population, raw replicas equal ceil(current replicas × current utilization / target utilization). The 10% tolerance avoids small changes. Minimum and maximum bounds, recent recommendation history and scaling policies are separate controls. Default upscale behavior allows the larger of four Pods or 100% over 15 seconds; this engagement had no earlier events in that period. Default downscale stabilization uses the highest recommendation in a 300-second window. HPA writes the scale subresource; Deployment and ReplicaSet controllers manage Pods.

Node fleet

Unschedulable Pods can trigger provisioning

The scheduler considers requests and placement constraints, not a promise from the HPA. A node autoscaler may provision capacity that fits pending Pods, subject to node policy, quota and provider availability. Cluster Autoscaler typically expands configured node groups; Karpenter can choose node configurations from NodePool constraints. The CPU-only diagram shows a packed target layout, not actual scheduler assignments. New nodes may still need initialization before Pods can start.

Ready path

Capacity is useful only after readiness

Desired, scheduled, Running and Ready are different states. Startup and readiness probes keep warming applications out of serving endpoints. After load falls, fewer replicas do not instantly reduce the bill: nodes must be eligible for consolidation, with feasible drains, placement and storage. PDBs can constrain voluntary node drains, but do not directly veto an HPA or Deployment replica reduction. Maintain service floors and verify latency, queue age and recovery before claiming an optimization.

03 / Keep the model honest

One observation, with its boundaries visible.

Transparent reasoning / Units and boundaries
utilization = observed_CPU / requested_CPU × 100
raw_replicas = ceil(current_replicas × utilization / target)
within_10_percent_tolerance → keep current recommendation
apply min=1, maxReplicas, recent downscale history
upscale cap = current + max(4,current), assuming no prior 15s events
required_nodes = ceil(total requested CPU / usable node CPU)
additional_nodes = min(permitted_extra, max(0,required−existing))
pending = desired Pods that still cannot fit
consolidation_candidates ≠ authorized node deletions
  • The HPA control loop normally runs periodically (15 seconds by default), not continuously. The brief visualization is a compressed explanation, not a clock or a cloud provisioning SLA.
  • Every selected Pod is initially Ready and has the same CPU request and observed usage. Actual missing or unready metrics dampen scaling; multiple metrics choose the largest recommendation and failed metrics can block downscaling.
  • The history assumption summarizes the highest recent recommendation. It does not simulate a rolling history buffer or earlier scale changes in the last 15 seconds.
  • Node capacity is usable after system reservations. Requests are powers-of-two multiples of 250m and divide the supported node sizes; the shown packing ignores all non-CPU constraints.
  • A missing CPU request reserves no CPU in this arithmetic but does not mean a Pod consumes no resources. It makes the HPA utilization metric unavailable.
  • Consolidation candidates remain billed in the displayed capacity cost until a separate drain/removal decision succeeds. The $2 per node-hour figure is the planning rate used in this engagement, not a provider price.
  • The small cluster does not generate the separate $2M portfolio bill. More traffic can increase instantaneous capacity cost even when good autoscaling reduces idle spend over a normalized month.

04 / Think it through

Questions behind the example.

  1. Why do three API Pods at 120% utilization and a 60% target produce six desired Pods, and why is one additional node required?

  2. Why would a zero extra-node allowance leave three target Pods unschedulable without changing the HPA recommendation?

  3. How does a larger CPU request change the utilization denominator without proving that the application is faster or cheaper?

  4. How does downscale stabilization differ from node consolidation and its disruption constraints?

  5. Why does one service’s HPA not solve a bottleneck in another microservice or dependency?

Primary documentation

References & further reading

Engineering notes

Read the project behind the model.

Real client engagements and the engineering behind them.

A useful next conversation

What needs to work better?

A system, a delivery bottleneck, or an engineering opportunity. Tell me what you are building and where you want to go.

Let’s talk

Technical glossary: definitions, connected ideas and further reading.

Optional analytics off. Contact works either way.

How measurement works