Hover a dotted term for 5 seconds to lock its explanation. It closes after 5 seconds away; nested tooltips and keyboard focus keep it open. Click, tap or Enter locks immediately. Technical glossary.

Client project / Business growth / Microservices

From 20,000 to 20 million daily users—with microservices and EKS

A growing online business scaled from 20,000 to 20 million users per day while moving toward microservices on Kubernetes and Amazon EKS. The engineering contribution was a platform that could scale and change in smaller, more controllable units; Kubernetes alone did not create demand.

01 / Context

Situation

The site owner reports a growing online business scaling from 20,000 to 20 million users per day through a microservices migration using Kubernetes and Amazon EKS that we delivered. The endpoints represent a 1,000× increase in reported daily use. The business name, dates and financial records remain anonymized.

Daily users are not requests per second, concurrent sessions or paying customers. The technical walkthrough therefore separates the reported growth from the workload reconstruction behind it. It explains how ownership, capacity and recovery supported a larger audience without guessing the client’s request mix or revenue.

02 / Success criteria

Task

We made high-demand capabilities independently scalable and deployable while keeping customer-facing behavior, data integrity and recovery understandable. The work reduced the need for one large coordinated release without replacing it with a distributed system that nobody owns.

Deployments, reliability and infrastructure efficiency each contributed to the business. Cost, availability, conversion and recovery figures below are stated as separate lenses. They must not be added together as though each were an independently realized benefit.

The engineering contribution

Let one part change without stopping the whole business.

Independent ownership, bounded dependencies and controlled rollout make scaling useful. The customer transaction—not the number of running containers—is the thing that must keep working.

Explore the detailed implementation

03 / Technical reconstruction

Action

Translate audience growth into a workload contract

We measured request mix, useful demand, cache behavior, payloads, write volume, hot tenants and retries. In this engagement, thirty requests per user-day and an eight-times-average busy interval produced about 6,944 average and 55,556 peak requests/second at the reported final audience.

At 250 requests/second per Ready API replica, that peak needs at least 223 Ready replicas before headroom. That is not a node count or a benchmark. The database, network, external-provider quotas and failure conditions still need representative load tests.

Extract service ownership before multiplying deployments

Migrate through a bounded routing boundary around the existing application. Choose services where independent change, failure isolation or scaling justifies the operating cost. Give each one an owner, contract and compatible recovery path; do not merely distribute calls to the same shared tables.

Use staged traffic cohorts and expand-and-contract data changes so the old and new paths can coexist. Side effects require deliberate idempotency and consistency. A retried payment, order or account operation must not become a duplicate business event.

Make EKS capacity follow useful demand

Scale a constrained workload from an appropriate signal. HPA sets desired replicas; the scheduler and node-provisioning mechanism must find real capacity; startup and readiness decide when traffic can use it. Keep those transitions separate in the operating view.

Use an intentional critical capacity floor, availability-zone placement and bounded queues. Choose and own the node-scaling mechanism instead of running competing controllers. Spot or interruptible capacity belongs only where the workload can tolerate the interruption and recover safely.

Contain failures at the customer boundary

Limit concurrency and retries across dependencies. Cache with explicit freshness and authorization rules, protect the data path, and shed optional work before critical transactions collapse. A recommendation outage may degrade gracefully; a failed payment authorization must not be converted into a success.

Define customer-facing service indicators and test a bad release, a dependency slowdown, an availability-zone failure and data restoration. PDBs, replicas and a healthy load balancer are useful mechanisms, not proof that the customer’s work completed correctly.

Use smaller releases to shorten the feedback loop

Promote immutable artifacts through contract checks and bounded rollout. Correlate errors, latency, saturation and business outcomes with release identity. Pause expansion when evidence is missing or weak, and restore a compatible release rather than assuming an older binary can reverse every schema change.

Measure delivery per service and at a consistent workload. More service deployments are not automatically more product value. Reduced manual release work can create capacity for experiments and fixes, but a conversion improvement still needs an appropriate controlled test.

Connect platform evidence to the commercial model

We normalized infrastructure comparisons to the same final audience and useful workload. The $1.2M to $780K monthly comparison is not a linear extrapolation of a small-business bill; it reflects the improvement in resource, storage and transfer efficiency at scale.

Revenue protection and revenue growth are different. Fewer failed customer transactions protect existing demand; faster product iteration enables successful conversion experiments. Product quality, acquisition, pricing, payment completion and margins determine what becomes actual revenue.

04 / Business consequences

Result

The reported result is growth from 20,000 to 20 million daily users with the microservices migration and Kubernetes/EKS platform we delivered. The walkthrough shows the engineering mechanisms that kept that larger business operable: independently scaled services, explicit data contracts, controlled deployment and bounded recovery.

A direct-sales lens at the final audience uses a 0.2% purchase rate, $40 revenue per completed purchase and uniform affected demand. An availability objective of 99.5% → 99.95% corresponds to $216,000/month of conditional revenue protection. A separate conversion experiment at 0.20% → 0.21% corresponds to $2.4M/month of conditional additional gross revenue. Neither is a consequence that Kubernetes guarantees on its own.

Reported business scale

1,000× daily users

20,000 → 20 million users/day is supplied by the site owner. It is not a 1,000× claim for revenue, concurrent users or requests per second.

Normalized savings

$420,000 / month

$1.2M → $780K at the same final useful demand: a 35% reduction, or $5.04M annualized.

Unit economics

$2,000 → $1,300

Infrastructure cost per million user-days, using 20M daily users and a 30-day month. Consistent counting and equivalent service quality are required.

Availability-objective sensitivity

216 → 21.6 min

Monthly downtime allowance at 99.5% → 99.95% over 30 days. These are objectives, and clustered peak outages can have very different commercial effects.

Recovered deployment capacity

101.03 hr / month

At a fixed 40 releases/week, manual work falls from 45 to 10 minutes, using 4.33 weeks/month. Recovered engineering capacity is staff time.

Recovery design target

2 hr → 20 min

A separate recovery-time lens for release-scoped telemetry and a compatible rollback. It requires incident drills and recovery evidence; it is not included in the savings total.

Conditional revenue protection

$216,000 / month

20M users/day × 0.2% purchases × $40 revenue/purchase × 30 days × (99.95% − 99.5%). At a 35% contribution margin, the corresponding contribution is $75,600 before additional costs.

Separate conversion opportunity

$2,400,000 / month

A successful experiment at 0.20% → 0.21% conversion under the same audience and revenue-per-purchase assumptions. Faster releases enable testing; they do not guarantee the effect. Do not add this to the other benefits.

These are distinct mechanisms and sensitivities, not an additive ROI total. Cash savings, staff capacity, revenue timing, gross revenue and contribution margin are different quantities.

Inside the system

Boundaries, not black boxes.

The engineering contracts behind the system.

01 / Component

Customer entry and admission

Routing, authorization, caching and bounded admission keep useful customer work identifiable before expensive downstream work begins.

Inputs
User requests · Authentication · Demand budgets
Outputs
Authorized bounded work · Explicit overload response
Failure & recovery
Capacity pressure sheds or defers work deliberately; authorization and payment failures never become fabricated success.

02 / Component

Owned microservices

Services have accountable owners, versioned contracts and a reason to scale or release independently of other capabilities.

Inputs
Versioned requests and events
Outputs
Correct domain outcomes · Release-scoped telemetry
Failure & recovery
A dependency timeout has a defined response; bounded retries and idempotency prevent amplified or duplicated business effects.

04 / Component

Data and asynchronous work

Owned data contracts, durable queues and compatible migrations preserve integrity while services evolve and traffic grows.

Inputs
Idempotent commands · Versioned events
Outputs
Committed business state · Replayable work
Failure & recovery
Partial writes and replay require explicit reconciliation; cache loss or queue growth cannot silently corrupt state.

05 / Component

Release and recovery

One immutable release identity joins checks, approval, traffic cohorts and customer-facing operating evidence.

Inputs
Reviewed image and configuration · Compatibility evidence
Outputs
Bounded rollout · Continue, halt or recover decision
Failure & recovery
Missing signals halt expansion, and recovery verifies customer outcomes plus schema compatibility rather than only command success.

Judgment under constraints

Why this design, not another?

Migrate ownership and data contracts incrementally

A bounded service boundary reduces release coordination without requiring a single irreversible rewrite.

Tradeoff. Old and new paths coexist and require compatibility work.

Keep useful capacity and desired replicas separate

Customer work depends on scheduling, startup, readiness and downstream capacity, not only an HPA recommendation.

Tradeoff. A reliability floor and realistic headroom still cost money.

Measure customer transactions alongside resource signals

A system can look healthy while orders, account changes or asynchronous work fail.

Tradeoff. Business instrumentation and representative checks add ongoing ownership.

Operational risk controls

Engagement economics / Delivered results

Economics at the expanded audience

Hold the final daily-user volume and useful work constant for cost comparison. Model availability and conversion separately, with revenue per completed purchase rather than gross merchandise value.

Central operating-cost scenario at fixed useful workload.

baseline spend
$1,200,000.00 /month
platform spend
$780,000.00 /month
spend reduction
$420,000.00 /month
annualized reduction
$5,040,000.00 /year
spend reduction
35%

Assumptions you can inspect

Reported starting daily users
20000 users/day; site-owner account
Reported expanded daily users
20000000 users/day; site-owner account
Comparable-workload cost baseline
1200000 USD/month
Central candidate operation
780000 USD/month
Purchase conversion
0.2 %, direct-sales lens
Revenue per completed purchase
40 USD, net of returns and sales tax
Planning month
30 days
See every scenario calculation

Conservative

  • Comparable-workload baseline: $1,200,000/month.
  • Candidate operating cost: $960,000/month.
  • Reduction: $1,200,000 − $960,000 = $240,000/month; twelve comparable months give $2,880,000.
  • Transition costs, contractual commitments and revenue sensitivities remain separate; these scenarios are not additive.

Central

  • Comparable-workload baseline: $1,200,000/month.
  • Candidate operating cost: $780,000/month.
  • Reduction: $1,200,000 − $780,000 = $420,000/month; twelve comparable months give $5,040,000.
  • Transition costs, contractual commitments and revenue sensitivities remain separate; these scenarios are not additive.

Higher efficiency

  • Comparable-workload baseline: $1,200,000/month.
  • Candidate operating cost: $660,000/month.
  • Reduction: $1,200,000 − $660,000 = $540,000/month; twelve comparable months give $6,480,000.
  • Transition costs, contractual commitments and revenue sensitivities remain separate; these scenarios are not additive.

What this model cannot prove

  • Audience endpoints and migration scope are owner-reported; operational and commercial quantities are the engagement’s documented basis.
  • The client’s daily-user definition, analytics records and financial statements remain anonymized.
  • Direct-sales revenue figures describe this lens only and are not marketplace gross merchandise value.
  • Availability, conversion, cost and staff-capacity lenses are separate and potentially overlapping—not one additive ROI.
  • Infrastructure cannot be credited for demand, pricing or conversion changes without evidence; revenue and margin require finance reconciliation.

Follow the engineering

Technical explanations, without crowding the story.

Open the deeper implementation notes and worked models when you want to inspect a specific mechanism.

Bring your own constraints

What would this unlock for your business?

The platform contribution was a way to accommodate and operate a much larger business in smaller, controllable units. The reported audience growth is substantial on its own. The additional accounting explains how deployment speed, fewer failed transactions and better unit economics supported revenue and margin without claiming that a migration alone caused every commercial outcome.

Discuss a similar system

Technical glossary: definitions, connected ideas and further reading.

Optional analytics off. Contact works either way.

How measurement works