Client project / Business growth / Microservices
From 20,000 to 20 million daily users—with microservices and EKS
A growing online business scaled from 20,000 to 20 million users per day while moving toward microservices on Kubernetes and Amazon EKS. The engineering contribution was a platform that could scale and change in smaller, more controllable units; Kubernetes alone did not create demand.
01 / Context
Situation
The site owner reports a growing online business scaling from 20,000 to 20 million users per day through a microservices migration using Kubernetes and Amazon EKS that we delivered. The endpoints represent a 1,000× increase in reported daily use. The business name, dates and financial records remain anonymized.
Daily users are not requests per second, concurrent sessions or paying customers. The technical walkthrough therefore separates the reported growth from the workload reconstruction behind it. It explains how ownership, capacity and recovery supported a larger audience without guessing the client’s request mix or revenue.
02 / Success criteria
Task
We made high-demand capabilities independently scalable and deployable while keeping customer-facing behavior, data integrity and recovery understandable. The work reduced the need for one large coordinated release without replacing it with a distributed system that nobody owns.
Deployments, reliability and infrastructure efficiency each contributed to the business. Cost, availability, conversion and recovery figures below are stated as separate lenses. They must not be added together as though each were an independently realized benefit.
The engineering contribution
Let one part change without stopping the whole business.
Independent ownership, bounded dependencies and controlled rollout make scaling useful. The customer transaction—not the number of running containers—is the thing that must keep working.
Explore the detailed implementation03 / Technical reconstruction
Action
Translate audience growth into a workload contract
We measured request mix, useful demand, cache behavior, payloads, write volume, hot tenants and retries. In this engagement, thirty requests per user-day and an eight-times-average busy interval produced about 6,944 average and 55,556 peak requests/second at the reported final audience.
At 250 requests/second per Ready API replica, that peak needs at least 223 Ready replicas before headroom. That is not a node count or a benchmark. The database, network, external-provider quotas and failure conditions still need representative load tests.
Extract service ownership before multiplying deployments
Migrate through a bounded routing boundary around the existing application. Choose services where independent change, failure isolation or scaling justifies the operating cost. Give each one an owner, contract and compatible recovery path; do not merely distribute calls to the same shared tables.
Use staged traffic cohorts and expand-and-contract data changes so the old and new paths can coexist. Side effects require deliberate idempotency and consistency. A retried payment, order or account operation must not become a duplicate business event.
Make EKS capacity follow useful demand
Scale a constrained workload from an appropriate signal. HPA sets desired replicas; the scheduler and node-provisioning mechanism must find real capacity; startup and readiness decide when traffic can use it. Keep those transitions separate in the operating view.
Use an intentional critical capacity floor, availability-zone placement and bounded queues. Choose and own the node-scaling mechanism instead of running competing controllers. Spot or interruptible capacity belongs only where the workload can tolerate the interruption and recover safely.
Contain failures at the customer boundary
Limit concurrency and retries across dependencies. Cache with explicit freshness and authorization rules, protect the data path, and shed optional work before critical transactions collapse. A recommendation outage may degrade gracefully; a failed payment authorization must not be converted into a success.
Define customer-facing service indicators and test a bad release, a dependency slowdown, an availability-zone failure and data restoration. PDBs, replicas and a healthy load balancer are useful mechanisms, not proof that the customer’s work completed correctly.
Use smaller releases to shorten the feedback loop
Promote immutable artifacts through contract checks and bounded rollout. Correlate errors, latency, saturation and business outcomes with release identity. Pause expansion when evidence is missing or weak, and restore a compatible release rather than assuming an older binary can reverse every schema change.
Measure delivery per service and at a consistent workload. More service deployments are not automatically more product value. Reduced manual release work can create capacity for experiments and fixes, but a conversion improvement still needs an appropriate controlled test.
Connect platform evidence to the commercial model
We normalized infrastructure comparisons to the same final audience and useful workload. The $1.2M to $780K monthly comparison is not a linear extrapolation of a small-business bill; it reflects the improvement in resource, storage and transfer efficiency at scale.
Revenue protection and revenue growth are different. Fewer failed customer transactions protect existing demand; faster product iteration enables successful conversion experiments. Product quality, acquisition, pricing, payment completion and margins determine what becomes actual revenue.
04 / Business consequences
Result
The reported result is growth from 20,000 to 20 million daily users with the microservices migration and Kubernetes/EKS platform we delivered. The walkthrough shows the engineering mechanisms that kept that larger business operable: independently scaled services, explicit data contracts, controlled deployment and bounded recovery.
A direct-sales lens at the final audience uses a 0.2% purchase rate, $40 revenue per completed purchase and uniform affected demand. An availability objective of 99.5% → 99.95% corresponds to $216,000/month of conditional revenue protection. A separate conversion experiment at 0.20% → 0.21% corresponds to $2.4M/month of conditional additional gross revenue. Neither is a consequence that Kubernetes guarantees on its own.
Reported business scale
1,000× daily users
20,000 → 20 million users/day is supplied by the site owner. It is not a 1,000× claim for revenue, concurrent users or requests per second.
Normalized savings
$420,000 / month
$1.2M → $780K at the same final useful demand: a 35% reduction, or $5.04M annualized.
Unit economics
$2,000 → $1,300
Infrastructure cost per million user-days, using 20M daily users and a 30-day month. Consistent counting and equivalent service quality are required.
Availability-objective sensitivity
216 → 21.6 min
Monthly downtime allowance at 99.5% → 99.95% over 30 days. These are objectives, and clustered peak outages can have very different commercial effects.
Recovered deployment capacity
101.03 hr / month
At a fixed 40 releases/week, manual work falls from 45 to 10 minutes, using 4.33 weeks/month. Recovered engineering capacity is staff time.
Recovery design target
2 hr → 20 min
A separate recovery-time lens for release-scoped telemetry and a compatible rollback. It requires incident drills and recovery evidence; it is not included in the savings total.
Conditional revenue protection
$216,000 / month
20M users/day × 0.2% purchases × $40 revenue/purchase × 30 days × (99.95% − 99.5%). At a 35% contribution margin, the corresponding contribution is $75,600 before additional costs.
Separate conversion opportunity
$2,400,000 / month
A successful experiment at 0.20% → 0.21% conversion under the same audience and revenue-per-purchase assumptions. Faster releases enable testing; they do not guarantee the effect. Do not add this to the other benefits.
These are distinct mechanisms and sensitivities, not an additive ROI total. Cash savings, staff capacity, revenue timing, gross revenue and contribution margin are different quantities.
Inside the system
Boundaries, not black boxes.
The engineering contracts behind the system.
01 / Component
Customer entry and admission
Routing, authorization, caching and bounded admission keep useful customer work identifiable before expensive downstream work begins.
- Inputs
- User requests · Authentication · Demand budgets
- Outputs
- Authorized bounded work · Explicit overload response
- Failure & recovery
- Capacity pressure sheds or defers work deliberately; authorization and payment failures never become fabricated success.
02 / Component
Owned microservices
Services have accountable owners, versioned contracts and a reason to scale or release independently of other capabilities.
- Inputs
- Versioned requests and events
- Outputs
- Correct domain outcomes · Release-scoped telemetry
- Failure & recovery
- A dependency timeout has a defined response; bounded retries and idempotency prevent amplified or duplicated business effects.
03 / Component
EKS workload and node capacity
Workload scaling, resource scheduling, node provisioning and readiness are observed as separate stages of becoming useful capacity.
- Inputs
- Demand signals · Resource requests · Placement and quota
- Outputs
- Scheduled and Ready replicas
- Failure & recovery
- Desired replicas do not count as served traffic; insufficient quota, startup failure and unschedulable Pods remain visible.
04 / Component
Data and asynchronous work
Owned data contracts, durable queues and compatible migrations preserve integrity while services evolve and traffic grows.
- Inputs
- Idempotent commands · Versioned events
- Outputs
- Committed business state · Replayable work
- Failure & recovery
- Partial writes and replay require explicit reconciliation; cache loss or queue growth cannot silently corrupt state.
05 / Component
Release and recovery
One immutable release identity joins checks, approval, traffic cohorts and customer-facing operating evidence.
- Inputs
- Reviewed image and configuration · Compatibility evidence
- Outputs
- Bounded rollout · Continue, halt or recover decision
- Failure & recovery
- Missing signals halt expansion, and recovery verifies customer outcomes plus schema compatibility rather than only command success.
Judgment under constraints
Why this design, not another?
Migrate ownership and data contracts incrementally
A bounded service boundary reduces release coordination without requiring a single irreversible rewrite.
Tradeoff. Old and new paths coexist and require compatibility work.
Keep useful capacity and desired replicas separate
Customer work depends on scheduling, startup, readiness and downstream capacity, not only an HPA recommendation.
Tradeoff. A reliability floor and realistic headroom still cost money.
Measure customer transactions alongside resource signals
A system can look healthy while orders, account changes or asynchronous work fail.
Tradeoff. Business instrumentation and representative checks add ongoing ownership.
Operational risk controls
- Load-test the actual request mix and failure cases.
- Bound fan-out, retries, concurrency and queue growth.
- Use compatible schema migration and explicit idempotency.
- Rehearse zone, dependency, release and data-recovery failures.
- Review normalized cost and commercial attribution with finance.
Engagement economics / Delivered results
Economics at the expanded audience
Hold the final daily-user volume and useful work constant for cost comparison. Model availability and conversion separately, with revenue per completed purchase rather than gross merchandise value.
Central operating-cost scenario at fixed useful workload.
- baseline spend
- $1,200,000.00 /month
- platform spend
- $780,000.00 /month
- spend reduction
- $420,000.00 /month
- annualized reduction
- $5,040,000.00 /year
- spend reduction
- 35%
Assumptions you can inspect
- Reported starting daily users
- 20000 users/day; site-owner account
- Reported expanded daily users
- 20000000 users/day; site-owner account
- Comparable-workload cost baseline
- 1200000 USD/month
- Central candidate operation
- 780000 USD/month
- Purchase conversion
- 0.2 %, direct-sales lens
- Planning month
- 30 days
See every scenario calculation
Conservative
- Comparable-workload baseline: $1,200,000/month.
- Candidate operating cost: $960,000/month.
- Reduction: $1,200,000 − $960,000 = $240,000/month; twelve comparable months give $2,880,000.
- Transition costs, contractual commitments and revenue sensitivities remain separate; these scenarios are not additive.
Central
- Comparable-workload baseline: $1,200,000/month.
- Candidate operating cost: $780,000/month.
- Reduction: $1,200,000 − $780,000 = $420,000/month; twelve comparable months give $5,040,000.
- Transition costs, contractual commitments and revenue sensitivities remain separate; these scenarios are not additive.
Higher efficiency
- Comparable-workload baseline: $1,200,000/month.
- Candidate operating cost: $660,000/month.
- Reduction: $1,200,000 − $660,000 = $540,000/month; twelve comparable months give $6,480,000.
- Transition costs, contractual commitments and revenue sensitivities remain separate; these scenarios are not additive.
What this model cannot prove
- Audience endpoints and migration scope are owner-reported; operational and commercial quantities are the engagement’s documented basis.
- The client’s daily-user definition, analytics records and financial statements remain anonymized.
- Direct-sales revenue figures describe this lens only and are not marketplace gross merchandise value.
- Availability, conversion, cost and staff-capacity lenses are separate and potentially overlapping—not one additive ROI.
- Infrastructure cannot be credited for demand, pricing or conversion changes without evidence; revenue and margin require finance reconciliation.
Follow the engineering
Technical explanations, without crowding the story.
Open the deeper implementation notes and worked models when you want to inspect a specific mechanism.
Platform engineering
Kubernetes Deployments: Readiness, Rollouts and Production Risk
A real order platform follows Pods, probes, rollout capacity and PodDisruptionBudgets to shorten release recovery without inventing uptime guarantees.
Cloud / Distributed systems
Monolith to Microservices on EKS: Ownership, Migration and Rollback
A real commerce migration connects EKS, service boundaries and rollback to release effort. The linked owner-reported audience growth remains separate.
Cloud architecture
Kubernetes Cost Optimization: Rightsizing, FinOps and Unit Economics
A real order platform cut cloud spend through Kubernetes rightsizing and autoscaling, with cost per successful order and reliability held explicit.
SRE and observability
OpenTelemetry for Microservices: Tracing, Incidents and Cost Control
A real order service uses OpenTelemetry to speed diagnosis and lower telemetry spend, with sampling, correlation and privacy boundaries kept visible.
Platform engineering
Terraform and GitOps: Ownership, Safe Releases and Recovery
A real platform team lowered release coordination effort through Terraform state boundaries, immutable artifacts and GitOps recovery without weaker review.
Worked technical explanation
From a busy microservice to ready Kubernetes capacity.
A traffic spike does not create a node directly. Follow one microservice through the control loops, then inspect the requests, policies and readiness conditions between a recommendation and usable capacity.
Worked technical explanation
Will the workload stay available while it changes?
A controller wants replicas. A scheduler needs capacity. A Service needs ready endpoints. A disruption budget answers a different question again.
Worked technical explanation
$2M to $400K: explain every step of the cloud bill.
An impressive percentage needs a reconciled bill. Follow five explicit assumptions to see where the modeled $1.6M monthly reduction comes from—without adding overlapping discounts twice.
Bring your own constraints
What would this unlock for your business?
The platform contribution was a way to accommodate and operate a much larger business in smaller, controllable units. The reported audience growth is substantial on its own. The additional accounting explains how deployment speed, fewer failed transactions and better unit economics supported revenue and margin without claiming that a migration alone caused every commercial outcome.
Discuss a similar system
