Client project / Data science / Multi-cloud platform
From Azure virtual machines to a data-science platform across two clouds
A data-science team moved from self-managed virtual machines in Azure to a custom platform using Kubernetes and MLflow across Azure and Google Cloud. The engineering story connects repeatable experiments, governed releases and cloud-local execution to the business work the team could take on.
01 / Context
Situation
The starting point was a data-science team working on self-managed Azure virtual machines. The destination was a custom Kubernetes platform spanning Azure and Google Cloud, with MLflow in the experiment and model lifecycle. Those platform facts come from the site owner’s account of the delivered engagement; the client and project dates are anonymized.
The practical constraint was not simply that virtual machines existed. It was that local environments, data paths, experiment history and deployment procedures had become tied to individual machines and the people who knew them. Moving that work into containers was useful only because the team gained a repeatable, supportable route from exploration to a usable business capability.
02 / Success criteria
Task
We gave the team a supported workflow for environments, training, artifacts, evaluation and model release without forcing every workload into one cloud or one operating mode. The distinction between interactive exploration, batch jobs and online serving stays visible in resource policy and ownership.
Productivity, cost and reliability were treated as testable outcomes. The operating metrics and financial figures below document the engagement’s planning and measurement basis; technical details are a documented walkthrough rather than a verbatim production runbook.
The engineering contribution
One supported workflow. Deliberate cloud boundaries.
The custom platform standardizes how work is requested and released. Cloud-local execution, identity, storage and recovery remain explicit rather than being hidden behind a multi-cloud label.
Explore the detailed implementation03 / Technical reconstruction
Action
Extract a reproducible contract from the VM workflow
Inventory dependencies, local datasets, scheduled work and the feature code that actually produces a model. Package a supported environment with an immutable image, dataset manifest, source revision and declared resource needs. Keep notebooks as consumers of reusable code instead of maintaining a second production implementation.
Move a representative workload through a reversible AKS cohort first. Compare outputs within a declared numerical tolerance and retain the original path until the workload owner accepts the replacement. The business effect is a smaller dependency on a particular machine—not an unsupported promise of faster model training.
Make the custom platform the ordinary path
Provide a project-aware interface for requesting an environment or run, selecting an approved execution location and inspecting status. Validate ownership, budget and identity before submitting a standard Kubernetes Job. Enforce resource requests, quotas, deadlines and bounded retries rather than hiding an unbounded job launcher behind a convenient button.
Use separate policies for notebooks, CPU preparation, GPU training and serving. Queue admission and cloud quotas affect useful completion time. A pending Pod is not productive capacity, and an accepted experiment is not a completed result.
Make MLflow evidence portable without pretending storage is shared
Use MLflow for run metadata, evaluation evidence, model lineage and version references. A conservative implementation keeps protected tracking/metadata and artifacts in the relevant cloud execution domain; the custom platform catalog records where a run lives and which digest was approved.
An Azure artifact does not become a GCP artifact when a label changes. Cross-cloud promotion needs an authorized copy, verified bytes, destination access and an egress budget. A model version is identified with its registry location and digest, not only a local version number or mutable alias.
Keep workload identity and release authority separate
Use Microsoft Entra Workload ID for AKS workloads and Workload Identity Federation for GKE instead of distributing durable cloud keys. Scope permissions to the project and storage resources actually required. The portal’s human login, the cluster-management identity and the job’s data-access identity are different boundaries.
A model registry entry does not authorize a production rollout. Promotion resolves an exact model version, image, feature schema and evaluation report. Retain a previous compatible release so that a service or batch output can be restored without blindly repointing an alias.
Turn the operating improvements into business evidence
Measure environment lead time, manual maintenance, successful run completion, release labor, recovery drills and cost per useful workload. Report queue and interruption losses alongside GPU utilization. Retire duplicate VM capacity only after data, artifact and recovery obligations are met.
The commercial mechanism is a team that can spend more attention on useful modeling and deliver a reviewed capability sooner. Earlier deployment can bring paid service activation forward, but demand, customer readiness, pricing and finance recognition still determine revenue. The platform does not manufacture model quality or a signed contract.
04 / Business consequences
Result
The delivered outcome was the move to Kubernetes and MLflow through a custom platform across Azure and GCP. The business accounting makes the contributions inspectable: a $75,000 monthly operating-cost reduction, 194.85 hours of monthly scientist maintenance capacity, and less manual release work.
A separate timing example: a paid capability launched eight business days earlier. At $150,000 of eligible monthly billable service over twenty business days, that represents $60,000 of service opportunity brought forward. It is not $60,000 of additional ARR, and it is zero if the customer’s activation date cannot move.
Platform savings
$75,000 / month
$180K → $105K at comparable useful work; $900K annualized. A $180K transition budget gives a 2.4-month simple payback before commitments and finance adjustments.
Recovered scientist capacity
194.85 hr / month
18 scientists × (3 − 0.5) maintenance hours/week × 4.33 weeks. At $110/hour that is $21,433.50 of capacity value.
Environment lead time
2 days → 20 min
A supported template avoids an ordinary machine-setup queue. This is elapsed waiting time, not two days of paid labor recovered per request.
Experiment reliability
36 fewer failed runs
At 600 runs/month, a success rate of 92% → 98% reduces failed runs from 48 to 12.
Recovered deployment capacity
87 hr / month
At a fixed 12 promotions/month, manual release work falls from eight hours to 45 minutes each. This is a separate capacity lens; it may overlap with other labor accounting.
Release lead time
10 days → 2 days
Reproducible packaging and evaluation remove handoffs while retaining approval. More promotion opportunities are useful only when candidate quality and customer readiness support them.
Conditional revenue timing
$60,000 brought forward
$150K eligible monthly service ÷ 20 business days × eight days earlier. Requires paid activation to advance; not new recurring revenue, and not additive with the cost or capacity accounting.
These are distinct mechanisms and sensitivities, not an additive ROI total. Cash savings, staff capacity, revenue timing, gross revenue and contribution margin are different quantities.
Inside the system
Boundaries, not black boxes.
The engineering contracts behind the system.
01 / Component
Custom platform
A project-aware interface validates run and environment requests, records ownership and exposes actual completion rather than only submission.
- Inputs
- Authorized project request · Approved environment contract
- Outputs
- Bounded Kubernetes workload · Run and release record
- Failure & recovery
- Missing authority or budget rejects the request; an accepted job is never reported as completed work.
02 / Component
AKS and GKE execution
Cloud-local notebooks, training Jobs and serving workloads have explicit resource requests, quotas, data placement and recovery policy.
- Inputs
- Image digest · Dataset manifest · Workload identity
- Outputs
- Completed run · Checkpoint and operational signals
- Failure & recovery
- Unschedulable or failed work remains visible with an owner; bounded retries do not fabricate successful completion.
03 / Component
MLflow evidence
Tracking and registry records join parameters, datasets, evaluation and artifact versions to an identifiable model release.
04 / Component
Cloud identity and storage
Short-lived cloud-specific workload credentials authorize the exact data and artifact operations required by the project.
- Inputs
- Projected service-account identity · Scoped cloud policy
- Outputs
- Authorized local storage access
- Failure & recovery
- Wrong project, revoked access or a missing cloud trust fails closed; cross-cloud copies require explicit authorization.
05 / Component
Reviewed serving release
An exact model, image, feature schema and evaluation record are promoted together and retain a compatible recovery route.
- Inputs
- Approved version and digest · Evaluation evidence
- Outputs
- Identifiable production or batch release
- Failure & recovery
- Failed evaluation, missing approval or incompatible rollback state stops promotion instead of broadening authority.
Judgment under constraints
Why this design, not another?
Standardize the workload contract, not every cloud service
The team needs a familiar path while Azure and GCP retain different identity, data and quota boundaries.
Tradeoff. Two execution domains require deliberate support and validation.
Keep bulk artifacts close to execution
Locality limits unnecessary transfer, data-residency surprises and remote storage dependencies.
Tradeoff. Cross-cloud promotion becomes an explicit copy and verification step.
Separate experiment completion from release approval
A successful training process is not proof that a model is fit for production use.
Tradeoff. Evaluation and accountable approval remain on the path.
Operational risk controls
Engagement economics / Delivered results
MLOps platform economics
Compare operating costs at the same useful workload. Staff capacity, run-success sensitivities and revenue timing are separate explanations, not one ROI total.
Central operating-cost scenario at fixed useful workload.
- baseline spend
- $180,000.00 /month
- platform spend
- $105,000.00 /month
- spend reduction
- $75,000.00 /month
- annualized reduction
- $900,000.00 /year
- spend reduction
- 41.67%
Assumptions you can inspect
- Comparable-workload baseline
- 180000 USD/month
- Central candidate operation
- 105000 USD/month
- Transition budget
- 180000 USD
- Scientist population
- 18 people
- Training workload
- 600 runs/month
- Eligible paid service
- 150000 USD/month, conditional
See every scenario calculation
Conservative
- Comparable-workload baseline: $180,000/month.
- Candidate operating cost: $140,000/month.
- Reduction: $180,000 − $140,000 = $40,000/month; twelve comparable months give $480,000.
- Transition costs, contractual commitments and revenue sensitivities remain separate; these scenarios are not additive.
Central
- Comparable-workload baseline: $180,000/month.
- Candidate operating cost: $105,000/month.
- Reduction: $180,000 − $105,000 = $75,000/month; twelve comparable months give $900,000.
- Transition costs, contractual commitments and revenue sensitivities remain separate; these scenarios are not additive.
Higher utilization
- Comparable-workload baseline: $180,000/month.
- Candidate operating cost: $90,000/month.
- Reduction: $180,000 − $90,000 = $90,000/month; twelve comparable months give $1,080,000.
- Transition costs, contractual commitments and revenue sensitivities remain separate; these scenarios are not additive.
What this model cannot prove
- Migration scope is owner-reported; price, staffing, duration and operating figures are the engagement’s documented basis.
- Client-identifying billing, contract and release records remain anonymized.
- Capacity is staff time, not booked cash. Revenue timing is not incremental ARR or guaranteed recognized revenue.
- Cloud-local artifact placement, duplicate migration capacity, data-transfer charges and recovery obligations can materially change the economics.
Follow the engineering
Technical explanations, without crowding the story.
Open the deeper implementation notes and worked models when you want to inspect a specific mechanism.
Platform engineering
Kubernetes Deployments: Readiness, Rollouts and Production Risk
A real order platform follows Pods, probes, rollout capacity and PodDisruptionBudgets to shorten release recovery without inventing uptime guarantees.
Cloud architecture
AWS vs Azure vs Google Cloud: Migration Costs and Workload Fit
A real SaaS migration compares AWS, Azure and Google Cloud through identity, egress and operating cost, with a transparent payback example.
AI infrastructure / Platform engineering
MLflow on Kubernetes: AKS-to-GCP Migration and Recovery Value
A real ML platform engagement delivered release handoff value while explaining MLflow, artifacts and AKS/GCP identity. Owner-reported project facts stay separate.
AI infrastructure
Production MLOps: Model Deployment, Evaluation and Release Value
A real forecasting business connects MLOps pipelines, model evaluation and rollback to release handoff effort without treating a notebook as production.
Platform engineering
Terraform and GitOps: Ownership, Safe Releases and Recovery
A real platform team lowered release coordination effort through Terraform state boundaries, immutable artifacts and GitOps recovery without weaker review.
Worked technical explanation
Will the workload stay available while it changes?
A controller wants replicas. A scheduler needs capacity. A Service needs ready endpoints. A disruption budget answers a different question again.
Worked technical explanation
What limits one warm decode step on Blackwell?
Start with fit within the total modeling budget. Then separate the cost of moving data from doing arithmetic. A low-precision label or a peak specification cannot answer either question on its own.
Worked technical explanation
Which cloud fits the workload—not just the headline price?
Normalize the workload and quote boundary first. Then inspect the modeled line items without pretending that cost alone selects an architecture.
Bring your own constraints
What would this unlock for your business?
The useful result is a data-science team with a repeatable route from an experiment to an accountable release across its chosen clouds. The reported platform migration stands on its own, and the operating and commercial accounting explains its delivered value. Financial reconciliation tied the accounting to bills, completed-work evidence and customer activation records.
Discuss a similar system
