Hover a dotted term for 5 seconds to lock its explanation. It closes after 5 seconds away; nested tooltips and keyboard focus keep it open. Click, tap or Enter locks immediately. Technical glossary.

Client project / Data science / Multi-cloud platform

From Azure virtual machines to a data-science platform across two clouds

A data-science team moved from self-managed virtual machines in Azure to a custom platform using Kubernetes and MLflow across Azure and Google Cloud. The engineering story connects repeatable experiments, governed releases and cloud-local execution to the business work the team could take on.

01 / Context

Situation

The starting point was a data-science team working on self-managed Azure virtual machines. The destination was a custom Kubernetes platform spanning Azure and Google Cloud, with MLflow in the experiment and model lifecycle. Those platform facts come from the site owner’s account of the delivered engagement; the client and project dates are anonymized.

The practical constraint was not simply that virtual machines existed. It was that local environments, data paths, experiment history and deployment procedures had become tied to individual machines and the people who knew them. Moving that work into containers was useful only because the team gained a repeatable, supportable route from exploration to a usable business capability.

02 / Success criteria

Task

We gave the team a supported workflow for environments, training, artifacts, evaluation and model release without forcing every workload into one cloud or one operating mode. The distinction between interactive exploration, batch jobs and online serving stays visible in resource policy and ownership.

Productivity, cost and reliability were treated as testable outcomes. The operating metrics and financial figures below document the engagement’s planning and measurement basis; technical details are a documented walkthrough rather than a verbatim production runbook.

The engineering contribution

One supported workflow. Deliberate cloud boundaries.

The custom platform standardizes how work is requested and released. Cloud-local execution, identity, storage and recovery remain explicit rather than being hidden behind a multi-cloud label.

Explore the detailed implementation

03 / Technical reconstruction

Action

Extract a reproducible contract from the VM workflow

Inventory dependencies, local datasets, scheduled work and the feature code that actually produces a model. Package a supported environment with an immutable image, dataset manifest, source revision and declared resource needs. Keep notebooks as consumers of reusable code instead of maintaining a second production implementation.

Move a representative workload through a reversible AKS cohort first. Compare outputs within a declared numerical tolerance and retain the original path until the workload owner accepts the replacement. The business effect is a smaller dependency on a particular machine—not an unsupported promise of faster model training.

Make the custom platform the ordinary path

Provide a project-aware interface for requesting an environment or run, selecting an approved execution location and inspecting status. Validate ownership, budget and identity before submitting a standard Kubernetes Job. Enforce resource requests, quotas, deadlines and bounded retries rather than hiding an unbounded job launcher behind a convenient button.

Use separate policies for notebooks, CPU preparation, GPU training and serving. Queue admission and cloud quotas affect useful completion time. A pending Pod is not productive capacity, and an accepted experiment is not a completed result.

Make MLflow evidence portable without pretending storage is shared

Use MLflow for run metadata, evaluation evidence, model lineage and version references. A conservative implementation keeps protected tracking/metadata and artifacts in the relevant cloud execution domain; the custom platform catalog records where a run lives and which digest was approved.

An Azure artifact does not become a GCP artifact when a label changes. Cross-cloud promotion needs an authorized copy, verified bytes, destination access and an egress budget. A model version is identified with its registry location and digest, not only a local version number or mutable alias.

Keep workload identity and release authority separate

Use Microsoft Entra Workload ID for AKS workloads and Workload Identity Federation for GKE instead of distributing durable cloud keys. Scope permissions to the project and storage resources actually required. The portal’s human login, the cluster-management identity and the job’s data-access identity are different boundaries.

A model registry entry does not authorize a production rollout. Promotion resolves an exact model version, image, feature schema and evaluation report. Retain a previous compatible release so that a service or batch output can be restored without blindly repointing an alias.

Turn the operating improvements into business evidence

Measure environment lead time, manual maintenance, successful run completion, release labor, recovery drills and cost per useful workload. Report queue and interruption losses alongside GPU utilization. Retire duplicate VM capacity only after data, artifact and recovery obligations are met.

The commercial mechanism is a team that can spend more attention on useful modeling and deliver a reviewed capability sooner. Earlier deployment can bring paid service activation forward, but demand, customer readiness, pricing and finance recognition still determine revenue. The platform does not manufacture model quality or a signed contract.

04 / Business consequences

Result

The delivered outcome was the move to Kubernetes and MLflow through a custom platform across Azure and GCP. The business accounting makes the contributions inspectable: a $75,000 monthly operating-cost reduction, 194.85 hours of monthly scientist maintenance capacity, and less manual release work.

A separate timing example: a paid capability launched eight business days earlier. At $150,000 of eligible monthly billable service over twenty business days, that represents $60,000 of service opportunity brought forward. It is not $60,000 of additional ARR, and it is zero if the customer’s activation date cannot move.

Platform savings

$75,000 / month

$180K → $105K at comparable useful work; $900K annualized. A $180K transition budget gives a 2.4-month simple payback before commitments and finance adjustments.

Recovered scientist capacity

194.85 hr / month

18 scientists × (3 − 0.5) maintenance hours/week × 4.33 weeks. At $110/hour that is $21,433.50 of capacity value.

Environment lead time

2 days → 20 min

A supported template avoids an ordinary machine-setup queue. This is elapsed waiting time, not two days of paid labor recovered per request.

Experiment reliability

36 fewer failed runs

At 600 runs/month, a success rate of 92% → 98% reduces failed runs from 48 to 12.

Recovered deployment capacity

87 hr / month

At a fixed 12 promotions/month, manual release work falls from eight hours to 45 minutes each. This is a separate capacity lens; it may overlap with other labor accounting.

Release lead time

10 days → 2 days

Reproducible packaging and evaluation remove handoffs while retaining approval. More promotion opportunities are useful only when candidate quality and customer readiness support them.

Conditional revenue timing

$60,000 brought forward

$150K eligible monthly service ÷ 20 business days × eight days earlier. Requires paid activation to advance; not new recurring revenue, and not additive with the cost or capacity accounting.

These are distinct mechanisms and sensitivities, not an additive ROI total. Cash savings, staff capacity, revenue timing, gross revenue and contribution margin are different quantities.

Inside the system

Boundaries, not black boxes.

The engineering contracts behind the system.

01 / Component

Custom platform

A project-aware interface validates run and environment requests, records ownership and exposes actual completion rather than only submission.

Inputs
Authorized project request · Approved environment contract
Outputs
Bounded Kubernetes workload · Run and release record
Failure & recovery
Missing authority or budget rejects the request; an accepted job is never reported as completed work.

03 / Component

MLflow evidence

Tracking and registry records join parameters, datasets, evaluation and artifact versions to an identifiable model release.

Inputs
Run metadata · Metrics · Artifact references
Outputs
Experiment history · Versioned model evidence
Failure & recovery
A missing artifact or unavailable metadata store blocks promotion; a healthy tracking endpoint is not proof of training success.

04 / Component

Cloud identity and storage

Short-lived cloud-specific workload credentials authorize the exact data and artifact operations required by the project.

Inputs
Projected service-account identity · Scoped cloud policy
Outputs
Authorized local storage access
Failure & recovery
Wrong project, revoked access or a missing cloud trust fails closed; cross-cloud copies require explicit authorization.

05 / Component

Reviewed serving release

An exact model, image, feature schema and evaluation record are promoted together and retain a compatible recovery route.

Inputs
Approved version and digest · Evaluation evidence
Outputs
Identifiable production or batch release
Failure & recovery
Failed evaluation, missing approval or incompatible rollback state stops promotion instead of broadening authority.

Judgment under constraints

Why this design, not another?

Standardize the workload contract, not every cloud service

The team needs a familiar path while Azure and GCP retain different identity, data and quota boundaries.

Tradeoff. Two execution domains require deliberate support and validation.

Keep bulk artifacts close to execution

Locality limits unnecessary transfer, data-residency surprises and remote storage dependencies.

Tradeoff. Cross-cloud promotion becomes an explicit copy and verification step.

Separate experiment completion from release approval

A successful training process is not proof that a model is fit for production use.

Tradeoff. Evaluation and accountable approval remain on the path.

Operational risk controls

  • Prove artifact and metadata restoration together.
  • Test revoked identity and wrong-project access.
  • Retain a reversible VM migration cohort until acceptance.
  • Bound jobs, retries and resource consumption.
  • Verify the full model, feature and runtime combination before rollback.

Engagement economics / Delivered results

MLOps platform economics

Compare operating costs at the same useful workload. Staff capacity, run-success sensitivities and revenue timing are separate explanations, not one ROI total.

Central operating-cost scenario at fixed useful workload.

baseline spend
$180,000.00 /month
platform spend
$105,000.00 /month
spend reduction
$75,000.00 /month
annualized reduction
$900,000.00 /year
spend reduction
41.67%

Assumptions you can inspect

Comparable-workload baseline
180000 USD/month
Central candidate operation
105000 USD/month
Transition budget
180000 USD
Scientist population
18 people
Training workload
600 runs/month
Eligible paid service
150000 USD/month, conditional
See every scenario calculation

Conservative

  • Comparable-workload baseline: $180,000/month.
  • Candidate operating cost: $140,000/month.
  • Reduction: $180,000 − $140,000 = $40,000/month; twelve comparable months give $480,000.
  • Transition costs, contractual commitments and revenue sensitivities remain separate; these scenarios are not additive.

Central

  • Comparable-workload baseline: $180,000/month.
  • Candidate operating cost: $105,000/month.
  • Reduction: $180,000 − $105,000 = $75,000/month; twelve comparable months give $900,000.
  • Transition costs, contractual commitments and revenue sensitivities remain separate; these scenarios are not additive.

Higher utilization

  • Comparable-workload baseline: $180,000/month.
  • Candidate operating cost: $90,000/month.
  • Reduction: $180,000 − $90,000 = $90,000/month; twelve comparable months give $1,080,000.
  • Transition costs, contractual commitments and revenue sensitivities remain separate; these scenarios are not additive.

What this model cannot prove

  • Migration scope is owner-reported; price, staffing, duration and operating figures are the engagement’s documented basis.
  • Client-identifying billing, contract and release records remain anonymized.
  • Capacity is staff time, not booked cash. Revenue timing is not incremental ARR or guaranteed recognized revenue.
  • Cloud-local artifact placement, duplicate migration capacity, data-transfer charges and recovery obligations can materially change the economics.

Follow the engineering

Technical explanations, without crowding the story.

Open the deeper implementation notes and worked models when you want to inspect a specific mechanism.

Bring your own constraints

What would this unlock for your business?

The useful result is a data-science team with a repeatable route from an experiment to an accountable release across its chosen clouds. The reported platform migration stands on its own, and the operating and commercial accounting explains its delivered value. Financial reconciliation tied the accounting to bills, completed-work evidence and customer activation records.

Discuss a similar system

Technical glossary: definitions, connected ideas and further reading.

Optional analytics off. Contact works either way.

How measurement works