AI infrastructure / Platform engineering
MLflow on Kubernetes: AKS-to-GCP Migration and Recovery Value
A real client engagement. The engineering and the results are described below.
A forecasting company we worked with loses time reconnecting experiments, artifacts and release permissions whenever a model moves toward production. This engagement ledger is separate from the linked owner-reported Azure/GCP project account, whose original facts remain unchanged.
Request a time through the inquiry form. A meeting is confirmed separately by email.
The business problem behind the technology
What business work resumes when the ML platform has one contract?
Each release needs the correct experiment metadata, artifacts, feature contract and permissions. Operators cannot keep recreating those links by hand whenever execution crosses a cloud boundary.
Read the client engagement ↓Client engagement / Delivered results
A reproducible handoff instead of another workstation rescue
A forecasting company we worked with makes twelve model releases each month. This engagement is separate from the linked owner-reported Azure/GCP account.
The constraint
Each release needs the correct experiment metadata, artifacts, feature contract and permissions. Operators cannot keep recreating those links by hand whenever execution crosses a cloud boundary.
The engineering decision
The client’s team inventories workstation dependencies, separates MLflow metadata from artifacts and production authority, and proves workload identity and recovery. Handoff effort fell from ten to four hours per release.
The delivered outcome
The monthly handoff effort fell from 120 to 48 hours. The 72-hour difference became available team capacity, separate from the reported facts of the linked project.
| Measure / unit | Before | After | Difference |
|---|---|---|---|
| Model release handoff effort hours/month | 120 | 48 | 72 |
Model release handoff effort. Twelve releases × ten hours versus twelve × four. All volume and effort values belong to this engagement.
The conditions behind the results
- The twelve releases retain equivalent evaluation, review and production acceptance.
- Ten/four-hour handoffs are this engagement’s values, separate from the owner-reported account.
- Cloud charges, migration work, ongoing platform support and retraining cost are excluded.
What this does not prove. The released staff capacity is this engagement’s result, separate from the owner-reported Azure/GCP project.
The linked client project account describes the same engagement from the project side.
Evidence to collect for your own decision
- Inventory tracking metadata, artifact references, environments and model-release identity.
- Restore metadata and artifacts together and verify actual artifact access after recovery.
- Measure handoff effort and exceptions while preserving evaluation and least privilege.
Key decisions
Give a data-science team one supported workflow without pretending that two clouds share a filesystem, identity boundary or failure domain. A custom platform can standardize the contract while execution and data stay deliberately placed.
- Inventory tracking metadata, artifact references, environments and model-release identity.
- Restore metadata and artifacts together and verify actual artifact access after recovery.
- Released staff capacity is not cash savings, extra customer revenue or an unreported outcome of the Azure/GCP project.
Follow the decision
Select a step to follow its reasoning, then continue into the technical chapters.
Problem → boundary → decision → evidence
Find the dependencies hidden in the original workstation.
Read every component and connection
- Inventory the experiment · Problem
- Find the dependencies hidden in the original workstation.
- Separate platform state · Boundary
- Keep metadata, artifacts and production authority distinct.
- Move identity deliberately · Decision
- Preserve least privilege across the cloud boundary.
- Prove useful recovery · Evidence
- Tie operational evidence to the outcome being claimed.
- Inventory the experiment → Separate platform state: identify the constraint
- Separate platform state → Move identity deliberately: choose a bounded change
- Move identity deliberately → Prove useful recovery: check the outcome
A conceptual decision map for this article, not a measured timeline, physical topology or a depiction of a specific client system.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
Inventory the workstation before packaging the notebook
The engineering reconstruction starts with workstation dependencies; the client’s team uses that inventory to examine handoff effort.
Capture the VM-dependent contract: operating-system packages, language environments, CUDA and driver expectations, local datasets, cron jobs, shared directories and credentials. Split interactive exploration, repeatable training, batch scoring and online serving. Their resource, isolation and recovery needs differ; wrapping the entire VM in one privileged container merely transports the original ambiguity.
Choose a representative workload and record a baseline for completed runs, queue time, useful GPU time, failed runs, operator intervention and cost. Record dataset and feature-schema versions. A seed alone does not promise bitwise reproducibility across different accelerators or libraries; define the numerical tolerance and business evaluation that actually matter.
Move a bounded workload to AKS first while keeping its data close to its Azure source. Use a reversible cohort cutover and output comparison. Add GKE execution only after the run contract is portable and the data-placement, residency and commercial case is explicit. Multi-cloud capability is not an instruction to copy every dataset everywhere.
Make the custom platform a supported contract
The platform becomes a service someone must support. That changes the meaning of a successful migration.
The platform presents a project, environment template, source revision, image digest, dataset manifest, requested resources, execution cloud and deadline. Its API validates ownership and budget, then submits a standard Kubernetes Job through a namespace-scoped identity. Kubernetes remains responsible for execution; the portal does not invent a competing scheduler or claim that an accepted request is a completed experiment.
Use separate workload classes for interactive notebooks, CPU preparation, GPU training and serving. ResourceQuota and LimitRange constrain accidental consumption; explicit CPU, memory and GPU requests make scheduling legible. Queue admission should reserve compatible capacity and respect project priority before a large parallel job starts. Namespace isolation is not a substitute for a hardened runtime when arbitrary untrusted code is involved.
Each job has a bounded retry policy, deadline and cleanup policy. Checkpoints go to durable cloud-local storage rather than a cross-cloud persistent volume. Re-running a scoring window must produce one logical published result rather than duplicate customer-facing side effects. Unrecoverable jobs retain their failure reason and useful logs instead of being silently resubmitted forever.
Submitted run contract
project + authorized owner
source revision + immutable training image
dataset manifest + feature schema
cloud + workload identity + artifact destination
CPU / memory / accelerator requests + deadline
Completed run evidence
Kubernetes Job outcome + MLflow run ID
artifact digests + evaluation results
cost attribution + explicit promotion decisionSeparate MLflow metadata, artifacts and production authority
MLflow tracking makes runs visible, but metadata, artifacts and deployment authority are different assets. Each needs a migration contract.
MLflow Tracking records parameters, metrics, run metadata and artifact references; it does not schedule the Kubernetes Job. The model registry records versions, lineage and aliases; it is not, by itself, a production approval system. A healthy tracking HTTP endpoint also does not prove that training jobs are completing or that artifacts can be recovered.
A conservative two-cloud design uses a protected MLflow service and database-backed metadata in each execution domain, with artifacts in Azure Blob Storage or Google Cloud Storage near the workload. The custom platform catalog keeps cloud, experiment/run identifiers and immutable release references. Local model-version numbers are scoped to their registry, so a release identity includes the registry location and artifact digest—not only version 17.
A cross-cloud promotion is an explicit, authorized transfer with a verified digest, destination access policy and accounted egress cost. Changing an experiment or deployment label does not relocate existing artifacts or make an Azure identity a GCP identity. MLflow artifact proxying changes which identity accesses the bucket; validate both the proxy and direct-access paths before allowing users to submit work.
Resolve a mutable alias to a specific version at promotion time. Pin the serving image, model artifacts, preprocessing code, schema and evaluation evidence in one release record. A running deployment must not silently switch model bytes because an alias changes later. Retain the previous compatible release independently of the current registry pointer.
Federate workload identity without broadening access
Identity is a migration responsibility in the reported account and a separate acceptance check in the engagement effort ledger.
On AKS, Microsoft Entra Workload ID exchanges a projected Kubernetes service-account token through the configured OIDC trust for access to authorized Azure resources. Use an SDK version that supports workload identity and scope the destination permissions to the required container, database or service. The pod identity is distinct from the human login and cluster-management identity.
On GKE, Workload Identity Federation maps Kubernetes principals to permitted Google Cloud resource access without distributing service-account key files. Be deliberate about identity sameness: clusters sharing a project pool and namespace/service-account names may represent the same principal. Separate trust domains or use appropriate conditions rather than assuming identical names imply isolation.
Authenticate the platform and MLflow endpoints, authorize project access, and enforce data access at the storage boundary. Do not give every experiment an artifact proxy with blanket access to every tenant. Test revoked access, a wrong namespace, an expired token and a missing artifact as real failure cases. No credential values belong in run parameters, model artifacts or logs.
Scale useful work, not just allocated GPUs
More allocated GPUs are not the operating objective. The platform needs to increase useful completed work under its stated constraints.
Scale serving replicas from a workload signal such as resource utilization with valid requests, queue depth or latency-related demand. HPA changes desired replicas; node provisioning supplies schedulable resources; readiness admits traffic. These are separate transitions. Training admission additionally needs compatible accelerators, quota, data locality and checkpoint policy.
Keep an intentional on-demand floor for interactive and deadline-sensitive work. Use interruptible capacity only where jobs can checkpoint and tolerate restart. Report completed useful work, interruption loss, idle reservation and queue age together. A higher utilization percentage is not a win if the business deadline is missed or expensive work is repeatedly discarded.
Coordinate retirement of the VM path with data retention, artifact readability and restore evidence. Stop paying for duplicate capacity only after the workload owner accepts the replacement. Model the overlap and implementation expense separately from the eventual monthly run-rate reduction.
Tie deployment and recovery evidence to commercial outcomes
Reported project facts stay unchanged while the client’s team tests whether its 72 hours of freed monthly capacity are credible.
A reproducible release path can reduce manual packaging, shorten model-promotion lead time and let product teams make a paid capability available earlier. Those are mechanisms, not proof of better model quality or new customer demand. Measure adoption, matched-cohort release time, correction burden and service activation dates before claiming realized revenue.
Test metadata recovery together with artifact access, identity restoration and a rebuilt serving release. State RPO and RTO as tested outcomes or as targets—never interchange them. For batch products, freshness and successful publication are customer-facing reliability metrics; for online products, latency, errors and compatible model loading also matter.
The linked case study separates the reported migration scope from its $75K monthly savings model, recovered staff capacity and conditional revenue-timing example. Finance needs actual bills and recognition rules; engineering needs completed-work and recovery evidence. Neither can be replaced by a successful Kubernetes rollout.
Questions behind the decision
What must be backed up for MLflow recovery?
Tracking metadata, artifact storage and the references connecting them must remain consistent and accessible. Also retain the identity, configuration and environment needed to use the restored model; a running tracking UI alone is not recovery proof.
Can moving MLflow to Kubernetes make model delivery reproducible?
Kubernetes can operate the service, but reproducibility still depends on captured data, code, environment and model artifacts. Explicit evaluation and release contracts are required across both the original and target cloud.
