Platform engineering
Terraform and GitOps: Ownership, Safe Releases and Recovery
A real client engagement. The engineering and the results are described below.
A software platform we worked with changes a deployment only to watch another controller change it back. Its team maps ownership before adding another tool, turning a noisy release process into a business decision about repeatable delivery.
Request a time through the inquiry form. A meeting is confirmed separately by email.
The business problem behind the technology
Can one owner per field remove the release tug-of-war?
The business needs lower operational friction without bypassing executable-plan review, credential custody or approval for production changes.
Read the client engagement ↓Client engagement / Delivered results
The release needed one desired state, not two competing successes
A platform team we worked with makes forty infrastructure-and-application releases a month. Manual coordination and resolving conflicting ownership consumed three staff-hours per release.
The constraint
The business needs lower operational friction without bypassing executable-plan review, credential custody or approval for production changes.
The engineering decision
The team gives Terraform and the GitOps reconciler separate resource/field ownership, binds releases to immutable artifacts and rehearses compatible recovery. Coordination effort fell to one hour per comparable release.
The delivered outcome
The effort ledger moved from 120 to 40 hours a month. The 80-hour difference bought capacity for engineering work.
| Measure / unit | Before | After | Difference |
|---|---|---|---|
| Release coordination effort hours/month | 120 | 40 | 80 |
Release coordination effort. Forty releases × three hours versus forty × one; elapsed rollout time and staff effort are not interchangeable.
The conditions behind the results
- The forty releases have comparable scope and unchanged review/approval requirements.
- Three/one-hour coordination values are this engagement’s operating figures.
- Platform operating costs, implementation effort and recovery incidents are excluded.
What this does not prove. The reduced coordination effort kept build and release authority separate.
Evidence to collect for your own decision
- Identify the single owner of each resource and desired-state field.
- Review the immutable artifact, Terraform plan and separate production approval boundary.
- Rehearse failed reconciliation and recovery with compatible application and data state.
Key decisions
Terraform and GitOps need explicit desired-state ownership, reviewed artifacts and compatible recovery. A successful command is not proof that production has converged safely.
Follow the decision
Select a step to follow its reasoning, then continue into the technical chapters.
Problem → boundary → decision → evidence
Separate infrastructure and application desired-state responsibilities.
Read every component and connection
- Assign ownership · Problem
- Separate infrastructure and application desired-state responsibilities.
- Review the change · Boundary
- Protect state and inspect the executable plan.
- Express release intent · Decision
- Keep artifacts, validation and deployment authorization distinct.
- Prove convergence · Evidence
- Test both the target state and a compatible recovery path.
- Assign ownership → Review the change: identify the constraint
- Review the change → Express release intent: choose a bounded change
- Express release intent → Prove convergence: check the outcome
A conceptual decision map for this article, not a measured timeline, physical topology or a depiction of a specific client system.
Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.
Draw an ownership boundary before writing modules
The client’s platform team begins by assigning each desired-state field a single owner.
Terraform can manage cloud resources and Kubernetes objects, while a GitOps controller can reconcile Kubernetes manifests. That overlap is convenient until both tools believe they own the same object. Competing desired states produce drift loops, surprising reversions, and reviews that do not describe the final cluster. Establish one authoritative controller per resource and document the handoff explicitly.
A practical split gives Terraform responsibility for networks, cluster foundations, workload identity, and managed data services. GitOps owns application workloads and their runtime configuration. Bootstrap resources need an equally clear choice: if Terraform installs the GitOps controller, decide whether Terraform continues to own that installation or whether ownership is transferred through a reviewed process.
Separate environments by credentials, access policy, and state boundaries appropriate to their risk. A directory name or Terraform workspace is not automatically a security boundary. Share reusable modules without sharing unrestricted production write access. Publish narrow outputs such as an approved endpoint or identity reference rather than allowing every application pipeline to read an entire infrastructure state snapshot.
Protect state and review the executable plan
The desired change becomes an executable plan. State custody and review determine whether that plan can be trusted.
Terraform state maps configuration to real resources and can contain sensitive values. Use an access-controlled remote backend with encryption, versioning or backups, and locking where the backend supports it. Locking protects against concurrent state writers; it is not a review gate or a backup. Do not disable locking to make a busy pipeline proceed, and only force-unlock after establishing that the original writer is no longer active.
Generate a plan using the intended source revision, dependency lock, variables, and target environment. Review replacements, deletions, identity changes, and data-service implications, not only the summary count. Apply the exact saved plan through the approved pipeline while it remains valid. If inputs or relevant state have changed, create and review a new plan instead of assuming earlier approval covers different actions.
Saved plans can also include sensitive information. Store them as protected short-lived artifacts, restrict who can apply them, and bind approval to their identity. Serialize applies for a state boundary. Cross-state dependencies require explicit sequencing and compatibility, but splitting every resource into its own state adds orchestration cost without necessarily improving isolation.
Let CI produce releases and Git express intent
The release still needs an immutable artifact and separate production authority even when coordination becomes faster.
CI should build and verify an immutable artifact, then propose a change to the environment repository. Pin application images by digest so the reviewed intent resolves to the same bytes later. Promote the same artifact between environments instead of rebuilding a tag and calling it the same release. Keep provenance and verification evidence connected to that digest.
The GitOps controller reconciles approved manifests into the cluster using its own scoped identity. Routine build jobs do not need broad cluster credentials. Repository protection, controller project boundaries, admission policy, and secret access must reinforce one another: pull-based delivery does not make every repository commit safe by definition.
Argo CD distinguishes automated synchronization, pruning, and self-healing. The example below is only the syncPolicy portion of an Application. It enables reconciliation and self-healing but leaves removal for a deliberate prune operation. Teams may enable automatic pruning once deletion policies and resource ownership are understood; the choice should reflect consequences, not a copied default.
syncPolicy:
automated:
enabled: true
selfHeal: true
prune: falseRollback desired state, not just the running pod
The team needs a way back. GitOps rollback must restore a compatible desired state, not only replace a running Pod.
With automated reconciliation, a manual image change can be reverted by the controller. Express an application rollback as a new reviewed Git change that restores the previous compatible image and configuration. Argo CD documents that its direct rollback operation is unavailable while automated sync is enabled. Changing the desired state keeps the recovery consistent with future reconciliation.
A Git revert does not reverse a database migration or reconstruct deleted infrastructure. Use backward-compatible schema expansion, support old and new application versions during transition, and delay destructive cleanup until the rollback window closes. Record which releases remain compatible with the current data shape. Recovery for irreversible changes may require a forward repair or a tested restore.
Terraform also has no universal undo button. Reverting configuration and planning again describes another infrastructure change, which may create replacements or fail because the world has moved on. Review that plan and preserve backup and restoration procedures for stateful services. Restoring an old state file is not equivalent to restoring the real resources it describes.
Prove convergence and recovery under failure
The team rehearses failure and recovery before treating freed coordination effort as a durable benefit.
Exercise a failed apply, unavailable image, failed readiness check, and interrupted reconciliation in a non-production environment. Confirm that the team can identify the intended revision, actual artifact, controller status, and next safe action. Health and synchronization are different: a manifest can be applied successfully while the application cannot serve requests.
Emergency changes need a bounded break-glass process with authorization, audit, and reconciliation back into source control. If reconciliation must be paused, identify the exact controller and scope, and define how it resumes. Otherwise, an urgent fix can disappear during the next sync or leave permanent undocumented drift.
Measure reviewed-change lead time, failed-release frequency, recovery time, drift age, and ownership conflicts. Validate restoration from protected state backups and application data backups separately. The success criterion is not simply that a pipeline turns green: approved intent must converge to a working service, and recovery must remain possible when one layer fails.
Questions behind the decision
How should Terraform and GitOps share responsibility?
Give each resource or desired-state field one authoritative owner, with explicit interfaces between infrastructure and application delivery. Overlapping reconciliation can undo valid changes and make recovery ambiguous.
Is reverting a Git commit enough to roll back production?
Only when the old desired state is compatible with current data, dependencies and external effects. Verify convergence and application behavior; an earlier manifest does not reverse every migration or side effect.
