
Operating on ArgoCD Drift
Last updated 2026-07-27 ·951c0c4
This is a debugging-focused companion to Operating on GitOps. That post covers day-to-day ArgoCD commands. This one covers what happens when OutOfSync stops being useful.
When I last ran argocd app list, 20 of 52 apps were OutOfSync. I’d been ignoring it — everything worked, dashboards were green. Then I investigated. Not one bug. Seven. And one of them was hiding a 21-day crashloop.
graph TB
subgraph git["Git Source"]
manifest["Deployment YAML<br/>no default values"]
end
subgraph api["API Server (apply)"]
crd["CRD Schema<br/>injects defaults"]
normalized["Normalized object<br/>with defaults set"]
end
subgraph argocd["ArgoCD Comparison"]
git_copy["Git version<br/>(no defaults)"]
live["Live version<br/>(with defaults)"]
diff["Three-way diff<br/>flags mismatch"]
end
manifest -->|"kubectl apply"| crd --> normalized
git_copy -.-> diff
live -.-> diff
diff -->|"OutOfSync"| dashboard["20 apps<br/>permanently red"]
diff -.->|"masked"| crashloop["21-day CrashLoopBackOff<br/>in argo-rollouts"]
What Healthy Looks Like
A healthy ArgoCD install isn’t one where every app is Synced — it’s one where OutOfSync actually means something is wrong. If the column is always red, you stop reading it.
Target: 52 apps, ≤2 OutOfSync, both with known and documented residuals.
Diagnose
Start with a bird’s-eye view of what’s drifting:
kubectl -n argocd get applications -o json \
| jq -r '.items[] | .metadata.name as $app
| .status.resources[]?
| select(.status != "Synced")
| "\($app)\t\(.kind)/\(.name)\t\(.namespace // "cluster")"' \
| sortThis gives you the shape of the drift: which app has which kind drifting, at which scope. Patterns jump out immediately.
On my cluster the output was dominated by three kinds: ExternalSecret (10 apps), Application (12 entries under root), and CustomResourceDefinition (12 entries).
Recover
Class A: CRD Schema Defaults
Symptom: every ExternalSecret drifts with fields like deletionPolicy: Retain, conversionStrategy: Default that the CustomResourceDefinitionThe object that teaches the Kubernetes API a new resource type. Install a CRD and the API server starts serving a kind it has never heard of, with validation and RBAC like any built-in. schema injects but git doesn’t have.
Fix: pin the defaults in git so the manifest matches what the CRD writes.
spec:
target:
creationPolicy: Owner
deletionPolicy: Retain
data:
- secretKey: ANTHROPIC_API_KEY
remoteRef:
key: ANTHROPIC_API_KEY
conversionStrategy: Default
decodingStrategy: None
metadataPolicy: NoneAfter pinning: 10 apps moved from OutOfSync to Synced.
Class B: Default-Value Phantom Diff
Symptom: root app shows child Applications drifting because prune: false is explicit in git but ArgoCD normalises it away (it’s the CRD schema default).
Fix: drop the explicit line.
for f in apps/root/templates/*.yaml; do
sed -i '/^ prune: false$/d' "$f"
doneSame pattern applies to group: "" in ignoreDifferences blocks — ArgoCD treats empty-string groups as unset.
Class C: Orphan CRDs
Symptom: Tekton and Argo Rollouts CRDs are OutOfSync — they were installed by a pre-ArgoCD bootstrap and have no argocd.argoproj.io/tracking-id annotation.
kubectl get crd rollouts.argoproj.io -o jsonpath='{.metadata.annotations}'Fix: annotate them, then use ignoreDifferences for the metadata noise:
for crd in analysisruns.argoproj.io analysistemplates.argoproj.io \
clusteranalysistemplates.argoproj.io experiments.argoproj.io \
rollouts.argoproj.io; do
kubectl annotate crd $crd \
"argocd.argoproj.io/tracking-id=argo-rollouts:apiextensions.k8s.io/CustomResourceDefinition:/$crd" \
--overwrite
doneignoreDifferences:
- group: apiextensions.k8s.io
kind: CustomResourceDefinition
jsonPointers:
- /metadata/labels
- /metadata/annotations
- /spec/preserveUnknownFieldsClass D: Zombie Sub-charts
Symptom: resources from disabled sub-charts persist because prune: false keeps them alive — redis-cluster Services, ingress-nginx ClusterRoles, ValidatingWebhookConfigurations.
Pre-delete verification:
# Are any Ingress resources still using the nginx IngressClass?
kubectl get ingress -A -o json \
| jq -r '.items[] | select(.spec.ingressClassName=="nginx")
| "\(.metadata.namespace)/\(.metadata.name)"'Rollback dump before deleting anything:
mkdir -p /tmp/argocd-drift
kubectl get clusterrole infisical-ingress-nginx -o yaml \
> /tmp/argocd-drift/rollback-infisical-clusterrole.yamlThen delete one at a time. Nothing broke in our case — every service uses Traefik now.
Class E: Two Apps Fighting for One Namespace
Symptom: sympozium-extras and sympozium both try to own the same namespace’s tracking-id annotation.
managedNamespaceMetadata won’t help — it only applies to namespaces ArgoCD auto-creates, not chart-rendered ones.
Fix: apply labels out-of-band, delete the duplicate manifest.
kubectl label ns sympozium-system \
pod-security.kubernetes.io/enforce=privileged \
pod-security.kubernetes.io/audit=privileged \
pod-security.kubernetes.io/warn=privileged --overwriteLabels are sticky — once applied, they survive chart re-renders. Document as a manual op.
Class F: Terminal Hook Noise
Symptom: completed Jobs (postgres-vk-init-electric, PipelineRun) show as OutOfSync — they complete, ArgoCD flags them because requiresPruning: true but prune: false refuses to clean up.
One-shot delete works but they’ll come back on the next sync (they’re supposed to). Permanent fix: hook-delete-policy: BeforeHookCreation,HookSucceeded on the hook definition.
Class G: Chart-Operator Disagreement
Symptom: apps where the chart or operator injects fields git doesn’t specify — timestamps, hashes, defaults.
| App | Field | Fix |
|---|---|---|
| victoria-metrics-grafana | checksum/config, checksum/secret | Narrow ignoreDifferences on those JSON pointers |
| gpu-operator ClusterPolicy | Dozens of unset sub-fields defaulted by webhook | ignoreDifferences on /spec wholesale |
| vcluster-experiments StatefulSet | vClusterConfigHash, whenScaled, revisionHistoryLimit | Per-pointer ignoreDifferences |
| infisical Deployment | updatedAt: "2026-04-04 UTC 21:31:24" (stamped every render) | ignoreDifferences |
| infisical-postgresql PodDisruptionBudgetA floor on how many replicas must stay available during voluntary disruption. It is what makes a node drain wait instead of taking the last healthy pod with it. | maxUnavailable: "" diverges from K8s default | pdb.create: false in values |
Class H: Stale Revision — the inverse failure (added 2026-07-27)
Classes A–G are all the same shape: OutOfSync that means nothing. This one is the mirror image — Synced that means nothing — and it is worse, because the first kind is merely noisy while this kind is silent.
Symptom: an Application reports Synced/Healthy minutes after a merge, while the live resource still holds pre-merge content. Seen twice in one session on 2026-07-27: longhorn’s rendered ConfigMap was missing a key the merged values add, and hermes-agent-shell’s PVC still requested the pre-merge size. argocd.argoproj.io/refresh: hard did not clear either.
There is no ignoreDifferences fix here, because nothing is drifting. Synced is a claim about live matching what ArgoCD fetched, not about live matching main.
Rule out your own manifest first — a chart silently dropping an unrecognised key looks identical from the cluster:
helm template <release> <chart> --version <ver> -f <values> | grep -A5 'name: <rendered-object>'If the key renders locally and is absent live while the app claims Synced, force a real sync operation:
kubectl -n argocd patch application <app> --type=merge \
-p '{"operation":{"sync":{"revision":"HEAD","syncOptions":["ServerSideApply=true","RespectIgnoreDifferences=true"]}}}'Both cleared within seconds. Multi-source apps (upstream chart + a $values ref into this repo) look the most exposed, though the underlying cause was never established — the recovery is known, the mechanism is not.
Distinguish it from ordinary poll lag before acting: an Application that has simply not polled the newest commit yet converges on its own within a few minutes and was never wrong, only behind. Compare .status.sync.revision against the commit you merged and wait one poll interval before forcing anything.
The wider pattern — four distinct ways a green tile diverges from reality, and the check for each — is Operating on Green.
The Unmasked Bug
The most important fix wasn’t about any of the seven drift classes.
When I resolved argo-rollouts’s drift, the controller logs became readable for the first time:
NAME READY STATUS RESTARTS AGE
argo-rollouts-6b4c4dfbd9-ghl9c 0/1 CrashLoopBackOff 1154 (2m18s ago) 21d1154 restarts. 21 days. The pod log:
level=fatal msg="Failed to download plugins: ...
response code Not Found"I had a trafficRouterPlugins entry in values.yaml pointing at a Cilium plugin URL that never existed — argoproj-labs has no such repo. The controller crashed on every startup. The drift noise had masked it completely.
Takeaways
- Classify before you fix. Seven drift classes. Each needed a different fix. Blanket
ignoreDifferenceswould have hidden the crashloop. - Pin schema defaults in git where possible. Preferred over
ignoreDifferencesbecause real changes still flag. - Dump before you delete.
kubectl get -o yaml > rollback.yamlsaves you from a bad Monday. - Read the logs of every app that moved from
Progressing. That’s where the hiding things are.
Missteps
| What we assumed | Why it was wrong | What it cost |
|---|---|---|
prune: false is harmless to have explicit in git | ArgoCD normalises it away (it’s the CRD schema default). The three-way diff flags the phantom gap on every child Application. | 12 templates edited to remove the explicit line. |
| Orphan CRDs will be adopted by ArgoCD when added to the Application | ArgoCD won’t adopt resources that predate its tracking annotations, even if the manifest is correct. | Manual annotation + ignoreDifferences for metadata noise. |
managedNamespaceMetadata fixes two-app namespace fights | That feature only works for auto-created namespaces, not chart-rendered ones. | Applied labels out-of-band via manual command. |
Disabled sub-charts are harmless with prune: false | The resources persist forever. Only a manual audit with rollback dumps can clean them safely. | Zombie ClusterRoles, ValidatingWebhookConfigurations, and IngressClasses blocked cleanup for weeks. |
| A Cilium traffic router plugin exists and works | The plugin URL 404s — it was never published. The config was aspirational. | 21 days of controller crashloop, masked by drift noise. |
Progressing in the health column means “mid-reconcile, probably fine” | When every second app is OutOfSync, you stop reading the column. Four apps had been Progressing for weeks. | Missed the crashloop signal until the noise was cleared. |
References
- Plan:
docs/superpowers/plans/2026-04-15--gitops--argocd-drift-cleanup.md - Operating on GitOps
- ArgoCD
ignoreDifferences - Kubernetes ServerSideApply
- Operating on Green — four ways a green tile diverges from reality
