Skip to content

Support idempotent and isolated ArgoCD bootstrap - #595

Merged
mdroll merged 2 commits into
developfrom
fix/idempotent-argocd-bootstrap
Sep 22, 2026
Merged

mdroll merged 2 commits into
developfrom
fix/idempotent-argocd-bootstrap

Conversation

@avetgit

@avetgit avetgit commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Make the Helm-based ArgoCD bootstrap idempotent and isolate prefixed ArgoCD installations.

Subsequent GOP runs now detect an existing bootstrap Application and skip the Helm bootstrap while keeping ArgoCD managed through GitOps.

Prefixed installations use their own Helm release name, which prevents ownership conflicts for cluster-scoped ArgoCD resources such as ClusterRoles. Existing ArgoCD CRDs are reused instead of being claimed by multiple Helm releases.

The Helm release cleanup and the self-managed ArgoCD Application now also use the installation-specific release name.

Verified with:

repeated Full profile deployments
repeated prefixed Full profile deployments
Full and prefixed Full ArgoCD installations running in parallel on the same cluster without Helm ownership conflicts
complete test suite

During parallel deployment testing, additional multi-instance issues in other tools were identified. These are independent of the ArgoCD bootstrap changes and are not part of this PR.

@mdroll
mdroll requested review from mdroll and a lite review from Copilot September 21, 2026 11:11

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Three moderate issues remain around legacy release cleanup and complete CRD ownership handling.

Review effort: Lite
Findings: None

What changed in this PR

This PR makes ArgoCD bootstrap idempotent and isolates prefixed installations through installation-specific Helm releases and CRD handling.

Changes:

  • Skips bootstrap when the Application already exists.
  • Adds resource and CRD ownership detection.
  • Aligns release names across deployment, cleanup, and self-managed Applications.
  • Expands coverage for idempotency, prefixes, CRDs, and resource lookup.

Review identified three moderate issues involving legacy release cleanup, CRD ownership during destruction, and incomplete CRD ownership checks.

File Summary
src/​test/​java/​com/​cloudogu/​gitops/​tools/​core/​argocd/​ArgoCDConfigurationTest.java Tests bootstrap idempotency, prefixes, CRDs, and cleanup.
src/​test/​java/​com/​cloudogu/​gitops/​infrastructure/​kubernetes/​api/​K8sClientTest.java Tests resource existence behavior.
src/​main/​java/​com/​cloudogu/​gitops/​tools/​core/​argocd/​ArgoCD.java Implements idempotent bootstrap and CRD ownership handling.
src/​main/​java/​com/​cloudogu/​gitops/​infrastructure/​kubernetes/​api/​K8sClientHelper.java Supports CRD resolution and aliases.
src/​main/​java/​com/​cloudogu/​gitops/​infrastructure/​kubernetes/​api/​K8sClient.java Adds resource existence checks.
src/​main/​java/​com/​cloudogu/​gitops/​destroy/​ArgoCDDestructionHandler.java Aligns Helm cleanup with installation-specific release names.
argocd/​cluster-resources/​apps/​argocd/​applications/​argocd.ftl.yaml Configures installation-specific Helm releases.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@mdroll mdroll left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

pr-595-full-prefix-run.txt

pr-595-full-prefix-run.txt
Code looks fine BUT:

As describe in your instructions:

  • 1st full-profil rollout - WORKS
  • 2nd full-profil rollout - WORKS
  • 1st full-prefix rollout - DOESNT WORKT

`3:55:26.998 [main] TRACE c.c.g.infrastructure.helm.HelmClient - Executing helm command: helm repo add registry https://twuni.github.io/docker-registry.helm
13:55:26.998 [main] TRACE c.c.gitops.utils.CommandExecutor - Executing command: '[helm, repo, add, registry, https://twuni.github.io/docker-registry.helm]'
"registry" already exists with the same configuration, skipping
13:55:27.049 [main] TRACE c.c.g.infrastructure.helm.HelmClient - Executing helm command: helm upgrade -i docker-registry registry/docker-registry --create-namespace --version 3.0.0 --values /tmp/gitops-playground-16655252673914750897 --namespace my-prefix-registry
13:55:27.049 [main] TRACE c.c.gitops.utils.CommandExecutor - Executing command: '[helm, upgrade, -i, docker-registry, registry/docker-registry, --create-namespace, --version, 3.0.0, --values, /tmp/gitops-playground-16655252673914750897, --namespace, my-prefix-registry]'
Release "docker-registry" does not exist. Installing it now.
Error: 1 error occurred:
* Service "docker-registry" is invalid: spec.ports[0].nodePort: Invalid value: 30000: provided port is already allocated

13:55:27.458 [main] ERROR c.c.gitops.utils.CommandExecutor - Executing command failed: helm upgrade -i docker-registry registry/docker-registry --create-namespace --version 3.0.0 --values /tmp/gitops-playground-16655252673914750897 --namespace my-prefix-registry
13:55:27.459 [main] ERROR c.c.gitops.utils.CommandExecutor - Stderr: Error: 1 error occurred:
* Service "docker-registry" is invalid: spec.ports[0].nodePort: Invalid value: 30000: provided port is already allocated
13:55:27.459 [main] ERROR c.c.gitops.utils.CommandExecutor - StdOut: Release "docker-registry" does not exist. Installing it now.
13:55:27.478 [main] ERROR c.c.g.cli.GitopsPlaygroundCliMain -
java.lang.IllegalStateException: Executing command failed: helm upgrade -i docker-registry registry/docker-registry --create-namespace --version 3.0.0 --values /tmp/gitops-playground-16655252673914750897 --namespace my-prefix-registry
at com.cloudogu.gitops.utils.CommandExecutor.getOutput(CommandExecutor.java:237)
at com.cloudogu.gitops.utils.CommandExecutor.execute(CommandExecutor.java:40)
at com.cloudogu.gitops.utils.CommandExecutor.execute(CommandExecutor.java:34)
at com.cloudogu.gitops.infrastructure.helm.HelmClient.helm(HelmClient.java:67)
at com.cloudogu.gitops.infrastructure.helm.HelmClient.upgrade(HelmClient.java:32)
at com.cloudogu.gitops.infrastructure.deployment.HelmStrategy.deployFeature(HelmStrategy.java:69)
at com.cloudogu.gitops.infrastructure.deployment.HelmStrategy.deployFeature(HelmStrategy.java:35)
at com.cloudogu.gitops.infrastructure.deployment.Deployer.deployFeature(Deployer.java:35)
at com.cloudogu.gitops.tools.common.AbstractTool.deployHelmChart(AbstractTool.java:227)
at com.cloudogu.gitops.tools.Registry.deployInternalRegistry(Registry.java:100)
at com.cloudogu.gitops.tools.Registry.deploy(Registry.java:70)
at com.cloudogu.gitops.tools.common.AbstractTool.execute(AbstractTool.java:61)
at com.cloudogu.gitops.application.orchestration.DeploymentOrchestrator.deployTools(DeploymentOrchestrator.java:30)
at com.cloudogu.gitops.application.Application.start(Application.java:73)
at com.cloudogu.gitops.cli.GitopsPlaygroundCli.run(GitopsPlaygroundCli.java:126)
at com.cloudogu.gitops.cli.GitopsPlaygroundCliMain.exec(GitopsPlaygroundCliMain.java:15)
at com.cloudogu.gitops.cli.GitopsPlaygroundCliMain.main(GitopsPlaygroundCliMain.java:9)

Process finished with exit code 2`

@mdroll mdroll left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

After --registry-url suggestion the previous error doesn't occur anymore.

Observed anomalies

Point 6 in ADR: Check helm release-name of the prefixed argocd Application

If I'm executing helm list -A - I just getting this. The my-prefix-argocd is missing
helm list -A
NAME NAMESPACE REVISION UPDATED STATUS CHART APP VERSION
docker-registry registry 2 2026-09-21 16:29:10.959193528 +0200 CEST deployed docker-registry-3.0.0 3.0.0
jenkins jenkins 2 2026-09-21 16:29:17.186630106 +0200 CEST deployed jenkins-5.9.56 2.568.3
jenkins my-prefix-jenkins 2 2026-09-21 16:54:02.328404105 +0200 CEST deployed jenkins-5.9.56 2.568.3
my-prefix-scmm my-prefix-scm-manager 2 2026-09-21 16:53:54.636492831 +0200 CEST deployed scm-manager-3.11.10 3.11.10
scmm scm-manager 2 2026-09-21 16:29:08.743065055 +0200 CEST deployed scm-manager-3.11.10 3.11.10

Cluster abnormalities

Triggering of ci pipelione of petclinic

Every time if the GOP is run the ci of the Jenkins is triggered, so that the Petclinic Jobs are executed multiple times

Multiple kube-system svclb-traefik-xy pods

`- kube-system svclb-traefik-xy running 2/2

  • kube-system svclb-traefik-xy running: 0/2

Warning FailedScheduling 6m11s default-scheduler 0/1 nodes are available: 1 node(s) didn't have free ports for the requested pod ports. no new claims to deallocate, preemption: 0/1 nodes are available: │
│ 1 node(s) didn't have free ports for the requested pod ports. │
│ Warning FailedScheduling 43s default-scheduler 0/1 nodes are available: 1 node(s) didn't have free ports for the requested pod ports. no new claims to deallocate, preemption: 0/1 nodes are available: │
│ 1 node(s) didn't have free ports for the requested pod ports.`

external-secrets-cert-controller Ready 0/1

│ Normal Pulling 26m kubelet Pulling image "ghcr.io/external-secrets/external-secrets:v0.9.16" │ │ Normal Pulled 26m kubelet Successfully pulled image "ghcr.io/external-secrets/external-secrets:v0.9.16" in 8.665s (8.665s including waiting). Image size: 61993465 bytes. │ │ Normal Created 26m kubelet Container created │ │ Normal Started 26m kubelet Container started │ │ Warning Unhealthy 32s (x56 over 5m35s) kubelet Readiness probe failed: HTTP probe failed with statuscode: 500

cert-manager CrashLoopBackOff Ready 0/1:

│ Type Reason Age From Message │ │ ---- ------ ---- ---- ------- │ │ Normal Scheduled 45m default-scheduler Successfully assigned cert-manager/cert-manager-cainjector-97fbb7746-7b5kc to k3d-gitops-playground-server-0 │ │ Normal Pulling 45m kubelet Pulling image "quay.io/jetstack/cert-manager-cainjector:v1.19.4" │ │ Normal Pulled 45m kubelet Successfully pulled image "quay.io/jetstack/cert-manager-cainjector:v1.19.4" in 7.533s (7.533s including waiting). Image size: 12514979 bytes. │ │ Normal Created 103s (x9 over 45m) kubelet Container created │ │ Normal Started 103s (x9 over 45m) kubelet Container started │ │ Normal Pulled 103s (x8 over 19m) kubelet Container image "quay.io/jetstack/cert-manager-cainjector:v1.19.4" already present on machine and can be accessed by the pod │ │ Warning BackOff 28s (x19 over 17m) kubelet Back-off restarting failed container cert-manager-cainjector in pod cert-manager-cainjector-97fbb7746-7b5kc_cert-manager(52977d08-05f2-4709-aa6d-a │ │ fcda312d0cf)

@avetgit

avetgit commented Sep 22, 2026

Copy link
Copy Markdown
Contributor Author

After --registry-url suggestion the previous error doesn't occur anymore.

Observed anomalies

Point 6 in ADR: Check helm release-name of the prefixed argocd Application

If I'm executing helm list -A - I just getting this. The my-prefix-argocd is missing helm list -A NAME NAMESPACE REVISION UPDATED STATUS CHART APP VERSION docker-registry registry 2 2026-09-21 16:29:10.959193528 +0200 CEST deployed docker-registry-3.0.0 3.0.0 jenkins jenkins 2 2026-09-21 16:29:17.186630106 +0200 CEST deployed jenkins-5.9.56 2.568.3 jenkins my-prefix-jenkins 2 2026-09-21 16:54:02.328404105 +0200 CEST deployed jenkins-5.9.56 2.568.3 my-prefix-scmm my-prefix-scm-manager 2 2026-09-21 16:53:54.636492831 +0200 CEST deployed scm-manager-3.11.10 3.11.10 scmm scm-manager 2 2026-09-21 16:29:08.743065055 +0200 CEST deployed scm-manager-3.11.10 3.11.10

Cluster abnormalities

Triggering of ci pipelione of petclinic

Every time if the GOP is run the ci of the Jenkins is triggered, so that the Petclinic Jobs are executed multiple times

Multiple kube-system svclb-traefik-xy pods

`- kube-system svclb-traefik-xy running 2/2

* kube-system svclb-traefik-xy running: 0/2

Warning FailedScheduling 6m11s default-scheduler 0/1 nodes are available: 1 node(s) didn't have free ports for the requested pod ports. no new claims to deallocate, preemption: 0/1 nodes are available: │ │ 1 node(s) didn't have free ports for the requested pod ports. │ │ Warning FailedScheduling 43s default-scheduler 0/1 nodes are available: 1 node(s) didn't have free ports for the requested pod ports. no new claims to deallocate, preemption: 0/1 nodes are available: │ │ 1 node(s) didn't have free ports for the requested pod ports.`

external-secrets-cert-controller Ready 0/1

│ Normal Pulling 26m kubelet Pulling image "ghcr.io/external-secrets/external-secrets:v0.9.16" │ │ Normal Pulled 26m kubelet Successfully pulled image "ghcr.io/external-secrets/external-secrets:v0.9.16" in 8.665s (8.665s including waiting). Image size: 61993465 bytes. │ │ Normal Created 26m kubelet Container created │ │ Normal Started 26m kubelet Container started │ │ Warning Unhealthy 32s (x56 over 5m35s) kubelet Readiness probe failed: HTTP probe failed with statuscode: 500

cert-manager CrashLoopBackOff Ready 0/1:

│ Type Reason Age From Message │ │ ---- ------ ---- ---- ------- │ │ Normal Scheduled 45m default-scheduler Successfully assigned cert-manager/cert-manager-cainjector-97fbb7746-7b5kc to k3d-gitops-playground-server-0 │ │ Normal Pulling 45m kubelet Pulling image "quay.io/jetstack/cert-manager-cainjector:v1.19.4" │ │ Normal Pulled 45m kubelet Successfully pulled image "quay.io/jetstack/cert-manager-cainjector:v1.19.4" in 7.533s (7.533s including waiting). Image size: 12514979 bytes. │ │ Normal Created 103s (x9 over 45m) kubelet Container created │ │ Normal Started 103s (x9 over 45m) kubelet Container started │ │ Normal Pulled 103s (x8 over 19m) kubelet Container image "quay.io/jetstack/cert-manager-cainjector:v1.19.4" already present on machine and can be accessed by the pod │ │ Warning BackOff 28s (x19 over 17m) kubelet Back-off restarting failed container cert-manager-cainjector in pod cert-manager-cainjector-97fbb7746-7b5kc_cert-manager(52977d08-05f2-4709-aa6d-a │ │ fcda312d0cf)

=========================================================================================
helm list -A does not show the ArgoCD release by design. GOP removes the ArgoCD Helm release metadata after the initial bootstrap, because ArgoCD manages itself via GitOps afterwards.
Point 6 refers to the helm.releaseName configured in the self-managed ArgoCD Application. Check it with:

kubectl -n my-prefix-argocd get application argocd
-o jsonpath='{.spec.source.helm.releaseName}{"\n"}'

Expected output:

my-prefix-argocd

@mdroll mdroll left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

After more clearification it now lgtm.

@mdroll
mdroll merged commit 2152a61 into develop Sep 22, 2026
3 checks passed
@mdroll
mdroll deleted the fix/idempotent-argocd-bootstrap branch September 22, 2026 12:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants