Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions .github/workflows/cd.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,12 @@ jobs:
with:
persist-credentials: false

- name: ⚙️ Setup Go
uses: actions/setup-go@b7ad1dad31e06c5925ef5d2fc7ad053ef454303e # v7.0.0
with:
go-version-file: go.mod
cache: false

- name: 💾 Validate two-stage PVC retirement
env:
PVC_PRUNE_BASE_SHA: ${{ github.sha }}
Expand Down
6 changes: 6 additions & 0 deletions .github/workflows/ci.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,12 @@ jobs:
go-version-file: go.mod
cache: false

- name: 💾 Validate Flux replacement safety
run: |
go test ./scripts/validate-flux-force-safety
shellcheck scripts/tests/test-flux-force-safety.sh
bash scripts/tests/test-flux-force-safety.sh

- name: 🧭 Validate GitHub bootstrap isolation
run: |
shellcheck scripts/guard-github-bootstrap-isolation.sh scripts/tests/test-github-bootstrap-isolation.sh
Expand Down
5 changes: 4 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -546,7 +546,10 @@ verify that deployment before treating the production lane as clean.
**Persistence retirement is always two-stage.** A merge-group artifact is speculative, but
Kubernetes PVC deletion is irreversible once `deletionTimestamp` is set: queue eviction and the
heal job cannot un-delete it. The production persistence-safety component therefore disables Flux
pruning on every PVC, HelmRelease, and Namespace, and disables Flux force replacement on PVCs.
pruning on every PVC, HelmRelease, and Namespace. Platform and generated tenant Flux layers set
`spec.force: false`; only individual Jobs that need recreation carry `force: enabled`.
`force: disabled` on a resource cannot override a forcing layer. The replacement-safety guard
checks layer defaults, both tenant template branches, and rendered patches before publication.
HelmRelease protection prevents chart uninstall from deleting chart-owned claims; Namespace
protection prevents cascading deletion from bypassing a claim's own annotation. To retire any of
these objects, first merge and deploy that protection in its own revision; only a later PR may
Expand Down
9 changes: 6 additions & 3 deletions docs/TENANTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -250,9 +250,12 @@ carry itself**:
only example in the repo: `wedding-app`'s CNPG `Cluster` `storage.storageClass: longhorn`
(local/docker has no longhorn, so the artifact stays class-agnostic and the overlay supplies
the class).
- **Operational-safety annotations** — platform-enforced guards such as
`kustomize.toolkit.fluxcd.io/{force,prune}: disabled` on a stateful resource to prevent a
Flux delete+recreate data-loss event.
- **Operational-safety settings** — `kustomize.toolkit.fluxcd.io/prune: disabled` preserves
a stateful resource when its manifest is removed. Platform and generated tenant Flux layers
set `spec.force: false` so immutable conflicts fail rather than replace resources. Jobs that
need recreation can opt in with `kustomize.toolkit.fluxcd.io/force: enabled`; persistent claims
and database clusters must not opt in. A resource's `force: disabled` annotation cannot
override a forcing layer.

**Everything a tenant can express in its own `deploy/` is tenant-owned and must NOT be patched
here** — **hostnames**, **`gethomepage.dev/*` dashboard annotations**, routes, and app config:
Expand Down
23 changes: 13 additions & 10 deletions docs/deletion-and-data-retention.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,11 +76,13 @@ Read from the charts this repository references and from the released source of
`Released` and the volume controller does nothing more with it.
- **CSI provisioner.** It deletes the storage behind a volume only when the volume is `Released`
and its policy is `Delete`.
- **Flux kustomize-controller 1.8.** All four layers set `force: true`, so an object whose change
the API server rejects as invalid, which is what an immutable field produces, is deleted and
created again. The annotation `kustomize.toolkit.fluxcd.io/force` is read only as an opt-in
(`enabled`): `force: disabled` does not exempt an object from a layer that forces. Within one
Kustomization, classes are applied in an earlier stage than HelmReleases.
- **Flux kustomize-controller 1.8.** Platform and generated tenant layers set `force: false`:
an immutable conflict fails without replacing the object. Three setup Jobs explicitly opt
into recreation with `kustomize.toolkit.fluxcd.io/force: enabled`. That annotation is only an
opt-in; `force: disabled` cannot exempt an object from a forcing layer. The repository guard
rejects forcing layers, including tenant templates and rendered patches, and force opt-ins
on manifest-owned claims or database clusters. Within one Kustomization, classes are applied
in an earlier stage than HelmReleases.
Comment thread
coderabbitai[bot] marked this conversation as resolved.
- **Flux helm-controller 1.5 (Helm 4.2).** Helm updates an object in place and never deletes it to
recreate it, so a changed `reclaimPolicy` fails the upgrade. Helm creates an object that is
missing, and deletes one that left the chart unless it carries `helm.sh/resource-policy: keep`.
Expand Down Expand Up @@ -154,11 +156,12 @@ would keep creating `Delete` volumes until the workload itself is recreated.
rebuilt cluster. It only ever moves `Delete` to `Retain`. It leaves Velero's temporary volumes
alone, which are the ones bound to claims in Velero's own namespace. And it leaves a deliberate
way to discard a single released volume.
5. **Forced replacement is a separate hazard with the same backstop.** All four layers force, and
`kustomize.toolkit.fluxcd.io/force: disabled` on a claim or a database cluster does not exempt
it. A change to such an object that the API server rejects is therefore answered by deleting
and recreating it. `Retain` turns that from data loss into an outage with recoverable data. It
does not prevent it. Preventing it is tracked in #4448 and is not part of this rollout.
5. **Forced replacement is a separate hazard with the same backstop.** Platform and generated
tenant layers set `force: false`, and the repository guard rejects forcing layers or force
opt-ins on manifest-owned claims and database clusters. Three setup Jobs explicitly opt into
recreation. `force: disabled` still cannot exempt an object from a forcing layer, while
`Retain` protects data but not service availability. The force-safety safeguard tracked by
#4448 is enforced separately from the storage-retention rollout described here.
6. **The order is fixed:** retain the data, prove it can be brought back, stop evicted candidates
deleting anything, and only then remove the opt-outs. Steps 2 and 3 of the order first written
on #3369 are swapped; the reason is under "Before the opt-outs are removed".
Expand Down
3 changes: 2 additions & 1 deletion k8s/bases/apps/ascoachingogvaner/flux-kustomization.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,8 @@ spec:
path: .
prune: true
wait: true
force: true
# Immutable conflicts require an operator plan; individual Jobs opt into recreation.
force: false
serviceAccountName: ascoachingogvaner
sourceRef:
kind: OCIRepository
Expand Down
7 changes: 2 additions & 5 deletions k8s/bases/apps/backstage/cluster.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -13,11 +13,8 @@ metadata:
name: backstage-db
namespace: backstage
annotations:
# Never let Flux force-delete+recreate this DB: the apps Flux Kustomization
# runs force + prune and longhorn PVCs are reclaimPolicy=Delete, so a
# force-recreate would wipe the data (the failure that hit coroot-db on
# 2026-06-18). Flux can still update it in place.
kustomize.toolkit.fluxcd.io/force: disabled
# The owning Flux layer has force=false, so immutable conflicts fail rather
# than replacing this database. Preserve it when its manifest is removed.
kustomize.toolkit.fluxcd.io/prune: disabled
spec:
# Pin the PostgreSQL operand image to clear fixable base-OS + PostgreSQL CVEs in
Expand Down
3 changes: 2 additions & 1 deletion k8s/bases/apps/github-config/flux-kustomization.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,8 @@ spec:
# operations (including manual repository bootstrap) cannot gate app delivery.
# Crossplane's Synced condition is monitored by crossplane-sync-alert instead.
wait: false
force: true
# Immutable conflicts require an operator plan; individual Jobs opt into recreation.
force: false
# Runs as the app SA: the GitHub managed resources are applied with only the
# namespace-scoped authority in rbac.yaml. This Kustomization is itself applied
# by the `apps` Flux Kustomization; it must NOT dependsOn it (deadlock) — it
Expand Down
11 changes: 4 additions & 7 deletions k8s/bases/apps/umami/cluster.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -14,12 +14,8 @@ metadata:
name: umami-db
namespace: umami
annotations:
# Never let Flux delete+recreate this database: the apps Flux Kustomization
# runs force: true + prune: true, and the longhorn PVCs are reclaimPolicy=
# Delete, so a force-recreate would wipe the data (the failure that hit
# coroot-db on 2026-06-18). Flux can still update it; immutable changes must
# be made by an operator with a backup/restore plan.
kustomize.toolkit.fluxcd.io/force: disabled
# The owning Flux layer has force=false: immutable conflicts require an
# operator's backup/restore plan. Keep this database when its manifest is removed.
kustomize.toolkit.fluxcd.io/prune: disabled
spec:
# Pin the PostgreSQL operand image explicitly instead of riding the CNPG operator
Expand All @@ -32,7 +28,8 @@ spec:
# 11 PostgreSQL CVEs incl. the RCE-class CVE-2026-6473. Renovate keeps this current
# via the cluster.yaml customManager in .github/renovate.json; CNPG rolls
# the image in place (replicas first, then a primary switchover), so the >=2-instance
# HA Clusters stay available with no data-loss path (force/prune are disabled above).
# HA Clusters remain available; Flux force=false and the prune annotation
# prevent automatic replacement and garbage collection of this Cluster.
# renovate: datasource=docker depName=ghcr.io/cloudnative-pg/postgresql
imageName: ghcr.io/cloudnative-pg/postgresql:18.4-system-trixie
# Run 3 instances (1 primary + 2 streaming replicas) for high availability.
Expand Down

This file was deleted.

Original file line number Diff line number Diff line change
Expand Up @@ -3,4 +3,3 @@ apiVersion: kustomize.config.k8s.io/v1alpha1
kind: Component
transformers:
- annotations-transformer-production-persistence-prune.yaml
- annotations-transformer-production-pvc-force.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -304,7 +304,7 @@ spec:
path: .
prune: true
wait: true
force: true
force: false
# A tenant that ships a CNPG database (wedding-app does) would
# otherwise fail this Kustomization for the whole of any rolling
# update — CNPG rolls every instance on an operator, barman-cloud
Expand Down Expand Up @@ -341,7 +341,7 @@ spec:
path: .
prune: true
wait: true
force: true
force: false
# A tenant that ships a CNPG database (wedding-app does) would
# otherwise fail this Kustomization for the whole of any rolling
# update — CNPG rolls every instance on an operator, barman-cloud
Expand Down
3 changes: 2 additions & 1 deletion k8s/clusters/base/flux-kustomization-apps.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -56,4 +56,5 @@ spec:
kind: Cluster
current: has(status.readyInstances) && status.readyInstances >= 1 && has(status.currentPrimary) && status.currentPrimary != ''
prune: true
force: true
# Immutable conflicts require an operator plan; individual Jobs opt into recreation.
force: false
3 changes: 2 additions & 1 deletion k8s/clusters/base/flux-kustomization-bootstrap.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -25,4 +25,5 @@ spec:
name: sops-age
wait: true
prune: true
force: true
# Immutable conflicts require an operator plan; individual Jobs opt into recreation.
force: false
Original file line number Diff line number Diff line change
Expand Up @@ -37,4 +37,5 @@ spec:
name: variables-cluster
wait: true
prune: true
force: true
# Immutable conflicts require an operator plan; individual Jobs opt into recreation.
force: false
3 changes: 2 additions & 1 deletion k8s/clusters/base/flux-kustomization-infrastructure.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -55,4 +55,5 @@ spec:
kind: Cluster
current: has(status.readyInstances) && status.readyInstances >= 1 && has(status.currentPrimary) && status.currentPrimary != ''
prune: true
force: true
# Immutable conflicts require an operator plan; individual Jobs opt into recreation.
force: false
Original file line number Diff line number Diff line change
Expand Up @@ -55,16 +55,8 @@ metadata:
namespace: aws
labels:
app.kubernetes.io/managed-by: ksail
annotations:
# The top-level `apps` Kustomization runs force: true, but this is a
# Crossplane managed resource with the default deletionPolicy: Delete — a
# force delete/recreate on immutable-field drift would cascade into DELETING
# the backing AWS IAM policy (the same class of failure that wiped coroot-db
# on 2026-06-18). The tenant Kustomization avoids force for exactly this
# reason; these platform-owned MRs opt out per-resource instead (#2631
# review). Flux can still update them in place; a genuine immutable conflict
# is an operator decision, not an automatic replace.
kustomize.toolkit.fluxcd.io/force: disabled
# Both platform and tenant Flux layers have force=false. An immutable conflict
# is an operator decision rather than recreation of this managed IAM policy.
spec:
forProvider:
description: >-
Expand Down
7 changes: 2 additions & 5 deletions k8s/providers/hetzner/apps/aws/role-eks-ci.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -47,11 +47,8 @@ metadata:
namespace: aws
labels:
app.kubernetes.io/managed-by: ksail
annotations:
# See policy-eks-ci-smoke-boundary.yaml: the `apps` Kustomization runs
# force: true, and a force-recreate of a Crossplane MR with the default
# deletionPolicy: Delete would DELETE the backing AWS IAM role. Opt out.
kustomize.toolkit.fluxcd.io/force: disabled
# The owning Flux layer has force=false; an immutable conflict cannot
# automatically recreate this managed resource and delete its backing IAM role.
spec:
forProvider:
description: >-
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -21,11 +21,8 @@ spec:
metadata:
name: wedding-db
annotations:
# Never let Flux delete+recreate this database (force-recreate +
# reclaimPolicy=Delete PVCs = data loss, the failure that hit
# coroot-db on 2026-06-18). Flux can still update it; immutable
# changes need a manual backup/restore by an operator.
kustomize.toolkit.fluxcd.io/force: disabled
# The tenant's Flux layer has force=false. Preserve the database
# if its manifest is removed; immutable changes need an operator plan.
kustomize.toolkit.fluxcd.io/prune: disabled
spec:
# Quorum-based SYNCHRONOUS replication, same rationale as umami-db: a
Expand Down
12 changes: 3 additions & 9 deletions k8s/providers/hetzner/infrastructure/coroot/cluster.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -31,15 +31,9 @@ metadata:
name: coroot-db
namespace: observability
annotations:
# Never let Flux delete+recreate this database. The infrastructure Flux
# Kustomization runs force: true + prune: true, so an immutable-field change
# or a live/Git drift on this Cluster would otherwise make kustomize-
# controller recreate it — and with the longhorn PVCs on reclaimPolicy=
# Delete that wipes the data via a fresh initdb (this happened 2026-06-18).
# force: disabled makes Flux surface the apply error for an operator to
# resolve deliberately (with a backup/restore plan) instead of recreating;
# prune: disabled stops GC if the Cluster is ever dropped from a kustomization.
kustomize.toolkit.fluxcd.io/force: disabled
# The owning Flux layer has force=false: immutable conflicts fail for an
# operator to resolve with a backup/restore plan. This annotation separately
# prevents garbage collection if the database manifest is removed.
kustomize.toolkit.fluxcd.io/prune: disabled
# CloudNativePG treats this supported annotation as a requested rolling
# restart. The timestamp is fixed declarative desired state: changing it is
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -25,8 +25,6 @@ kind: PersistentVolumeClaim
metadata:
name: vault-snapshots
namespace: openbao
annotations:
kustomize.toolkit.fluxcd.io/force: disabled
spec:
# Hetzner Cloud Volumes attach reliably via the API to any node in the
# location, unlike a Longhorn RWO volume mounted by node-mobile Jobs.
Expand Down
10 changes: 5 additions & 5 deletions scripts/rgd-template-static-scan-baseline.tsv
Original file line number Diff line number Diff line change
Expand Up @@ -12,15 +12,15 @@
# change; a new finding is fixed, not baselined.
# Measured with Trivy v0.74.0 and the policies embedded in that binary. CI pins the binary on pull
# requests, merge groups, and direct main pushes; the gate rejects other versions and policy updates.
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml RGD-CONTENT SHA256 01953c88886b926f0fdbfa2063cecd0139f63cd801ebc5bd59ce57e2e11eb92e
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml RGD-CONTENT SHA256 164441249612faccef24576f8cec7c7c70a527b81f45574f13dec917439c26d1
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/defaultDeny.yaml KSV-0039 LOW CAUSE-SHA256:6246c5e7ede528806f1f84a2b466b1a22c061b7d7f2ad38e18dc4c722ea7b32b
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/defaultDeny.yaml KSV-0040 LOW CAUSE-SHA256:6246c5e7ede528806f1f84a2b466b1a22c061b7d7f2ad38e18dc4c722ea7b32b
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/ghcrAuth.yaml KSV-0039 LOW CAUSE-SHA256:f3c7ced815c76aaba6407cf691461a7e46b0037f888f32d760286b0fde1bd286
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/ghcrAuth.yaml KSV-0040 LOW CAUSE-SHA256:f3c7ced815c76aaba6407cf691461a7e46b0037f888f32d760286b0fde1bd286
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/kustomization.yaml KSV-0039 LOW CAUSE-SHA256:1ebd8487222e03b649f95b96fb502578bd5f3fee3a775f5b54b450ec12ea9f81
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/kustomization.yaml KSV-0040 LOW CAUSE-SHA256:1ebd8487222e03b649f95b96fb502578bd5f3fee3a775f5b54b450ec12ea9f81
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/kustomizationSops.yaml KSV-0039 LOW CAUSE-SHA256:62bdc59f0e694bd4baccf8f6c2523696b1afd6ad47c6a53be670419a1e42038d
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/kustomizationSops.yaml KSV-0040 LOW CAUSE-SHA256:62bdc59f0e694bd4baccf8f6c2523696b1afd6ad47c6a53be670419a1e42038d
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/kustomization.yaml KSV-0039 LOW CAUSE-SHA256:2b3834d0822048e2aaefbd8076a28785cf8781b65230ea8ea37eb854e1169bc4
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/kustomization.yaml KSV-0040 LOW CAUSE-SHA256:2b3834d0822048e2aaefbd8076a28785cf8781b65230ea8ea37eb854e1169bc4
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/kustomizationSops.yaml KSV-0039 LOW CAUSE-SHA256:59d21ae5decd83ec9c1e66036a71db7aa29ca574b4e3ef0fb803a8b45d7fac61
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/kustomizationSops.yaml KSV-0040 LOW CAUSE-SHA256:59d21ae5decd83ec9c1e66036a71db7aa29ca574b4e3ef0fb803a8b45d7fac61
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/limitRange.yaml KSV-0039 LOW CAUSE-SHA256:c37b17db55af250264e52c9089e25decadde834c9ebae540db73774a4f2b5702
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/limitRange.yaml KSV-0040 LOW CAUSE-SHA256:c37b17db55af250264e52c9089e25decadde834c9ebae540db73774a4f2b5702
1 k8s/bases/infrastructure/resource-graph-definitions/tenant/resource-graph-definition.yaml.resources/ociRepository.yaml KSV-0039 LOW CAUSE-SHA256:7655439f41ac63530bbbacc158dbf9a0ea3b63d14d1e7135ce6b25e2218adb42
Expand Down
16 changes: 16 additions & 0 deletions scripts/tests/test-flux-force-safety.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
#!/usr/bin/env bash
set -euo pipefail

root_dir="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
readonly root_dir
render_dir="$(mktemp -d)"
readonly render_dir
trap 'rm -rf "${render_dir}"' EXIT

cd "${root_dir}"
# Source checks cover both KRO tenant branches; rendered checks cover patches.
kubectl kustomize k8s/clusters/prod >"${render_dir}/cluster.yaml"
for overlay in bootstrap infrastructure/controllers infrastructure apps; do
kubectl kustomize "k8s/providers/hetzner/${overlay}" >"${render_dir}/${overlay//\//-}.yaml"
done
go run ./scripts/validate-flux-force-safety k8s "${render_dir}"
Loading
Loading