Schedule the deletion drain safely - #7827
Conversation
d0c2af1 to
6549909
Compare
6549909 to
b79bd19
Compare
wesbillman
left a comment
There was a problem hiding this comment.
Carl, an automated reviewer, commenting via Wes’s GitHub account.
No new blocking defect found in the incremental scheduler at b79bd19c73b188d9fb785cbcd5ba50ccad74b295, against acf561219f4223bc6090ce5c82555ee294d11ea4. This is a source review, not an approval.
The disabled-by-default, fixed-command job preserves the existing approval boundary and PostgreSQL lease/checkpoint authority without exposing a generic command or signing-secret environment. Two non-blocking runbook corrections are inline; no engine redesign is requested.
Validation: existing chart CI passed all 11 suites / 54 tests, including the deletion-drain suite. Its checkout 39edc1151804cc7fcece70e82d32000f5f66eb47 has the same tree as this head. Reviewers did not check out, render, build, or execute PR code. Live Kubernetes interruption, workload identity, and rollout behavior were not exercised.
This does not clear the separate base-layer #7818 recovery blocker.
wpfleger96
left a comment
There was a problem hiding this comment.
🤖 Thanks for this. I reviewed head b79bd19c against acf56121 (the head of #7818), covering only this PR's two commits. I found no blockers.
The job is scoped tightly:
- It runs one typed command,
buzz-admin deletions drain. - It gets only the eight env entries the deletion engine reads.
- It has no
envFromand no relay signing key. - It runs with
Forbid,restartPolicy: NeverandbackoffLimit: 0. additionalProperties: falsekeeps arbitrary command/env settings out of the values schema.
PostgreSQL keeps the retry and approval authority. Claim requires an approved request, and submitted/inventoried requests are never claimable, so the note that this can't make a new owner request destructive is accurate.
We also ran it locally:
- The chart suite passed 54/54 with Helm 3.16.4 and helm-unittest 0.8.2. The enabled quickstart render passes strict kubeconform for Kubernetes 1.31 (18/18 resources).
- We broke the chart five ways and the suite failed each time: removing the reserved-label
fail, settingconcurrencyPolicy: Allow, settingbackoffLimit: 6, dropping the Redis validation, and adding an extra secret env. - A local build of
buzz-admin deletions drainran against fresh Postgres/Redis with exactly the chart's eight env vars and an unwritableTMPDIR. On an empty queue it exited 0 with no output.
Non-blocking:
- CronJob name length.
{{ include "buzz.fullname" }}-deletion-drainadds 15 characters to a name capped at 63. Kubernetes rejects CronJob names over 52, so any release whose fullname is over 37 characters fails at install once the job is enabled. The failure is loud, the job is opt-in, and the existing-storage-accountingCronJob has the same limit with an even longer suffix. So I'm not blocking on this, but I think one small helper that truncates the base to fit before adding the suffix would fix both CronJobs. A long-release-name test would pin it. - Service account and cloud IAM. When
serviceAccountNameis empty, the job uses the relay's service account and so inherits the relay's cloud role.automountServiceAccountToken: falseonly removes the Kubernetes API token; admission-injected workload identity, such as the EKS pod-identity webhook's projected token and AWS env, still gets added. So the README line saying the job doesn't receive a service-account token is broader than what actually happens. I'd say explicitly that the default inherits the relay's IAM, and recommend a pre-created dedicated account when IAM isolation matters. - Deadline and shutdown wording.
activeDeadlineSeconds: 3600is more than the 60-second renewable lease, but it isn't a limit on total work, since a large queue can outrun it. SIGTERM cancels stage execution and tries to release the lease, but not every claim/heartbeat/release await is bounded by cancellation, andstatement_timeoutdefaults to zero. So exiting within the 30-second grace period isn't guaranteed, and a SIGKILL leaves the lease until it expires and gets reclaimed. That recovery path works; it would just be good for the runbook to describe it that way rather than promise a graceful finish. It should also mention that a stage that can't fit inside the deadline gets restarted by each successive Job without a recorded error. - Test coverage gaps.
- The reserved-label test only covers
app.kubernetes.io/component. - Removing the
failis caught by a YAML duplicate-key error, not by the guard's message, sonameandinstancehave no coverage of their own. - No test asserts the pod's own labels (only the CronJob metadata), the default
sidecar.istio.io/inject: "false"annotation, thes3.endpointrequired, the unsupported-typefail, or schema rejection of unknown keys.
- The reserved-label test only covers
CI is green at this head: 11 pass, 28 skipped.
b79bd19 to
494d574
Compare
|
Addressed the runbook feedback at head
No chart behavior changed. The pre-existing shared CronJob name-length convention and broader optional chart-test expansion remain out of this review-fix scope. Helm validation passed 54/54 plus lint and the render matrix. Both inline threads have substantive replies and are resolved. AI-generated implementation reply (Elrond). |
wesbillman
left a comment
There was a problem hiding this comment.
Carl, an automated reviewer, commenting via Wes’s GitHub account.
Review clear: no new blocking defect found. The prior runbook findings are addressed: deadline interruption is distinguished from durable dependency retries, non-resumable-stage recovery is documented, and manual runs discover the actual CronJob name. The separately reviewed base #7818 recovery fix is present. The opt-in typed job preserves the existing approval and PostgreSQL lease/checkpoint boundaries; no generic command/environment mechanism is introduced.
Reviewed HEAD 494d5744360f2d39c464e8ac8b83ae2276ccf0f6 against BASE 8ebe3b881c98f19cfc93da7c1d00d4b0618a356e, including chart/credential profiles, CLI/image/env concordance, shutdown/reclaim, and operator guidance. All assigned lanes are complete. Source-only on pinned Blox; no checkout, render, build, or execution by reviewers.
Existing chart CI passed 11 suites / 54 tests plus its fixture render matrix. It tested merge 2086b18b84fe22e67ce1eb16d15619c83d78d623, whose tree equals this head. The enabled job is exercised by helm-unittest, not the fixture matrix; the inline coverage follow-up is non-blocking. Live Kubernetes interruption/IAM behavior remains untested. Broad CI has a Desktop Smoke E2E (4) failure, not attributed here; required CI remains a merge gate. This comment is not approval.
wpfleger96
left a comment
There was a problem hiding this comment.
🤖 Re-review at 494d5744, which sits on the current #7818 head (8ebe3b88). git range-diff shows your two existing commits are unchanged from b79bd19c, so the only new change is the runbook/values commit. I don't see anything blocking in it.
I checked the new runbook claims against source and they hold:
- Shutdown releases the claim without touching
retry_count/last_error, while a claim incrementsattempts. So "rising attempts, flat retry count, never blocked" is the right thing to tell operators to look for. freeze_destructive_manifestclears its partial chunks and re-enumerates from scratch, so only the object drain resumes mid-stage.DEFAULT_LEASE_DURATIONis 60s andHEARTBEAT_INTERVALis 10s, which matches the SIGKILL recovery text.jobTemplatehas no metadata labels while the pod template does, so finding attempts via pod labels is correct.- Discovering the CronJob by component/instance label fixes the
buzz-buzz-deletion-drainname confusion. - The deadline and termination sections now describe recovery honestly instead of promising a clean handoff.
One small wording nit: the service-account paragraph says IRSA bindings are "resolved by the node/metadata path". That's accurate for GKE Workload Identity, but on EKS it's the pod-identity webhook injecting its own projected token volume plus AWS_* env, which automountServiceAccountToken: false doesn't suppress. The conclusion that the pod inherits the relay's cloud role is right either way, so maybe just say "injected by the platform's workload-identity mechanism".
Still open from my last review, both non-blocking:
- CronJob name length.
<fullname>-deletion-drainstill fails install for fullnames over 37 characters, same as-storage-accounting. - Chart test gaps.
name/instancereserved labels, pod labels, the istio annotation, thes3.endpointrequired, and schema rejection of unknown keys. This overlaps with Carl's note about adding an enabled fixture underci/.
CI: everything passes except Desktop Smoke E2E (4) and its two aggregate checks. The failing specs are sidebar-snapshot, video-attachment and workflows, and this PR only touches the chart, docs and ARCHITECTURE.md. Since this is stacked on #7818, it also inherits the lock-order blocker I just raised there.
|
Desktop-smoke CI investigation pinned to The test uses A throwaway test-only experiment located each player under |
|
Published follow-ups at Changes: cover reserved-label/schema/S3 guards and enabled deployment fixtures; share a suffix-preserving 52-character CronJob-name helper between the two jobs; correct workload-identity documentation; scope both video-menu test players to the exact emitted event IDs. Production Desktop behavior and forced-right-click interactions are unchanged. Complexity improves by sharing the bounded-name construction; no new scheduler or deletion engine. Blox evidence (not a merge-ready claim):
Report and logs remain on |
wesbillman
left a comment
There was a problem hiding this comment.
Carl, an automated reviewer, commenting via Wes’s GitHub account.
No new blocking source findings. Reviewed HEAD dfd5db4e087b77215d6a81ab81342d025e28c3e4 against BASE 8ebe3b881c98f19cfc93da7c1d00d4b0618a356e, focused on the four commits since the prior clear review.
The shared naming helper keeps both CronJobs within the 52-character limit while preserving previously valid short names. Enabled bundled/existing-Secret fixtures and the new guard/schema/identity assertions address the earlier chart-coverage follow-up. Workload-identity wording now correctly distinguishes cloud credentials from the Kubernetes API token. The video-test selectors bind to the exact emitted message without weakening the product assertions. The disabled-by-default typed command and PostgreSQL approval/lease/retry boundaries are unchanged.
Validation: existing chart CI passed 11 suites / 60 tests and rendered both enabled credential-profile fixtures. Its checkout was synthetic merge 8221a84576020a7d596ffd2e1ec135bfcd4b855b (this head plus f856a337), not the standalone head/pinned-base tree. At the captured snapshot, Desktop Smoke E2E was unfinished, so the video-test change has no completed result claimed here; Security Review authorization failed and remains a workflow gate. Neither is attributed to a source defect in this review.
This is a source-review comment, not approval or merge-readiness certification. No reviewer checkout, render, build, test or PR-code execution; live Kubernetes interruption and workload-identity behavior remain unexercised.
|
Desktop residual deferred to #7945: at dfd5db4 versus base 8ebe3b8, Desktop production code is unchanged; the only Desktop change is the ID-scoped test locator, so this unchanged-path flake is treated as inherited for scope disposition, not as a proven root cause. The issue preserves the 18/20 and two distinct 19/20 failures plus the finite instrumented 20/20 non-reproduction. This path remains known-flaky, not fixed or certified stable. Terminal current-head CI, downstream stack updates, and live CronJob evidence/explicit exception remain separate gates. |
wpfleger96
left a comment
There was a problem hiding this comment.
🤖 Re-review at dfd5db4e. It's a fast-forward from 494d5744, still based on the #7818 branch at 8ebe3b88. Nothing blocking. The new commits pick up the earlier non-blocking notes: CronJob name length, chart test gaps, and the workload-identity wording.
buzz.cronJobName truncates only the fullname part so the result fits in 52 characters. The deletion drain keeps 37 characters of fullname plus -deletion-drain, and storage accounting keeps 33 plus -storage-accounting. trimSuffix drops a separator left dangling at the cut. Names that already fit don't change. We rendered base and head side by side: release-name, buzz, and buzz-prod give identical metadata.name for both CronJobs, so an upgrade doesn't rename a deployed job. A long release now renders very-long-release-name-for-operator-j-deletion-drain and very-long-release-name-for-operat-storage-accounting, both exactly 52 characters. At base those were 60 and 64. With only the trunc removed from the helper, the suite goes 58/60: both new long-name tests fail, and nothing else does.
The new tests check what they say they check. The istio annotation is now asserted as the inherited default rather than supplied by the test. Each reserved-label case has to hit its specific guard message. The existing-Secret case reaches the s3.endpoint required check, and the unknown-key case exercises the schema's additionalProperties: false. Both enabled credential fixtures go through the render matrix in helm-chart.yml. The README and runbook wording now matches what _operator-jobs.tpl does. It turns off the normal service-account token mount and service links. It doesn't stop a workload-identity webhook from injecting its own credentials.
A few optional notes:
- Truncation isn't unique.
very-long-release-name-for-operator-jobs-aand-bboth rendervery-long-release-name-for-operator-j-deletion-drain, so two long releases in one namespace would fight over the same CronJob. That's the same propertybuzz.fullname'strunc 63already has, and it only affects the new overlength case. If you want it closed, add a short hash of the untruncated fullname when truncating, plus a two-release assertion. Otherwise, a line in the runbook would do. - The bundled fixture only proves the chart renders. Nothing checks the pod's secret references against the generated Secret or the MinIO endpoint, and the unsupported-type branch in
_operator-jobs.tplhas no test. - The
video-attachment.spec.tschange is a good fix:.last()could pick the wrong video when two messages land in the same second. It doesn't belong to the deletion scheduler, though, and would read more cleanly as its own PR. - The PR description doesn't mention the storage-accounting CronJob now going through the shared name helper, or the desktop spec change. The Testing section also still cites
e68cb753.
CI at this head: the chart job passed 60/60 tests across 11 suites and rendered every fixture, and all four Desktop Smoke shards are green. Authorize Security Review fails with Pull request #7827 is not an open PR targeting main. That's because the PR is stacked, not a problem in the code. The base branch has since moved to c237fd46, so this needs a rebase onto current #7818 before it can land.
Signed-off-by: Codex <noreply@openai.com>
wpfleger96
left a comment
There was a problem hiding this comment.
🤖 Re-review at 464a4e60. Nothing blocking.
The new head is two merges on top of dfd5db4e, which I last reviewed clear. The first merge picks up #7818 at 1f1b5f4d, and the second picks it up at 9f6f7201. Across both, the PR's own diff hasn't changed. Once you strip blob IDs and hunk coordinates it's byte-identical to the one at dfd5db4e. The only difference is the surrounding context in the chart README, where #7341's readiness wording landed right above the new "Community deletion operator job" section. Between 8956cf83 and 464a4e60, nothing changed under deploy/charts/buzz in templates, values, the schema, tests, ci/ or Chart.yaml, and nothing changed in crates/buzz-admin, buzz-deletion, buzz-db, buzz-media, buzz-pubsub, migrations/ or the runbook. We started our checks at 8956cf83, and the head moved to 464a4e60 while we were working. The second merge only brings in main's relay readiness change, relay-crate code and docs, so nothing below depends on which of those two SHAs it ran at.
What I looked at is whether the base moving from 8ebe3b88 to 9f6f7201 makes anything in the unchanged runbook, README, ARCHITECTURE.md or _operator-jobs.tpl wrong. I don't think it does:
- The abort boundary still lines up. #7818 now says abort stays open through
fencedand closes fromdrainedonward "when tenant-state destruction may have begun". The runbook doesn't claim anything about abort. Its statement that owner admission stops atsubmitted, and that the engine won't claimsubmitted/inventoried, still matches the claim guard indeletion.rs. - The new owner-convergence serialization takes the shared community advisory lock before owner row locks, the same order abort uses. It changes lock ordering, not what the drain can claim.
- The drain doesn't need any new env or credentials.
buzz-admin,buzz-deletion,buzz-mediaandbuzz-pubsubare unchanged across the base move, andbuzz-admin deletionsdelegates straight to the deletion crate without building relay config. So the new relay listener, NIP-FI and readiness settings aren't missing inputs to the CronJob. The env set in_operator-jobs.tplis still exactly what the executor reads. - No migration reference went stale. Owner admission was renamed from
0050to0051with byte-identical contents, and the runbook only says "confirm database migrations are current" without pinning a number. The new0050_operator_listener_mentionstables are deployment-global and excluded from tenant fencing. - The lease/retry/checkpoint wording still holds, since the deletion executor is untouched. Shutdown releases the claim without recording a retry, and only the object drain resumes mid-stage.
On the chart side, with the CI-pinned Helm 3.16.4 and helm-unittest 0.8.2, we rendered ci/quickstart-values.yaml and tests/fixtures/deletion-drain-bundled-values.yaml at dfd5db4e and 8956cf83. The only differences were eight randomly generated quickstart Secret values and the relay checksum/secret, and re-rendering the same head changes exactly those nine lines. The enabled deletion-drain CronJob is byte-identical without any normalization. helm unittest is 60/60 across 11 suites.
A few non-blocking notes are still open from last time: CronJob name truncation isn't unique across two long releases in one namespace, the bundled fixture only proves the chart renders, and the unsupported-type branch has no test. The PR description also still cites e68cb753 in Testing and doesn't mention the storage-accounting CronJob moving to the shared name helper or the video-attachment.spec.ts change.
CI at 464a4e60: the chart job passed. Authorize Security Review fails with "Pull request #7827 is not an open PR targeting main", which comes from the stack, not the code. Desktop Core and the four Smoke shards were still running when I checked.
|
At exact head Proposed named exception: Blox private-cluster cgroupv2 startup blocker. The chart remains disabled by default ( The bounded setup, failed launch and cleanup are retained in |
Co-authored-by: Codex <noreply@openai.com> Signed-off-by: tornquist <tornquist@squareup.com>
Co-authored-by: Codex <noreply@openai.com> Signed-off-by: tornquist <tornquist@squareup.com>
The runbook told operators to target `cronjob/<release>-buzz-deletion-drain`. `buzz.fullname` collapses to the release name when it already contains the chart name, so the documented name is wrong for the documented install: `helm install buzz ...` renders `buzz-deletion-drain`. Discover the CronJob by its component label instead, and describe the name rule rather than a single guessed spelling. Three more corrections: `activeDeadlineSeconds` was presented as if a timed-out run were just another retry. It is not recorded as one — shutdown releases the claim without recording a retry, and only the object-store drain resumes mid-stage — so a deadline landing repeatedly inside a non-resumable stage loops forever with a rising `attempts`, a flat `retry_count`, and no block. Document how to spot that from both the request and Kubernetes, how to size the deadline, and how to recover. `terminationGracePeriodSeconds` was presented as a clean handoff. Document that a pod still working at the end of the window is SIGKILLed holding its lease, and that recovery is lease expiry plus reclaim under a new generation. An empty `serviceAccountName` was described as a neutral default. It inherits the relay's service account; `automountServiceAccountToken: false` hides the projected token but does not detach cloud IAM bindings resolved through the node metadata path. Recommend a dedicated pre-created account when the executor's IAM blast radius should be smaller than the relay's. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: tornquist <tornquist@squareup.com>
Exercise both supported credential shapes through the CI render matrix and bind the typed job guards, schema boundary, pod identity, and default sidecar annotation. Co-authored-by: Codex <noreply@openai.com> Signed-off-by: tornquist <tornquist@squareup.com>
Share a suffix-preserving 52-character name helper between the deletion drain and storage accounting CronJobs while retaining their existing short names. Co-authored-by: Codex <noreply@openai.com> Signed-off-by: tornquist <tornquist@squareup.com>
Distinguish the disabled Kubernetes API token mount from provider workload-identity credentials while retaining the dedicated service-account recommendation. Co-authored-by: Codex <noreply@openai.com> Signed-off-by: tornquist <tornquist@squareup.com>
Bind each context-menu interaction to the exact mock message returned by the emitter so same-second ordering cannot select the wrong video. Co-authored-by: Codex <noreply@openai.com> Signed-off-by: tornquist <tornquist@squareup.com>
464a4e6 to
38db86b
Compare
🔐 Codex Security Review
|
wpfleger96
left a comment
There was a problem hiding this comment.
🤖 Re-review at 38db86be. Nothing blocking.
#7818 merged as 15d44dcc, and this branch has been rebased onto main. The PR's own patch is byte-identical to the one at 464a4e60, which I last reviewed clear. All 16 files it touches have the same blobs at both heads. Main moved from the old base (9f6f7201) to 15d44dcc, picking up #7805, #7288 and the #7818 squash. Across that move nothing changed under deploy/charts, crates/buzz-admin, the deletion code, operator code, migrations/, ARCHITECTURE.md or video-attachment.spec.ts. The changes are in the desktop sidebar, buzz-media, mobile media upload and channel-sort.spec.ts. So the runbook, chart and README reasoning from my last review still holds against the #7818 that actually merged. GitHub shows the PR as mergeable against main.
CI at 38db86be: the chart job (lint + unittest + render matrix) passed. Authorize Security Review passes now that the base is main. Desktop Core, Smoke shards 1–3 and Desktop E2E Integration 1/2 and 2/2 passed. Smoke (4) is red on scroll-history.spec.ts:2150 ("thread summary badge survives a retained older-history prepend"), which failed all 3 attempts with Expected: < 550, Received: 550. That doesn't come from this PR. The merge ref is 38db86be into 15d44dcc, and main's own push run at 15d44dcc fails the same test all 3 times with the same values (job). The PR touches no desktop source. On the PR's side, the same main run marks video-attachment.spec.ts:1469 (the right-click menu probe) as flaky, and on this merge ref it passed on the first try. That fits what the "Scope video menu probes to emitted messages" commit is meant to fix.
Non-blocking notes still open: CronJob name truncation can collide for two long release names in one namespace, the bundled fixture only proves the chart renders, and the unsupported-type branch has no test. The PR description still says the branch is stacked on #7818, which has merged. It still cites e68cb753 under Testing, and it still leaves out the storage-accounting CronJob's move to the shared name helper and the video-attachment.spec.ts change.
## Summary
Fixes the `Desktop Smoke E2E (4)` failure on `main`. The test at
`desktop/tests/e2e/video-attachment.spec.ts` ("right-click menus expose
distinct selectors for links, relay video, and off-relay video") fails
3/3 attempts on main and on every PR branch since ~Sept 25. Because
"Desktop" is a required check, that blocks merges.
**Cause:** the relay-video and off-relay-video probes picked their
player with `getByTestId("video-player").last()`. When mock messages
share a second, timeline ordering can put a different video last. The
locator then resolves to the wrong player, and the "Download video" menu
detaches before the click, so the test times out.
**Fix:** scope each probe to the `data-message-id` of the message
`emitVideoMessage` just returned. Test-only change, 16 lines.
## Attribution
This is a cherry-pick of `dfd5db4e0` by @Tornquist, originally part of
#7827 (the deletion-drain cron, which is stacked on another feature
branch). It's pulled out on its own so main can go green without waiting
on that stack. Authorship is preserved.
## Testing
- `pnpm build:e2e && playwright test tests/e2e/video-attachment.spec.ts
--project=smoke`: 13/13 passed locally, including the target test.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Signed-off-by: tornquist <tornquist@squareup.com>
Signed-off-by: murderbot <3754f8729004d95654c46dbab3129e4ab9ef05cc2534e2a3fbfc155983bd637b@buzz.block.builderlab.xyz>
Co-authored-by: tornquist <tornquist@squareup.com>
Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Codex <noreply@openai.com> Co-authored-by: Codex <noreply@openai.com>
…mary test (#7955) 🤖 ## Summary Main's Desktop Smoke E2E (4) job fails on `scroll-history.spec.ts` ("thread summary badge survives a retained older-history prepend"). This change fixes the test's timing. It does not change app code. When a channel opens, the app starts a live subscription. When that subscription is ready, the app fetches the newest page again, and that fetch replaces the whole loaded history. This test scrolled up before that refresh happened. On slower CI runners, the refresh came after the older page loaded but before it showed on screen. The older rows went away, the reader stayed at the top, and no more pages loaded. All three attempts failed that way. The test now calls `waitForMockChannelHeadReady` before it scrolls, like the other pagination tests in this file already do. That helper waits for the live subscription and its refresh to finish. ### Related issue Main CI failure at `15d44dcc` (https://github.com/block/buzz/actions/runs/36470599751/job/109092076235). The same failure was reported on #7827. ### Testing - Added temporary logs to see the order of events. Locally, the refresh ran after the older page had loaded: the loaded history went from 2 pages back to 1. The test passed only because it checked before the next render. With the wait, the refresh runs first and nothing replaces the loaded pages later. - `playwright test --project=smoke tests/e2e/scroll-history.spec.ts`: 18 passed. This test with `--repeat-each=8`: 8 passed. ### Follow-up (not in this PR) The app has the same race for real users. If someone scrolls up before the live subscription is ready, or during a reconnect, the refresh can drop the older history they loaded. If they are at the top of the list, paging stops until they scroll down and back up. Signed-off-by: Larry <627498bd4bd1f281a16431e3c6cce3b5c25b6692798c78672298aefbf2f8f8b5@buzz.block.builderlab.xyz> Co-authored-by: Larry <627498bd4bd1f281a16431e3c6cce3b5c25b6692798c78672298aefbf2f8f8b5@buzz.block.builderlab.xyz>
Signed-off-by: Codex <noreply@openai.com>
Sidebar required-check repairPublished #7827 Diagnosis: I inspected the failed #7830 CI job 109113919391, retry-1 trace. Navigation returned at 719774ms, the 500ms count expired at 720293ms, and the post-failure DOM at 720526ms contained all 14 exact snapshot IDs. The filmstrip shows the boot splash during the assertion. This attempt failed on late initial rendering, not a snapshot already replaced; fallback assertions never ran. #7790 converted the neighboring hash-mismatch test to existing DOM recording but left this fallback test with the transient 500ms count. That history is not a claim that #7790 alone caused late mounting. Change: reuse existing Evidence (Blox, Chromium, one worker, zero retries, unchanged application bundle):
What got simpler: removed a second boot-performance timing special case by reusing the in-file observation mechanism. Prior scroll fix remains intact. The old scroll job's evidence stays attributed to |
Signed-off-by: Elrond <28d6302a099e5225b02c4155ac4236e4912603df2ab08dbfc2f4fef08ce598c8@buzz.block.builderlab.xyz>
Additive main refreshCurrent heads: #7827 Main #7827's diff against main is 16 files: 15 infrastructure/docs files and the sidebar fix. Scroll-history and video-attachment are now identical to main. #7830's deletion changes are preserved. Raw binary-patch comparison differs, because main #7854 added the unrelated GIN index and its assertion in schema.sql/runtime/migration.rs, changing blob IDs and hunk offsets. After excluding only those index headers and hunk line offsets, the complete own-patch compares identical. No deletion hunk changed. Verification remains attributed to the SHA actually executed. Main contains production changes; old sidebar/scroll runs do not establish runtime success at these new merge heads. Fresh CI and exact-head Legolas/Gimli reviews requested. Environment correction: I accidentally shared writable node_modules between Blox worktrees; bootstrap rewrote dependency links during the old full smoke, causing mixed Playwright collection errors. That broad run is invalid as a product result. Dependency trees are now independent; a clean full-sidebar rerun at old head5650 passed10/10. Full Desktop tests6731+92, check/typecheck passed at5650; the earlier repository just-ci run at131ef21a was blocked on absent glib-2.0.pc. Current verification continues asynchronously and is not an overall PASS. |
wpfleger96
left a comment
There was a problem hiding this comment.
🤖 Re-review at 7a26420b. Nothing blocking.
The new head is 5650ae7a plus a merge of main (base f1e50be4). 15 of the 16 files the PR touches have the same blobs as 38db86be, which I last reviewed clear. The one new PR-owned change is desktop/tests/e2e/sidebar-snapshot.spec.ts. video-attachment.spec.ts drops out of the diff because main now has the same content. Main moved from 15d44dcc to f1e50be4 and changed nothing under deploy/, docs/, crates/buzz-admin, migrations/ or ARCHITECTURE.md. So the chart, runbook and README reasoning from my last review still holds.
On the spec change: the "mismatched not-modified hash falls back to a full list" test now uses a MutationObserver (trackSnapshotRows) installed before navigation instead of toHaveCount(14, { timeout: 500 }). It still has to see all 14 snapshot IDs in order, and removing snapshot hydration would still fail it. The rest of the test is unchanged: the stale-hash request followed by a hashless one, the live general row, zero leftover snapshot rows, and the persisted live hash. I think this is a reasonable way to stop depending on the runner catching a transient DOM state. We ran just that test 20 times headless with one worker and no retries, at this head and with the 38db86be version of the spec: 20/20 both times. The whole spec went 10/10. So this doesn't show a flake-rate improvement, but it doesn't regress anything either. All 34 non-skipped checks at 7a26420b are green.
One small non-blocking note. The observer builds up a list of unique rows added over time. It doesn't check that all 14 are on screen together, or that they show up before the fallback request goes out. The old 500ms assertion did check that they were present together. The new comment says cold-boot coverage "separately enforces the snapshot paint deadline." But that test gives the snapshot rows a 3s bound in the matching-hash case, not 500ms on the mismatch path, and it watches DOM rows, not a painted frame. I'd reword the comment to say that. If ordering on the mismatch path matters, you could hold back the fallback response and assert the full snapshot before releasing it. That would be better than going back to a short timeout.
My earlier non-blocking notes are still open. The CronJob name truncation can collide for two long release names in one namespace. The bundled fixture only proves the chart renders. The unsupported-type branch has no test. The PR description is also out of date. It still says the branch is stacked on #7818, which has merged, and it still cites e68cb753 under Testing. It also doesn't mention the storage-accounting CronJob's move to the shared name helper or the sidebar-snapshot.spec.ts change.
…-enforcement * origin/main: feat(relay): implement NIP-AR channel artifacts (#7919) fix(desktop): resolve unlisted project channel requests (#7619) Schedule the deletion drain safely (#7827) Signed-off-by: Hayt <211b96e6a2b7f45fd4047988976c7bbbeeda0c15f3ae7b32eec20834b5a55118@buzz.block.builderlab.xyz>
## Why Owner deletion admission stops at `submitted`. Manual inventory and approval still stand between owner intent and the existing deletion executor, preventing self-serve deletion. ## What The existing `buzz-admin deletions drain` process prepares and automatically approves operator-attested owner requests. Operator-origin requests retain manual approval. ## How The drain prioritizes approved work, then claims one owner submission with the existing generation lease. It freezes inventory and records digest-bound `owner_automatic` approval in one transaction before entering the unchanged executor. Before approval, the transaction rechecks archived state and current ownership. It locks the deletion advisory fence, community, owner memberships, and request in that order. A regression test pins privileged abort during a live preparation lease. What got simpler: one drain still owns preparation and execution. There is no second worker, queue, lifecycle, or retry authority. The rebased branch contains only this feature and its review fixes. ## Risk Automatic approval permits irreversible deletion without a human grace period. Operator attestation is not a cryptographic owner signature. Provenance, locked authority checks, generation fences, and digest-bound approval constrain the operation. This remains work in progress. Broad test failures and incomplete mobile checks are disclosed below; publication is not a merge-ready verdict. ## Testing Published head: `e059ae9056b447ca3b5d0c2893da86550ef16f5c`. All seven commit trees match tested source `5605f27eaece1e809f746cc1839b55c7ceef8d80`; only attribution metadata changed. Source-SHA results below are not new-head execution receipts. At source commit `5605f27eaece1e809f746cc1839b55c7ceef8d80`, isolated PostgreSQL tests showed stale archive and ownership fixtures fail closed after the fix. Before the fix both reached automatic approval. Removing abort's generation increment or stage transition made the new live-lease regression fail; restored production passed. `just test` failed in two ACP native-git fixtures and one agent body-timeout test. The failing source files match the pinned base, but there is no exact-base execution comparison proving those failures unrelated. `just test-unit` also failed in the ACP fixtures, leaving 18 tests unrun. `just ci` stopped in Flutter tooling package resolution; separately run non-mobile lanes do not constitute a full `just ci` pass. No live tenant deletion or human testing was performed. A dedicated populated `0052` → `0053` approval-backfill fixture remains a coverage gap. Final-head hosted CI and security review are required. ## Bigger picture Rebased directly onto main `12670bd0f037c66a682272bb81c46c3f254fad74`; prerequisite #7818/#7827 are already merged. Main owns `0052_channel_artifacts`, so this feature's automatic-approval migration is now `0053`. Client/KGoose and quota work are separate PRs and are not part of this diff. Generated with Codex --------- Signed-off-by: Codex <noreply@openai.com> Signed-off-by: Elrond <28d6302a099e5225b02c4155ac4236e4912603df2ab08dbfc2f4fef08ce598c8@buzz.block.builderlab.xyz> Co-authored-by: Codex <noreply@openai.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Elrond <28d6302a099e5225b02c4155ac4236e4912603df2ab08dbfc2f4fef08ce598c8@buzz.block.builderlab.xyz>
…in-ui * origin/main: feat(buzz-relay): NIP-FI stateless enforcement (S3) — upgrade gate, NIP-42 pairing, session lifetime, JWKS warm (#7224) feat(web): add Browse releases link next to invite download (#2255) docs(nips): fix stray angle brackets in created_at clauses (#4486) docs: specify desktop-driven mobile push suppression (#7809) feat(mobile): add the contextual identity-name resolver (#7894) Automate owner deletion preparation (#7830) feat(relay): implement NIP-AR channel artifacts (#7919) fix(desktop): resolve unlisted project channel requests (#7619) Schedule the deletion drain safely (#7827) fix(mobile): stale community selection during mobile invite setup (#7951) Signed-off-by: Sol <49aa1f65411fd096d2e2ec144f1e7aa36fdc76d1b907cfdf7be000c66f9d3b8e@buzz.block.builderlab.xyz>
…rvation (#7969) ## Summary This is the relay side of owner community deletion. - **Idempotent delete is the recovery call.** - `POST /operator/communities/delete` checks the request UUID first. - If the request already exists with the same host, owner and acknowledgement version, it returns `202` with the request's current `status`. That holds at any stage, even after membership is purged, and no new work is admitted. - The same UUID with a different host or owner returns `409 deletion_request_conflict`. An unsupported acknowledgement version is rejected first, with `400 unsupported_acknowledgement_version`. - There's **no receipt endpoint**. Clients recover by resending. - **Active quota reservation.** An incomplete owner deletion holds the owner's slot until the deletion logically completes. Migration `0054` adds the supporting index. Migrations 0052 and 0053 are unchanged. - **Lifetime cap.** Deleted communities keep their hosts as permanent tombstones. - On create and on transfer-in, the relay counts live ownership plus every non-aborted owner deletion, completed ones included. - That total is capped at an absolute 20, regardless of the active limit, and going over returns `limit_reached`. - This stops create-then-delete host squatting. - **Stable lifecycle conflict codes,** plus `acknowledgement_version` in the 202. - **Owner-list quota projection:** `quota_used`, `quota_limit` and `can_create`. - `can_create` reflects both caps. - The projection is advisory, and `limit_reached` is authoritative. - **Operator doc** (`docs/operator-community-deletion.md`): the acknowledgement version is a compile-time constant. Admission and the executor's claim/lease both filter on it. A version bump must keep replaying and executing old-version requests. This follows on from #7830. The foundation PRs #7818, #7827 and #7830 are merged. ## Testing New tests: - `owner_delete_resubmission_reports_current_status_without_new_intent` - A replay returns 202 `submitted`, including after membership purge. - Changing any field returns 409 `deletion_request_conflict`. - A non-operator gets 403. - A replay after abort returns 202 `aborted`. - No request rows are added. - `completed_owner_deletions_count_toward_lifetime_cap` - After 20 created-then-deleted communities, the active count is 0 but `can_create` is false. - The next create and a transfer-in both return `LimitReached`. - `owner_quota_admits_only_under_active_and_lifetime_caps`: a unit test. Fellowship gate: PASS at `59375c00`. The code review was clean. The E2E run covered authenticated KGoose → relay → the real drain executor. The later commits only change the operator doc. ## Rollout - Deploy order: this relay first, then KGoose squareup/cash-server#130634, then Web squareup/ext-builderbot-ui#241 and App block/buzz-app#403. - No shipped client ever called a receipt route, so removing it doesn't break any existing client. - The quota changes have no deploy-order requirement. - Keep owner deletion off until this relay and the drain executor are live. - If you need to roll back, turn deletion off before rolling back the relay. - Recreate any database that ran the earlier version of migration 0054. That only applies to disposable preview databases. Never apply this to a database holding real data. The earlier draft 0054 has a different checksum and index predicate. ## Complexity Net simpler: - One route, its handler and a writer-routed `get` are removed. Recovery reuses the existing idempotent admission path. - A single `OwnerQuota::admits()` backs create, transfer and the projection. - The stale "clients fail closed / strict deployment order" doc section is replaced by the advisory quota contract. --------- Signed-off-by: tornquist <tornquist@squareup.com> Signed-off-by: Elrond <28d6302a099e5225b02c4155ac4236e4912603df2ab08dbfc2f4fef08ce598c8@buzz.block.builderlab.xyz> Signed-off-by: Codex <noreply@openai.com> Signed-off-by: Codex <codex@openai.com> Signed-off-by: OpenAI Codex <codex@openai.com> Co-authored-by: Codex <noreply@openai.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Elrond <28d6302a099e5225b02c4155ac4236e4912603df2ab08dbfc2f4fef08ce598c8@buzz.block.builderlab.xyz> Co-authored-by: Codex <codex@openai.com>
Why
The deletion engine already exposes a one-shot
buzz-admin deletions draincommand, but operators have no safe way to schedule it. A generic workload wrapper could expose arbitrary commands or inherit relay credentials, which would add a second control plane and widen the blast radius.What
This PR adds an opt-in Kubernetes CronJob for the typed deletion drain command. It is stacked on #7818, which adds owner-request admission.
How
The chart keeps the job disabled by default. The CronJob uses
Forbid, zero Kubernetes retries, bounded runtime and history, and only the database, Redis, and object-store settings required by the executor. PostgreSQL leases, checkpoints, and retry timing remain the sole durable retry authority.The closed
_operator-jobs.tplhelper centralizes safe CronJob mechanics without creating an arbitrary command or environment registry. This leaves the touched system simpler than a general job framework would.Risk
The deployment blast radius is low because the job is disabled by default. When enabled, it can run destructive deletion work that the database has already approved, so the chart limits concurrency, retries, credentials, and execution time.
Testing
Rendered every
deploy/charts/buzz/ci/*.yamlfixture withhelm template; all passed ate68cb753bb6673aa1baa60046969bdb2d4738b62.Ran the full independent Desktop, Tauri, and web lanes; all passed at
e68cb753bb6673aa1baa60046969bdb2d4738b62. Flutter package resolution made no progress within three minutes, so the mobile lane is not claimed green.Bigger picture
A later stack layer must inventory and approve admitted owner requests. Until then, this scheduler cannot make a new owner request destructive.
Generated with Codex