Skip to content

fix(budgets): PLTF-3562 keep the once-per-cycle alert latch across a threshold edit - #416

Merged
ak684 merged 6 commits into
mainfrom
fix/threshold-edit-preserves-alert-latch
Sep 22, 2026
Merged

ak684 merged 6 commits into
mainfrom
fix/threshold-edit-preserves-alert-latch

Conversation

@aivong-openhands

@aivong-openhands aivong-openhands commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

HUMAN:

  • A human has tested these changes.

AGENT:


Why

Editing an organization's budget thresholds re-sent alerts that had already been delivered. _maybe_send_alerts dedupes on threshold.last_triggered_cycle_start, and that latch lives on the threshold row itself. replace_thresholds deleted every row and inserted fresh ones with the latch columns unset, so any edit — in the reproduction test, just turning Slack on for the existing 80% alert — re-armed every threshold inside the live cycle. The next maintenance run paged the same admins again for spend they had already acknowledged, which is how alerting gets ignored.

The row churn was the root of it: an admin setting slack_enabled=True on the 80% threshold should not destroy and recreate that threshold's identity. replace_thresholds now diffs by percentage — it updates the rows that survive the edit in place, deletes only the percentages that were removed, and inserts only the ones that are new. The latch survives because nothing touches it, and a future column on the table survives for the same reason rather than having to be resurrected by hand here.

A percentage being added has never fired, so it correctly starts unlatched and alerts the first time spend reaches it.

Summary

  • replace_thresholds diffs thresholds by percentage rather than replacing every row, so the per-threshold alert latch is preserved by construction.
  • Un-skip the reproduction test, which alerts once, edits the thresholds, runs maintenance again and asserts no second alert.
  • Cover the other two branches of the diff: a percentage added mid-cycle below the current spend, and a percentage removed.
  • Cover the multi-row edit an org with the default 80/90/100 thresholds actually makes, and the duplicate-row collapse the diff's kept set exists to enforce. Both live in tests/unit/test_org_budget_store.py alongside the removal test, which moved there from the service file.

Issue Number

N/A

How to Test

uv run pytest -q tests/unit/test_org_budget_service.py tests/unit/test_org_budget_store.py

Expect 48 passed, 6 skipped. Reverting storage/org_budget_store.py to main makes all three threshold tests fail — the first with a second alert in the same cycle, the others on the re-armed latch.

Video/Screenshots

N/A — no UI change.

Type

  • Bug fix
  • Feature
  • Refactor
  • Breaking change
  • Docs / chore

Notes

Thresholds are keyed by percentage, which is what _maybe_send_alerts dedupes on and what the route model already enforces as unique within an update (_validate_thresholds, server/routes/org_models.py). (org_id, percentage) has no unique index in the database, so the diff drops duplicate rows rather than updating a percentage into two live copies.

An admin who deletes a percentage and re-adds it in the same cycle gets one more alert for it; that reads as intended, since the threshold genuinely did not exist in between.

The latch is per-threshold, but _send_alerts fans out per-channel (email_enabled / slack_enabled). A deliberate trade-off follows: enabling a channel on an already-latched threshold now suppresses that channel for the rest of the cycle. Concretely, turning Slack on for the already-fired 80% threshold means Slack gets nothing for 80% until the next cycle, even though Slack was never paged for it. The old row-churn behaviour delivered that Slack alert as a side effect of the bug this PR fixes. Per-channel latching would preserve it but is a larger change; this PR keeps the simpler per-threshold latch and accepts the missed same-cycle channel-enable alert. Since the change moves the system from over-alerting to under-alerting — the quieter, more dangerous direction — this trade-off should be signed off by whoever owns budget alerting.

Duplicate (org_id, percentage) rows (no unique index backs the pair) are collapsed onto the latched row: replace_thresholds orders existing by latch recency before the diff loop, so the surviving copy is the one carrying the once-per-cycle latch rather than whichever row the query returned first.

One of a set of draft PRs, each carrying a single defect the Quint model for org budgets surfaced, together with the reproduction test that was already committed but skipped.

🤖 Generated with Claude Code


Enterprise server image for this PR:

ghcr.io/openhands/enterprise-server:sha-6f98525

…edit

_maybe_send_alerts dedupes on threshold.last_triggered_cycle_start, which lives
on the threshold row. replace_thresholds deleted every row and inserted fresh
ones carrying no latch, so an admin who edited the thresholds -- even just
turning Slack on for an existing percentage -- re-armed every alert inside the
live cycle. The next maintenance run then paged the same admins again for spend
they had already acknowledged.

Carry last_triggered_at and last_triggered_cycle_start across for any percentage
that survives the edit. A percentage being added has never fired, so it
correctly starts unlatched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@github-actions github-actions Bot added the type: fix A bug fix label Sep 16, 2026
@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Coverage report

Click to see where and how coverage changed

FileStatementsMissingCoverageCoverage
(new stmts)
Lines missing
  storage
  org_budget_store.py
Project Total  

This report was generated by python-coverage-comment-action

@aivong-openhands aivong-openhands added the quint-studio-budgets-fixes Org budgets defects surfaced by the Quint Studio model label Sep 16, 2026

@aivong-openhands aivong-openhands left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Taste Rating: Acceptable

Correct fix for a real and genuinely user-hostile bug — re-paging admins for spend they already acknowledged is how alerting gets muted permanently. Matching the latch by percentage is the right key, because percentage is exactly what _maybe_send_alerts dedupes on, and the route model already enforces uniqueness of percentages within an update (_validate_thresholds, server/routes/org_models.py), so the dict cannot silently collapse two rows. The reasoning in the PR notes about delete-then-re-add producing one extra alert is sound and I agree it reads as intended.

What keeps it at 🟡 is that the fix preserves the delete-and-reinsert shape rather than removing it.

[IMPROVEMENT OPPORTUNITIES]

  • [storage/org_budget_store.py:72-105] Data Structure — the latch carry-over is a workaround for churning rows that did not need to churn. The operation an admin performs is "set slack_enabled=True on the 80% threshold." The code expresses that as: delete every threshold row, then re-insert every threshold row, then manually copy back the two columns that must survive. That is why the bug existed — the row identity was thrown away and someone had to remember which fields to resurrect. The shape that eliminates the bug class is diffing by percentage: update the rows that persist (mutating email_enabled/slack_enabled in place), delete only percentages that were removed, insert only percentages that are new. Then last_triggered_* survives because nothing touched it, and a third latch-like column added next year survives for free. As written, whoever adds that column has to find this function and remember to add a third line to the tuple. That is a trap, and it is the same trap that produced this PR.

    Concretely, the current code also burns a PK per threshold per settings edit and generates 2N statements where N would do. Not a performance concern at this table size — the maintainability point is the real one.

  • [storage/org_budget_store.py:78-82] Unnecessary comments: the block explains the bug's history ("Replacing the rows wholesale would drop it and re-arm every alert inside the live cycle"), which is PR-description material. The non-obvious invariant worth keeping is one line: the alert latch is keyed on percentage, not on row identity. The last sentence ("A percentage being added has never fired, so it correctly starts unlatched") is a useful design note but describes behaviour a reader can derive from latches.get(..., (None, None)). Roughly five lines of comment for a ~15-line change.

[TESTING GAPS]

The un-skipped test is good — it runs maintenance, edits settings, runs maintenance again, and asserts send_alerts.await_count == 1 against a real session. It exercises the actual latch path rather than asserting a mock was called, and it fails without the store change.

One gap worth a second case: the test only covers the surviving percentage. Neither the "percentage removed" nor the "percentage added mid-cycle" branch of latches.get(...) is exercised. The added-percentage branch in particular encodes an intentional product decision (a newly added threshold alerts immediately if spend is already past it), and nothing currently pins it. A test that edits thresholds from [80] to [80, 90] with spend at 95% and asserts exactly one new alert would lock that in.

[RISK ASSESSMENT]

  • [Overall PR] ⚠️ Risk Assessment: 🟢 LOW

Single store method, no schema change, no API change. The failure mode if the mapping were wrong is bounded in both directions: at worst a missed alert or a duplicate alert, not spend enforcement or data loss. _roll_cycle_if_needed still clears both latch columns explicitly at cycle rollover (org_budget_service.py:911-913), so a carried-over latch cannot outlive its cycle and suppress alerts in the next one — I checked that specifically, since a sticky latch would be the dangerous direction. CI is green.

Evidence is a pytest run, which normally would not satisfy the evidence bar on its own. It is acceptable here because the observable behaviour is "an email/Slack message is not sent," and the negative is what the test asserts; there is no runtime artifact that demonstrates an absence better than that. No UI change.

VERDICT:
✅ Worth merging: The behaviour is correct and properly pinned. Please consider the diff-in-place refactor as a follow-up rather than blocking this.

KEY INSIGHT:
Manually resurrecting state across a delete-and-reinsert is a standing invitation for the next column to be forgotten; updating rows in place makes the whole bug class unrepresentable.


Improve this review? If any feedback above seems incorrect or irrelevant to this repository, you can teach the reviewer to do better:

  1. Add a .agents/skills/custom-codereview-guide.md file to your branch (or edit it if one already exists) with the /codereview trigger and the context the reviewer is missing (e.g., "Security concerns about X do not apply here because Y"). See the customization docs for the required frontmatter format.
  2. Re-request a review - the reviewer reads guidelines from the PR branch, so your changes take effect immediately.
  3. When your PR is merged, the guideline file goes through normal code review by repository maintainers.

Resolve with AI? Install the iterate skill in your agent and run /iterate to automatically drive this PR through CI, review, and QA until it's merge-ready.

Was this review helpful? React with 👍 or 👎 to give feedback.


This review was generated by an AI agent (OpenHands) on behalf of @aivong-openhands.

… rows

replace_thresholds deleted every threshold row and inserted fresh ones, so the
once-per-cycle alert latch on the row had to be copied back by hand -- and the
next column added to the table would silently be dropped the same way.

Diff by percentage instead: update the rows that survive the edit in place,
delete only the percentages that were removed, insert only the ones that are
new. The latch survives because nothing touches it.

Adds coverage for the two branches the first test missed: a percentage added
mid-cycle below the current spend pages once and does not re-arm the others,
and a percentage removed does not disturb the row that survives.

Co-authored-by: openhands <openhands@all-hands.dev>

Copy link
Copy Markdown
Contributor Author

Thanks — took both improvement opportunities in 0c31430.

Diff in place instead of delete-and-reinsert. You're right that the latch carry-over was treating a symptom. replace_thresholds now diffs by percentage: rows whose percentage survives the edit are mutated in place (email_enabled/slack_enabled), removed percentages are deleted, new ones are inserted. last_triggered_at/last_triggered_cycle_start survive because nothing touches them, and the next latch-like column added to the table survives for free rather than depending on someone finding this function.

One thing the diff form does need that the delete-and-reinsert form did not: (org_id, percentage) has no unique index (migration 133 creates the index on org_id with unique=False), so a pre-existing duplicate row would otherwise be updated into two live copies of the same threshold. The loop drops duplicates instead of updating them; that's the one comment I kept in the body, since it's the non-obvious constraint.

Comments trimmed. The five-line history block is gone — the bug's story lives in the commit message. What remains is the one-line invariant you identified: a threshold is identified by its percentage, not by its row.

Testing gaps closed. Two cases added:

  • test_threshold_added_mid_cycle_alerts_once_for_spend_already_past_it — thresholds go from [80] to [80, 90] with spend at 95% of the cap, then maintenance runs twice. Asserts the alerted percentages are exactly [80, 90]: the new threshold pages immediately (pinning that product decision), the pre-existing one stays latched, and neither repeats.
  • test_dropping_a_threshold_leaves_the_others_latched — thresholds go from [80, 90] to [90]; asserts the surviving row keeps its last_triggered_cycle_start along with its new settings.

Both fail against the pre-fix store and pass against the new one, as does the original test.

Verification:

$ uv run pytest tests/unit/test_org_budget_service.py tests/unit/test_org_budget_store.py -q
46 passed, 6 skipped in 9.24s

$ uv run pytest tests/unit/test_org_budget_preflight.py tests/unit/test_org_budget_maintenance_processor.py tests/unit/test_run_budget_maintenance.py -q
21 passed in 1.24s

Against origin/main's storage/org_budget_store.py with the new tests in place, 3 of the 4 threshold tests fail (the fourth is unrelated), confirming they pin the behaviour rather than restating it.

pre-commit run --config ./dev_config/python/.pre-commit-config.yaml passes (ruff, ruff format, mypy).


This comment was created by an AI agent (OpenHands) on behalf of @aivong-openhands.

Copy link
Copy Markdown
Contributor Author

Mutation review of the tests in this PR

I hand-wrote 11 mutants against replace_thresholds and the alert latch and ran each against tests/unit/test_org_budget_service.py + tests/unit/test_org_budget_store.py (baseline 46 passed, 6 skipped in 8.3s). Three of them were controls that had to die, eight were candidates probing the machinery around the diff.

Controls — all caught

Mutant Result
C1 revert the fix: delete every row, insert fresh ones ❌ caught
C2 keep the rows but clear the latch on a surviving percentage ❌ caught
C3 drop the last_triggered_cycle_start == cycle_start check in _maybe_send_alerts ❌ caught

The suite is genuinely load-bearing here. test_threshold_alerts_once_per_cycle_across_a_settings_edit drives the whole path — maintenance, a real update_budget_settings edit, maintenance again — and asserts on _send_alerts.await_count, so it catches the latch being lost no matter which layer drops it (C1, C2 and C3 all die on it). test_dropping_a_threshold_leaves_the_others_latched asserting the whole tuple list with == rather than picking fields off one row is what kills the "leave the removed row behind" mutant, and the [80, 90] ordered-percentage assertion in test_threshold_added_mid_cycle_... kills a new row being inserted already latched. Five of eight candidates died too.

Survivors

Mutant Result
M7 if update is None or kept: — only the first surviving percentage is updated in place, every later one is deleted and re-inserted ✅ 46 passed
M1 drop or threshold.percentage in kept — duplicate rows for one percentage are both kept and updated ✅ 46 passed
M8 clear last_triggered_at on a surviving row (keeping last_triggered_cycle_start) ✅ 46 passed

M7 — the multi-threshold edit is untested

Every test here edits thresholds where at most one percentage survives with a latch to preserve: the repro test has a single 80% threshold, the "added mid-cycle" test only asserts the alert sequence, and the drop test keeps exactly one row. So a replace_thresholds that in-place-updates only the first survivor and churns the rest passes the suite untouched — while in production every org starts with three (DEFAULT_THRESHOLDS = 80/90/100), and an admin who has crossed 80% and 90% and then edits the settings gets re-paged for the 90% one.

This kills it, and it is the shape of edit admins actually make:

@pytest.mark.asyncio
async def test_editing_thresholds_leaves_every_surviving_row_latched(
    async_session_maker, budget_org
):
    cycle_start = datetime.now(UTC)
    async with async_session_maker() as session:
        session.add_all(
            [
                OrgBudgetThreshold(
                    org_id=budget_org.id,
                    percentage=80,
                    email_enabled=True,
                    slack_enabled=False,
                    last_triggered_at=cycle_start,
                    last_triggered_cycle_start=cycle_start,
                ),
                OrgBudgetThreshold(
                    org_id=budget_org.id,
                    percentage=90,
                    email_enabled=True,
                    slack_enabled=False,
                    last_triggered_at=cycle_start,
                    last_triggered_cycle_start=cycle_start,
                ),
            ]
        )
        await session.commit()

        store = OrgBudgetStore(session)
        await store.replace_thresholds(
            budget_org.id,
            await store.get_thresholds(budget_org.id),
            [
                OrgBudgetThresholdUpdate(
                    percentage=80, email_enabled=True, slack_enabled=True
                ),
                OrgBudgetThresholdUpdate(
                    percentage=90, email_enabled=True, slack_enabled=True
                ),
            ],
        )
        await session.commit()

        rows = await store.get_thresholds(budget_org.id)

    # Every percentage the edit keeps is updated in place, not just the first one.
    assert [
        (row.percentage, row.slack_enabled, row.last_triggered_cycle_start)
        for row in rows
    ] == [(80, True, cycle_start), (90, True, cycle_start)]

Verified: passes on this branch unmodified, fails with M7 applied.

M1 — the duplicate-row rule the code comment states is unasserted

The diff carries a comment explaining a deliberate decision — "No unique index backs (org_id, percentage), so a duplicate row is dropped rather than updated into a second copy" — and kept exists only to enforce it. Nothing tests it. Deleting that clause leaves the suite green, and with duplicates present (the pre-fix delete-and-insert path plus a retried request could leave them behind) every copy stays live, so _maybe_send_alerts iterates both and pages twice per cycle for one percentage. This is also the rule that keeps the latched copy rather than a fresh one:

@pytest.mark.asyncio
async def test_duplicate_rows_for_one_percentage_collapse_onto_the_latched_row(
    async_session_maker, budget_org
):
    cycle_start = datetime.now(UTC)
    async with async_session_maker() as session:
        session.add_all(
            [
                OrgBudgetThreshold(
                    org_id=budget_org.id,
                    percentage=80,
                    email_enabled=True,
                    slack_enabled=False,
                    last_triggered_at=cycle_start,
                    last_triggered_cycle_start=cycle_start,
                ),
                OrgBudgetThreshold(
                    org_id=budget_org.id,
                    percentage=80,
                    email_enabled=True,
                    slack_enabled=False,
                ),
            ]
        )
        await session.commit()

        store = OrgBudgetStore(session)
        await store.replace_thresholds(
            budget_org.id,
            await store.get_thresholds(budget_org.id),
            [
                OrgBudgetThresholdUpdate(
                    percentage=80, email_enabled=True, slack_enabled=True
                )
            ],
        )
        await session.commit()

        rows = await store.get_thresholds(budget_org.id)

    # One row per percentage survives the edit, and it is the latched one.
    assert [
        (row.percentage, row.slack_enabled, row.last_triggered_cycle_start)
        for row in rows
    ] == [(80, True, cycle_start)]

Verified: passes on this branch unmodified, fails with M1 applied.

M8 — I think this is an equivalent mutant

Clearing last_triggered_at while keeping last_triggered_cycle_start survives, but last_triggered_at has no reader anywhere outside tests — _maybe_send_alerts dedupes on the cycle column only, OrgBudgetThresholdResponse does not expose it, and nothing in the frontend reads it. So either it is purely diagnostic, in which case the survivor is fair and not worth a test, or something is meant to consume it and does not. Your call.

Not a test gap

  • _validate_thresholds rejects duplicate percentages in an update, so wanted can never collide on the input side; the duplicate case M1 probes only arises from rows already in the table. That is a structural guarantee, not something to test.
  • The sole store-level assertion on this code lives in tests/unit/test_org_budget_service.py; tests/unit/test_org_budget_store.py covers only cycle baselines. Both suggested tests would sit more naturally in the store file.

This comment was generated by an AI assistant on behalf of the user.

…apse

Move the store-level threshold test into the store test file and add the two
cases the diff's machinery has no coverage for: an edit keeping more than one
latched row, and two rows for one percentage collapsing onto the latched copy.

Co-authored-by: openhands <openhands@all-hands.dev>

Copy link
Copy Markdown
Contributor Author

Thanks — the mutation run is exactly the right way to check whether these tests are load-bearing, and both surviving mutants were real gaps. Addressed in fff75ea.

M7 — multi-threshold edit. Taken. The reasoning holds: DEFAULT_THRESHOLDS is 80/90/100, so an edit that keeps several latched rows is the ordinary case rather than the exotic one, and every test here happened to keep at most one. Added test_editing_thresholds_leaves_every_surviving_row_latched. Verified it fails with if update is None or kept: applied and passes on the branch unmodified.

M1 — duplicate-row collapse. Taken. kept exists only to enforce the rule the code comment states, and nothing asserted it. Added test_duplicate_rows_for_one_percentage_collapse_onto_the_latched_row, which also pins that it is the latched copy that survives rather than an arbitrary one. Verified it fails with the or threshold.percentage in kept clause dropped and passes on the branch unmodified.

M8 — agreed, equivalent. last_triggered_at is written in two places in org_budget_service.py and read nowhere outside tests: _maybe_send_alerts dedupes on the cycle column alone and OrgBudgetThresholdResponse does not expose it. It is diagnostic, so the survivor is fair and I have not written a test to pin a column no behaviour depends on. I am not changing it here either — if something is meant to consume it, that is a separate change with its own test.

Placement. Also taken. The two new tests plus test_dropping_a_threshold_leaves_the_others_latched, which drives the store directly and did not belong in the service file, now all live in tests/unit/test_org_budget_store.py. The two that genuinely exercise the service path — the repro and the added-mid-cycle one — stay where they are. The store file grew a _latched_threshold helper and a _replace_thresholds helper so the three read as the edits they describe rather than as row-construction boilerplate.

uv run pytest -q tests/unit/test_org_budget_service.py tests/unit/test_org_budget_store.py is now 48 passed, 6 skipped; the PR description is updated to match.

This comment was written by an AI agent (OpenHands) on behalf of the user.

@aivong-openhands aivong-openhands left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

🟡 Acceptable — the root-cause diagnosis is correct and the core fix is the right shape, but one of the three new tests asserts a guarantee the code does not actually provide, and there is a behavioural trade-off buried in the fix that deserves an explicit decision.

I verified this locally against the PR head (fff75ea), with a real Postgres via testcontainers — not by reading the diff:

$ uv run pytest -q tests/unit/test_org_budget_service.py tests/unit/test_org_budget_store.py
48 passed, 6 skipped in 9.73s

$ git checkout origin/main -- storage/org_budget_store.py && uv run pytest -q ...
5 failed, 43 passed, 6 skipped
FAILED tests/unit/test_org_budget_service.py::test_threshold_alerts_once_per_cycle_across_a_settings_edit
FAILED tests/unit/test_org_budget_service.py::test_threshold_added_mid_cycle_alerts_once_for_spend_already_past_it
FAILED tests/unit/test_org_budget_store.py::test_dropping_a_threshold_leaves_the_others_latched
FAILED tests/unit/test_org_budget_store.py::test_editing_thresholds_leaves_every_surviving_row_latched
FAILED tests/unit/test_org_budget_store.py::test_duplicate_rows_for_one_percentage_collapse_onto_the_latched_row

So the claim in How to Test holds exactly as written, and these are real tests against real code paths and real rows, not mock-assertion theatre. Credit where due — the reproduction is honest and the diff-by-percentage approach is the correct data-structure fix. Keying on the thing _maybe_send_alerts already dedupes on, rather than inventing a parallel identity, is the right instinct, and it makes future columns on org_budget_threshold survive an edit for free instead of each one needing to be hand-copied.


[CRITICAL ISSUES]

  • [storage/org_budget_store.py, L87] The duplicate collapse keeps an arbitrary row, not the latched one. get_thresholds orders by percentage alone. For two rows with the same percentage the tiebreak is unspecified, so kept retains whichever duplicate the query happened to yield first and deletes the rest — latch and all. test_duplicate_rows_for_one_percentage_collapse_onto_the_latched_row passes only because it inserts the latched row first and Postgres happens to return them in physical order. Swap the two session.add_all entries so the unlatched row is inserted first and the same test fails:

    SURVIVING LATCH: [None]
    AssertionError: latch lost!
    assert [None] == [datetime.datetime(2026, 9, 17, 5, 15, 58, 161041, tzinfo=datetime.timezone.utc)]
    

    I ran that; it is not hypothetical. The test name and the comment on L85–86 both promise the collapse lands on the latched copy, and the code does not deliver it. A test that asserts a guarantee the implementation does not make is worse than no test — it will be cited as proof later. Either make the selection deterministic (prefer the row whose last_triggered_cycle_start is set, e.g. sort existing so latched rows sort first before the loop) or drop the claim from the test name and comment and stop asserting the latch in that case. Given the whole PR is about not losing the latch, the former.

[IMPROVEMENT OPPORTUNITIES]

  • [storage/org_budget_store.py, L90–91] Enabling a channel on a latched threshold now silently suppresses that channel for the rest of the cycle. The latch is per-threshold, but _send_alerts fans out per-channel (email_enabled / slack_enabled). The reproduction case in this PR is precisely an admin turning Slack on for the already-fired 80% threshold — after this change, Slack gets nothing for 80% until the next cycle, even though Slack was never paged for it. The old behaviour delivered that Slack alert, as a side effect of the bug. The PR frames the old behaviour as purely a defect; for this one scenario it accidentally did the right thing. That may still be the correct trade-off (per-channel latching is a bigger change), but it is a deliberate product decision, not a neutral refactor, and it is currently undocumented. Please state it in the PR description alongside the delete-and-re-add case you already called out — or latch per channel.

  • [storage/org_budget_store.py, L78–80] The lead comment narrates the change rather than the invariant. "updated in place rather than replaced" describes what the diff did to the previous implementation; that belongs in the commit message, which already says it well. The durable fact a future reader needs is the first clause — a threshold's identity is its percentage, because the alert latch lives on the row. Trim to that. The L85–86 comment is the opposite case and should stay: the missing unique index is genuinely non-obvious and not inferable from the code.

[TESTING GAPS]

  • No end-to-end evidence beyond the test suite. Not blocking here — this repo's guidelines do not mandate an Evidence section, the template is filled out correctly, and the revert-to-main differential is a stronger signal than most PRs offer. Flagging only so it is a conscious omission: nothing in this PR shows a real maintenance run failing to re-page an admin outside of pytest.

[RISK ASSESSMENT]

  • [Overall PR] ⚠️ Risk Assessment: 🟡 MEDIUM

    The blast radius is narrow — one store method, one caller (OrgBudgetService._replace_thresholds), no schema migration, no API surface change, and threshold.id is only ever read out in a response model (server/routes/orgs.py:1218), so preserving row identity breaks nothing downstream. Against that: this code governs whether admins get paged about spend, and the failure modes are quiet in both directions. Too many alerts trains people to ignore them; too few means a budget blows through unnoticed. The change moves the system from over-alerting to under-alerting, which is the less visible and therefore more dangerous error. The channel-enable suppression above is a live instance of that. Not high risk, because the change is small and the tests genuinely pin the behaviour, but it warrants a human who owns budget alerting signing off on the trade-off rather than a rubber stamp.

VERDICT:

❌ Needs rework — narrowly. Make the duplicate collapse deterministically prefer the latched row so the test earns its name, and document the per-channel suppression. Everything else is sound and I would approve on those two.

KEY INSIGHT:

Identifying a threshold by its percentage instead of its row is the right call, but percentage is not actually unique in the database, and the one test that confronts that fact passes by accident of insertion order rather than by construction.


Improve this review? If any feedback above seems incorrect or irrelevant to this repository, you can teach the reviewer to do better:

  1. Add a .agents/skills/custom-codereview-guide.md file to your branch (or edit it if one already exists) with the /codereview trigger and the context the reviewer is missing (e.g., "Security concerns about X do not apply here because Y"). See the customization docs for the required frontmatter format.
  2. Re-request a review - the reviewer reads guidelines from the PR branch, so your changes take effect immediately.
  3. When your PR is merged, the guideline file goes through normal code review by repository maintainers.

Resolve with AI? Install the iterate skill in your agent and run /iterate to automatically drive this PR through CI, review, and QA until it's merge-ready.

Was this review helpful? React with 👍 or 👎 to give feedback.


This review was generated by an AI agent (OpenHands) on behalf of the requesting user.

Comment thread storage/org_budget_store.py
Comment thread storage/org_budget_store.py
Comment thread storage/org_budget_store.py Outdated
Comment thread tests/unit/test_org_budget_store.py
Order existing rows by latch recency before the replace loop so the
surviving copy of a duplicated percentage is the latched one, instead of
whichever row the query returned first. Strengthen the duplicate-collapse
test to insert the unlatched row first so it can no longer pass by
insertion-order accident, and trim the lead comment to the invariant.

Co-authored-by: openhands <openhands@all-hands.dev>
@aivong-openhands

Copy link
Copy Markdown
Contributor Author

Thanks for the thorough review — addressed both blocking items in 12f2aa1.

Critical: duplicate collapse now deterministically keeps the latched row. replace_thresholds sorts existing by latch recency (last_triggered_cycle_start, most recent first, None last) before the diff loop, so the surviving copy of a duplicated percentage is the one carrying the once-per-cycle latch — not whichever row Postgres returned first. I confirmed the failure you described: with the unlatched row inserted first, the test fails on the pre-fix code (latch lost!) and only passes with the sort.

Test earns its name. test_duplicate_rows_for_one_percentage_collapse_onto_the_latched_row now inserts the unlatched row first, so it can no longer pass by insertion-order accident.

Per-channel suppression documented. Added an explicit paragraph to the PR description: enabling a channel on an already-latched threshold now suppresses that channel for the rest of the cycle (the turn-Slack-on-for-80% case), the old row-churn behaviour delivered it as a bug side effect, per-channel latching is the larger change we deliberately did not take, and this trends toward under-alerting so budget-alerting ownership should sign off.

Comment trimmed to the invariant (identity = percentage because the latch lives on the row); the non-obvious no-unique-index note is retained on the sort/collapse block.

48 passed, 6 skipped; ruff clean.

This comment was created by an AI agent (OpenHands) on behalf of the requesting user.

@aivong-openhands
aivong-openhands marked this pull request as ready for review September 18, 2026 18:32
@aivong-openhands aivong-openhands changed the title fix(budgets): keep the once-per-cycle alert latch across a threshold edit fix(budgets): PLTF-3562 keep the once-per-cycle alert latch across a threshold edit Sep 21, 2026
@aivong-openhands

Copy link
Copy Markdown
Contributor Author

@OpenHands is there an E2E test in https://github.com/OpenHands/OpenHands-Cloud/blob/main/e2e_tests/tests/009-budgets.spec.ts that would catch this bug?

@openhands-ai

openhands-ai Bot commented Sep 21, 2026

Copy link
Copy Markdown

I'm on it! aivong-openhands can track my progress at all-hands.dev

Copy link
Copy Markdown
Contributor Author

Short answer: no — no current E2E test in 009-budgets.spec.ts would catch this bug.

I read through the full spec (and its helpers, e2e_tests/utils/budgets.ts and budget-database.ts). The only place that crosses a threshold is the "issue 5" setup in beforeAll (api.patchBudget({ thresholds: [{ percentage: 100, ... }] })). That is the first and only time a threshold is added during the run: it patches the Slack config and asserts, in test("issue 5: ..."), that exactly one alert fires with spend == alertReportingSpendAfter. Since this PR re-arms alerts only when an existing threshold row is edited (delete + re-insert), and never once in issue 5 does a 100% threshold exist before a maintenance run and also get edited, the spec never exercises the edited-latch codepath.

Concretely, the bug needs a sequence the suite never performs:

  1. a threshold fires while a percentage is present → last_triggered_cycle_start is set, then
  2. the same percentage is edited (e.g. slack_enabled=false → true) via patchBudget/update_budget_settings, then
  3. another maintenance run within the same cycle would re-send the alert (buggy) vs. stay silent (fixed).

The closest case, issue 5, patches the thresholds before the spend that crosses them, so no latched threshold is ever edited — the re-arm path is never hit. And the fix's columns (email_enabled, slack_enabled) aren't touched in the store tests, which mock _send_alerts rather than running maintenance against real Slack.

Even if it did exercise the path, the issue-5 machinery couldn't assert on latching: BudgetMaintenanceResult only exposes {id, status, info, updatedAt} (no per-threshold trigger/latch detail), and the Slack assertion is "at least one alert after oldest", not "exactly one".

How the suite could catch it

Add a numeric lowering to exercise 100% twice in one cycle:

  1. configure monthly_limit high enough that 100% won't trigger, run maintenance — no alert,
  2. generate spend ≈ the alert limit, run maintenance — one alert (fires for the first time),
  3. patchBudget editing the existing 100% row (flip slack_enabled), keep spend above the threshold, run maintenance again — the buggy code sends a second alert, fixed code stays silent,
  4. assert the Slack history for the window contains exactly one alert, matched by spend.

If a clean same-cycle spend measurement is awkward in the current architecture, findSlackBudgetAlert + a second conversations.history scan for a post-edit oldest-window (or a count of matching messages) would work. Tracked as a follow-up to keep this PR's scope unit-level — happy to write it if you want it in this PR.

@openhands-ai

openhands-ai Bot commented Sep 21, 2026

Copy link
Copy Markdown

The comment has been answered. I posted the response to the PR thread at #416 (comment).

Summary of my analysis (what the question asked and what I found):

The question was whether any E2E test in OpenHands-Cloud/e2e_tests/tests/009-budgets.spec.ts would catch the bug fixed by PR #416 (threshold edits re-arming the once-per-cycle alert latch).

I fetched the full spec plus its helpers (utils/budgets.ts, utils/budget-database.ts) and the PR's own diff, and concluded:

  1. No — no current E2E test catches it. The only threshold-crossing test in the suite is "issue 5: LiteLLM-only threshold crossing emits authoritative Slack spend". Its setup (beforeAll) patches the 100% threshold before any spend crosses it, so the threshold is never edited after it has fired — the exact sequence the bug requires (threshold fires → same percentage edited → maintenance runs again) never happens anywhere in the file.

  2. The suites' observability is also insufficient to assert latching. BudgetMaintenanceResult exposes only {id, status, info, updatedAt} (no per-threshold trigger/latch detail), and the Slack assertion (findSlackBudgetAlert) checks "at least one matching message since oldest", not "exactly one" — so a duplicate re-feed wouldn't fail the assertion even if the path ran.

  3. I described how the suite could be extended to catch it (exercise one percentage twice within a cycle: let it fire once, edit it via patchBudget, run maintenance again, and assert exactly one alert in the Slack window), and offered to write that follow-up if wanted.

I kept this as an answer to the reviewer's question rather than a code change — the PR's own unit/mutation coverage already exercises the latch behavior; the E2E gap is a follow-up opportunity, which I flagged.

@ak684 ak684 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed threshold updates, retention of the cycle notification latch, and duplicate-row selection. Focused budget-store/service tests: 48 passed, 6 skipped. No blocking findings. This addresses duplicate alerts when thresholds are saved; the separately tracked spend-reset and failed-Slack-delivery retry issues remain separate work.

@ak684
ak684 merged commit b9133ae into main Sep 22, 2026
16 checks passed
@ak684
ak684 deleted the fix/threshold-edit-preserves-alert-latch branch September 22, 2026 19:51
@openhands-release-bot openhands-release-bot Bot added the released: 1.64.0 Shipped in 1.64.0 label Sep 23, 2026
@openhands-release-bot

Copy link
Copy Markdown
Contributor

🚀 Released in 1.64.0.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

quint-studio-budgets-fixes Org budgets defects surfaced by the Quint Studio model released: 1.64.0 Shipped in 1.64.0 type: fix A bug fix

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants