Skip to content

fix(handlers): attribute partial account-deletion failures and emit retryable codes - #677

Merged
dmeiser merged 11 commits into
mainfrom
fm/KW-673-AUDITABILITY
Oct 4, 2026
Merged

dmeiser merged 11 commits into
mainfrom
fm/KW-673-AUDITABILITY

Conversation

@dmeiser

@dmeiser dmeiser commented Oct 3, 2026

Copy link
Copy Markdown
Owner

Intent

Make the account-deletion partial-state failure report honestly and attribute it to the actor. On the admin path the new logger.error omits actor_sub, which every other line in admin_purge_user_account carries - so the single most audit-worthy event in an admin-triggered destructive path is the one with no actor attribution. The Cognito helper already logs the failure with the real AWS code, so the handler logs it twice and neither line identifies the operator. The removed outer except ClientError was the only place a sweep failure was logged with account_id, so sweep failures now reach the decorator whose log line does not say which account was affected. The message instructs the caller to retry but emits INTERNAL_ERROR, the catch-all the frontend maps to a generic failure, when this project mandates RESOURCE_BUSY for a transient or throttled Cognito failure. The error_code field is a constant that can never carry anything else - it is either the fallback or the helper own INTERNAL_ERROR - and it collides with four other sites that use error_code for the AWS code. And the message says the account data was deleted, which is true for self-service but false for the admin path where the accounts row is deleted after the Cognito step and survives the raise, so an operator could conclude the record is gone and skip the retry, orphaning it permanently. Log once with full context, attribute every line to the actor and the account, emit the code that matches the retry guidance, and name the resources that actually survive per path.

What Changed

  • Added a shared run_deletion_steps helper in src/handlers/deletion_cascade.py that both delete_my_account and admin_purge_user_account use to run the data sweep then the Cognito delete, logging each failure exactly once with per-path context (account_id on the self-service path, account_id plus actor_sub on the admin path) and naming the resources that actually survive each failure phase.
  • Transient Cognito faults surviving the retry wrapper now raise RESOURCE_BUSY with retry-guidance messages instead of INTERNAL_ERROR, using a single-home COGNITO_TRANSIENT_ERROR_CODES set and is_transient_cognito_error classifier in src/utils/cognito.py; sweep failures are likewise classified via is_transient_client_error, and the admin Cognito delete propagates its real AWS error code to the one attributed log line.
  • Split the admin purge's Cognito user lookup into its own failure handling (retryable on throttle) and extracted the residue sweep into _sweep_purge_residue, with new unit tests in test_account_operations.py and test_admin_operations.py covering attribution, single-log, transient classification, and per-path surviving-resource messages; AGENTS.md's transient-classification entry was corrected to reflect the separate Cognito classifier.

Risk Assessment

⚠️ Medium: The change correctly fixes the attribution, single-logging, code/message agreement, and per-path resource-naming the intent demands, but leaves a misaligned transient classification (InternalErrorException and un-split sweep/lookup catches mapping retryable failures to INTERNAL_ERROR) and an AppError-typed sweep-failure path that still reaches the decorator without account attribution — both user-visible in destructive flows and mechanically fixable.

Testing

I stood up the handlers the way AppSync invokes them (real decorated entry points, AppSync-shaped events) against a disposable local AWS instance — moto's HTTP server fronted by per-service fault-injecting proxies so botocore parses genuine AWS error responses — pointed there exclusively through the product's own endpoint overrides, and drove 10 scenarios covering every failure path the intent enumerates (attribution, single log line, retryable code vs. guidance, per-path survival naming) with post-state read back from the local instance and retry counts observed at the proxy; a base-commit copy of src/ reproduced the reported defect under the same driver, proving sensitivity. All target scenarios passed, the regression reproduction matched the reported failure, the change's own unit tests (287) pass, and the worktree was left clean.

  • Live validation: ✅ go - 11 of 11 scenarios driven live against the product
Scenario Result Live Evidence
Admin purge: transient Cognito delete fault (InternalErrorException) -> retryable RESOURCE_BUSY, logged exactly once with actor_sub, account_id and the real AWS code ✅ pass live target_results.md — scenario 1 (payload errorCode RESOURCE_BUSY; single ERROR line with actor_sub=admin-operator-live-sub, aws_error_code=InternalErrorException; accounts row + Cognito user verified s…
Admin purge: permanent Cognito delete fault -> INTERNAL_ERROR with manual-completion message, no false retry promise, survivor named honestly ✅ pass live target_results.md — scenario 2 (message contains 'complete the deletion manually' and 'the accounts record still exists', no 'retry'; state verified)
Admin purge: throttled data sweep -> RESOURCE_BUSY with 'Cognito user and the accounts record are untouched', both verified untouched, Cognito delete never attempted ✅ pass live target_results.md — scenario 3 (4 injected ProvisionedThroughputExceededException Query faults observed at proxy; state checks true; no AdminDeleteUser request)
Admin purge: sweep raises typed AppError (S3 fault in QR purge) -> re-raised unchanged with an attributed sweep-failure log line ✅ pass live target_results.md — scenario 4 (payload message exactly 'Failed to purge payment QR codes from S3'; sweep line carries actor_sub + account_id + error_code)
Admin purge: transient Cognito lookup fault -> RESOURCE_BUSY retry guidance before any deletion (R3 sibling fix) ✅ pass live target_results.md — scenario 5 (RESOURCE_BUSY 'Retry the purge'; accounts row + user intact; no AdminDeleteUser)
Admin purge success path: returns True and actually deletes the accounts row and the Cognito user, final audit line attributed ✅ pass live target_results.md — scenario 6 (payload true; accounts_row_exists=False, cognito_user_exists=False; 'User account purged' line carries actor_sub + account_id)
Self-service delete: transient Cognito fault after sweep -> RESOURCE_BUSY, 'Account data was deleted' verified TRUE (accounts row gone), Cognito user survives, 3 real retries observed ✅ pass live target_results.md — scenario 7 (proxy rule hits=3; accounts row deleted, Cognito user present; single ERROR line with account_id + aws_error_code)
Self-service delete: permanent Cognito fault -> INTERNAL_ERROR 'complete the deletion manually in Cognito', no retry promise, zero pointless retries ✅ pass live target_results.md — scenario 8 (rule hits=1; state accounts row gone / Cognito user survives)
Self-service delete: throttled sweep -> RESOURCE_BUSY, 'the Cognito user is untouched' verified against persisted state ✅ pass live target_results.md — scenario 9 (4 throttled profiles-Query attempts; Cognito user + accounts row present)
Self-service delete: transient lookup fault retries via the project wrapper then surfaces RESOURCE_BUSY ✅ pass live target_results.md — scenario 10 (exactly 3 ListUsers attempts at proxy; RESOURCE_BUSY 'Failed to delete account. Retry to complete the deletion.')
Regression sensitivity: base-commit src under the same driver exhibits the reported defect (INTERNAL_ERROR for a transient code, unattributed helper log line) ✅ pass live baseline_results.md — payload {'errorCode': 'INTERNAL_ERROR', 'message': 'Failed to delete user from Cognito'}; error line 'Cognito admin_delete_user failed' with error_code=InternalErrorException but…
Evidence: Live driver transcript — target commit (10 scenarios, payloads, log lines, state, checks)
# Live driver results — target (target commit 15b8f44)

## PASS: admin purge: transient Cognito delete fault is retryable and logged once with actor+account+AWS code

Returned payload:
`` `json
{
  "__isError": true,
  "errorCode": "RESOURCE_BUSY",
  "message": "Account data was swept but the Cognito user could not be confirmed deleted; the accounts record still exists. Retry the purge to complete the deletion."
}
`` `

Post-state: {"accounts_row_exists": true, "cognito_user_exists": true}

Error-level structured log lines:
`` `json
[
  {
    "timestamp": "2026-10-03T18:05:15.246424+00:00",
    "level": "ERROR",
    "message": "Account data swept but the Cognito delete did not report success; the account's Cognito state is unknown and needs a retry or manual completion",
    "correlationId": "f09100f2-5f4c-4c87-8735-410bd6c8930e",
    "account_id": "e28e8d36-4bde-48ca-b607-852f8043597a",
    "actor_sub": "admin-operator-live-sub",
    "error": "An error occurred (InternalErrorException) when calling the AdminDeleteUser operation: Cognito backend unstable",
    "aws_error_code": "InternalErrorException",
    "traceback": "Traceback (most recent call last):\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/handlers/admin_operations.py\", line 893, in _delete_cognito_user_after_sweep\n    _delete_user_from_cognito(cognito, user_pool_id, username, email, logger, actor_sub)\n    ~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/handlers/admin_operations.py\", line 850, in _delete_user_from_cognito\n    cognito.admin_delete_user(UserPoolId=user_pool_id, Username=username)\n    ~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/botocore/client.py\", line 602, in _api_call\n    return self._make_api_call(operation_name, kwargs)\n           ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/botocore/context.py\", line 123, in wrapper\n    return func(*args, **kwargs)\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/botocore/client.py\", line 1078, in _make_api_call\n    raise error_class(parsed_response, operation_name)\nbotocore.errorfactory.InternalErrorException: An error occurred (InternalErrorException) when calling the AdminDeleteUser operation: Cognito backend unstable\n"
  }
]
`` `

Checks:
- [x] payload is the structured GraphQL error payload
- [x] errorCode is retryable RESOURCE_BUSY (not the INTERNAL_ERROR catch-all)
- [x] message carries retry guidance
- [x] message names the resource that actually survives on the admin path
- [x] exactly ONE handler error line for the failure (no double logging)
- [x] that line is the only ERROR the handler emitted
- [x] failure line attributes the actor (actor_sub)
- [x] failure line names the affected account (account_id)
- [x] field carries the real AWS code (not a constant)
- [x] no legacy helper line 'Cognito admin_delete_user failed'
- [x] accounts row really still exists (message claim verified)
- [x] Cognito user really still exists (message claim verified)
- [x] fault actually exercised through the real client stack
- [x] completes in sane wall time

Proxy request counts: {"cognito-idp": 2, "dynamodb": 2, "s3": 1}

## PASS: admin purge: permanent Cognito delete fault is INTERNAL_ERROR without a false retry promise

Returned payload:
`` `json
{
  "__isError": true,
  "errorCode": "INTERNAL_ERROR",
  "message": "Account data was swept but the Cognito user could not be confirmed deleted; the accounts record still exists and the Cognito user may also survive \u2014 complete the deletion manually."
}
`` `

Post-state: {"accounts_row_exists": true, "cognito_user_exists": true}

Error-level structured log lines:
`` `json
[
  {
    "timestamp": "2026-10-03T18:05:15.304980+00:00",
    "level": "ERROR",
    "message": "Account data swept but the Cognito delete did not report success; the account's Cognito state is unknown and needs a retry or manual completion",
    "correlationId": "f41a3acc-7c18-47b8-abd0-f2e4df345fc0",
    "account_id": "7a127fe7-37fd-4dc2-9e96-76778dbe13eb",
    "actor_sub": "admin-operator-live-sub",
    "error": "An error occurred (NotAuthorizedException) when calling the AdminDeleteUser operation: access denied by Cognito",
    "aws_error_code": "NotAuthorizedException",
    "traceback": "Traceback (most recent call last):\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/handlers/admin_operations.py\", line 893, in _delete_cognito_user_after_sweep\n    _delete_user_from_cognito(cognito, user_pool_id, username, email, logger, actor_sub)\n    ~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/handlers/admin_operations.py\", line 850, in _delete_user_from_cognito\n    cognito.admin_delete_user(UserPoolId=user_pool_id, Username=username)\n    ~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/botocore/client.py\", line 602, in _api_call\n    return self._make_api_call(operation_name, kwargs)\n           ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/botocore/context.py\", line 123, in wrapper\n    return func(*args, **kwargs)\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/botocore/client.py\", line 1078, in _make_api_call\n    raise error_class(parsed_response, operation_name)\nbotocore.errorfactory.NotAuthorizedException: An error occurred (NotAuthorizedException) when calling the AdminDeleteUser operation: access denied by Cognito\n"
  }
]
`` `

Checks:
- [x] payload is the structured GraphQL error payload
- [x] errorCode is the permanent INTERNAL_ERROR
- [x] message does NOT promise a retry
- [x] message points at manual completion
- [x] message names the surviving accounts record
- [x] exactly ONE handler error line
- [x] failure line attributes the actor
- [x] failure line names the account
- [x] failure line carries the real AWS code
- [x] accounts row survives (message claim verified)
- [x] Cognito user survives (message claim verified)
- [x] fault actually exercised

Proxy request counts: {"cognito-idp": 2, "dynamodb": 2, "s3": 1}

## PASS: admin purge: throttled sweep call is retryable and attributed; untouched-state claim verified

Returned payload:
`` `json
{
  "__isError": true,
  "errorCode": "RESOURCE_BUSY",
  "message": "Data sweep failed before the Cognito delete; the Cognito user and the accounts record are untouched. Retry the purge to complete the deletion."
}
`` `

Post-state: {"accounts_row_exists": true, "cognito_user_exists": true}

Error-level structured log lines:
`` `json
[
  {
    "timestamp": "2026-10-03T18:05:16.049645+00:00",
    "level": "ERROR",
    "message": "Data sweep failed before the Cognito delete; the Cognito user and the accounts record are untouched",
    "correlationId": "d8539767-df95-4f08-af6f-47556d3ee277",
    "account_id": "53b21e8a-4a13-4ae0-b319-41557b589c80",
    "actor_sub": "admin-operator-live-sub",
    "error": "An error occurred (ProvisionedThroughputExceededException) when calling the Query operation (reached max retries: 0): sweep throttled",
    "aws_error_code": "ProvisionedThroughputExceededException",
    "traceback": "Traceback (most recent call last):\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/handlers/admin_operations.py\", line 1094, in admin_purge_user_account\n    delete_inbound_shares(account_id, logger)\n    ~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^

... [17629 bytes truncated] ...

)\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/botocore/client.py\", line 1078, in _make_api_call\n    raise error_class(parsed_response, operation_name)\nbotocore.errorfactory.NotAuthorizedException: An error occurred (NotAuthorizedException) when calling the AdminDeleteUser operation: access denied by Cognito\n"
  }
]
`` `

Checks:
- [x] payload is the structured GraphQL error payload
- [x] errorCode is the permanent INTERNAL_ERROR
- [x] message does NOT promise a retry
- [x] message points at manual completion
- [x] exactly ONE handler error line
- [x] failure line names the account
- [x] failure line carries the real AWS code
- [x] accounts row deleted by the sweep (claim verified)
- [x] Cognito user survives
- [x] no pointless retries for a permanent fault

Proxy request counts: {"cognito-idp": 2, "dynamodb": 3, "s3": 1}

## PASS: self-service delete: throttled sweep -> RESOURCE_BUSY; Cognito user verified untouched

Returned payload:
`` `json
{
  "__isError": true,
  "errorCode": "RESOURCE_BUSY",
  "message": "Data sweep failed before the Cognito delete; the Cognito user is untouched. Retry to complete the deletion."
}
`` `

Post-state: {"accounts_row_exists": true, "cognito_user_exists": true}

Error-level structured log lines:
`` `json
[
  {
    "timestamp": "2026-10-03T18:05:17.337293+00:00",
    "level": "ERROR",
    "message": "Data sweep failed before the Cognito delete; the Cognito user is untouched",
    "correlationId": "656dd0fe-5f72-4be9-bacf-1a0291f94a81",
    "account_id": "d07f6520-5987-4c76-9532-820bf720daa3",
    "error": "An error occurred (ProvisionedThroughputExceededException) when calling the Query operation (reached max retries: 0): sweep throttled",
    "aws_error_code": "ProvisionedThroughputExceededException",
    "traceback": "Traceback (most recent call last):\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/handlers/account_operations.py\", line 141, in delete_my_account\n    delete_all_user_data(account_id, logger)\n    ~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/handlers/deletion_cascade.py\", line 253, in delete_all_user_data\n    profiles = get_user_profiles(normalize_account_id(account_id))\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/handlers/deletion_cascade.py\", line 61, in get_user_profiles\n    return query_all_items(\n        tables.profiles,\n    ...<3 lines>...\n        },\n    )\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/utils/pagination.py\", line 129, in query_all_items\n    return list(_paginated_query(table, query_kwargs, max_items=max_items))\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/utils/pagination.py\", line 107, in _paginated_query\n    yield from _paginated(table, \"query\", query_kwargs, max_items=max_items, log_label=\"query\")\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/utils/pagination.py\", line 87, in _paginated\n    response = _call_with_retry(method, query_kwargs, log_label)\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/utils/pagination.py\", line 54, in _call_with_retry\n    return method(**kwargs)\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/boto3/resources/factory.py\", line 581, in do_action\n    response = action(self, *args, **kwargs)\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/boto3/resources/action.py\", line 88, in __call__\n    response = getattr(parent.meta.client, operation_name)(*args, **params)\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/botocore/client.py\", line 602, in _api_call\n    return self._make_api_call(operation_name, kwargs)\n           ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/botocore/context.py\", line 123, in wrapper\n    return func(*args, **kwargs)\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/botocore/client.py\", line 1078, in _make_api_call\n    raise error_class(parsed_response, operation_name)\nbotocore.errorfactory.ProvisionedThroughputExceededException: An error occurred (ProvisionedThroughputExceededException) when calling the Query operation (reached max retries: 0): sweep throttled\n"
  }
]
`` `

Checks:
- [x] payload is the structured GraphQL error payload
- [x] errorCode is RESOURCE_BUSY
- [x] message carries retry guidance
- [x] message states the Cognito user is untouched
- [x] exactly ONE sweep-failure line
- [x] sweep-failure line names the account
- [x] sweep-failure line carries the real AWS code
- [x] accounts row survives the failed sweep
- [x] Cognito user really untouched (claim verified)
- [x] Cognito delete never attempted
- [x] sweep throttle exercised incl. wrapper retries
- [x] completes in sane wall time

Proxy request counts: {"cognito-idp": 1, "dynamodb": 4, "s3": 0}

## PASS: self-service delete: transient lookup fault retries then surfaces RESOURCE_BUSY

Returned payload:
`` `json
{
  "__isError": true,
  "errorCode": "RESOURCE_BUSY",
  "message": "Failed to delete account. Retry to complete the deletion."
}
`` `

Post-state: {"accounts_row_exists": true, "cognito_user_exists": true}

Error-level structured log lines:
`` `json
[
  {
    "timestamp": "2026-10-03T18:05:17.672101+00:00",
    "level": "ERROR",
    "message": "Cognito lookup failed before deletion",
    "correlationId": "656dd0fe-5f72-4be9-bacf-1a0291f94a81",
    "account_id": "68437015-d708-4e02-8d09-14ebfd00b5ca",
    "error": "An error occurred (TooManyRequestsException) when calling the ListUsers operation: lookup throttled",
    "traceback": "Traceback (most recent call last):\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/handlers/account_operations.py\", line 121, in delete_my_account\n    username = _lookup_cognito_user_with_retry(cognito, user_pool_id, account_id)\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/handlers/account_operations.py\", line 54, in _lookup_cognito_user_with_retry\n    users_response = retry_on_transient_errors(\n        cognito.list_users, UserPoolId=user_pool_id, Filter=filter_expression, Limit=1\n    )\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/src/utils/cognito.py\", line 56, in retry_on_transient_errors\n    return method(*args, **kwargs)\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/botocore/client.py\", line 602, in _api_call\n    return self._make_api_call(operation_name, kwargs)\n           ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/botocore/context.py\", line 123, in wrapper\n    return func(*args, **kwargs)\n  File \"~/.no-mistakes/worktrees/0e3f105a97cb/01M419PA797T446TEJZTAN0NDS/.venv/lib/python3.14/site-packages/botocore/client.py\", line 1078, in _make_api_call\n    raise error_class(parsed_response, operation_name)\nbotocore.errorfactory.TooManyRequestsException: An error occurred (TooManyRequestsException) when calling the ListUsers operation: lookup throttled\n"
  }
]
`` `

Checks:
- [x] payload is the structured GraphQL error payload
- [x] errorCode is RESOURCE_BUSY after retry exhaustion
- [x] message carries retry guidance
- [x] lookup failure logged with the account
- [x] the project's own retry wrapper really retried (3 attempts)
- [x] nothing deleted before the abort
- [x] no Cognito delete attempted
- [x] completes in sane wall time

Proxy request counts: {"cognito-idp": 3, "dynamodb": 0, "s3": 0}
  • Evidence: Live driver raw results — target commit (local file: ~/.no-mistakes/evidence/01M419PA797T446TEJZTAN0NDS/target_results.json)
Evidence: Live driver transcript — base-commit regression reproduction
# Live driver results — baseline (base commit b4134fd)

## PASS: BASELINE (pre-fix src): same transient Cognito fault

Returned payload:
`` `json
{
  "__isError": true,
  "errorCode": "INTERNAL_ERROR",
  "message": "Failed to delete user from Cognito"
}
`` `

Post-state: {"accounts_row_exists": true, "cognito_user_exists": true}

Error-level structured log lines:
`` `json
[
  {
    "timestamp": "2026-10-03T18:05:48.893553+00:00",
    "level": "ERROR",
    "message": "Cognito admin_delete_user failed",
    "correlationId": "efaea3de-ff4a-4111-a970-245b410e5135",
    "error": "An error occurred (InternalErrorException) when calling the AdminDeleteUser operation: Cognito backend unstable",
    "error_code": "InternalErrorException"
  }
]
`` `

Checks:
- [x] PRE-FIX: transient fault surfaced as INTERNAL_ERROR (the reported defect)
- [x] PRE-FIX: helper emitted its own unattributed error line
- [x] PRE-FIX: that line has no actor_sub (the reported defect)
- [x] PRE-FIX: that line has no account_id (the reported defect)
- [x] PRE-FIX: no narration line naming surviving resources
- [x] accounts row survived anyway
- [x] fault exercised

Proxy request counts: {"cognito-idp": 2, "dynamodb": 2, "s3": 1}
Evidence: Live driver raw results — base commit
{
  "mode": "baseline",
  "src": "base commit b4134fd",
  "generated_at": "2026-10-03T14:05:49",
  "scenarios": [
    {
      "name": "BASELINE (pre-fix src): same transient Cognito fault",
      "passed": true,
      "payload": {
        "__isError": true,
        "errorCode": "INTERNAL_ERROR",
        "message": "Failed to delete user from Cognito"
      },
      "state": {
        "accounts_row_exists": true,
        "cognito_user_exists": true
      },
      "error_logs": [
        {
          "timestamp": "2026-10-03T18:05:48.893553+00:00",
          "level": "ERROR",
          "message": "Cognito admin_delete_user failed",
          "correlationId": "efaea3de-ff4a-4111-a970-245b410e5135",
          "error": "An error occurred (InternalErrorException) when calling the AdminDeleteUser operation: Cognito backend unstable",
          "error_code": "InternalErrorException"
        }
      ],
      "total_log_lines": 0,
      "proxy": {
        "cognito-idp": {
          "requests": 2,
          "faulted": 1,
          "ops": [
            "ListUsers",
            "AdminDeleteUser"
          ],
          "rules": [
            {
              "service": "cognito-idp",
              "op": "AdminDeleteUser",
              "table": null,
              "code": "InternalErrorException",
              "hits": 1
            }
          ]
        },
        "dynamodb": {
          "requests": 2,
          "faulted": 0,
          "ops": [
            "GetItem",
            "Query"
          ],
          "rules": []
        },
        "s3": {
          "requests": 1,
          "faulted": 0,
          "ops": [
            "ListObjectVersions"
          ],
          "rules": []
        }
      },
      "checks": [
        {
          "desc": "PRE-FIX: transient fault surfaced as INTERNAL_ERROR (the reported defect)",
          "ok": true,
          "detail": {
            "__isError": true,
            "errorCode": "INTERNAL_ERROR",
            "message": "Failed to delete user from Cognito"
          }
        },
        {
          "desc": "PRE-FIX: helper emitted its own unattributed error line",
          "ok": true,
          "detail": [
            "Cognito admin_delete_user failed"
          ]
        },
        {
          "desc": "PRE-FIX: that line has no actor_sub (the reported defect)",
          "ok": true,
          "detail": null
        },
        {
          "desc": "PRE-FIX: that line has no account_id (the reported defect)",
          "ok": true,
          "detail": null
        },
        {
          "desc": "PRE-FIX: no narration line naming surviving resources",
          "ok": true,
          "detail": [
            "Cognito admin_delete_user failed"
          ]
        },
        {
          "desc": "accounts row survived anyway",
          "ok": true,
          "detail": ""
        },
        {
          "desc": "fault exercised",
          "ok": true,
          "detail": 1
        }
      ]
    }
  ]
}
  • Evidence: The live driver itself (moto server + fault-injecting proxies + 11 scenarios) (local file: ~/.no-mistakes/evidence/01M419PA797T446TEJZTAN0NDS/live_driver.py)

Pipeline

Updates from git push no-mistakes

... (5 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)

⚠️ **Review** - 7 issues (3 warnings, 4 infos)

🔧 Fix applied.
6 issues (3 warnings, 3 infos) still open:

  • ⚠️ src/handlers/admin_operations.py:896 - Transient classification of Cognito failures is misaligned with the repo's own retry policy, so a genuinely transient Cognito code still surfaces as INTERNAL_ERROR with 'complete the deletion manually' instead of the mandated retryable RESOURCE_BUSY. Concrete sequence (admin path): Cognito backend unstable → admin_delete_user raises ClientError code InternalErrorException → is_transient_client_error(e) checks TRANSIENT_ERROR_CODES (src/utils/dynamodb.py:54-56 = {ProvisionedThroughputExceededException, ThrottlingException, TooManyRequestsException}) which omits InternalErrorException → falls into the INTERNAL_ERROR branch with 'the accounts record still exists and the Cognito user may also survive — complete the deletion manually', telling an operator to do manual console work for a failure the project itself defines as retryable (_RETRYABLE_CODES in src/utils/cognito.py:12-14 includes InternalErrorException). Self path (src/handlers/account_operations.py:158) is worse: the internal retry_on_transient_errors (account_operations.py:74) retries InternalErrorException twice, then the catch classifies it as permanent — same code, two different transient-definitions on one call chain. This is the same code-vs-retry-guidance contradiction the change was written to eliminate, surviving for one member of the Cognito transient set. Same invariant is violated at: account_operations.py:158 (same check, same gap for InternalErrorException); account_operations.py:133-140 (sweep catch maps ANY ClientError to INTERNAL_ERROR with no transient split — a throttled query_all_items/accounts-row delete_item from delete_all_user_data (raw ClientError re-raised at src/handlers/deletion_cascade.py:268-272) gets 'Failed to delete account data' / INTERNAL_ERROR, while the identical throttle raised by batch_delete_keys on the same sweep becomes retryable RESOURCE_BUSY via _raise_delete_error, src/handlers/campaign_operations.py:90-98 — retry converges for both, so RESOURCE_BUSY matches the guidance in both); account_operations.py:124-126 (lookup catch, same no-split shape). Mechanical remedy contained to this change: classify Cognito delete failures against the Cognito retryable set (union with the canonical set or a shared is_transient_cognito_error), and apply the same transient→RESOURCE_BUSY split in the sweep/lookup catches instead of blanket INTERNAL_ERROR.
  • ℹ️ src/handlers/admin_operations.py:1098 - Residual attribution gap inside the same destructive purge path, unchanged by this diff: the accounts-record deletion failure logs 'Failed to delete account record' with error and account_id but no actor_sub (src/handlers/admin_operations.py:1098-1099), breaking the admin_operations.py never logs the acting admin; _validate_admin_and_get_account_id returns the target, not the actor #507 invariant ('every admin audit line downstream records both') that the newly added lines now honor; the shared cascade helper's failure line 'Failed to delete account from DynamoDB' (src/handlers/deletion_cascade.py:271) likewise lacks actor_sub. Noting for a follow-up pass; not a blocker for this change.
  • ℹ️ src/handlers/account_operations.py:170 - The unconditional success narration 'Deleted account from DynamoDB' inside delete_all_user_data (src/handlers/deletion_cascade.py:269, unchanged) is ordered last in the sweep, so a sweep failure before that point combined with the change's honest 'retry converges' framing is consistent — the narration is only emitted on actual success. No defect found on the ordering; noting the two paths' resource-naming claims were verified: self-service deletes the accounts row inside the sweep (deletion_cascade.py:268) so 'Account data was deleted' is true; admin deletes it after the Cognito step (admin_operations.py:1090-1096) so 'the accounts record still exists' is true. No action needed.
  • ⚠️ src/utils/dynamodb_exceptions.py:15 - Fix-round change: COGNITO_TRANSIENT_ERROR_CODES duplicates _RETRYABLE_CODES (src/utils/cognito.py:12-14) with only a comment claiming parity — a second hand-maintained definition of the same concept, unenforced. If a code is later added to the retry wrapper but not here (or vice versa), the exhausted-retry classification (is_transient_cognito_error at account_operations.py:130/185, admin_operations.py:900) diverges from what actually gets retried — the same code-vs-retry-guidance contradiction F1 targeted, reintroduced structurally. Remedy contained to the change: make the set single-homed (export a public name from utils.cognito and build COGNITO_TRANSIENT_ERROR_CODES from it, or vice versa). This is also the smallest honest remedy under the simplification pass: the parallel copy of the rule is not required by the intent, which only requires one consistent classification.
  • ⚠️ src/handlers/admin_operations.py:828 - Sibling left behind by the fix round: _find_cognito_user_by_sub (called by admin_purge_user_account at admin_operations.py:1058) maps ANY ClientError — transient Cognito codes included — to INTERNAL_ERROR 'Failed to look up Cognito user' with no retry wrapper, while the self path's identical lookup now emits retryable RESOURCE_BUSY when the same fault exhausts its retries (src/handlers/account_operations.py:126-138, the fix round's own work). Concrete sequence on the purge path: Cognito list_users throttled (TooManyRequestsException/InternalErrorException) → aborts the purge with the catch-all the frontend maps to a generic failure, contradicting the intent mandate that a transient or throttled Cognito failure carry retry guidance. Note the purge path uses this lookup with NO retry wrapper at all, so the throttle aborts on the first attempt. Remedy contained: classify against is_transient_cognito_error here like the sibling sites (the missing actor_sub in this log line is the F3 gap the user chose to ignore — not part of this finding).
  • ℹ️ src/handlers/admin_operations.py:885 - Round-2 fix (e23e66f) copied the self-service path's failure-provenance comment into the admin wrapper's docstring: '_delete_cognito_user_after_sweep' claims a 'failed re-lookup of an absent user lands here too', but the admin path's _delete_user_from_cognito (admin_operations.py:839) takes an already-resolved username from _find_cognito_user_by_sub and performs no re-lookup — that clause describes only the self path (account_operations.py:174, where the fallback lookup inside the account _delete_user_from_cognito can fail). Minor doc inaccuracy on the destructive purge path; drop the clause.

🔧 Fix applied.
7 issues (3 warnings, 4 infos) still open:

  • ⚠️ src/handlers/admin_operations.py:896 - Transient classification of Cognito failures is misaligned with the repo's own retry policy, so a genuinely transient Cognito code still surfaces as INTERNAL_ERROR with 'complete the deletion manually' instead of the mandated retryable RESOURCE_BUSY. Concrete sequence (admin path): Cognito backend unstable → admin_delete_user raises ClientError code InternalErrorException → is_transient_client_error(e) checks TRANSIENT_ERROR_CODES (src/utils/dynamodb.py:54-56 = {ProvisionedThroughputExceededException, ThrottlingException, TooManyRequestsException}) which omits InternalErrorException → falls into the INTERNAL_ERROR branch with 'the accounts record still exists and the Cognito user may also survive — complete the deletion manually', telling an operator to do manual console work for a failure the project itself defines as retryable (_RETRYABLE_CODES in src/utils/cognito.py:12-14 includes InternalErrorException). Self path (src/handlers/account_operations.py:158) is worse: the internal retry_on_transient_errors (account_operations.py:74) retries InternalErrorException twice, then the catch classifies it as permanent — same code, two different transient-definitions on one call chain. This is the same code-vs-retry-guidance contradiction the change was written to eliminate, surviving for one member of the Cognito transient set. Same invariant is violated at: account_operations.py:158 (same check, same gap for InternalErrorException); account_operations.py:133-140 (sweep catch maps ANY ClientError to INTERNAL_ERROR with no transient split — a throttled query_all_items/accounts-row delete_item from delete_all_user_data (raw ClientError re-raised at src/handlers/deletion_cascade.py:268-272) gets 'Failed to delete account data' / INTERNAL_ERROR, while the identical throttle raised by batch_delete_keys on the same sweep becomes retryable RESOURCE_BUSY via _raise_delete_error, src/handlers/campaign_operations.py:90-98 — retry converges for both, so RESOURCE_BUSY matches the guidance in both); account_operations.py:124-126 (lookup catch, same no-split shape). Mechanical remedy contained to this change: classify Cognito delete failures against the Cognito retryable set (union with the canonical set or a shared is_transient_cognito_error), and apply the same transient→RESOURCE_BUSY split in the sweep/lookup catches instead of blanket INTERNAL_ERROR.
  • ℹ️ src/handlers/admin_operations.py:1098 - Residual attribution gap inside the same destructive purge path, unchanged by this diff: the accounts-record deletion failure logs 'Failed to delete account record' with error and account_id but no actor_sub (src/handlers/admin_operations.py:1098-1099), breaking the admin_operations.py never logs the acting admin; _validate_admin_and_get_account_id returns the target, not the actor #507 invariant ('every admin audit line downstream records both') that the newly added lines now honor; the shared cascade helper's failure line 'Failed to delete account from DynamoDB' (src/handlers/deletion_cascade.py:271) likewise lacks actor_sub. Noting for a follow-up pass; not a blocker for this change.
  • ℹ️ src/handlers/account_operations.py:170 - The unconditional success narration 'Deleted account from DynamoDB' inside delete_all_user_data (src/handlers/deletion_cascade.py:269, unchanged) is ordered last in the sweep, so a sweep failure before that point combined with the change's honest 'retry converges' framing is consistent — the narration is only emitted on actual success. No defect found on the ordering; noting the two paths' resource-naming claims were verified: self-service deletes the accounts row inside the sweep (deletion_cascade.py:268) so 'Account data was deleted' is true; admin deletes it after the Cognito step (admin_operations.py:1090-1096) so 'the accounts record still exists' is true. No action needed.
  • ⚠️ src/utils/dynamodb_exceptions.py:15 - Fix-round change: COGNITO_TRANSIENT_ERROR_CODES duplicates _RETRYABLE_CODES (src/utils/cognito.py:12-14) with only a comment claiming parity — a second hand-maintained definition of the same concept, unenforced. If a code is later added to the retry wrapper but not here (or vice versa), the exhausted-retry classification (is_transient_cognito_error at account_operations.py:130/185, admin_operations.py:900) diverges from what actually gets retried — the same code-vs-retry-guidance contradiction F1 targeted, reintroduced structurally. Remedy contained to the change: make the set single-homed (export a public name from utils.cognito and build COGNITO_TRANSIENT_ERROR_CODES from it, or vice versa). This is also the smallest honest remedy under the simplification pass: the parallel copy of the rule is not required by the intent, which only requires one consistent classification.
  • ⚠️ src/handlers/admin_operations.py:828 - Sibling left behind by the fix round: _find_cognito_user_by_sub (called by admin_purge_user_account at admin_operations.py:1058) maps ANY ClientError — transient Cognito codes included — to INTERNAL_ERROR 'Failed to look up Cognito user' with no retry wrapper, while the self path's identical lookup now emits retryable RESOURCE_BUSY when the same fault exhausts its retries (src/handlers/account_operations.py:126-138, the fix round's own work). Concrete sequence on the purge path: Cognito list_users throttled (TooManyRequestsException/InternalErrorException) → aborts the purge with the catch-all the frontend maps to a generic failure, contradicting the intent mandate that a transient or throttled Cognito failure carry retry guidance. Note the purge path uses this lookup with NO retry wrapper at all, so the throttle aborts on the first attempt. Remedy contained: classify against is_transient_cognito_error here like the sibling sites (the missing actor_sub in this log line is the F3 gap the user chose to ignore — not part of this finding).
  • ℹ️ src/handlers/admin_operations.py:885 - Round-2 fix (e23e66f) copied the self-service path's failure-provenance comment into the admin wrapper's docstring: '_delete_cognito_user_after_sweep' claims a 'failed re-lookup of an absent user lands here too', but the admin path's _delete_user_from_cognito (admin_operations.py:839) takes an already-resolved username from _find_cognito_user_by_sub and performs no re-lookup — that clause describes only the self path (account_operations.py:174, where the fallback lookup inside the account _delete_user_from_cognito can fail). Minor doc inaccuracy on the destructive purge path; drop the clause.
  • ℹ️ src/handlers/admin_operations.py:919 - Cosmetic path-dependent wording mismatch within this change: the admin partial-state message reads 'Account data was swept but the Cognito user could not be confirmed deleted' (singular 'Account data') while its log narration (admin_operations.py:894) and the analogous self-path message ('Account data was deleted', account_operations.py:190) use different phrasings for the same distinction. Every intent requirement (honest survival naming, actor attribution, code/retry alignment) is met on both paths and verified against the source: the sweep (deletion_cascade.py:268) deletes the accounts row on the self path before the Cognito delete; the admin path deletes it after (admin_operations.py:1129-1136) so 'the accounts record still exists' is accurate there. Purely a narration-wording nuance; no action needed.

🔧 No changes applied.
7 issues (3 warnings, 4 infos) still open:

  • ⚠️ src/handlers/admin_operations.py:896 - Transient classification of Cognito failures is misaligned with the repo's own retry policy, so a genuinely transient Cognito code still surfaces as INTERNAL_ERROR with 'complete the deletion manually' instead of the mandated retryable RESOURCE_BUSY. Concrete sequence (admin path): Cognito backend unstable → admin_delete_user raises ClientError code InternalErrorException → is_transient_client_error(e) checks TRANSIENT_ERROR_CODES (src/utils/dynamodb.py:54-56 = {ProvisionedThroughputExceededException, ThrottlingException, TooManyRequestsException}) which omits InternalErrorException → falls into the INTERNAL_ERROR branch with 'the accounts record still exists and the Cognito user may also survive — complete the deletion manually', telling an operator to do manual console work for a failure the project itself defines as retryable (_RETRYABLE_CODES in src/utils/cognito.py:12-14 includes InternalErrorException). Self path (src/handlers/account_operations.py:158) is worse: the internal retry_on_transient_errors (account_operations.py:74) retries InternalErrorException twice, then the catch classifies it as permanent — same code, two different transient-definitions on one call chain. This is the same code-vs-retry-guidance contradiction the change was written to eliminate, surviving for one member of the Cognito transient set. Same invariant is violated at: account_operations.py:158 (same check, same gap for InternalErrorException); account_operations.py:133-140 (sweep catch maps ANY ClientError to INTERNAL_ERROR with no transient split — a throttled query_all_items/accounts-row delete_item from delete_all_user_data (raw ClientError re-raised at src/handlers/deletion_cascade.py:268-272) gets 'Failed to delete account data' / INTERNAL_ERROR, while the identical throttle raised by batch_delete_keys on the same sweep becomes retryable RESOURCE_BUSY via _raise_delete_error, src/handlers/campaign_operations.py:90-98 — retry converges for both, so RESOURCE_BUSY matches the guidance in both); account_operations.py:124-126 (lookup catch, same no-split shape). Mechanical remedy contained to this change: classify Cognito delete failures against the Cognito retryable set (union with the canonical set or a shared is_transient_cognito_error), and apply the same transient→RESOURCE_BUSY split in the sweep/lookup catches instead of blanket INTERNAL_ERROR.
  • ℹ️ src/handlers/admin_operations.py:1098 - Residual attribution gap inside the same destructive purge path, unchanged by this diff: the accounts-record deletion failure logs 'Failed to delete account record' with error and account_id but no actor_sub (src/handlers/admin_operations.py:1098-1099), breaking the admin_operations.py never logs the acting admin; _validate_admin_and_get_account_id returns the target, not the actor #507 invariant ('every admin audit line downstream records both') that the newly added lines now honor; the shared cascade helper's failure line 'Failed to delete account from DynamoDB' (src/handlers/deletion_cascade.py:271) likewise lacks actor_sub. Noting for a follow-up pass; not a blocker for this change.
  • ℹ️ src/handlers/account_operations.py:170 - The unconditional success narration 'Deleted account from DynamoDB' inside delete_all_user_data (src/handlers/deletion_cascade.py:269, unchanged) is ordered last in the sweep, so a sweep failure before that point combined with the change's honest 'retry converges' framing is consistent — the narration is only emitted on actual success. No defect found on the ordering; noting the two paths' resource-naming claims were verified: self-service deletes the accounts row inside the sweep (deletion_cascade.py:268) so 'Account data was deleted' is true; admin deletes it after the Cognito step (admin_operations.py:1090-1096) so 'the accounts record still exists' is true. No action needed.
  • ⚠️ src/utils/dynamodb_exceptions.py:15 - Fix-round change: COGNITO_TRANSIENT_ERROR_CODES duplicates _RETRYABLE_CODES (src/utils/cognito.py:12-14) with only a comment claiming parity — a second hand-maintained definition of the same concept, unenforced. If a code is later added to the retry wrapper but not here (or vice versa), the exhausted-retry classification (is_transient_cognito_error at account_operations.py:130/185, admin_operations.py:900) diverges from what actually gets retried — the same code-vs-retry-guidance contradiction F1 targeted, reintroduced structurally. Remedy contained to the change: make the set single-homed (export a public name from utils.cognito and build COGNITO_TRANSIENT_ERROR_CODES from it, or vice versa). This is also the smallest honest remedy under the simplification pass: the parallel copy of the rule is not required by the intent, which only requires one consistent classification.
  • ⚠️ src/handlers/admin_operations.py:828 - Sibling left behind by the fix round: _find_cognito_user_by_sub (called by admin_purge_user_account at admin_operations.py:1058) maps ANY ClientError — transient Cognito codes included — to INTERNAL_ERROR 'Failed to look up Cognito user' with no retry wrapper, while the self path's identical lookup now emits retryable RESOURCE_BUSY when the same fault exhausts its retries (src/handlers/account_operations.py:126-138, the fix round's own work). Concrete sequence on the purge path: Cognito list_users throttled (TooManyRequestsException/InternalErrorException) → aborts the purge with the catch-all the frontend maps to a generic failure, contradicting the intent mandate that a transient or throttled Cognito failure carry retry guidance. Note the purge path uses this lookup with NO retry wrapper at all, so the throttle aborts on the first attempt. Remedy contained: classify against is_transient_cognito_error here like the sibling sites (the missing actor_sub in this log line is the F3 gap the user chose to ignore — not part of this finding).
  • ℹ️ src/handlers/admin_operations.py:885 - Round-2 fix (e23e66f) copied the self-service path's failure-provenance comment into the admin wrapper's docstring: '_delete_cognito_user_after_sweep' claims a 'failed re-lookup of an absent user lands here too', but the admin path's _delete_user_from_cognito (admin_operations.py:839) takes an already-resolved username from _find_cognito_user_by_sub and performs no re-lookup — that clause describes only the self path (account_operations.py:174, where the fallback lookup inside the account _delete_user_from_cognito can fail). Minor doc inaccuracy on the destructive purge path; drop the clause.
  • ℹ️ src/handlers/admin_operations.py:919 - Cosmetic path-dependent wording mismatch within this change: the admin partial-state message reads 'Account data was swept but the Cognito user could not be confirmed deleted' (singular 'Account data') while its log narration (admin_operations.py:894) and the analogous self-path message ('Account data was deleted', account_operations.py:190) use different phrasings for the same distinction. Every intent requirement (honest survival naming, actor attribution, code/retry alignment) is met on both paths and verified against the source: the sweep (deletion_cascade.py:268) deletes the accounts row on the self path before the Cognito delete; the admin path deletes it after (admin_operations.py:1129-1136) so 'the accounts record still exists' is accurate there. Purely a narration-wording nuance; no action needed.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 11 of 11 scenarios driven live against the product
Scenario Result Live Evidence
Admin purge: transient Cognito delete fault (InternalErrorException) -> retryable RESOURCE_BUSY, logged exactly once with actor_sub, account_id and the real AWS code ✅ pass live target_results.md — scenario 1 (payload errorCode RESOURCE_BUSY; single ERROR line with actor_sub=admin-operator-live-sub, aws_error_code=InternalErrorException; accounts row + Cognito user verified s…
Admin purge: permanent Cognito delete fault -> INTERNAL_ERROR with manual-completion message, no false retry promise, survivor named honestly ✅ pass live target_results.md — scenario 2 (message contains 'complete the deletion manually' and 'the accounts record still exists', no 'retry'; state verified)
Admin purge: throttled data sweep -> RESOURCE_BUSY with 'Cognito user and the accounts record are untouched', both verified untouched, Cognito delete never attempted ✅ pass live target_results.md — scenario 3 (4 injected ProvisionedThroughputExceededException Query faults observed at proxy; state checks true; no AdminDeleteUser request)
Admin purge: sweep raises typed AppError (S3 fault in QR purge) -> re-raised unchanged with an attributed sweep-failure log line ✅ pass live target_results.md — scenario 4 (payload message exactly 'Failed to purge payment QR codes from S3'; sweep line carries actor_sub + account_id + error_code)
Admin purge: transient Cognito lookup fault -> RESOURCE_BUSY retry guidance before any deletion (R3 sibling fix) ✅ pass live target_results.md — scenario 5 (RESOURCE_BUSY 'Retry the purge'; accounts row + user intact; no AdminDeleteUser)
Admin purge success path: returns True and actually deletes the accounts row and the Cognito user, final audit line attributed ✅ pass live target_results.md — scenario 6 (payload true; accounts_row_exists=False, cognito_user_exists=False; 'User account purged' line carries actor_sub + account_id)
Self-service delete: transient Cognito fault after sweep -> RESOURCE_BUSY, 'Account data was deleted' verified TRUE (accounts row gone), Cognito user survives, 3 real retries observed ✅ pass live target_results.md — scenario 7 (proxy rule hits=3; accounts row deleted, Cognito user present; single ERROR line with account_id + aws_error_code)
Self-service delete: permanent Cognito fault -> INTERNAL_ERROR 'complete the deletion manually in Cognito', no retry promise, zero pointless retries ✅ pass live target_results.md — scenario 8 (rule hits=1; state accounts row gone / Cognito user survives)
Self-service delete: throttled sweep -> RESOURCE_BUSY, 'the Cognito user is untouched' verified against persisted state ✅ pass live target_results.md — scenario 9 (4 throttled profiles-Query attempts; Cognito user + accounts row present)
Self-service delete: transient lookup fault retries via the project wrapper then surfaces RESOURCE_BUSY ✅ pass live target_results.md — scenario 10 (exactly 3 ListUsers attempts at proxy; RESOURCE_BUSY 'Failed to delete account. Retry to complete the deletion.')
Regression sensitivity: base-commit src under the same driver exhibits the reported defect (INTERNAL_ERROR for a transient code, unattributed helper log line) ✅ pass live baseline_results.md — payload {'errorCode': 'INTERNAL_ERROR', 'message': 'Failed to delete user from Cognito'}; error line 'Cognito admin_delete_user failed' with error_code=InternalErrorException but…
  • live_driver.py --mode target: 10 scenarios driving the real handlers end-to-end against a local moto ThreadedMotoServer with fault-injecting HTTP proxies (target_results.json/.md)
  • admin purge, transient Cognito AdminDeleteUser fault (InternalErrorException): RESOURCE_BUSY payload, one ERROR line with actor_sub + account_id + aws_error_code, accounts row + Cognito user verified surviving, no legacy helper line
  • admin purge, permanent Cognito fault (NotAuthorizedException): INTERNAL_ERROR with 'complete the deletion manually' and no retry promise, single attributed line, state verified
  • admin purge, DynamoDB sweep throttle (Query on shares table): RESOURCE_BUSY with 'Cognito user and the accounts record are untouched' — both verified untouched, AdminDeleteUser never attempted
  • admin purge, S3 AccessDenied during QR sweep: AppError re-raised unchanged ('Failed to purge payment QR codes from S3') with attributed sweep-failure line (error_code field), nothing deleted
  • admin purge, transient ListUsers lookup throttle: RESOURCE_BUSY 'Retry the purge', nothing deleted
  • admin purge success: returns True, accounts row and Cognito user actually gone, 'User account purged' audit line carries actor_sub + account_id
  • self-service delete, transient Cognito fault after sweep: RESOURCE_BUSY, 'Account data was deleted' verified true (accounts row gone), Cognito user survives, 3 real retry attempts observed at the proxy
  • self-service delete, permanent Cognito fault: INTERNAL_ERROR 'complete the deletion manually in Cognito', zero pointless retries, accounts row gone / Cognito user survives
  • self-service delete, sweep throttle: RESOURCE_BUSY, 'the Cognito user is untouched' verified, 4 throttled Query attempts observed
  • self-service delete, transient lookup throttle: RESOURCE_BUSY after exactly 3 project-wrapper attempts
  • live_driver.py --mode baseline (git archive b4134fd src/): reproduced the pre-fix defect — InternalErrorException surfaced as INTERNAL_ERROR 'Failed to delete user from Cognito' with an unattributed helper error line (baseline_results.json/.md)
  • targeted unit validation: .venv/bin/python -m pytest tests/unit/test_admin_operations.py tests/unit/test_account_operations.py -q (287 passed)
✅ **Document** - passed

✅ No issues found.

🔧 **Lint** - 2 issues found → auto-fixed ✅
  • 🚨 src/handlers/account_operations.py:84 - The CI complexity gate now fails on this change: uv run xenon --max-average A --max-absolute B --exclude &#39;*payment_methods.py,*logging.py&#39; src/ exits 1 with block &#34;src/handlers/account_operations.py:84 delete_my_account&#34; has a rank of C. delete_my_account grew from cyclomatic complexity 6 (rank B) at base commit b4134fd to 14 (rank C, >10) after the change's three new try/except classification blocks. AGENTS.md documents that this xenon command is the enforced CI gate (Grade A average, no block above Grade B), so the backend CI job will fail. Unresolved here because the fix requires restructuring functional code on the destructive deletion path (extracting the lookup/sweep/Cognito-delete classification blocks into helpers), which cannot be safely verified without running the test suite, and tests must not be changed in this phase.
  • 🚨 src/handlers/admin_operations.py:1027 - Same CI complexity gate breach as L1: block &#34;src/handlers/admin_operations.py:1027 admin_purge_user_account&#34; has a rank of C. The function grew from complexity 8 (rank B) at base to 11 (rank C, just over the B ceiling of 10) after the change added the sweep try/except AppError/except ClientError classification and the _delete_cognito_user_after_sweep call site. Same constraint as L1: a behavior-preserving extraction of the sweep-failure classification into a helper would clear it, but that is a functional refactor needing test verification, which this documentation/lint phase may not run or may not apply to tests.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Push** - passed

✅ No issues found.

…on both paths

A failure after the data sweep leaves a half-deleted account; every log
line and error on that path must say who acted, which account, what
survives, and whether a retry can help.

admin_purge_user_account:
- The partial-failure error log now carries actor_sub like every other
  line in the handler.
- The Cognito helper no longer logs the failure itself (its line and the
  handler's line were double noise, neither with the actor); it
  propagates the raw ClientError so the single log line where the
  context is added can extract the real AWS error code (aws_error_code)
  instead of the constant error_code field that could only ever be
  INTERNAL_ERROR.
- A throttled AdminDeleteUser is classified with the canonical transient
  classifier and raised as retryable RESOURCE_BUSY (#291), paired with a
  message that promises the retry; other failures stay INTERNAL_ERROR
  with a manual-completion message.
- The message names what survives on the admin path: the accounts
  record (deleted after the Cognito step) and possibly the Cognito user,
  so an operator cannot conclude the record is gone and skip the retry.

delete_my_account:
- The Cognito delete after the sweep gets the same loud partial-failure
  handling: one error line with the account id, transient -> RESOURCE_BUSY
  with a retry message, otherwise INTERNAL_ERROR with manual completion.
- A data-sweep failure is logged with the account id (the removed outer
  except was the only attributed sweep log) and leaves the Cognito user
  untouched.

Tests assert actor attribution, single logging, real AWS codes, message/
code agreement, and per-path surviving resources.
@dmeiser
dmeiser deployed to ephemeral October 3, 2026 18:44 — with GitHub Actions Active
…t comments at it (#669)

* docs(scripts): correct run-id provenance list and point script comments at the single contract (#595)

* no-mistakes(document): Run-id contract docs verified; shellcheck warning fixed

* no-mistakes(ci): Ephemeral tests for PR failed on one integration test: resolvers/campaignQueries.integration.test.ts 'should return zero for totalOrders and totalRevenue when campaign has no orders' threw TypeError on data.getCampaign.totalOrders because getCampaign was read immediately after createCampaign. getCampaign's pipeline resolves the campaign via the eventually-consistent campaignId-index GSI (query_campaign_fn.js), so a straight read can return null for an existing campaign (Bug #21); the test was the only positive read-after-create site in the file with no consistency poll, unlike its siblings. Fix: that test and its four sibling sites in the same file (campaignId/all-fields, shared-user-READ, no-endDate, both-dates) now poll getCampaign through the file's existing waitForGSIConsistency(query, len, 10, 1000) helper before asserting, exactly as the four call sites that already did so. Assertions were rewritten against campaigns[0] (visibleCampaigns in the no-endDate test, which already declares campaigns for its polled create) with identical field/value expectations; cleanups and timeouts unchanged. Negative tests (expect null) and the totalOrders=2 test (guarded by two successful createOrder calls that read the same GSI) were correctly left alone. The PR's own diff (docs/scripts/README.md + shell comments) is unrelated to the failure. Verified the edited file parses/links (module load reaches only ERR_MODULE_NOT_FOUND for uninstalled deps; the method was proven to catch duplicate declarations, which it caught before the shadowing rename); the integration suite cannot run locally (no .env/AWS credentials; it provisions cloud resources), so CI re-run is the definitive check
…ree-wide (#670)

* refactor(appsync): route ID-prefix normalization through lib/ids.js tree-wide (#534)

Every hand-rolled value.startsWith('PREFIX#') ternary and unconditional
'PREFIX#' + id concatenation in the js-resolvers now routes through the
shared lib/ids.js helpers (normalizeId / normalizeIdOrPrefix, plus a new
stripIdPrefix inverse for the API-boundary prefix strips). The three
verify_profile_owner_for_* resolvers - which had drifted into two
different owner-key implementations, one of them double-prefixing an
already-prefixed sub - now share lib/owner_key.js (expectedOwnerKey +
ownerGetItemRequest) and differ only in their FORBIDDEN message, per the

Behavior-preserving for every schema-valid input. Two deliberate
tolerances are kept rather than normalized away: update_campaign_fn still
passes catalogId null through as null (a deliberate clear-catalog
update), and check_existing_share_fn still passes a missing profileId /
targetAccountId through untouched. The share owner-verifier no longer
double-prefixes a prefixed sub - the divergence the issue flags as the
rot hazard - and identity-sub key building is now idempotent everywhere
(a Cognito sub is a bare UUID today; that assumption is no longer baked
into 20+ resolvers).

tests/unit/check_id_prefix_normalization.test.ts now sweeps every
non-test resolver for all four table prefixes (ACCOUNT/PROFILE/CATALOG/
CAMPAIGN), with a documented autoId() minting exemption; a re-inlined
ternary fails the guard. lib/verify_profile_owner_cases.js asserts the
consolidated owner-key behavior once for all three families, and
id_prefix_contract.test.js pins emitted DynamoDB requests for prefixed,
unprefixed, foreign-prefixed, and null inputs across the converted
patterns.

* no-mistakes(review): Add missing normalizeIdOrPrefix import to listMySharedCampaigns resolver

* test(appsync): pin verify_profile_read_access GetItem key bytes after the #508/#534 merge

The rebase onto main merged #508's consistent-GetItem rewrite of
verify_profile_read_access_fn.js with #534's routing of its two ID-prefix
constructions through lib/ids.js. The composite GetItem key IS the
authorization signal, so pin it: a bare Cognito sub must still produce
exactly ACCOUNT#<sub> and PROFILE#<id>, byte-identical to the pre-#534
hand-rolled ternaries.

* test(appsync): align update_campaign catalogId null contract with #659 in the #534 sweep

The #534 id_prefix_contract test pinned the pre-#659 tolerance that an
explicit null catalogId passes through as a NULL attribute. The base this
branch rebases onto (#659) made Campaign.catalogId ID! non-nullable and
rejects an explicit null with INVALID_INPUT, so the rebased contract test
now pins the rejection instead. Also correct the #534 comment in
normalizeCatalogId that claimed null must survive as null.
…t profileIds input (#672)

* fix(frontend): regenerate GraphQL types to match adminPurgeUserAccount profileIds list input

The committed generated types were stale from the client-side account-purge
work: the schema accepts a single ID or a list for profileIds, but
GqlAdminPurgeUserAccountMutationVariables still declared only the array form.
Regenerated with the lockfile-pinned toolchain (@graphql-codegen/cli 6.3.1,
@graphql-codegen/typescript 5.0.10) via npm run codegen.

* ci(frontend): fail when the generated GraphQL types drift from the schema

frontend/src/types/graphql-generated.ts is committed generated output that no
build or test step regenerates, so it can go stale against
tofu/application/schema/schema.graphql without anything noticing: the
adminPurgeUserAccount profileIds argument landed in #641 and the generated file
stayed behind until it was regenerated later.

Add a frontend CI step that runs the repo's lockfile-pinned codegen
(npm run codegen) and fails on any diff in the generated file, with an error
annotation naming the fix command.

Verified locally: the step passes on the current tree, and fails (exit 1) when
the generated file is hand-edited and committed - the regenerated union line
shows up in the diff.

* no-mistakes(document): Documented the codegen sync CI gate in owner docs
…ls at Cognito (#673)

* fix(handlers): make Cognito the commit point in user deletion (#551)

Delete the Cognito user before running the shared data cascade in both
admin_delete_user and delete_my_account. A Cognito failure now leaves every
record intact, and a data-phase failure afterwards leaves only inert records
the user can no longer reach (no sign-in, so no post_authentication
re-bootstrap); a re-run of the data phase is safe in both directions.

A data-phase failure after the Cognito commit is logged loudly with an
error_code-tagged field so a partial delete is detectable in post-incident
review rather than inferred from a client error.

* no-mistakes(review): Fix delete_my_account data-phase error attribution and update AGENTS.md #521 purge order

* no-mistakes(review): Revert user deletion to data-sweep-first order with loud partial-failure handling

* no-mistakes(document): Verified deletion docs current; reformatted one log string
Both KW-673 branches were cut from b4134fd and reworked the same deletion
paths, so they collided with the merged #673 (6ae5f4c). Resolved in favour
of this branch's implementation, which supersedes #673's classification:

- account_operations / admin_operations: the sweep -> Cognito delete sequence
  now goes through the shared `run_deletion_steps`, so both paths classify a
  failure once and identically. #673's `except Exception` around the Cognito
  delete is replaced by that helper.
- #521 sweep-first ordering is preserved and verified: `run_deletion_steps`
  runs the sweep before the Cognito delete, and the admin purge still deletes
  the accounts record last, after `run_deletion_steps`.
- utils/cognito.py single-homes the Cognito transient set as the public
  `COGNITO_TRANSIENT_ERROR_CODES` and adds `is_transient_cognito_error`, so
  the retry wrapper and the exhausted-retry classification cannot diverge.

Three #673-era tests asserted the messages this branch replaces. Updated to
the new contract (a permanent fault no longer promises a retry; the sweep
failure names what survived). One of them, the partial-state post-state
assertion, was orphaned onto the wrong test by the merge; moved back to
`test_delete_account_cognito_admin_delete_error`, which owns the seed.

Gates: `pytest tests/unit` 1563 passed / 2 skipped at 100% coverage,
xenon (max-average A, max-absolute B), ruff, and mypy all clean.
@dmeiser
dmeiser deployed to ephemeral October 4, 2026 02:59 — with GitHub Actions Active
@dmeiser
dmeiser merged commit c62b028 into main Oct 4, 2026
11 checks passed
@dmeiser
dmeiser deleted the fm/KW-673-AUDITABILITY branch October 4, 2026 03:55
@dmeiser
dmeiser deployed to ephemeral October 4, 2026 03:55 — with GitHub Actions Active
dmeiser added a commit that referenced this pull request Oct 4, 2026
…he deletion lookups (#678)

* fix(backend): harden get_caller_id and classify transport faults on the deletion lookups

Carries forward the two non-duplicate pieces of the closed #671 and #676,
neither of which made it into the merged work.

get_caller_id (from #671, whose rest is a duplicate of merged #662):
`event.get("identity") or {}` then `.get("sub")` assumed identity is a
mapping. A present-but-non-mapping identity (a string, a number, a list)
is truthy, so the read raised AttributeError instead of returning None,
and the caller surfaced as the decorator's generic INTERNAL_ERROR rather
than the typed UNAUTHORIZED the handlers raise for an absent caller. Guard
with isinstance against Mapping.

Cognito pre-check lookups (from #676, whose rest is merged #673 and #677):
both `_lookup_cognito_user_for_deletion` and `_find_cognito_user_by_sub`
caught ClientError only, so a BotoCoreError during the lookup escaped to
the decorator with no attributed log line and no retry guidance -- even
though the lookup runs before the sweep and the Cognito delete, so nothing
has been mutated and a retry converges. Widen both to
`(ClientError, BotoCoreError)`.

To classify that without adding a second definition of "transient",
`is_transient_cognito_error` now takes a BaseException and counts a bare
BotoCoreError as retryable; the pre-delete catches just call it. The
post-sweep Cognito delete in `run_deletion_steps` keeps its narrower
`isinstance(error, ClientError)` guard, and the asymmetry is now documented
on both sides: a transport fault *after* the delete request was sent leaves
the Cognito state genuinely unknown, so it must not promise a retry, while
the same fault *before* anything is mutated is plainly retryable. This
keeps #677's tested unknown-state contract intact.

Gates: `pytest tests/unit` 1575 passed / 2 skipped at 100% coverage, xenon,
ruff, and mypy all clean.

* chore(spelling): add 'isinstance' to the project cspell dictionary

The AGENTS.md #291 entry now names the isinstance narrowing that distinguishes
the post-sweep Cognito delete from the pre-delete lookups; the word was not in
the dictionary and failed the frontend job's Spell check step.

This branch was successfully deployed

1 active deployment
ephemeral — f20c2410 Deployed Oct 4, 2026 by dmeiser via Tear down ephemeral environment for PR #272
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant