Skip to content

Send an alert webhook when a connector keeps failing - #976

Merged
helpcodeai merged 2 commits into
HelpCode-ai:mainfrom
rangarajan19:feat/857-connector-failure-webhook
Oct 10, 2026
Merged

helpcodeai merged 2 commits into
HelpCode-ai:mainfrom
rangarajan19:feat/857-connector-failure-webhook

Conversation

@rangarajan19

Copy link
Copy Markdown
Contributor

Summary

Backend half of #857 (PR 1 of the planned split, UI card to follow in PR 2): detects a connector failing repeatedly within a sliding window and dispatches a signed webhook (or Slack message), with a cooldown so a sustained outage sends one alert, not one per call.

Changes

  • New AlertsModule: per-organization webhook config stored via OrgSettingsService, signing secret generated server-side and encrypted (AES-256-GCM, AAD-bound to the org), shown once. Admin endpoints: GET/PUT/DELETE + POST .../test at /api/admin/settings/alert-webhook.
  • Detection hooks into AuditService.logInvocation, firing before the repeat-failure dedup check added since this plan was discussed, so a fast-repeating identical failure (the common case for a real outage) still counts toward the threshold instead of being silently skipped. Fire-and-forget: a dispatch failure is caught and logged, never affects the MCP call path.
  • Sliding-window failure counter and cooldown via RedisService (incr/expire), with an in-memory fallback when Redis isn't configured. Dispatch fires only when the counter hits the threshold exactly, and the cooldown key is written before the dispatch goes out, so two failures crossing the threshold at the same moment can't both fire.
  • Dispatch is HMAC-SHA256 signed (X-AnythingMCP-Signature) and goes through the existing SSRF-guarded outbound path. The webhook URL deliberately ignores both the env (SSRF_ALLOWED_HOSTS) and the admin-editable DB allowlist — ssrf.util.ts/outbound-http.ts gain an opt-in skipAllowlists option for this, since those allowlists exist for connectors to reach a specific internal host for their own reason, and a webhook target must not inherit that trust.
  • Docs: new section in docs/operations/observability.md — payload shape, headers, a Node verification snippet.

PR 2 (the Alerts card on /settings/admin) will follow once this is reviewed.

Testing

  • Added/updated unit tests
  • Tested manually (describe steps below)
  • All existing tests pass (npm test in packages/backend)

New: src/alerts/alerts.service.spec.ts, src/alerts/alerts.controller.spec.ts, and additions to src/audit/audit.service.spec.ts covering threshold→one dispatch, cooldown suppressing the next, cross-org isolation, signature verification, dispatch errors never affecting logInvocation, SSRF (private IP and the cloud metadata address rejected even when allowlisted), and non-admin → 403.

Ran from packages/backend:

npx tsc --noEmit -p tsconfig.json
npx eslint src/alerts src/audit/audit.module.ts src/audit/audit.service.ts src/common/ssrf.util.ts src/common/outbound-http.ts
npx jest src/alerts src/audit/audit.service.spec.ts src/audit/audit.breakdowns.spec.ts src/common/ssrf.util.spec.ts src/common/outbound-http.spec.ts src/common/outbound-fetch.util.spec.ts src/settings

All clean; 747 tests passing in the scoped run above (the files this PR touches and everything that depends on them).

Related Issues

Relates to #857

Detects a connector failing repeatedly (configurable threshold over a
sliding window) and dispatches a signed webhook or Slack message, with
a cooldown so a sustained outage sends one alert, not one per call.

- New AlertsModule: per-org webhook config (OrgSettings), encrypted
  secret shown once, GET/PUT/DELETE + POST .../test admin endpoints.
- Detection hooks into AuditService.logInvocation ahead of the
  repeat-failure dedup check, so a fast-repeating identical failure
  still counts toward the threshold. Fire-and-forget; never affects
  the MCP call path.
- Sliding-window counter and cooldown via Redis (incr/expire), with
  an in-memory fallback when Redis isn't configured.
- Dispatch is HMAC-signed and SSRF-guarded; the webhook URL ignores
  the env/DB connector allowlists so a workspace can't point it at a
  host allowed for an unrelated reason (ssrf.util.ts/outbound-http.ts
  gain an opt-in skipAllowlists flag for this).
- Docs: new section in docs/operations/observability.md.

Relates to HelpCode-ai#857.
Comment thread packages/backend/src/alerts/alerts.service.ts Dismissed
@rangarajan19

Copy link
Copy Markdown
Contributor Author

Hi @keysersoft, PR 1 for the connector-failure alert webhook is up here. While working on it I noticed #857 (and the whole #840-867 range, including the October Challenge issue #846) is no longer accessible — just a 404, not sure if that was intentional. Do you still want this feature added, or has priority changed? Happy to keep going as-is if so.

@rangarajan19

Copy link
Copy Markdown
Contributor Author

Hi @mirkopoloni , do we need this, because the issue is not showing , should I delete this or ?

@helpcodeai

Copy link
Copy Markdown
Contributor

@rangarajan19 yes, we still want it, please keep going. The issue is now #1006, and the October Challenge list is #997. I'll review this PR in the next days. The CodeQL warning on alerts.service.ts is a false positive: that HMAC signs the webhook payload with a random secret, it isn't a password hash.

- Cache the webhook config per organization for a minute, so failed calls of
  organizations without a webhook cost no database read.
- Heal a failure counter whose EXPIRE was lost instead of never alerting again.
- On Cloud, require https and skip the SSRF allowlists; self-hosted keeps the
  operator's allowlist, so an internal chat server can receive alerts.
- Escape the connector name and vendor error in Slack messages.
- Look the connector up inside the failing organization only.
- Log undelivered alerts (without the URL), validate the URL and bound the
  numbers, throttle the test endpoint.
- Docs: restore the 'Correlating across pipelines' heading, replay-safe
  verification snippet, describe the window as starting at the first failure.
@helpcodeai
helpcodeai self-requested a review as a code owner October 10, 2026 11:51
@helpcodeai
helpcodeai merged commit eb7c5c3 into HelpCode-ai:main Oct 10, 2026
10 of 11 checks passed
@helpcodeai

Copy link
Copy Markdown
Contributor

Merged, thanks @rangarajan19! I pushed a few review fixes on top so this could go in today:

  • the webhook config is cached per organization for a minute (failed calls of orgs without a webhook no longer cost a DB read);
  • a failure counter whose EXPIRE was lost now heals instead of never alerting again;
  • on Cloud the URL must be https and the SSRF allowlists are skipped, while self-hosted keeps the operator's allowlist (as 🎃 Send an alert webhook when a connector keeps failing #1006 asked, for internal Mattermost or Slack proxies);
  • Slack text escaping for the connector name and the vendor error;
  • the connector lookup is scoped to the organization;
  • undelivered alerts are logged, the DTO validates the URL and bounds the numbers, and the test endpoint is throttled;
  • docs: the 'Correlating across pipelines' heading is back, and the verify snippet is replay-safe.

The UI card (PR 2) is next whenever you're ready. Please keep the field names and defaults from this PR.

@github-actions github-actions Bot locked and limited conversation to collaborators Oct 10, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants