Skip to content

Ransomware guard, heartbeat, Prometheus metrics, one-click repair, skip next run, disk power - #28

Merged
firsttris merged 2 commits into
masterfrom
claude/daemon-parity-features
Oct 9, 2026
Merged

firsttris merged 2 commits into
masterfrom
claude/daemon-parity-features

Conversation

@firsttris

@firsttris firsttris commented Oct 9, 2026 •

Copy link
Copy Markdown
Owner

This PR adds the six features snapraid-daemon 2.0 has and SnapRAID UI lacked.

Features

  1. Sync guard for changed files (ransomware)
    • A scheduled sync is skipped when diff reports more updated files than Max. changed files. The default is 100 for new schedules, the same as the daemon's sync_threshold_updates.
    • Why: ransomware encrypts files in place, and the next sync would overwrite their parity, which is the copy you could restore them from.
    • The schedule row shows the limit. The skip reason and the notification name the risk.
    • Existing schedules have no limit until you edit them.
  2. Heartbeat
    • After every successful scheduled run, the UI sends a GET to a healthchecks.io, Uptime Kuma or Dead Man's Snitch URL.
    • If the pings stop (server down, runs failing or skipped), the monitoring service raises the alarm.
    • It sits as a card under Notifications, with a Send ping button.
  3. Prometheus metrics at /api/metrics
    • Switched on under Automation; a token is generated, and the card shows a ready scrape_configs entry.
    • Covers array health, last sync and scrub, bad blocks, disk usage, SMART temperatures and sectors, and schedules.
    • Values come only from cached data (logs, last status read, histories), so a scrape never wakes a disk.
    • Not behind the login; it checks its own bearer token instead.
  4. Repair and verify
    • One button runs fix -e and then scrub -p bad (POST /api/snapraid/heal).
    • The scrub only runs after a successful fix and only when no other job started in between.
    • Verify only stays available as a second button.
  5. Skip next run
    • A button per schedule skips the next timed run once, and a second click takes it back.
    • The skipped run is recorded with "Skipped once, as asked." and sends no notification. Run now ignores the flag.
  6. Disk power and temperature
    • A ⋯ menu per disk and Spin up/down in the header for the whole array (snapraid up/down, POST /api/snapraid/power).
    • SnapRAID's error message is shown when a disk can't be controlled.
    • The status column shows the temperature from SMART with a sparkline of the last 30 reads.

Other changes

Tests

  • Backend: 135 passed. New tests:
    • update guard and skip-next in the scheduler
    • heartbeat logic, and the ping against a local server
    • metrics rendering
    • the heal sequence
    • disk power arguments and errors
  • Frontend: Biome, typecheck, Vitest 27 (new: temperature trend).
  • E2E: 15/15, locally and against the Docker image. New E2E cases:
    • one-click repair of bit rot
    • ransomware-like mass changes stopping the sync
    • skip next and taking it back
    • heartbeat on Send ping and after a scheduled run
    • metrics refused without the token and correct with it
    • spin up/down plumbing
  • Docs: mkdocs build --strict passes. Updated pages:
    • scheduling.md, notifications.md (Heartbeat), automation.md (Prometheus metrics with metric list and alert examples)
    • usage.md, architecture.md (API), development.md
    • README feature list

🤖 Generated with Claude Code

https://claude.ai/code/session_01JuXnquCPdbG5sBbDyin7yF

claude added 2 commits October 9, 2026 10:17
…ip next run, disk power

What snapraid-daemon offers and SnapRAID UI lacked:

- Sync guard for changed files: a scheduled sync is skipped when diff
  reports more updated files than allowed (default 100 for new
  schedules), as ransomware leaves them; the notification says so
- Heartbeat: a GET to a healthchecks.io or Uptime Kuma URL after every
  successful scheduled run, so a dead server or failing runs alert
- Prometheus metrics at /api/metrics with a bearer token, from cached
  values only, so a scrape never wakes a disk
- Repair and verify: fix -e, then scrub -p bad, with one click
- Skip next run of a schedule, once, without a notification
- Spin disks up or down by hand, and the temperature with a 30-day
  trend in the disk table

With unit and end-to-end tests and documentation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JuXnquCPdbG5sBbDyin7yF
@firsttris
firsttris merged commit 9dda952 into master Oct 9, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants