Skip to content

Release expired domain reservations daily - #2965

Open
OlegPhenomenon wants to merge 3 commits into
masterfrom
feature/2963-release-expired-reservations
Open

OlegPhenomenon wants to merge 3 commits into
masterfrom
feature/2963-release-expired-reservations

Conversation

@OlegPhenomenon

@OlegPhenomenon OlegPhenomenon commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Closes #2963

What

Expired domain reservations are now removed by a daily cron job instead of relying only on lazy removal during Business Registry availability checks.

  • ReservedDomain.release_expired removes every reservation with expire_at < now.
  • Permanent reservations (expire_at IS NULL, e.g. created manually in admin) are never removed.
  • Every removal leaves a full audit trail in log_reserved_domains (PaperTrail): process, reason and timestamp.
  • Index on reserved_domains.expire_at for the cleanup query.

How it works

For every expired reservation (find_each, batches of 1000):

  1. The row is locked (SELECT ... FOR UPDATE) and expiry is re-checked, so a reservation extended in the meantime is not removed.
  2. Inside PaperTrail.request(whodunnit: audit):
    • update!(updator_str: audit) — creates an update version, persists the reason;
    • destroy! — creates a destroy version whose object['updator_str'] and whodunnit contain the reason.
  3. After commit, UpdateWhoisRecordJob.perform_later(name, 'reserved') refreshes WHOIS (after_destroy → after_destroy_commit, so the job never sees the row before it is deleted).

Audit string format (built in one place, ReservedDomain#release_audit_message, so #2962 can switch it to a structured format later):

Automated daily cleanup job - Expired reservation deadline reached - 2026-10-02T00:35:00+03:00
Business registry availability check - Expired reservation deadline reached - 2026-10-02T14:12:03+03:00

The second form is written by the existing lazy removal (destroy_if_expired in BusinessRegistry::DomainAvailabilityCheckerService), which now goes through the same code path.

Error handling: a failing record is logged (stdout → log/cron.log, and Rails.logger.error → application log) with its id, name and exception, its transaction is rolled back and the run continues with the next record. A record already deleted by a concurrent process is skipped silently.

Deployment / setup

1. Migration

bundle exec rails db:migrate

20261002120000_add_index_to_reserved_domains_on_expire_at creates the index with algorithm: :concurrently (no table lock, disable_ddl_transaction!). Rollback: bundle exec rails db:rollback (drops the index).

2. Crontab

The job is defined in config/schedule.rb (whenever) inside the @cron_group == 'registry' block:

every :day, at: '12:35am' do
  runner 'ReservedDomain.release_expired'
end

The crontab is not updated automatically by the code change — it must be regenerated on the registry (admin) server after deploy:

# via mina (see doc/application_build_doc.md#cron)
mina pr cron:setup      # production, cron_group=registry
mina cron:setup         # alpha/staging

# or directly on the server, from the current release dir
bundle exec whenever --update-crontab registry_production \
  --set 'environment=production&path=/path/to/registry/current&cron_group=registry'

Only servers deployed with cron_group=registry get this job; epp, registrar, registrant groups do not.

Preview what will be written to crontab without installing it:

bundle exec whenever --set 'environment=production&cron_group=registry' | grep release_expired

Expected line:

35 0 * * * /bin/bash -l -c '...; cd <app path> && bin/rails r -e production "ReservedDomain.release_expired" >> log/cron.log 2>&1'

3. Timezone

whenever uses the server system time; expire_at is calculated in Time.zone (Tallinn). With #2961 / #2964 reservations expire at 23:59:59 Tallinn time, so on a server running in Tallinn time the job removes them ~35 minutes after expiry. If the server runs in UTC, the job runs at 00:35 UTC (03:35 Tallinn) — still correct, only the window is longer. Change the time in config/schedule.rb if needed.

Manual run / troubleshooting

# how many reservations are expired right now (dry run, nothing is deleted)
bundle exec rails r 'puts ReservedDomain.expired.count'

# run the cleanup manually (same as cron)
bundle exec rails r 'ReservedDomain.release_expired'

# with a custom cutoff, e.g. everything expired before yesterday midnight
bundle exec rails r 'ReservedDomain.release_expired(at: Time.zone.yesterday.beginning_of_day)'

Local Docker environment:

docker compose exec registry bash -c "cd /opt/webapps/app && rails r 'ReservedDomain.release_expired'"

Output (stdout → log/cron.log when run by cron):

2026-10-02 00:35:01 UTC - Released 12 expired reserved domains (failed: 0)
2026-10-02 00:35:01 UTC - Failed to release reserved domain 42 (example.ee): ActiveRecord::RecordInvalid - ...

failed > 0 means some records stayed in the table; details are in log/cron.log and the Rails application log. They will be retried on the next run.

Audit trail queries

# reservations removed by the daily job
Version::ReservedDomainVersion
  .where(event: 'destroy')
  .where('whodunnit LIKE ?', 'Automated daily cleanup job%')
  .order(created_at: :desc)
  .limit(20)
  .map { |v| [v.object['name'], v.object['expire_at'], v.whodunnit] }

# full history of one domain (create version has no `object`, so match object_changes too)
ids = Version::ReservedDomainVersion
  .where("object->>'name' = :n OR object_changes->'name'->>1 = :n", n: 'example.ee')
  .distinct.pluck(:item_id)
Version::ReservedDomainVersion.where(item_id: ids).order(:id)
  .map { |v| [v.event, v.whodunnit, v.created_at] }

Tests

docker compose exec registry bash -c "cd /opt/webapps/app && rails test test/models/reserved_domain_test.rb"

Covered in test/models/reserved_domain_test.rb:

  • expired rows removed; future and expire_at = NULL rows kept; returned count;
  • boundary: expire_at == cutoff is not removed;
  • update + destroy versions with reason/process/timestamp in object['updator_str'] and whodunnit;
  • reservation extended after loading is not removed (row lock + re-check);
  • one failing record does not stop the run;
  • WHOIS refresh job enqueued for released domains;
  • lazy removal marks the process as Business registry availability check;
  • no expired rows → returns 0.

Also ran: test/services/business_registry/, test/integration/api/business_registry/, test/models/free_domain_reservation_holder_test.rb, test/models/reserve_domain_invoice_test.rb, test/models/domain_test.rb — all green.

Related

Files

  • app/models/reserved_domain.rb — expired scope, release_expired, release_if_expired, after_destroy_commit
  • config/schedule.rb — daily job at 00:35
  • db/migrate/20261002120000_add_index_to_reserved_domains_on_expire_at.rb, db/structure.sql — index
  • test/models/reserved_domain_test.rb — tests
  • CHANGELOG.md

Add ReservedDomain.release_expired, run by cron every day at 00:35, which
removes reservations whose expire_at has passed. Permanent reservations
(expire_at NULL) are kept.

Each record is locked and re-checked before removal. The reason is stored
in updator_str and whodunnit before destroy, so the PaperTrail destroy
version in log_reserved_domains carries the process, reason and timestamp.
Lazy removal on availability check reuses the same path.

A failing record is logged and reported to Airbrake without stopping the
run. WHOIS refresh is enqueued after commit. Add index on expire_at.

Closes #2963
The expired reservation was created one day before the real current
time but released at a fixed past moment, so the test failed once the
real date moved past it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add daily cron job to remove expired reservations

4 participants