Skip to content

CAS: retry the mount lease renewal until the store answers - #2474

Draft
filimonov wants to merge 31 commits into
antalya-26.6from
fix/antalya-26.6/cas-mount-renewal-no-deadline
Draft

filimonov wants to merge 31 commits into
antalya-26.6from
fix/antalya-26.6/cas-mount-renewal-no-deadline

Conversation

@filimonov

Copy link
Copy Markdown
Member

The background renewal of the CAS mount lease gave up 4 s after it was due. The give-up fenced the mount and forced a remount under a fresh epoch (36.5 s or more), so an 8-second burst of 503 cost a full remount cycle. The renewal now retries until the store gives a definitive answer; the holder's own timer alone decides when writes stop, and they resume under the same writer_epoch when a renewal succeeds and leaves enough lease.

Changes

  • The background renewal retries one immutable request about once a second until the store answers, a stop or a remount request. It runs on its own request plane, so a latched fence cannot refuse the renewal that restores the lease.
  • An expired lease is visible: lifecycle_reason = 'lease_expired' in system.cas_mounts, the event CASMountLeaseExpired, expired_ms in the watermark_renew event, one WARNING per expiry.
  • GC dates a sighting of a mount token after the read that returned it. Before, a slow round could fence a holder that still had authority.
  • A memory-limit exception no longer ends the lease thread; a conditional PUT runs on the calling thread with a buffer of the body size.

Risks / notes

  • Write admission is unchanged: a conditional write still needs 2 × envelope + margin (16 s with defaults) of remaining lease, so writes are refused from about 4 s into a renewal outage, with a transient NETWORK_ERROR, until a renewal lands.
  • A GC leader on another pool member still fences a silent server after its observation threshold, and a definitive answer of the store still ends the mount.
  • One control-plane attempt is not bounded by a wall-clock deadline (per-socket-operation timeouts only). Not changed here.

Testing

  • Unit: CAS* gate, 2605 tests in release, 2609 under ASan with no error.
  • Integration: test_cas_mount_renewal_retry (test_short_put_outage_does_not_fence, test_long_put_outage_resumes_under_the_same_epoch), both red on the base.

Closes: #2431
Closes: #2403

Changelog category (leave one):

  • Bug Fix (user-visible misbehavior in an official stable release)

Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):

Fix a CAS mount lease being fenced, with a remount under a fresh epoch and a long wait, after a short object-store outage on the lease renewal. The renewal now retries until the store answers. Writes are refused while the remaining lease cannot cover them or has expired, and resume under the same writer_epoch once a renewal succeeds and leaves enough lease. A GC leader on another member or a definitive store answer still ends the mount. Add the CASMountLeaseExpired event and the lease_expired lifecycle reason in system.cas_mounts. Fix GC counting its wait for a mount token from before it saw the token, and the memory tracker failing the lease thread.

Documentation entry for user-facing changes

  • Documentation is written (mandatory for new features)

Updated docs/en/antalya/cas/architecture/mounts-and-leases.md, docs/en/antalya/cas/configuration.md, docs/en/antalya/cas/operations/troubleshooting.md, docs/en/antalya/cas/operations/monitoring.md, docs/en/antalya/cas/operations/debugging.md and docs/en/operations/system-tables/cas_mounts.md.

CI/CD Options

Exclude tests:

  • Fast test
  • Integration Tests
  • Stateless tests
  • Stateful tests
  • Unit tests
  • Performance tests
  • Aarch64 tests
  • All with ASAN
  • All with TSAN
  • All with MSAN
  • All with UBSAN
  • All with Coverage
  • All Regression
  • Disable CI Cache

Regression jobs to run:

  • Fast suites (mostly <1h)
  • Aggregate Functions (2h)
  • Alter (1.5h)
  • Benchmark (30m)
  • CAS (content-addressed storage; Antalya only)
  • ClickHouse Keeper (1h)
  • Iceberg (2h)
  • LDAP (1h)
  • OAuth (5m)
  • Parquet (1.5h)
  • RBAC (1.5h)
  • SSL Server (1h)
  • S3 (2h)
  • S3 Export (2h)
  • Swarms (30m)
  • Tiered Storage (2h)

filimonov and others added 30 commits October 2, 2026 01:27
`Retry::untilDefinitive` has no window and no lease bound; `bind` saturates, so it ends only on an answer, the fence or liveness. `attempt_spacing_ms` carries the spacing the engine will apply to its reissues.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
`claimMountAwaitingExpiry` and `computeHeartbeatFloor` each kept their own token and first-seen pair. Both now use `TokenWatch`, whose stability test is false for a sample older than the sighting instead of wrapping. `computeHeartbeatFloor` takes the observer clock as a function; it still samples it once per call.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The tracker still counts the renewal thread's allocations but no longer throws on it. An allocation failure outside the request used to end the thread; one inside it already was, and stays, a retried transient failure.

Related: #2403

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A spaced policy reissues a `PUT` (ordinary, connect-hint and credential-refresh), a conflict and a read no sooner than a spacing draw after the failed request started; a request that took longer is reissued at once. The first-attempt fuse and the read after an unclear `PUT` stay immediate. Policies without spacing keep their pauses and their clock reads.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
`computeHeartbeatFloor` dated a first sighting with the clock sample taken before the `LIST`, so a slow walk counted as time the token was watched. A second round could then fence a slot whose holder still had authority. The sighting is now sampled after the read, in the first decision and in the re-decision after a refused fence-out; the stability test keeps the round-start sample.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A conditional PUT ran on the remote-FS writer pool, outside the memory guard of the mount-lease thread, with a 1 MiB buffer for an object of a few hundred bytes. It now uploads inline and its buffer starts at the body size, one byte for an empty body.

Related: #2403

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A caller that renews the mount lease needs to see requests and failures while a write is still retrying. `CasOperation::setRequestObserver` is told of every physical `PUT` when sent, of every `PUT` that throws, and of every failed resolve read. A refused precondition, a successful read and the caller's own reads are not reported, and an observer that throws is ignored.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`Pool` owns `lease_requests` (open fence, the mount plane's interruptible sleep) and hands it through `CasMountRuntime` to `MountLeaseRenewer`, which routes an `UntilDefinitive` renewal there. Adds `MountRenewPolicy`, `kMountRenewRetrySpacingMs`, `MountRenewRequestEvent` and the `policy`/`on_request` environment fields. Nothing sets `UntilDefinitive` yet, so no renewal changes plane or policy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tive

`MountLeaseRenewer::renew` with `MountRenewPolicy::UntilDefinitive` retries under `Retry::untilDefinitive(kMountRenewRetrySpacingMs)` on the worker plane, with no lease bound, and reports each `PUT` and each failure through `MountRenewOperationEnvironment::on_request`. Classification is unchanged. `LeaseBound` keeps `Retry::untilLeaseSafe`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The background worker renews with `MountRenewPolicy::UntilDefinitive`: it retries one second apart on the lease plane until the store answers, and a stop, a park or a terminal lifecycle still ends it. Startup, remount and direct renewals stay `LeaseBound`. Three worker tests that ended a renewal by moving the clock to the deadline now end it with a fenced slot, which is still terminal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The lease is expired while the pool is Live, the fence is not lost and the confirmed deadline has passed. Add the accessor and the CASMountLeaseExpired counter that the restoring renewal will bump.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The worker's renewal counts its requests as they are sent and keeps the last failure text. A renewal that restores an expired lease bumps CASMountLeaseExpired, logs a warning with the duration and the last failure, and puts expired_ms into its watermark_renew event. A renewal that commits with a start already more than a TTL ago keeps the expiry and when it began.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
While the lease is expired and the pool is Live, the row reads lifecycle = 'not_live', lifecycle_reason = 'lease_expired', lifecycle_since = the expired deadline as wall time and lifecycle_detail = the last failed renewal request. The PoolLifecycle enum and the state column are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
checkFenceOrThrow asks the mount runtime whether the lease expired with the fence still held and the generation unchanged. Other refusals keep their text; admission is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The attempt counters now advance as requests are sent, so replaying one line per request when the renewal ends adds nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The text describes the outage that expired the lease; once a renewal ends it, a later expiry must not show it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Any committed renewal whose deadline is still in the future clears the stored failure text, so an earlier retried failure never explains a later expiry. A renewal that commits with its deadline already past keeps the text: the lease is still expired. The restore edge is one case of this rule and shares its clearing site.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A 6 s outage that covers a renewal start must not fence; an outage longer than the TTL expires the lease, refuses writes and resumes under the same writer_epoch with no remount.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The background renewal retries until the store answers; an expired lease refuses writes and is reported in system.cas_mounts, CASMountLeaseExpired and expired_ms.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
- Store the carried expiry instant before the new deadline when a renewal commits but leaves the lease expired, so a concurrent snapshot never shows a later `lifecycle_since`.
- Derive `worker_call` from the driver state in `renewRenewerOnce` instead of passing the same fact twice.
- AND `renewal_live_for_test` with the ordinary liveness predicate, so a stop still ends a renewal under the hook.
- Under a spaced policy, skip the backoff draw that the spacing replaces; the reissue counter still advances.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ored lease row

A failing test left the mount path at 503 and broke the node restart of the next test. The long outage test now also checks the detail text, the restored lifecycle row and that the refused row was not written.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…cing

State when writes stop, that GC or a definitive store answer still ends a mount, and bring the monitoring and debugging pages in line with the renewal counters and classifications.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…adline

Conditional writes reserve two attempt envelopes, a retried sentinel probe one. Also reflow the touched lines and replace history wording in the renewal log pages.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
- Spaced conflict pause of `readModifyWriteOnPresence` and spaced retries of a direct `op.read`.
- Count one restore per expiry across two expiries; refused writes and a fence-out count nothing.
- Measure the over-limit allocation on the calling thread's tracker.
- Hold the first retry wait of the stop test for 30 s, so only a stop can end it.
- Bound the definitive-answer and fence tests by the clock, so a regression fails instead of hanging.
- Declare hook-captured locals before the backend in the pool tests.
- Assert a positive line in the debug-level log capture.
- Drop the spec references from the spacing test comments.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
With a regression that zeroes every spaced pause the test clock stops, and the renewal loops forever. The tests now end the renewal after 500 requests and fail with a message instead of hanging.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…cleSnapshot

The sleep wakes only on a stop of the workers; a park or a remount request does not wake it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The background renewal retries past the lease; only the bounded renewals stop at it. A successful renewal resumes writes only if it leaves enough lease. The lease keeper logs nothing until a renewal ends.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A first-attempt success logs nothing; the recovery `INFO` needs a retry, a read-settled commit or a restored lease. The background renewal ends on any lifecycle other than `Live`, not only a terminal one.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
While the object store is down the lease keeper logged nothing until a renewal restored the lease. It now writes one WARNING at the first renewal request after the expiry. The expiry instant identifies the expiry, so retries, refused writes and a renewal that commits past its own deadline stay silent, and a second expiry warns again.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Workflow [PR], commit [e1c38ca]

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant