Conversation
`Retry::untilDefinitive` has no window and no lease bound; `bind` saturates, so it ends only on an answer, the fence or liveness. `attempt_spacing_ms` carries the spacing the engine will apply to its reissues. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
`claimMountAwaitingExpiry` and `computeHeartbeatFloor` each kept their own token and first-seen pair. Both now use `TokenWatch`, whose stability test is false for a sample older than the sighting instead of wrapping. `computeHeartbeatFloor` takes the observer clock as a function; it still samples it once per call. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The tracker still counts the renewal thread's allocations but no longer throws on it. An allocation failure outside the request used to end the thread; one inside it already was, and stays, a retried transient failure. Related: #2403 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A spaced policy reissues a `PUT` (ordinary, connect-hint and credential-refresh), a conflict and a read no sooner than a spacing draw after the failed request started; a request that took longer is reissued at once. The first-attempt fuse and the read after an unclear `PUT` stay immediate. Policies without spacing keep their pauses and their clock reads. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
`computeHeartbeatFloor` dated a first sighting with the clock sample taken before the `LIST`, so a slow walk counted as time the token was watched. A second round could then fence a slot whose holder still had authority. The sighting is now sampled after the read, in the first decision and in the re-decision after a refused fence-out; the stability test keeps the round-start sample. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A conditional PUT ran on the remote-FS writer pool, outside the memory guard of the mount-lease thread, with a 1 MiB buffer for an object of a few hundred bytes. It now uploads inline and its buffer starts at the body size, one byte for an empty body. Related: #2403 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A caller that renews the mount lease needs to see requests and failures while a write is still retrying. `CasOperation::setRequestObserver` is told of every physical `PUT` when sent, of every `PUT` that throws, and of every failed resolve read. A refused precondition, a successful read and the caller's own reads are not reported, and an observer that throws is ignored. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`Pool` owns `lease_requests` (open fence, the mount plane's interruptible sleep) and hands it through `CasMountRuntime` to `MountLeaseRenewer`, which routes an `UntilDefinitive` renewal there. Adds `MountRenewPolicy`, `kMountRenewRetrySpacingMs`, `MountRenewRequestEvent` and the `policy`/`on_request` environment fields. Nothing sets `UntilDefinitive` yet, so no renewal changes plane or policy. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tive `MountLeaseRenewer::renew` with `MountRenewPolicy::UntilDefinitive` retries under `Retry::untilDefinitive(kMountRenewRetrySpacingMs)` on the worker plane, with no lease bound, and reports each `PUT` and each failure through `MountRenewOperationEnvironment::on_request`. Classification is unchanged. `LeaseBound` keeps `Retry::untilLeaseSafe`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The background worker renews with `MountRenewPolicy::UntilDefinitive`: it retries one second apart on the lease plane until the store answers, and a stop, a park or a terminal lifecycle still ends it. Startup, remount and direct renewals stay `LeaseBound`. Three worker tests that ended a renewal by moving the clock to the deadline now end it with a fenced slot, which is still terminal. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The lease is expired while the pool is Live, the fence is not lost and the confirmed deadline has passed. Add the accessor and the CASMountLeaseExpired counter that the restoring renewal will bump. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The worker's renewal counts its requests as they are sent and keeps the last failure text. A renewal that restores an expired lease bumps CASMountLeaseExpired, logs a warning with the duration and the last failure, and puts expired_ms into its watermark_renew event. A renewal that commits with a start already more than a TTL ago keeps the expiry and when it began. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
While the lease is expired and the pool is Live, the row reads lifecycle = 'not_live', lifecycle_reason = 'lease_expired', lifecycle_since = the expired deadline as wall time and lifecycle_detail = the last failed renewal request. The PoolLifecycle enum and the state column are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
checkFenceOrThrow asks the mount runtime whether the lease expired with the fence still held and the generation unchanged. Other refusals keep their text; admission is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The attempt counters now advance as requests are sent, so replaying one line per request when the renewal ends adds nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The text describes the outage that expired the lease; once a renewal ends it, a later expiry must not show it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Any committed renewal whose deadline is still in the future clears the stored failure text, so an earlier retried failure never explains a later expiry. A renewal that commits with its deadline already past keeps the text: the lease is still expired. The restore edge is one case of this rule and shares its clearing site. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A 6 s outage that covers a renewal start must not fence; an outage longer than the TTL expires the lease, refuses writes and resumes under the same writer_epoch with no remount. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The background renewal retries until the store answers; an expired lease refuses writes and is reported in system.cas_mounts, CASMountLeaseExpired and expired_ms. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
- Store the carried expiry instant before the new deadline when a renewal commits but leaves the lease expired, so a concurrent snapshot never shows a later `lifecycle_since`. - Derive `worker_call` from the driver state in `renewRenewerOnce` instead of passing the same fact twice. - AND `renewal_live_for_test` with the ordinary liveness predicate, so a stop still ends a renewal under the hook. - Under a spaced policy, skip the backoff draw that the spacing replaces; the reissue counter still advances. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ored lease row A failing test left the mount path at 503 and broke the node restart of the next test. The long outage test now also checks the detail text, the restored lifecycle row and that the refused row was not written. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…cing State when writes stop, that GC or a definitive store answer still ends a mount, and bring the monitoring and debugging pages in line with the renewal counters and classifications. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…adline Conditional writes reserve two attempt envelopes, a retried sentinel probe one. Also reflow the touched lines and replace history wording in the renewal log pages. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
- Spaced conflict pause of `readModifyWriteOnPresence` and spaced retries of a direct `op.read`. - Count one restore per expiry across two expiries; refused writes and a fence-out count nothing. - Measure the over-limit allocation on the calling thread's tracker. - Hold the first retry wait of the stop test for 30 s, so only a stop can end it. - Bound the definitive-answer and fence tests by the clock, so a regression fails instead of hanging. - Declare hook-captured locals before the backend in the pool tests. - Assert a positive line in the debug-level log capture. - Drop the spec references from the spacing test comments. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
With a regression that zeroes every spaced pause the test clock stops, and the renewal loops forever. The tests now end the renewal after 500 requests and fail with a message instead of hanging. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…cleSnapshot The sleep wakes only on a stop of the workers; a park or a remount request does not wake it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The background renewal retries past the lease; only the bounded renewals stop at it. A successful renewal resumes writes only if it leaves enough lease. The lease keeper logs nothing until a renewal ends. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A first-attempt success logs nothing; the recovery `INFO` needs a retry, a read-settled commit or a restored lease. The background renewal ends on any lifecycle other than `Live`, not only a terminal one. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
While the object store is down the lease keeper logged nothing until a renewal restored the lease. It now writes one WARNING at the first renewal request after the expiry. The expiry instant identifies the expiry, so retries, refused writes and a renewal that commits past its own deadline stay silent, and a second expiry warns again. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
10 of 31 tasks
11 of 30 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The background renewal of the CAS mount lease gave up 4 s after it was due. The give-up fenced the mount and forced a remount under a fresh epoch (36.5 s or more), so an 8-second burst of
503cost a full remount cycle. The renewal now retries until the store gives a definitive answer; the holder's own timer alone decides when writes stop, and they resume under the samewriter_epochwhen a renewal succeeds and leaves enough lease.Changes
lifecycle_reason = 'lease_expired'insystem.cas_mounts, the eventCASMountLeaseExpired,expired_msin thewatermark_renewevent, oneWARNINGper expiry.PUTruns on the calling thread with a buffer of the body size.Risks / notes
2 × envelope + margin(16 s with defaults) of remaining lease, so writes are refused from about 4 s into a renewal outage, with a transientNETWORK_ERROR, until a renewal lands.Testing
CAS*gate, 2605 tests in release, 2609 under ASan with no error.test_cas_mount_renewal_retry(test_short_put_outage_does_not_fence,test_long_put_outage_resumes_under_the_same_epoch), both red on the base.Closes: #2431
Closes: #2403
Changelog category (leave one):
Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):
Fix a CAS mount lease being fenced, with a remount under a fresh epoch and a long wait, after a short object-store outage on the lease renewal. The renewal now retries until the store answers. Writes are refused while the remaining lease cannot cover them or has expired, and resume under the same
writer_epochonce a renewal succeeds and leaves enough lease. A GC leader on another member or a definitive store answer still ends the mount. Add theCASMountLeaseExpiredevent and thelease_expiredlifecycle reason insystem.cas_mounts. Fix GC counting its wait for a mount token from before it saw the token, and the memory tracker failing the lease thread.Documentation entry for user-facing changes
Updated
docs/en/antalya/cas/architecture/mounts-and-leases.md,docs/en/antalya/cas/configuration.md,docs/en/antalya/cas/operations/troubleshooting.md,docs/en/antalya/cas/operations/monitoring.md,docs/en/antalya/cas/operations/debugging.mdanddocs/en/operations/system-tables/cas_mounts.md.CI/CD Options
Exclude tests:
Regression jobs to run: