Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
31 commits
Select commit Hold shift + click to select a range
c0e112d
Add an unbounded retry policy with attempt spacing to `Retry`
filimonov Oct 1, 2026
703b9ab
Watch mount tokens through one TokenWatch helper
filimonov Oct 1, 2026
0ec8214
Keep memory-limit exceptions off the CAS mount-lease thread
filimonov Oct 1, 2026
26de2d5
Space engine reissues under `Retry::attempt_spacing_ms`
filimonov Oct 1, 2026
38c5ac5
Date a GC token sighting after the read that produced it
filimonov Oct 1, 2026
817426d
Run CAS conditional PUTs on the calling thread with a body-sized buffer
filimonov Oct 1, 2026
4804c47
Report each CAS write attempt to an optional observer
filimonov Oct 2, 2026
2b6ab36
Give the CAS mount renewer a lease plane with no fence
filimonov Oct 2, 2026
ff39bb1
Renew the CAS mount lease until a definitive answer under UntilDefini…
filimonov Oct 2, 2026
01ffe71
Run the worker's CAS mount renewal without a lease deadline
filimonov Oct 2, 2026
73ed84b
Derive the expired state of a CAS mount lease
filimonov Oct 2, 2026
08e8242
Report a CAS mount lease expiry and its restore
filimonov Oct 2, 2026
45381f7
Show an expired CAS mount lease in system.cas_mounts
filimonov Oct 2, 2026
eb5bb4e
Say that writes resume when a CAS lease expiry is the only refusal
filimonov Oct 2, 2026
aaaa57b
Drop the per-request DEBUG lines of a CAS mount renewal
filimonov Oct 2, 2026
c08fea9
Clear the last CAS renewal failure when a renewal restores the lease
filimonov Oct 2, 2026
98776bf
End the CAS renewal failure text with the run of trouble it explains
filimonov Oct 2, 2026
1f7d2e7
Test CAS mount renewal across a short and a long mount PUT outage
filimonov Oct 2, 2026
5b21ea3
Document CAS mount renewal without a lease deadline
filimonov Oct 2, 2026
66dea1f
Tidy the mount renewal and the spaced reissue pause
filimonov Oct 2, 2026
016de00
Disarm the S3 fault proxy after every renewal test and check the rest…
filimonov Oct 2, 2026
204390a
Make the CAS mount renewal docs precise about write admission and fen…
filimonov Oct 2, 2026
f0b2275
State which CAS writes reserve how many envelopes before the lease de…
filimonov Oct 2, 2026
13a9b92
Test the spaced pauses, the restore count and bounded waits
filimonov Oct 2, 2026
ee58abb
Correct comments on the renewal and its observer
filimonov Oct 2, 2026
0c2ba2a
Bound the unbounded-renewal tests by request count
filimonov Oct 2, 2026
8dc0af6
Correct the mount-plane sleep comments and drop spec tags from Lifecy…
filimonov Oct 2, 2026
7443ccd
Fix CAS mount renewal docs that contradict the code
filimonov Oct 2, 2026
7e24350
Say which CAS mount renewals log and what ends the background renewal
filimonov Oct 2, 2026
2712455
Log one WARNING when the CAS mount lease expires
filimonov Oct 2, 2026
e1c38ca
Document the CAS mount lease expiry WARNING
filimonov Oct 2, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
116 changes: 79 additions & 37 deletions docs/en/antalya/cas/architecture/mounts-and-leases.md
Original file line number Diff line number Diff line change
Expand Up @@ -83,12 +83,14 @@ watermark — there is no separate watermark object. `MountLease` fields: `serve
equals its immutable request. If the predecessor token is still current, another identical `PUT`
may follow bounded backoff. A same-pair twin, GC-fenced body, successor, foreign holder, or absent
body is never treated as this renewal.
- **Absolute deadline.** Renewal uses `CLOCK_BOOTTIME`, not `CLOCK_MONOTONIC`, so a VM resumed from
suspend correctly observes itself expired. Its absolute deadline is the minimum of the existing
request-operation budget and the last confirmed lease deadline minus the safety margin. The
controller checks that one attempt envelope still fits before each backend `PUT` or resolving
`GET`, after each interruptible backoff, and before accepting success. A retry, `GET`, response
timestamp, or wall-clock step never extends authority.
- **Lease deadline.** The lease is valid for `cas_mount_lease_ttl_ms` from the start of the last
confirmed renewal. It is measured on `CLOCK_BOOTTIME`, not `CLOCK_MONOTONIC`, so a VM resumed from
suspend correctly observes itself expired. The deadline decides about writes: none is admitted past
it, and none is admitted once the remaining lease cannot cover the requests it may send plus the
safety margin. The background renewal does not stop at the deadline. It retries timeouts, `5xx`
answers and connection errors about a second apart until the store answers. The renewals at startup,
after a remount and the direct renewal stay bounded: they stop at the last confirmed deadline minus
the safety margin. A retry, `GET`, response timestamp, or wall-clock step never extends authority.
- **Cadence.** The runtime normally starts a logical renewal every `cas_mount_renew_period_ms` (default
10 s), with TTL `cas_mount_lease_ttl_ms` (default 30 s, TTL/3 renewal ratio). The next beat is anchored
at the committed body's pre-I/O BOOTTIME start. A slow recovery therefore causes an immediate
Expand All @@ -101,43 +103,79 @@ watermark — there is no separate watermark object. `MountLease` fields: `serve
`2 × envelope + safety_margin` fits inside the remaining lease (a write and its settlement read),
rejecting with `BAD_ARGUMENTS` at request-admission time rather than mid-flight.

**Losing the lease is neither read-only mode nor a process abort.** `MountLeaseRenewer` is a
synchronous durable-slot state machine. A committed result advances its token, sequence, confirmed
BOOTTIME deadline, and cadence anchor. Any admitted deterministic failure, confirmed conflict, or
ambiguity left at the deadline/attempt limit moves it to `RenewalTerminal`; it cannot mint another
body or publish a clean farewell. Owner cancellation before any request is the only
`NotAttempted` result and leaves clean release possible. Cancellation after a request was sent is
terminal because that request may still land.
**A transient renewal failure does not end the mount.** `MountLeaseRenewer` is a synchronous
durable-slot state machine. A committed result advances its token, sequence, confirmed BOOTTIME
deadline, and cadence anchor. The background renewal ends, and moves the renewer to `RenewalTerminal`,
only on one of these:

- a definitive answer from the store: a confirmed foreign, successor or same-pair body, `gc_fenced`, an
absent object, or a request the store refuses on a clear attempt;
- a stop or a remount request;
- a deterministic local failure;
- a lifecycle other than `Live`;
- a lost fence.

A terminal renewer cannot mint another body or publish a clean farewell. Owner cancellation before
any request is the only `NotAttempted` result and leaves clean release possible. Cancellation after a
request was sent is terminal because that request may still land.

While retries continue past the lease deadline, the lease is *expired*. New durable writes are refused
with a transient `NETWORK_ERROR` that names the lease, reads are not gated, and the pool does not
remount. Writes are admitted again, under the same `writer_epoch`, when a renewal succeeds and leaves
enough lease for a write's reservation. A renewal that succeeds after its own deadline (its start plus
the TTL has already passed) leaves the lease expired, and the next renewal follows at once.
The log carries one `WARNING` when the lease expires, written at the first request of the renewal
after the expiry (a request that hangs delays it by up to one attempt timeout), and the restore
`WARNING` when a renewal restores it.
`system.cas_mounts` shows the expired period as `lifecycle = 'not_live'`,
`lifecycle_reason = 'lease_expired'` (see [`system.cas_mounts`](#mounts-table)); it shows
`lifecycle = 'live'` until the deadline passes, although with the defaults conditional writes stop
16 s earlier. The `CASMountLeaseExpired` event advances by one when a renewal restores an expired
lease, not when the lease expires. Failing renewals alone do not fence or remount the mount. A GC
leader on another member still fences a slot whose token has not changed for
`TTL + floor(TTL/20) + period`, and a definitive answer from the store that the slot holds something
else ends it; in both cases the server remounts under a new `writer_epoch`.

After the renewer call returns, `CasMountRuntime` consumes the result. A terminal result trips the
local fence (latches `lost`, bumps the fence generation, moves the in-process runtime to
`TransientNotLive`) and latches one self-remount generation. A confirmed foreign/successor or
same-pair conflict remains a typed fail-closed error; it is never adopted. A real fence still costs
only an epoch: recovery reclaims with a fresh one, bounded at three whole-chain attempts. This is the
general CAS posture: doubt about the source fails closed, while transport ambiguity may retry only
inside authority already proved by the last confirmed lease.

GC's own view of a dead server is symmetric and clock-skew-immune: a slot becomes fence-eligible
only after the leader observes the *same* renewal token hold stable, on its own monotonic clock,
for `TTL + floor(TTL/20) + period` — close to, but not identical to, the threshold a re-mounting
server uses to wait out a predecessor, which observes `TTL + floor(TTL/20) + max(1,
floor(period/2))`. Both thresholds are evaluated purely on the observer's own clock and its own
configured `TTL`/`period`; nothing about the writer's timing travels on the wire. The stamped
`expires_at_ms` never participates in either decision — it is a writer-stamped diagnostic used by
`system.cas_mounts` and by the non-authoritative decommission epoch-recovery precheck, never an
authorization; local fencing is derived instead from the confirmed request's pre-I/O `BOOTTIME`
anchor plus the TTL, and wall-clock `now` stays audit-only.
general CAS posture: doubt about the source fails closed. Fenced mutations and the bounded renewals
(startup, remount and direct) retry transport ambiguity only inside authority already proved by the last
confirmed lease. The background renewal may keep retrying after that lease has expired; write authority
stays bounded by the start of the renewal that last succeeded plus the TTL.

GC's own view of a dead server is symmetric and clock-skew-immune: a slot becomes fence-eligible only
after the leader observes the *same* renewal token hold stable, on its own monotonic clock, for `TTL +
floor(TTL/20) + period` — close to, but not identical to, the threshold a re-mounting server uses to
wait out a predecessor, which observes `TTL + floor(TTL/20) + max(1, floor(period/2))`. GC starts
counting from a clock sample taken after the read that returned the token, so time the round spent on a
slow `LIST` or `GET` before that read is not credited as time spent watching. Both thresholds are
evaluated purely on the observer's own clock and its own configured `TTL`/`period`; nothing about the
writer's timing travels on the wire. The stamped `expires_at_ms` never participates in either decision —
it is a writer-stamped diagnostic used by `system.cas_mounts` and by the non-authoritative decommission
epoch-recovery precheck, never an authorization; local fencing is derived instead from the confirmed
request's pre-I/O `BOOTTIME` anchor plus the TTL, and wall-clock `now` stays audit-only.

Every server sharing a pool must therefore run the identical `cas_mount_lease_ttl_ms` and
`cas_mount_renew_period_ms`: a member or GC leader configured with a shorter threshold than its
peers can fence out a healthy peer whose token-update gap merely exceeds that shorter threshold —
a peer renewing frequently stays live, one that missed a renewal does not. Change these values only
with every member of the pool stopped; a graceful restart removes only that member's own startup
observation and does not make mixed thresholds safe. With the defaults (TTL 30 s, period 10 s,
margin 2 s), `TTL − margin − period − 2 × envelope = 4 s` is the scheduling-lateness budget before
the first renewal attempt of a period can begin, where `envelope = attempt_timeout + 2 × cap` and
`cap` is `attempt_timeout` when the disk's `connect_timeout_ms` is `0`, else
`min(connect_timeout_ms, attempt_timeout)` (7 s with defaults).
`cas_mount_renew_period_ms`: a member or GC leader configured with a shorter threshold than its peers
can fence out a healthy peer whose token-update gap merely exceeds that shorter threshold — a peer
renewing frequently stays live, one that missed a renewal does not. Change these values only with every
member of the pool stopped; a graceful restart removes only that member's own startup observation and
does not make mixed thresholds safe. With the defaults (TTL 30 s, period 10 s, margin 2 s), the startup
check `period + 2 × envelope + margin < TTL` leaves `TTL − margin − period − 2 × envelope = 4 s` of
slack: a renewal that starts up to 4 s late still leaves time to admit a write before the next one. A
write is admitted only while the remaining lease exceeds the time its requests may still take plus the
margin. A conditional write reserves two envelopes on every attempt (the attempt and the read that
settles it; a ref-log append reserves the same). A removal that re-observes the key after a mismatch
reserves two plus its pause, and a retried sentinel probe reserves one plus its pause. With the defaults
a conditional write is refused once less than 16 s of the lease remains, a retried sentinel probe once
less than 9 s remains. Renewals start every 10 s, so the lease has 30 s left after one and 20 s just
before the next; if renewals keep failing, conditional writes stop 14 s after the last confirmed renewal
started, about 4 s after the next one was due. Here `envelope = attempt_timeout + 2 × cap` and `cap` is
`attempt_timeout` when the disk's `connect_timeout_ms` is `0`, else `min(connect_timeout_ms,
attempt_timeout)` (7 s with defaults).

## The two monotone counters {#counters}

Expand Down Expand Up @@ -205,7 +243,7 @@ The in-process `PoolLifecycle` runtime, by contrast, is a literal enum (`CasMoun
stateDiagram-v2
[*] --> Live: Pool constructed, fence unarmed
Live --> Live: mountWritable arms the fence
Live --> TransientNotLive: renewal failure, tripMountLost, lost=true
Live --> TransientNotLive: terminal renewal result, tripMountLost, lost=true
TransientNotLive --> Live: self-remount succeeds with a fresh epoch
TransientNotLive --> TransientNotLive: probe inconclusive, retry with backoff
TransientNotLive --> IdentityLost: pool meta and owner both authoritatively absent
Expand Down Expand Up @@ -268,7 +306,11 @@ becomes `state = 'corrupt'`, never an exception). Shows every `server_root_id` i
| `writer_epoch`, `renewal_sequence`, `started_at`, `expires_at`, `min_active_build_sequence`, `gc_fenced` | lease state (`DateTime64(3)` columns; the millisecond-integer field names live only in the internal `MountLease` struct and the on-disk body) |
| `state` | one of `live`, `expired`, `terminated`, `fenced`, `corrupt` |
| `is_leader`, `pending_reclaim`, `last_success_age_seconds`, `wedged_namespace_count` | GC health, process-local; **`NULL` on every peer row** — a process-local fact must never be stamped onto another server's row |
| `lifecycle`, `lifecycle_reason`, `lifecycle_detail`, `lifecycle_since` | the SQL surface for the in-process `PoolLifecycle` runtime above: `lifecycle` is one of `live`, `not_live`, `identity_lost`, `vanished`, `constructing`, `shutdown`; `lifecycle_reason` distinguishes `replaced` from `forgotten` for a `vanished` disk; `lifecycle_detail` carries the full diagnosis text; `lifecycle_since` is when the current non-live state began (`NULL` while live) |
| `lifecycle`, `lifecycle_reason`, `lifecycle_detail`, `lifecycle_since` | the SQL surface for the in-process `PoolLifecycle` runtime above: `lifecycle` is one of `live`, `not_live`, `identity_lost`, `vanished`, `constructing`, `shutdown`; `lifecycle_reason` distinguishes `replaced` from `forgotten` for a `vanished` disk, and is `lease_expired` for a `not_live` disk whose lease deadline has passed while renewals keep failing; `lifecycle_detail` carries the full diagnosis text (for `lease_expired`, the text of the last failed renewal request, empty if none failed); `lifecycle_since` is when the current non-live state began (for `lease_expired`, the passed deadline; `NULL` while live) |

`lifecycle_reason = 'lease_expired'` reports this server's own confirmed lease deadline. The `state`
column is a different signal: it is derived from the body's wall-clock `expires_at` plus a skew
allowance. The two need not agree.

The lifecycle snapshot is I/O-free and ungated, so a not-live, never-started, or vanished disk
still produces a row instead of silently disappearing from the table.
32 changes: 22 additions & 10 deletions docs/en/antalya/cas/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -96,7 +96,7 @@ entirely before release. Treat this table as a snapshot of the current build, no
| `cas_blob_hash` | `cityhash128` | Pool blob content-hash function (`cityhash128` \| `xxh3-128` \| `sha256`). Recorded in the pool at creation; a mismatching config is refused at mount |
| `cas_blob_hash_allow_new` | `false` | Explicit opt-in to admit a new hash algorithm into an existing pool. One-way: once admitted, the pool carries both algorithms permanently |
| `skip_access_check` | `false` | Skip the boot-time capability probe (start now, fix later). Only the preflight probe is skipped — the conditional-write correctness check still runs on every writable mount. **Not available on a writable generation-token (GCS) disk**, which refuses to mount with it: there, the probe battery is the only proof that a token-exact delete carries its generation precondition. Mount such a disk read-only if you need to defer the check |
| `cas_mount_lease_ttl_ms` | `30000` | Milliseconds for which a mount lease remains valid after a successful claim or renewal (≥ 1). Lower values shorten stale-mount recovery but reduce tolerance for object-storage and scheduling delays |
| `cas_mount_lease_ttl_ms` | `30000` | Milliseconds for which a mount lease remains valid after the start of a successful claim or renewal (≥ 1). Writes are never admitted past it, and stop earlier (see below). The background renewal keeps retrying after it until the store answers, so a failing or timing-out renewal does not by itself fence or remount the mount; GC on another pool member can still fence it after the observation threshold below, and a definitive answer from the store (`gc_fenced`, a foreign or successor body, an absent object) ends it. Lower values shorten stale-mount recovery but shorten the outage that writes ride out |
| `cas_mount_renew_period_ms` | `10000` | Milliseconds between background mount-lease renewals (≥ 1). It must leave enough time for two attempt envelopes (a renewal write and its settlement read) and the lease safety margin before the TTL expires: `period + 2 × envelope + margin < TTL` |
| `cas_gc_snapshot_generations_to_keep` | `3` | GC snapshot generations retained |
| `cas_gc_shards` | `1` | Blob-hash-prefix reducer shards (≥ 1). Recorded in the pool at creation; a mismatching config is refused at mount |
Expand All @@ -117,21 +117,33 @@ Startup reclaim and GC's fence-out both judge liveness by the mount slot's write
on the observer's own `CLOCK_BOOTTIME`, using the observer's own threshold — nothing about a writer's
timing travels on the wire. Startup observes `cas_mount_lease_ttl_ms + floor(cas_mount_lease_ttl_ms /
20) + max(1, floor(cas_mount_renew_period_ms / 2))`; GC observes `cas_mount_lease_ttl_ms +
floor(cas_mount_lease_ttl_ms / 20) + cas_mount_renew_period_ms`. A pool member or GC leader
floor(cas_mount_lease_ttl_ms / 20) + cas_mount_renew_period_ms`. GC counts the observation from a clock sample taken after the read that returned the token, not from the start of its round. A pool member or GC leader
configured with a shorter threshold than its peers can therefore fence out a healthy peer whose
token-update gap exceeds that shorter threshold — a peer renewing frequently stays live, one that
missed a renewal does not. Change these values only with every member of the pool stopped: a
graceful restart removes only that member's own startup observation and does not make mixed
thresholds safe.

A shorter TTL reduces the tolerance for object-storage delays; a shorter renewal period increases it
(renewal starts earlier) at the cost of more background traffic. With the defaults,
`cas_mount_lease_ttl_ms − cas_lease_safety_margin_ms − cas_mount_renew_period_ms − 2 × envelope =
4000` ms is the scheduling-lateness budget before the first renewal attempt of a period can begin,
where `envelope = cas_attempt_timeout_ms + 2 × cap` (7000 ms with defaults) and `cap` is
`cas_attempt_timeout_ms` when the disk's `connect_timeout_ms` is `0`, else
`min(connect_timeout_ms, cas_attempt_timeout_ms)` (1000 ms with defaults); the renewal then keeps
retrying until `confirmed deadline − cas_lease_safety_margin_ms`.
A longer TTL lengthens the outage that writes ride out without a refusal. A shorter renewal period starts
each renewal earlier, so more of the TTL is left when one fails, at the cost of more background traffic.
With the defaults, `cas_mount_lease_ttl_ms − cas_lease_safety_margin_ms − cas_mount_renew_period_ms −
2 × envelope = 4000` ms is the slack the startup check leaves: a renewal that starts up to that late
still leaves time to admit a write before the next one. Here `envelope = cas_attempt_timeout_ms + 2 × cap`
(7000 ms with defaults) and `cap` is `cas_attempt_timeout_ms` when the disk's `connect_timeout_ms` is `0`,
else `min(connect_timeout_ms, cas_attempt_timeout_ms)` (1000 ms with defaults). The background renewal
retries, about a second apart, until the store answers. It does not stop at the lease deadline.

A write is admitted only while the remaining lease exceeds the time its requests may still take plus
`cas_lease_safety_margin_ms`. A conditional write (create, replace, read-modify-write) reserves two
attempt envelopes on every attempt: the attempt and the read that settles it. A removal that re-observes
the key after a mismatch reserves two plus its pause, and a sentinel probe that is retried reserves one
plus its pause. With the defaults (envelope 7000 ms, margin 2000 ms) a conditional write is therefore
refused once less than 16000 ms of the lease remains, and a retried sentinel probe once less than 9000
ms remains. The lease is renewed every `cas_mount_renew_period_ms` (10000 ms) from the start of the last
confirmed renewal, so it has 30000 ms left after a renewal and 20000 ms just before the next one. If
renewals keep failing, conditional writes are refused from 14000 ms after the last confirmed renewal
started, about 4000 ms after the next one was due. `system.cas_mounts` shows `lifecycle = 'live'` until
the deadline itself passes; from then it shows `lifecycle_reason = 'lease_expired'`.

The `expires_at_ms` stamped into the mount object is a writer-stamped diagnostic used by
`system.cas_mounts` and by the non-authoritative decommission epoch-recovery precheck; it never
Expand Down
Loading
Loading