Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
38 commits
Select commit Hold shift + click to select a range
2835741
Route the CAS reclaim's lease-loss step and fence arm through the run…
filimonov Oct 2, 2026
76979a4
Latch the fence at the start of a CAS reclaim and arm only when admis…
filimonov Oct 2, 2026
07ecdee
Run the CAS mount lease and its reclaim on one thread
filimonov Oct 2, 2026
ecd7d59
Let only the CAS lease thread drive the renewer while it runs
filimonov Oct 2, 2026
fd39d11
Describe the single CAS lease thread in the mounts and leases doc
filimonov Oct 2, 2026
bb38179
Trip and request the remount of a CAS interference report in one step
filimonov Oct 2, 2026
6897788
Pin the one-step CAS interference report with a deterministic test
filimonov Oct 2, 2026
a9af278
Arm a latched CAS mount fence from a renewal with room for a ref append
filimonov Oct 2, 2026
ccda27b
Say that writes resume when a renewal restores the CAS mount lease
filimonov Oct 2, 2026
123d67c
Make a CAS mount ready through the lease thread instead of a re-anchor
filimonov Oct 2, 2026
78ae6eb
Describe CAS mount readiness through the lease thread
filimonov Oct 2, 2026
30ff8c8
Run the CAS readiness death test in a fresh process and pin the arm o…
filimonov Oct 2, 2026
e5e1dcb
Describe when a CAS remount counts as succeeded or failed
filimonov Oct 2, 2026
f535696
Merge the CAS lease expiry warning from the fix branch
filimonov Oct 2, 2026
a127ae0
Renew the CAS mount lease through one step, also from the test seam
filimonov Oct 2, 2026
1facfbc
Remove MountLeaseRenewer::renewForRemount
filimonov Oct 2, 2026
59c93ef
Run every CAS mount renewal on the lease plane without a lease bound
filimonov Oct 2, 2026
8ba79c8
Remove the CAS mount renewal deadline counter and classifications
filimonov Oct 2, 2026
cbada51
Return the identity, ending and duration of a CAS mount renewal
filimonov Oct 2, 2026
0145e1c
Report CAS mount renewals from their result, without a thread-local s…
filimonov Oct 2, 2026
68462f8
Pass the lease deadline to the CAS mount farewell
filimonov Oct 2, 2026
531f6cb
Define the two-envelope write cost once in CasRequestBudget
filimonov Oct 2, 2026
dabc19a
Name the default safety margin and the default write window
filimonov Oct 2, 2026
48fd237
Correct the lease margin, grace, FORGET trip and renewal descriptions
filimonov Oct 2, 2026
7a3d9f6
Make warnOnceIfLeaseExpired private
filimonov Oct 2, 2026
45aac31
Delete process_epoch and the epoch accessors nothing reads
filimonov Oct 2, 2026
97953a2
Test the Vanished lifecycle instead of a shadow flag in enterVanished
filimonov Oct 2, 2026
e488f97
Read the renewal period and the boot clock from one place
filimonov Oct 2, 2026
a9a02fc
Derive allow_steal from the round trigger in CasGcScheduler
filimonov Oct 2, 2026
e6cabaf
Use remountTerminal in the GC scheduler loops
filimonov Oct 2, 2026
1152bc8
Remove the write-only fence, runtime and renewal result fields
filimonov Oct 2, 2026
8ffad69
Log the stale-mount observation once per window start
filimonov Oct 2, 2026
8ddd799
Assert GC quiescence over a parked real round
filimonov Oct 2, 2026
04b4e91
Name the renewer's claim and farewell plane for what it is
filimonov Oct 2, 2026
6bdde54
Name the lease keeper's two recurring conditions; drop dead code
filimonov Oct 2, 2026
8920c76
Bound the lease-keeper tests that could hang; pin two untested paths
filimonov Oct 2, 2026
98ca1ad
Make lease-keeper comments, counters and docs say what the code does
filimonov Oct 2, 2026
56897c5
Fix four lease-keeper comments that named the wrong callers or order
filimonov Oct 2, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
71 changes: 43 additions & 28 deletions docs/en/antalya/cas/architecture/mounts-and-leases.md
Original file line number Diff line number Diff line change
Expand Up @@ -88,9 +88,7 @@ watermark — there is no separate watermark object. `MountLease` fields: `serve
suspend correctly observes itself expired. The deadline decides about writes: none is admitted past
it, and none is admitted once the remaining lease cannot cover the requests it may send plus the
safety margin. The background renewal does not stop at the deadline. It retries timeouts, `5xx`
answers and connection errors about a second apart until the store answers. The renewals at startup,
after a remount and the direct renewal stay bounded: they stop at the last confirmed deadline minus
the safety margin. A retry, `GET`, response timestamp, or wall-clock step never extends authority.
answers and connection errors about a second apart until the store answers. A retry, `GET`, response timestamp, or wall-clock step never extends authority.
- **Cadence.** The runtime normally starts a logical renewal every `cas_mount_renew_period_ms` (default
10 s), with TTL `cas_mount_lease_ttl_ms` (default 30 s, TTL/3 renewal ratio). The next beat is anchored
at the committed body's pre-I/O BOOTTIME start. A slow recovery therefore causes an immediate
Expand All @@ -112,8 +110,9 @@ only on one of these:
absent object, or a request the store refuses on a clear attempt;
- a stop or a remount request;
- a deterministic local failure;
- a lifecycle other than `Live`;
- a lost fence.
- a terminal lifecycle, which includes a published `FORGET` intent.

A lost fence alone does not end it: the renewal that makes a mount ready runs under a latched fence.

A terminal renewer cannot mint another body or publish a clean farewell. Owner cancellation before
any request is the only `NotAttempted` result and leaves clean release possible. Cancellation after a
Expand Down Expand Up @@ -141,10 +140,10 @@ local fence (latches `lost`, bumps the fence generation, moves the in-process ru
`TransientNotLive`) and latches one self-remount generation. A confirmed foreign/successor or
same-pair conflict remains a typed fail-closed error; it is never adopted. A real fence still costs
only an epoch: recovery reclaims with a fresh one, bounded at three whole-chain attempts. This is the
general CAS posture: doubt about the source fails closed. Fenced mutations and the bounded renewals
(startup, remount and direct) retry transport ambiguity only inside authority already proved by the last
confirmed lease. The background renewal may keep retrying after that lease has expired; write authority
stays bounded by the start of the renewal that last succeeded plus the TTL.
general CAS posture: doubt about the source fails closed. Fenced mutations retry transport ambiguity
only inside authority already proved by the last confirmed lease. The renewal may keep retrying after
that lease has expired; write authority stays bounded by the start of the renewal that last succeeded
plus the TTL.

GC's own view of a dead server is symmetric and clock-skew-immune: a slot becomes fence-eligible only
after the leader observes the *same* renewal token hold stable, on its own monotonic clock, for `TTL +
Expand Down Expand Up @@ -242,9 +241,9 @@ The in-process `PoolLifecycle` runtime, by contrast, is a literal enum (`CasMoun
```mermaid
stateDiagram-v2
[*] --> Live: Pool constructed, fence unarmed
Live --> Live: mountWritable arms the fence
Live --> Live: the open's claim or first renewal arms the fence
Live --> TransientNotLive: terminal renewal result, tripMountLost, lost=true
TransientNotLive --> Live: self-remount succeeds with a fresh epoch
TransientNotLive --> Live: the reclaim or the next renewal arms the fence under a fresh epoch
TransientNotLive --> TransientNotLive: probe inconclusive, retry with backoff
TransientNotLive --> IdentityLost: pool meta and owner both authoritatively absent
TransientNotLive --> VanishedReplaced: foreign pool_id observed
Expand All @@ -254,29 +253,45 @@ stateDiagram-v2
VanishedForgotten --> [*]
```

`IdentityLost`, `VanishedReplaced` and `VanishedForgotten` are terminal and absorbing: the remount
and GC threads self-exit, and there is deliberately no auto-revive — an identity disappearing
`IdentityLost`, `VanishedReplaced` and `VanishedForgotten` are terminal and absorbing: the lease
thread and the GC threads exit, and there is deliberately no auto-revive — an identity disappearing
under a live mount is an operator-level event.

## Mount, unmount, crash {#mount-lifecycle}

**Writable open** runs in a strict order: bootstrap-residual proof, capability probe under a
random per-mount prefix, pool-meta create-or-validate, `validateServerRootId`, owner claim,
`allocateWriterEpoch`, mount claim and synchronous renewer start, arm the fence, then create and
release the runtime-owned renewal and remount workers before the writable pool becomes externally
visible. If the claim consumed the TTL, one fresh synchronous renewal re-anchors the deadline
before the fence is armed.
Failure to construct either worker joins the partial pair, closes the fence, and fails the writable
open. No incident path constructs a thread.

The renewal and remount workers are separate and long-lived under one stable `CasMountRuntime`.
`scheduleRemount` increments a requested-generation latch and wakes the persistent remount worker,
including while an older generation is active. Before renewer replacement, remount requests
`ParkRequested` and waits for the renewal driver to report `Parked`, which proves that no renewer call
is in flight. A successful remount handles only its snapshotted generation; a newer request is
processed before renewal resumes.

**Clean unmount:** request stop and join both persistent workers, drain the ref lanes, and only if
`allocateWriterEpoch`, mount claim and synchronous renewer start, arm the fence, then start the
runtime-owned lease thread before the writable pool becomes externally visible.
Failure to start the lease thread closes the fence and fails the writable open. No incident path
constructs a thread.

**Readiness.** An open arms the fence from its claim only when the claim's deadline, its start plus the
TTL, still leaves room for a ref append: `2 × envelope + margin` (16 s with the defaults). Otherwise the
fence stays latched, the lease thread renews at once, and the first renewal whose own deadline leaves that
room arms the fence; then the open returns. A renewal that commits with less room arms nothing, and the
next one follows at once. If no renewal arms the fence within one TTL, the open stops and joins the lease
thread and fails with `ABORTED`, naming the last failed renewal request. A reclaim whose arming conditions
(below) do not hold keeps the pool `TransientNotLive` with the fence latched. Then the next renewal arms
the fence and reports `Live` when no request is pending, a newer request gets another reclaim, and a stop
or a terminal lifecycle ends the thread. The `mount_remount` row of such a reclaim has `outcome = 'ok'`
and `step = 'claimed_not_armed'`.

One lease thread per writable mount renews the lease and runs the self-remount, one after the other.
An interference report (`tripAndRequestRemount`) and a terminal renewal trip the fence, raise the
requested remount generation unless a request no reclaim has started serving is already pending or the
pool is stopping or terminal, and wake the thread; a pending request also ends a renewal in progress. While the thread runs, only it replaces, starts or resets the renewer.

- A reclaim latches the fence first.
- It arms the fence, and reports `Live`, only when no newer request is pending, no stop is requested
and the lifecycle is not terminal. The check and the arm are one step under the runtime's mutex,
which a stop, a request and a FORGET intent also take.
- A reclaim acknowledges only the generation it served, so a request raised during a reclaim is served
by the next reclaim before renewal resumes.
- A renewal's result is published before the thread looks at the requests again, so a renewal that
finished before a request cannot overwrite the deadline of the reclaim that follows.

**Clean unmount:** request stop and join the lease thread, drain the ref lanes, and only if
the drain *certified* quiescence call `MountLeaseRenewer::release` on an `Active` renewer to write the
terminal farewell (`expires_at_ms` already expired, `min_active_build_sequence = UINT64_MAX`). That sentinel is what
lets a successor reclaim instantly. A `RenewalTerminal` renewer, an unresolved ref write, or a sent
Expand Down
4 changes: 2 additions & 2 deletions docs/en/antalya/cas/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -108,8 +108,8 @@ entirely before release. Treat this table as a snapshot of the current build, no
| `cas_gc_meta_pool_size` | `16` | Bounded pool size for GC per-hash freshness-meta writes |
| `cas_gc_io_concurrency` | `16` | Bounded pool size for GC object-storage requests that run in parallel: the fold's read-ahead (checkpoints, ref logs, manifests, zero-candidate HEADs), the orphan-manifest sweep planning reads, the `SYSTEM CAS GC REBUILD` read-ahead, and the `pending_deletes` blob `HEAD` + conditional `DELETE` fan-out. Not covered: meta writes (`cas_gc_meta_pool_size`) and all other GC requests, which run on the round thread. `1` runs the covered requests sequentially. `cas_gc_read_concurrency` is rejected without an alias; use `cas_gc_io_concurrency` instead |
| `cas_attempt_timeout_ms` | `5000` | Budget for one HTTP attempt of a writable Native mount's control-plane requests (read, head, list, remove, conditional write), at least 1. Together with the connect cap it forms the attempt envelope (`cas_attempt_timeout_ms + 2 × cap`; the cap is `cas_attempt_timeout_ms` itself when the disk's `connect_timeout_ms` is `0`, else `min(connect_timeout_ms, cas_attempt_timeout_ms)`) that the lease arithmetic reserves: one TCP connect and one TLS handshake under the cap each, send/receive bounded per socket operation by `cas_attempt_timeout_ms`. With background renewal the cadence check requires `cas_mount_renew_period_ms + 2 × envelope + cas_lease_safety_margin_ms < cas_mount_lease_ttl_ms`, which puts an effective ceiling on the frozen connect cap: under the defaults (TTL 30000, period 10000, margin 2000) the envelope must stay under 9000, so a disk `connect_timeout_ms` of 2000 ms or more refuses to open writable — lower the connect timeout or raise the TTL if you hit this |
| `cas_lease_safety_margin_ms` | `2000` | Startup-only margin validated against the mount lease TTL: the attempt envelope + `cas_lease_safety_margin_ms` must be strictly less than the mount lease TTL, and `cas_mount_renew_period_ms` + 2 × envelope + `cas_lease_safety_margin_ms` too, or the disk refuses to open writable |
| `cas_unsafe_remount_no_delay` | `0` | Reclaim a mount slot that carries this server's own uuid at once after a hard restart, without observing the slot's token for the lease TTL. Unsafe whenever two processes can hold the same `server_uuid` (a copied uuid file, a stalled predecessor). After such a reclaim the predecessor can still start conditional writes until its own cutoff (`confirmed deadline − cas_lease_safety_margin_ms − 2 × envelope`) or until its next renewal meets the token guard, and a request it already sent may still materialize later. That is not a data hazard: ref-log keys carry `(writer_epoch, sequence)` and creates are conditional, so two writers can never commit different bodies to one key, and recovery's epoch seal settles any straggler (recovery fails closed after 64 successive seal-create attempts displaced by newly materializing old-epoch transactions). The exposure is availability, not data. Intended for test stands and deployments that guarantee one process per uuid |
| `cas_lease_safety_margin_ms` | `2000` | Room kept between a request and the mount lease deadline: a write is admitted only while the requests it may send plus this margin fit in the remaining lease. Validated against the mount lease TTL when a writable mount opens: the attempt envelope + `cas_lease_safety_margin_ms` must be strictly less than the mount lease TTL, and `cas_mount_renew_period_ms` + 2 × envelope + `cas_lease_safety_margin_ms` too, or the disk refuses to open writable |
| `cas_unsafe_remount_no_delay` | `0` | Reclaim a mount slot that carries this server's own uuid at once after a hard restart, without observing the slot's token for the lease TTL. Unsafe whenever two processes can hold the same `server_uuid` (a copied uuid file, a stalled predecessor). After such a reclaim, unless its next renewal meets the token guard first, the predecessor can still start a conditional write, a ref-log append included, until its own cutoff (`confirmed deadline − cas_lease_safety_margin_ms − 2 × envelope`: the write and the read that settles it), and a single-envelope request such as a conditional delete until one envelope later. A request it already sent may still materialize later. That is not a data hazard: ref-log keys carry `(writer_epoch, sequence)` and creates are conditional, so two writers can never commit different bodies to one key, and recovery's epoch seal settles any straggler (recovery fails closed after 64 successive seal-create attempts displaced by newly materializing old-epoch transactions). The exposure is availability, not data. Intended for test stands and deployments that guarantee one process per uuid |
| `cas_staging_backend` | `local` | Blob staging backend (`local` \| `s3`); `s3` is opt-in and requires native same-store copy on writable mount |

All servers sharing a pool must run the same `cas_mount_lease_ttl_ms` and `cas_mount_renew_period_ms`.
Expand Down
22 changes: 11 additions & 11 deletions docs/en/antalya/cas/operations/debugging.md
Original file line number Diff line number Diff line change
Expand Up @@ -133,25 +133,25 @@ follows:
request; `committed_after_retry` means a later identical physical `PUT` completed and the response
itself proved it; `committed_after_expiry` means the renewal restored a lease that had expired (the
row also has `expired_ms`). When more than one applies, the first of those three in that order wins.
- `outcome = 'failed'` carries the decisive `classification`: `external_lease_deadline` (the
confirmed lease's own safety margin, not the request policy, ran out first — check object-store
latency or `BOOTTIME` advancement before anything else), `request_deadline` (the ninety-second
request policy exhausted first), `unresolved` (every attempt was ambiguous and never settled by
- `outcome = 'failed'` carries the decisive `classification`: `unresolved` (every attempt was ambiguous and never settled by
the time the operation gave up), `conflict` (an exact resolve read found another body — a
same-pair twin, a GC-fenced body, a successor epoch, or a foreign holder), `cancelled` (a
renewal in flight was cancelled by shutdown or a remount park request; expected during graceful
shutdown), `fence_or_lifecycle_lost` (another local fence loss or a terminal lifecycle transition
closed admission while the operation was active), `deterministic_failure` (the store's own
renewal in flight was cancelled by shutdown; expected during graceful shutdown),
`fence_or_lifecycle_lost` (a pending remount request or a terminal lifecycle, including a published
FORGET intent, ended the renewal), `deterministic_failure` (the store's own
answer proved the write never applied), and `vanished` (an exact resolve read proved the mount
slot absent — the pool directory was removed or renamed out of band, or a decommission raced the
renewal). `terminal_unclassified` means the renewal terminated through a path that assigned no
classification; that is a defect to report together with the surrounding rows, not an operator
condition. Do not collapse these into a generic timeout — the action
differs by classification, and only `external_lease_deadline` and `request_deadline` are about a
deadline at all.
differs by classification.
- A following `mount_remount` row names the whole-chain `attempt_no` and final `step`. An `ok` row
restored `Live` under the reported fresh `writer_epoch`; a `failed` row's `step` and optional
`error` identify where that whole-chain attempt stopped.
with `step = 'publish_live'` restored `Live` under the reported fresh `writer_epoch`; with
`step = 'claimed_not_armed'` the claim succeeded but its arming conditions did not hold (too little
lease left, a newer remount request, a stop, or a terminal lifecycle), and the fence stays latched. The
next renewal arms it and reports `Live` when no request is pending; a newer request gets another
reclaim; a stop or a terminal lifecycle ends the lease thread. A `failed` row's `step` and optional `error` identify where
that whole-chain attempt stopped.

Use deltas of the mount counters from
[monitoring](/antalya/cas/operations/monitoring#mount-renewal-remount-counters) to check completeness:
Expand Down
Loading