Skip to content

Research lightweight handling of panicking wakers #342

Description

@tisonkun

Motivation

Follow-up to #341, with the earlier discussion in #335. Asyncband currently expects executor waker operations to be non-panicking and does not promise recovery from panicking waker operations. Removing the previous recovery machinery made notification and permit-distribution paths substantially easier to follow.

A wake callback can still panic after a batch of waiters has been detached, interrupting the remaining notifications. Investigate whether we can attempt those remaining wakes with a small, clear implementation, without restoring the complexity removed in #341.

Research questions

  • Which real-world executor or custom-waker scenarios benefit from recovery, and what behavior should callers observe after catching the panic?
  • Can an ownership or RAII-based approach provide useful recovery? Distinguish cleaning up remaining wakers from actually waking them, and account for a later wake also panicking or cleanup running during an existing unwind.
  • What state must be committed before invoking callbacks, especially for permit distribution and cancellation handoff?
  • Where should the recovery boundary be? Focus on wake and wake_by_ref; assess clone and drop where they affect the proposed approach, rather than assuming that every waker operation needs a general recovery framework.
  • What are the costs in code complexity, allocation, and the normal notification path?

Existing examples

The initial source review found several different policies rather than one ecosystem-wide convention:

  • Tokio 1.53.1 WakeList uses a drop guard to destroy remaining wakers after a wake panic; it does not attempt the remaining wakes. Its internal AtomicWaker separately restores registration state after a clone panic.
  • thingbuf 0.1.6 WaitCell follows Tokio's registration-recovery strategy, while its batch queue notifications directly invoke callbacks.
  • event-listener 5.4.2 directly invokes wake callbacks in its notification loop, without per-callback panic isolation. async-lock, async-channel, and async-broadcast build on this notification mechanism.
  • embassy-sync 0.8.0 MultiWakerRegistration clears the stored length before waking to preserve memory safety during unwinding; a panic stops the loop and can leak the remaining wakers.
  • asupersync 0.5.0 Notify catches each wake panic, retains the first payload, attempts the remaining wakes, and then resumes the first panic. This is a concrete comparison point for the stronger behavior and its implementation cost.

Desired outcome

An evidence-backed proposal with a small prototype and targeted validation, or a documented conclusion explaining why the available approaches do not justify their complexity. Keep callbacks and replaced or cancelled waker destruction outside primitive locks, preserve basic ownership and unwind safety, introduce no new dependencies, and keep the normal path simple.

A successful approach is intended to be a non-breaking robustness improvement under the contract established by #341. Broader guarantees for arbitrary panicking waker operations should be justified separately.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    help wantedExtra attention is needed

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions