refactor: [SDK-4995] bootstrap dispatchers off the caller thread - #2766
fadi-george wants to merge 1 commit into
Conversation
Co-authored-by: Cursor <cursoragent@cursor.com>
There was a problem hiding this comment.
Multi-model review (Opus 5.5, GPT 5.6 Sol, Grok 4.7)
Intent: first launchOn* / dispatch must queue and return without building a pool on the caller (often main).
Act on
failPendingreopensCLOSED(3/3).close()is not terminal if in-flightinitialize()then fails — the old gate goes back toCOLDand can leak a new executor afterresetForTesthas already dropped that generation.createTargetleaks workers on prestart failure (3/3).prestartAllCoreThreads()can throw after some cores are alive;allowCoreThreadTimeOut(false)and noshutdownNow()on that path, so retries leak more threads.Errorfrom thread creation skips fallback (2/3, Opus + Grok). Onlycatch (Exception)triesDispatchers.IO/Default.prestartAllCoreThreads()throwingOutOfMemoryErroris the realistic primary-failure case.
Consider
- Cold
pendingis unbounded; drain then burst-submits into the 200-slot IO/Default queues, so work queued during a slow bootstrap can be cancelled at birth (3/3). prewarm()tests no longer assert off-caller construction; a synchronousprewarmwould still pass (3/3).- Coordinator
running+ConcurrentLinkedQueueexit vsrequest()can leave a lane stuck inSTARTING(Grok). prewarmStartedstays true if the bootstrap thread fails to start, so laterprewarm()no-ops (Opus, Grok).
Noted
cancelAndCompleteonDispatchers.IOcan run SerialIOfinallyoff the serial thread (Opus).close()duringDRAINING/READYshutdownNow()s without completing those Jobs (Opus).
Dismissed
- Public API contract is intact. SerialIO remaining unbounded is intentional.
Sent by Cursor Automation: PR Reviews
| synchronized(lock) { | ||
| val copy = pending.toList() | ||
| pending.clear() | ||
| state = LaneState.COLD |
There was a problem hiding this comment.
Act on (3/3): failPending always sets COLD, including after close() has set CLOSED. If resetForTest / shutdown races an in-flight initialize(), the discarded generation can bootstrap again and leak an executor that later resets will not tear down. Keep CLOSED terminal.
| OptimizedThreadFactory(config.threadName, config.priority), | ||
| ) | ||
| executor.allowCoreThreadTimeOut(false) | ||
| executor.prestartAllCoreThreads() |
There was a problem hiding this comment.
Act on (3/3): If prestartAllCoreThreads() throws after starting some workers, this executor is never shutdownNow()’d (allowCoreThreadTimeOut(false)). Also 2/3 (Opus + Grok): that throw is typically OutOfMemoryError, so createTargetOrFallback skips the Dispatchers.* fallback and only catch (Exception) would have used it.
abdulraqeeb33
left a comment
There was a problem hiding this comment.
Requesting changes. CI is still red (NotificationGenerationProcessorTests > processNotificationData should not display notification when external callback indicates not to).
prestartAllCoreThreads throwing Error skips the Dispatchers fallback and leaks the half-built executor. That is the failure the fallback is for.
| failPending("OneSignal $lane dispatcher fallback failed") | ||
| null | ||
| } | ||
| } catch (t: Throwable) { |
There was a problem hiding this comment.
prestartAllCoreThreads (line 365) throws OutOfMemoryError after some core threads exist. That misses catch (Exception) and lands here, which cancels queued work and never calls createFallbackTarget. Shut the partial executor down in a finally, then fall back. Only failPending if the fallback also fails. failPending also sets the lane back to COLD, which reopens a CLOSED generation during resetForTest.


Description
One Line Summary
Build dispatcher executors on a background bootstrap thread so the first
launchOn*call never constructs a pool on the caller (often main) thread.Details
Motivation
Split out of #2712.
prewarm()is best-effort: a cold-start entry point that dispatches right after it can still win the lazy-init race and pay executor construction on the main thread (SDK-4995 Play Vitals data showed no improvement after the 5.9.5 prewarm).Scope
GateDispatcherthat queues work while cold and requests bootstrap from a singleBootstrapCoordinatorthread.Dispatchers.IO/Default/IO.limitedParallelism(1). If fallback also fails, queued jobs are cancelled and later lanes are not wedged.Dispatchers.IO, never inline on the caller thread.newSingleThreadExecutor.prewarm()now just requests warmup for every lane. No public API changes.Testing
Unit testing
Manual testing
Not tested on device separately; covered by the full core and notifications unit suites.
Affected code checklist
Checklist
Overview
Testing
Final pass