Skip to content

fix: preserve CPU headroom during background indexing - #2022

Open
mvanhorn wants to merge 3 commits into
DeusData:mainfrom
mvanhorn:fix/1084-background-index-cpu
Open

mvanhorn wants to merge 3 commits into
DeusData:mainfrom
mvanhorn:fix/1084-background-index-cpu

Conversation

@mvanhorn

@mvanhorn mvanhorn commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

What does this PR do?

Carry an internal background-execution marker on the index requests created by both session auto-index and watcher re-index paths, preserving it through daemon coordination and supervised-worker serialization. When handle_index_repository builds the pipeline, translate that marker into pipeline execution context rather than changing the public MCP tool schema or adding a user-facing configuration key. Make full and incremental pipeline worker-count selection consult that context and use the existing background/incremental default that leaves CPU headroom; retain the current CBM_WORKERS precedence, single-thread crash-recovery behavior, and all-core default for explicit/manual indexing.

Enabling auto_index has repeatedly caused high CPU usage and severe Windows UI stutter, including a fresh confirmation on v0.10.8. Main already contains the maintainer-identified non-Git auto_index_limit guard, dirty-state watcher deduplication, and subprocess RSS isolation, so reimplementing those fixes would be a no-op. The remaining production path still treats automatic first indexing like a foreground full index: the pipeline selects the initial=true worker policy, whose documented behavior is to use every detected core because “the user is waiting.” Automatic session and watcher jobs run in the background, so they should instead use the repository's existing headroom-preserving worker policy while explicit indexing retains its current throughput.

Fixes #1084

Checklist

  • Every commit is signed off (git commit -s) — required, CI rejects
    Not run: no test command resolved in this workspace, so nothing was executed to pass.
    unsigned commits (DCO, see CONTRIBUTING.md)
  • Tests pass locally (make -f Makefile.cbm test)
    Not run: no test command resolved in this workspace, so nothing was executed to pass.
  • Lint passes (make -f Makefile.cbm lint-ci)
    Not run: no test command resolved in this workspace, so nothing was executed to pass.
  • New behavior is covered by a test (reproduce-first for bug fixes)

@mvanhorn
mvanhorn requested a review from DeusData as a code owner September 3, 2026 07:21
@github-actions

github-actions Bot commented Sep 3, 2026

Copy link
Copy Markdown

Thanks for opening this — it has been seen, and it is queued.

This note is automated, but it is not a brush-off: it exists so you know where your PR stands instead of having to guess from silence.

Current review status: working through a backlog. 0.9.1-rc.1 is out, so the release freeze that held reviews is over — but it left a large queue of open pull requests behind it, and we are reading through them oldest-first. The background is in discussion #1144.

What that means for this PR, concretely:

  • It will not be closed for inactivity. No stale bot touches pull requests here.
  • It may still sit a while before a human reads it. That is on us, not on you.
  • Older PRs are read first, so a recent one is not being skipped — it is behind a queue.

Things that will genuinely speed it up whenever review does happen:

  • Keep it rebased on main — the tree is moving quickly right now, and a conflicting branch cannot be reviewed as the diff you intended.
  • Get CI green, or say which failures you believe are pre-existing.
  • Keep the change to one claim. Bundled features and refactors get split before they get merged, which costs you a round trip.
  • Every commit needs a sign-off (git commit -s) — CI enforces DCO.

If this fixes a bug, a reproduction we can run is worth more than a description of the symptom.

Thanks for contributing, and sorry in advance for the wait.

@DeusData

DeusData commented Sep 3, 2026

Copy link
Copy Markdown
Owner

Welcome, and thank you for the PR 👋 — one quick, mechanical thing before review, so you are not left guessing at the red.

dco is failing because your commit is not signed off. c846ed7b ("fix: preserve CPU headroom during background indexing") carries no Signed-off-by trailer, and this repository enforces DCO on push.

The fix is to amend the commit with a sign-off and force-push your branch:

  • git commit --amend -s --no-edit — -s appends the trailer using your configured user.name and user.email
  • then force-push your own PR branch (with lease), naming the branch explicitly

If you end up with more than one commit, git rebase --signoff <base> does the whole branch at once.

Two things that will save you time when the rest of CI reports:

  • test-msan is currently failing for everyone — it dies at the Docker image build step, before any test runs. If you see it red, it is ours, not yours.
  • test-windows-guards has a standing red on test_daemon_stability.py this week, likewise unrelated to any diff.

I will review the change itself properly once it is signed off — background-indexing CPU headroom is a good thing to be looking at, and +209/-12 over 8 files is a reviewable size.

@DeusData

DeusData commented Sep 3, 2026

Copy link
Copy Markdown
Owner

Attributed your reds so you are not chasing nine separate things — there are really only two, and one of them is a genuinely interesting test-design point.

1. dco — still unsigned

c846ed7b still has no Signed-off-by trailer. Amend with -s and force-push your own branch (see my earlier note).

2. The sanitizer legs — your own new test, and worth thinking about

The failure is not spread across nine problems. On the macOS TSan leg:

mcp_auto_index_in_process_uses_background_worker_policy
  FAIL tests/test_mcp.c:13035: selected_workers == 2, expected cbm_default_worker_count(false) == 4

That is the test this PR adds, asserting:

ASSERT_EQ(selected_workers, cbm_default_worker_count(false));

against your production change:

workers = cbm_default_worker_count(!p || !p->background);

The assertion re-evaluates the helper rather than naming an expected number, so it silently assumes nothing between cbm_default_worker_count(false) and the value the pipeline actually selects can change the answer. The run says otherwise: the helper returned 4 while the pipeline chose 2. Something downstream — a clamp, a cap, or an environment-sensitive input — is intervening.

That is worth pinning down rather than papering over, because it is the same shape as a class of bug we have been clearing out of this suite all week: assertions whose verdict depends on the machine rather than the code. A test that compares a captured value against a live re-evaluation of a helper will agree on your laptop and disagree on a constrained runner, and neither result tells you whether the policy is right.

Two questions that should settle it:

  • Is 2 the value your change intends on that runner, and the helper's 4 simply not the right comparison? Then assert the policy relationship directly, or assert against whatever the pipeline is actually given.
  • Or is the pipeline clamping below your intended policy? Then the production change is not yet doing what the description says, and the test has correctly caught it.

Either way the test is doing its job — it failed rather than passing vacuously, which is more than a lot of new tests manage.

Not everything red is that

For completeness: the test-unix (ubuntu-latest) leg reports 2216 passed, 0 failed and failed at a later step, so that one is environmental rather than yours. And test-msan fails at the Docker image build with the suite skipped — a DNS resolution failure inside buildkit, entirely ours.

I have not reviewed the change on merit yet — I will once it is signed off and the worker-count question is resolved. Preserving CPU headroom during background indexing is a good thing to be fixing.

@adfjadfj16-a11y

This comment has been minimized.

@DeusData

DeusData commented Sep 3, 2026

Copy link
Copy Markdown
Owner

Maintainer notice: please disregard comments from @adfjadfj16-a11y on this thread

@adfjadfj16-a11y is not a maintainer of this project and does not speak for it. That account has posted replies on 17 threads here written in the project's voice — promising merges, announcing that a case has been "escalated to the development team", asking to close issues, and in some threads replying as though it were the author of someone else's pull request. None of those were maintainer decisions, and none of them carried any weight.

@DeusData is the only account that gives a maintainer response on this repository. If a comment about the fate of your issue or pull request did not come from @DeusData, it is not a decision, however official it reads.

If you were waiting on something because of one of those comments — a promised merge, a review "immediately", a request to close your ticket — I am sorry. That was noise you had no way to identify as noise, and it should not have been on your thread. Your issue or PR is judged on its own merits, and I will answer it here myself.

Nothing in this notice reflects on your contribution. Thank you for your patience, and thank you for the work.

@DeusData DeusData added bug Something isn't working stability/performance Server crashes, OOM, hangs, high CPU/memory priority/high Needs near-term maintainer attention; high-impact bug, regression, safety issue, or release blocker. labels Sep 3, 2026
@mvanhorn
mvanhorn force-pushed the fix/1084-background-index-cpu branch from c846ed7 to 84fffef Compare September 7, 2026 06:39

@DeusData DeusData left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The attribution I owed you on 3 September, with apologies that it took this long. Every red on your run 34091819685 is now accounted for, and none of them is a mystery:

job what failed whose
test-diag, test-lsan-macos, test-unix 2/3 (x64 + arm64) suites fully green (7643 passed, 0 failed etc.), then LSan: 360 bytes in 9 allocations, stack autoindex_thread → cbm_pipeline_run → … → cbm_kind_in_set ours, stale base: the in-process auto-index thread never freed its extraction bitset cache; fixed on main by e45d7051 (#2116) after your merge base. Your new test is simply the second thing to exercise that thread. Disappears on rebase.
test-tsan (macos-14, ubuntu x64, ubuntu arm64) mcp_auto_index_in_process_uses_background_worker_policy: selected_workers == 2, expected cbm_default_worker_count(false) == 4 (3 vs 4 on the 4-core runners) the PR's own test, assertion order: the test unsets CBM_WORKERS for the run, then restores the environment, and only then computes the expected value — which now sees the lane's CBM_WORKERS=4. Deterministic; here it fails the other way under CBM_WORKERS=1 (5, expected 1).

No infrastructure flake among them, and the Windows guard was green on your run.

On the change itself, which I read in full on your head and built and ran locally: I like it. Routing every automatic index — in-process auto-index, the supervised worker, the daemon's session auto-index and the watcher re-index — onto the existing "leave headroom" policy through one flag, while an explicit index_repository keeps every core and CBM_WORKERS / CBM_INDEX_SINGLE_THREAD keep their precedence, is the smallest shape that does the job, and reusing the incremental policy rather than inventing a second one is the right instinct. It changes only the worker count fed to the scheduler, never merge order or a work cap, so the graph is unaffected; no delay, no raw fopen, no new allocation site. pipeline 266/0 and daemon_application 52/0 pass here, and with the production hunks neutered your preserves_overrides and the two daemon_application assertions go red, so the tests bind.

Four things before it can merge:

  1. Rebase onto main. The four LSan reds vanish with it. The conflicts are all same-anchor insertion drift against the opt-in resource-policy work (21ae0215) and #713's tests: keep both sides in pipeline.h / pipeline.c and the two test files, and in the three args builders (application_auto_index_args, application_background_index, index_run_supervised_path) chain your _background bool after main's policy call inside its guarded form.
  2. Fix the assertion order in mcp_auto_index_in_process_uses_background_worker_policy: capture int expected = cbm_default_worker_count(false); inside the window where the variables are unset, before mcp_test_restore_env, and assert against that. Your pipeline_background_worker_policy_preserves_overrides already does it this way.
  3. pipeline_background_policy_preserves_index_results is inert. setup_test_repo() writes about three files, below MIN_FILES_FOR_PARALLEL (50), so foreground and background both take the sequential path and the test stays green with the policy removed. Either give it a fixture past 50 files so it really compares an all-core index against a headroom index, or drop it and name the existing determinism test that covers the property.
  4. The description overstates the coverage. Two all-core sites are still unrouted on a background run: the LSP-surface pass (lsp_surface.c, cbm_parallel_for with auto-detected workers) and incremental hashing (pipeline_incremental.c, hash_workers = cbm_default_worker_count(true)). Both have the pipeline in reach; route them through cbm_pipeline_worker_count too, or narrow the description to what is covered. Small naming nit while you are there: _cbm_background matches the existing _cbm_index_policy convention for hidden keys.

One product question I want to be explicit about rather than bury: with this, a user who connects with auto_index=true gets their first graph built on perf_cores-1 workers, so the first index takes a little longer in exchange for a responsive machine. I think that is the right trade and it is the thesis of the PR, but it is a visible change on first connect, and I will confirm it on our side before merging. Thank you for the careful set of tests and for waiting out nineteen days of silence you did nothing to earn.

@mvanhorn
mvanhorn force-pushed the fix/1084-background-index-cpu branch from 84fffef to 7498e2b Compare September 22, 2026 06:56
@mvanhorn

Copy link
Copy Markdown
Contributor Author

Thank you for the attribution, and for taking the time to sort every red into whose it was. That saved me chasing the LSan leak.

All four are done, rebased onto main in 01fba2c6 and the fixes in 7498e2b7.

  1. Rebase. Resolved as you described: kept both sides in pipeline.h / pipeline.c and the two test files, and chained the background bool after main's policy call inside its guarded form in all three args builders. The LSan reds are gone with it.

  2. Assertion order. expected is now captured inside the unset window, before mcp_test_restore_env, the way pipeline_background_worker_policy_preserves_overrides already did it.

  3. The inert test. Gave it a real fixture rather than dropping it: 64 padding files put it past MIN_FILES_FOR_PARALLEL, and I also unset any inherited CBM_WORKERS / CBM_INDEX_SINGLE_THREAD, since a pin would have made both runs pick the same count and left the test inert for a second reason.

  4. Coverage. Routed both sites you named, lsp_surface.c's cbm_parallel_for and pipeline_incremental.c's hash_workers, through cbm_pipeline_worker_count rather than narrowing the description; the pipeline is plumbed into cbm_pipeline_build_semantic_manifest for that. Took the naming nit too, so the hidden key is _cbm_background now.

One thing worth your eye in 4: closure_probe_surfaces previously passed NULL as the pipeline ctx deliberately, so file errors were not recorded twice. I set probe_ctx.pipeline around the call and clear it after, so the surface pass can read the worker policy without changing the error-recording behaviour. If you would rather it stayed NULL and the probe path kept all cores, say so and I will pull that hunk.

On the product question: I agree that is the right trade, and I think it is worth saying in the release note rather than only here, since a slower first index on auto_index=true is the kind of thing people notice and file.

666 passed, 4 skipped locally across the pipeline, mcp and daemon_application suites.

@DeusData DeusData left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for turning all four around so quickly. The assertion order, the 64-file fixture that finally makes the invariance test compare two different worker counts, and routing the two remaining all-core sites are all right, and I verified them on the merge result with today's main (e783f73d): build clean; pipeline 289, daemon_application 56 and mcp 321 passed; memory-core linter unchanged.

Your question about closure_probe_surfaces first, because the answer is that you were right and the hunk should stay. The surface builder reads the context for exactly one thing, the worker count, and the only function it forwards the context to, cbm_pipeline_result_acquire, reads nothing but ctx->spill, which probe_ctx never sets. So that call takes the same branch it took when you passed NULL, and no file error can be recorded through it, twice or at all. What would break the property is leaving probe_ctx.pipeline set for the struct's whole lifetime, because then cbm_parallel_extract above it would see a pipeline. Your set-and-clear around the one call is exactly the shape that avoids that.

Two things before it merges, and the first is the product question you and I discussed:

  1. A project's first index runs at full width. We have decided the trade the other way round for the one case where it matters most. When someone connects and the project has no index yet, they are waiting for a graph that does not exist, and a first build one worker short is the slowest possible moment to make them wait. Once an index exists, a refresh really is background work, and leaving a core free is right. So _cbm_background should take effect only when a committed index for the project already exists; a project's very first index keeps every core even when it was started automatically, and watcher re-indexes and auto-indexes of an already-indexed project keep the headroom. The flag is honoured in two places today, handle_index_repository (reads _cbm_background, then cbm_pipeline_set_background) and the in-process autoindex_thread (sets it directly), so the check belongs there, or in one shared helper both call, as long as the two paths agree. For tests, a pair makes the rule visible: the in-process auto-index of a fresh project asserts full width, and the same auto-index of a project that is already indexed asserts headroom.
  2. The lint job. src/pipeline/pipeline.c:207 still has two spaces before the trailing comment on the new bool background; field where the block wants one. I reproduced it with our own Homebrew clang-format 22.1.8, so it is a genuine violation and not the standalone clang-format-20 drift we have been caught by before; clang-format -i on that file fixes it and touches nothing else.

Everything else stands as reviewed. After those two, this merges. If the Windows guard job goes red on your next run, that is the daemon endpoint cold-start race on our side (#2275), not your change.

@mvanhorn

Copy link
Copy Markdown
Contributor Author

Thanks, that trade makes sense. Both are done in 7a307ab.

  1. _cbm_background now only applies when the project already has a committed index. There's one helper, project_has_committed_index(), and both handle_index_repository (checked after the final project name is set) and autoindex_thread go through it, so the two paths can't drift. A first index keeps every core even when auto-started. The tests come in pairs as you described: fresh in-process auto-index asserts full width and the already-indexed one asserts headroom, plus the same pair for the _cbm_background tool path.
  2. The bool background; comment spacing at pipeline.c:207 is fixed. Homebrew clang-format 22.1.8 is clean on both source files.

make test-focused for mcp and pipeline: 613 passed.

@DeusData

DeusData commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Thank you, @mvanhorn, for four quick and careful rounds on this. The first index keeping full width is exactly what was decided, and the review of 7a307ab is clean.

One last step before it can merge: main has moved a long way (about 215 commits), and the branch now conflicts in two places. Both are same-spot insertions:

Keeping both sides resolves the text in each case. One thing is worth a look while you're there: #2344 changed how handle_index_repository picks the project name for an unnamed re-index (it now reuses the project that owns the root). Your headroom decision reads the pipeline's project name at that same site, so project_has_committed_index() should be asked about the name that is actually being indexed.

Could you rebase onto current main? We'll re-verify the merge result and merge right after. Thanks again for your patience with this one!

mvanhorn and others added 3 commits October 2, 2026 14:03
Fixes DeusData#1084

Signed-off-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Addresses the four blockers from review:

- Capture the expected worker count while CBM_WORKERS is still unset in
  mcp_auto_index_in_process_uses_background_worker_policy, so the assertion no
  longer reads the lane override restored just before it.

- Give pipeline_background_policy_preserves_index_results a fixture past
  MIN_FILES_FOR_PARALLEL and drop any inherited CBM_WORKERS /
  CBM_INDEX_SINGLE_THREAD pin, so it really compares an all-core index against
  a headroom index instead of two sequential runs.

- Route the two remaining all-core sites through cbm_pipeline_worker_count:
  the LSP surface pass and incremental manifest hashing. The pipeline is
  plumbed into cbm_pipeline_build_semantic_manifest for that.

- Rename the hidden key _background to _cbm_background, matching the existing
  _cbm_index_policy convention, across every producer and consumer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Du315CaKLufAYPEdEwcocq
Signed-off-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Signed-off-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
@mvanhorn
mvanhorn force-pushed the fix/1084-background-index-cpu branch from 7a307ab to 26eece5 Compare October 2, 2026 23:16
@mvanhorn

mvanhorn commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

Rebased onto current main and pushed (26eece5). Both conflicts kept both sides: cbm_pipeline_export_error() and cbm_pipeline_set_background() now sit one after the other in pipeline.c, and index_root_owner_append() / index_root_owner_resolve() sit ahead of project_has_committed_index() in mcp.c.

On the unnamed re-index: handle_index_repository swaps in the root owner's name before cbm_pipeline_set_project_name(), and the headroom check reads cbm_pipeline_project_name(p) after that, so project_has_committed_index() is asked about the reused project, not the path-derived one. No change was needed there.

I ran make -f Makefile.cbm test-par on the rebased branch: 8327 passed, 0 failed, 10 skipped.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working priority/high Needs near-term maintainer attention; high-impact bug, regression, safety issue, or release blocker. stability/performance Server crashes, OOM, hangs, high CPU/memory

Projects

None yet

Development

Successfully merging this pull request may close these issues.

CPU memory usage is too high

3 participants