Skip to content

fix(pipeline): expose unresolved call coverage - #2305

Open
pcristin wants to merge 9 commits into
DeusData:mainfrom
pcristin:fix/issue-2265-unresolved-call-coverage
Open

pcristin wants to merge 9 commits into
DeusData:mainfrom
pcristin:fix/issue-2265-unresolved-call-coverage

Conversation

@pcristin

@pcristin pcristin commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

What does this PR do?

Fixes #2265.

Records unresolved call sites as per-file index coverage, including caller, method leaf, source byte span, one-based line, reason, and the resolver's considered candidate QN. check_index_coverage exposes those records. Incremental indexing preserves them, and CALLS traces report unknown when relevant unresolved sites or older metadata prevent an exact total.

The review amendment reconciles duplicate records by exact call site and removes a diagnostic when that site actually emitted a CALLS edge through any resolver strategy, including route registration. Coverage capture happens while the resolved result is still live, including spilled results. Inbound matching uses the resolver's candidate QN; nested outbound calls use the extractor's enclosing caller. Indexed unresolved files stay out of skipped[]. The JSON collector uses the core allocator.

Nested callers remain distinct from their enclosing functions; the final indexed fixture reports inner unknown and outer eq when the unresolved name matches an in-project method. Existing indexes need reindexing for coverage metadata version 5. The deeper TypeScript LSP attribution fix remains separate.

The CLI fixture lifetime fix was submitted separately in #2304 and has merged.

Outbound project-symbol amendment

The latest maintainer-requested change limits outbound unknown to unresolved callee names matching non-container symbols in the same project. File, Folder, Project, Module, Package and Section name collisions do not count. Class/constructor symbols remain eligible. Inbound candidate matching and recorded coverage are unchanged; failed lookups remain conservative.

Regressions cover external-only names, another project's name collision, case-sensitive matching, constructors, all six containers, and nested caller attribution with Client.buscar defined but an untyped receiver.

The branch is rebased onto upstream 80eb92a7; both conflicting pipeline suite registrations were retained. PR #2294 remains open. The maintainer owns the separate qualified-name fallback leak fix and full-kernel measurement; the branch can pick up the leak fix once it lands.

Current validation:

  • Final MCP suite: 324 passed, 0 failed, 4 platform skips, exit 0. Pipeline, incremental, store_nodes and store_search: 609 passed, with no assertion failures in those suites. These are 933 passing tests across focused runs; the corrected MCP fixture was rerun separately. Local ASan/UBSan enabled, detect_leaks=0.
  • Final changed-line clang-tidy 20 and full CI lint with clang-format 20 / cppcheck 2.20 passed.
  • Complete make -f Makefile.cbm security passed, including MCP robustness 32/32.
  • Independent review accepted the final production patch and corrected nested fixture.
  • Default scripts/test.sh stopped at the unchanged Step 0x Scoop version-metadata contract. A clean complete default gate is not claimed. The prior full-harness result below applies to the previous published snapshot.

Refreshed Prometheus v3.0.0 (c5d009d57fcccb7247e1191a0b10d74b06295388), full mode, 2 workers, 4096 MiB budget, no forced spill: 23.396 s, 556016 KiB peak worker RSS; 505 unresolved rows / 19006 sites / 3758665 JSON bytes. All 400 sampled function/direction pairs match the previous sample exactly: 43/400 unknown, comprising 1/200 inbound and 42/200 outbound (previously 1 inbound / 121 outbound), with zero errors. The upstream rebase also differs between runs, so no isolated timing improvement is claimed.

Current production binary SHA256: 726abe090c53728a603d4c1a60ed7ca72be372a09369c7a741da5f31f727ee09.

Previous reconciliation validation (published 853192aa)

  • Baseline regressions failed before the fix: 612 passed, 4 failed, 4 platform skips. Failures covered unrelated same-name inbound matching and the sequential, parallel, and spill pipeline cases.
  • Final ASAN_OPTIONS=detect_leaks=0 scripts/run-tests-parallel.sh build/c/test-runner 2: 8222 passed, 0 failed, 8 platform skips, all 144 suites accounted for by the union guard. ASan and UBSan enabled; LeakSanitizer disabled.
  • Focused MCP, pipeline, and incremental suites: 780 passed, 4 platform skips, exit 0. Regressions include 61-file parallel indexing, forced spill, repeated calls on one line, nested caller attribution, direct imports, and route registration.
  • make -f Makefile.cbm lint-tidy-diff CC=clang-20 CLANG_TIDY=clang-tidy-20: exit 0. Full scripts/lint.sh --ci with clang-format 20 and cppcheck 2.20.0, two cppcheck jobs and a build cache: exit 0. Formatting, NOLINT, no-skips policy, and memory-core checks passed.
  • Security static, binary-string, UI, install, network, and vendored integrity stages passed. The initial fuzz run had startup timeouts under concurrent load; its serialized retest passed 32/32. Parent/worker watchdog, TLS budget, worker error transport, watcher-disabled behavior, and worker request-scope guards passed.
  • Independent review identified the route registration emission gap; both emitters were corrected. Follow-up review found no actionable defects.

Gate limitations

The default scripts/test.sh stops at the unchanged Scoop version-metadata contract: its 0.11.0 pin conflicts with the newest stable v0.11.0 tag. The full C harness was run separately as described above. Existing UBSan null-argument warnings remain in the ObjectScript scanners, semantic-edge code, and project-listing qsort(records, 0, ...) path. A clean complete default gate is not claimed.

Previous pinned corpus measurements (853192aa)

Host: Linux x86_64, 4 reported CPUs, 7941 MiB RAM, no swap. Peak worker RSS uses /proc VmHWM sampled every 100 ms, rather than the CLI process's /usr/bin/time RSS. No baseline overhead comparison is claimed.

Prometheus v3.0.0, c5d009d57fcccb7247e1191a0b10d74b06295388, full mode, 2 workers, 4096 MiB budget: 31.052 s, worker RSS 552900 KiB; 505 unresolved rows / 19006 sites, 3758665 JSON bytes (3782469 logical path/kind/detail bytes excluding SQLite overhead and project keys). A sample of 200 equally spaced Function/Method QNs in lexical order, depth 1, CALLS only, both directions: 122/400 unknown (1 inbound, 121 outbound), zero query errors. This final-binary run followed the C harness.

Full Linux v6.12, adc218676eef25575469234709c2d87185ca223a, 81924 files: the initial 2-worker/4096 MiB run was OOM-killed after 1881.23 s, at 3522632 KiB worker RSS. Forced-spill runs with 2 workers and 2048/3072 MiB budgets stopped at the memory guard after 173.280/848.803 s, at 2939752/4539200 KiB. The forced-spill runs published no partial graph. A completed full-kernel index, coverage row sizes, and trace share remain unavailable; that measurement needs a larger machine. Linux timings overlapped verification jobs.

Final production binary SHA256 for Prometheus and the forced-spill Linux runs: b3035597e242263093c82d41177e3c2419f797537b67dd3805f82f3b2e0d23e9.

Disclosure and checklist

OpenAI Codex assisted with investigation, implementation, testing, and PR preparation. Pcristin reviewed and personally certified the signed-off contributions under DCO 1.1.

  • DCO checker passes for the contribution range
  • New behavior covered by MCP and pipeline regressions
  • Previous published snapshot passed the full 144-suite C harness
  • Current affected suites pass: 324 MCP tests plus 609 pipeline/incremental/store tests
  • Full 144-suite harness repeated for the current rebased amendment
  • Full CI lint and changed-line clang-tidy pass
  • Complete default test script passes (unchanged packaging contract blocks it)
  • Completed full-kernel coverage size and trace measurements (larger machine needed)

@pcristin
pcristin requested a review from DeusData as a code owner September 24, 2026 20:45
@github-actions

Copy link
Copy Markdown

Thanks for opening this — it has been seen, and it is queued.

This note is automated, but it is not a brush-off: it exists so you know where your PR stands instead of having to guess from silence.

Current review status: working through a backlog. 0.9.1-rc.1 is out, so the release freeze that held reviews is over — but it left a large queue of open pull requests behind it, and we are reading through them oldest-first. The background is in discussion #1144.

What that means for this PR, concretely:

  • It will not be closed for inactivity. No stale bot touches pull requests here.
  • It may still sit a while before a human reads it. That is on us, not on you.
  • Older PRs are read first, so a recent one is not being skipped — it is behind a queue.

Things that will genuinely speed it up whenever review does happen:

  • Keep it rebased on main — the tree is moving quickly right now, and a conflicting branch cannot be reviewed as the diff you intended.
  • Get CI green, or say which failures you believe are pre-existing.
  • Keep the change to one claim. Bundled features and refactors get split before they get merged, which costs you a round trip.
  • Every commit needs a sign-off (git commit -s) — CI enforces DCO.

If this fixes a bug, a reproduction we can run is worth more than a description of the symptom.

Thanks for contributing, and sorry in advance for the wait.

@BryanQuiceno

Copy link
Copy Markdown

Thank you @pcristin, both for picking this up so quickly and for keeping it focused on the misleading zero. That was exactly the part that hurt us. I built the PR head and ran it against the repro from #2265, plus one extra case. The inbound side works well. I found three rough edges and traced each one to the source, so I hope this saves you some time. Please take whatever is useful.

All line links below point to the PR head 69b84ca.

How I tested

Linux x86_64, gcc, scripts/build.sh on the PR branch. I used a throwaway HOME so it wouldn't touch my regular index:

export HOME=$(mktemp -d)
B=./build/c/codebase-memory-mcp
R=/path/to/repro            # the 9 files from #2265 (+ anidada.js below)
echo "{\"repo_path\":\"$R\",\"mode\":\"full\"}" | $B cli index_repository
P=<project name printed above>
q() { echo "{\"project\":\"$P\",\"function_name\":\"$1\",\"direction\":\"$2\",\"format\":\"json\"}" | $B cli trace_path; }

What works

query before with #2305
q buscar inbound 0, eq 0, unknown ✅
q buscarPlano inbound 0, eq 0, unknown ✅
q procesarConClase outbound (class method) 0, eq 0, unknown ✅

check_index_coverage with paths: ["servicio.js"] returns status: "unresolved_calls" and recommended_action: "read_source_and_verify_calls", plus caller, leaf, span and reason. For an agent, that is the right nudge: go read the source before trusting the count. 🙌


1. Outbound: nested functions inside a factory still report an exact zero

Repro (anidada.js, added next to the original files):

import { ayudante } from "./ayudante.js";

export function crearAnidada({ cliente }) {
  function interna(id) {
    ayudante(id);              // resolved → CALLS edge from `interna`
    return cliente.buscar(id); // unresolved → recorded with caller `crearAnidada`
  }
  return { interna };
}
q interna outbound      → callees_total: 1, callees_total_relation: "eq"       ❌ (buscar is missing, but it says exact)
q crearAnidada outbound → callees_total: 0, callees_total_relation: "unknown"  (the site isn't really its own)

The original repro shows the same thing: procesar, procesarTipado and procesarPlano (outbound) all stay at 0 / eq.

Why: the two passes disagree on who the caller is.

  • The extractor attributes a call to the innermost function: cbm_enclosing_func_qn → cbm_find_enclosing_func (helpers.c#L1187-L1233). That gives …anidada.interna, which is the node trace_path looks up.
  • The TS LSP walk only sets enclosing_func_qn in process_function_body (ts_lsp.c#L3652), i.e. for top-level functions and class methods. For nested function_declaration / function_expression / arrow_function, the branch at ts_lsp.c#L3551-L3574 pushes a fresh scope but keeps the outer enclosing_func_qn ("Nested functions just get a fresh scope"). So ts_emit_unresolved_call_at (#L290) stores caller = …anidada.crearAnidada.
  • The new outbound check in trace_path matches on that caller QN (mcp.c#L9426). interna is never found, so the relation stays eq. Inbound works because it matches on leaf, which doesn't depend on the caller.

I checked this with a temporary fprintf inside ts_emit_unresolved_call_at: every site in the repro is emitted with the outer factory as caller.

Possible directions (you know the codebase much better, so these are only ideas):

  • (a) In the nested-function branch, set enclosing_func_qn to the same QN the def walk produces for that node, and restore it on exit. That would also help the resolved ts_emit_resolved_call_at paths (#L254, #L274) join with the extractor's caller, the same kind of mismatch the comment at helpers.c#L1198-L1205 describes for nested classes. Wider blast radius, though.
  • (b) Keep this PR narrow: in trace_path, match an unresolved site to the traced function by file + source range (the site's span inside the function's range) instead of by caller QN.

This is the case that made me open #2265. Our backend builds services with factory functions and injected dependencies, so almost every method we'd trace outbound is a nested function.

2. A resolved call is also recorded as unresolved, twice

In directo.js, usarDirecto → ayudante resolves (heuristic, 0.95), but coverage also lists it as unresolved, twice, with the same span:

[{"caller":"…directo.usarDirecto","leaf":"ayudante","start_byte":84,"end_byte":95,"reason":"import_symbol_not_in_registry"},
 {"caller":"…directo.usarDirecto","leaf":"ayudante","start_byte":84,"end_byte":95,"reason":"import_symbol_not_in_registry"}]

As a result, q ayudante inbound and q usarDirecto outbound report unknown even though the count is complete.

Why: the TS LSP runs twice per file, and each run gives a different callee QN for the same site. My trace shows:

per-file  (cbm.c#L2324 cbm_run_ts_lsp)            callee=./ayudante.js.ayudante                    reason=import_symbol_not_in_registry
cross-file (pass_lsp_cross.c#L1356 cbm_run_ts_lsp_cross) callee=<project>.ayudante.ayudante        reason=import_symbol_not_in_registry

pxc_append_results dedupes on kind + caller + callee + span + origin (pass_lsp_cross.c#L1090). The callee strings differ, so both records survive. cbm_pipeline_record_unresolved_calls then reduces the callee to its leaf (pipeline.c#L466-L470), and the two become identical rows. The method-call sites (buscar) are emitted twice too, but with the same callee string, so they are deduped correctly.

Possible fixes:

  • Dedupe by (caller, leaf, start_byte, end_byte) when building the coverage rows. Cheap and local to this PR.
  • Skip unresolved sites whose span already produced a CALLS edge by another strategy. Otherwise, in a real TS/JS repo nearly every file with an import may end up flagged, and the signal loses its value.

(Separately, and maybe worth its own issue: the cross-file pass has the correct module QN <project>.ayudante.ayudante and still reports import_symbol_not_in_registry for a plain one-argument exported function. I didn't dig into why.)

3. index_repository lists these files as skipped

"skipped_count": 5,
"skipped": {"files":[{"path":"clase.js","reason":"[{\"caller\":…}]","phase":"unresolved_calls"}, …]}

These files were indexed. add_skipped_summary only filters parse_partial/parse_unusable through is_parse_coverage (mcp.c#L10195), so unresolved_calls falls into skipped[]. The comment right above that function says it well: "a reader who sees a file there believes it is absent from the graph entirely." Adding unresolved_calls to that predicate, and perhaps an unresolved_calls_count in the summary, would match what you already did in the coverage summary.

Minor

A 1-based line next to start_byte/end_byte would save agents a conversion step when they go to read the source.


A proposal, if it works for you

I'd like to help with code, not only with reports:

  • (2) and (3): both are small and live inside this PR's scope. If you're OK with it, I'll open a PR against your branch (fix/issue-2265-unresolved-call-coverage) with those two fixes and tests, signed off. You can take them, adapt them or ignore them, so this PR stays yours and stays one claim.
  • (1): the nested-function attribution lives in the TS LSP walk and predates this PR, and it likely affects resolved LSP calls too. To keep fix(pipeline): expose unresolved call coverage #2305 focused, I'd open a separate issue for it with this repro and follow with its own PR, unless you'd rather handle it here.

Either way, I'm happy to re-run the same repro on any follow-up commit. Thanks again for the careful work on this!

@DeusData DeusData left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you. This is the follow-up we agreed on in #1682, and persisting unresolved call sites so coverage and trace_path can admit what they do not know is exactly the direction we want. We accept the "unknown" relation value.

@BryanQuiceno, thank you too for that test report. You traced each rough edge to the line that causes it, and that saved us real time. Your three findings overlap with what we were going to ask for, so here they are in one list.

Before it merges:

  1. Remove the false "unknown"s. Two sources drain the signal:

    • Bryan's finding 2: one call site that resolved (usarDirecto → ayudante) is also recorded as unresolved, twice, because the per-file and cross-file TS passes use different callee strings.
    • Inbound matching on any short name: any unresolved site whose leaf matches a visited node's short name marks that trace "unknown". Common names (get, run, init) would then almost always read "unknown".

    Please dedupe coverage rows on (caller, leaf, start_byte, end_byte), drop an unresolved site whose span already produced a CALLS edge, and link an unresolved site to a traced node only through the caller or a leaf the resolver would actually have considered for it.

  2. Outbound on nested functions must not claim an exact zero (Bryan's finding 1). A factory-built function (crearAnidada → interna) still reports callees_total_relation: "eq" while a call is missing. Correcting that "eq" is what this PR is for. Matching the site to the traced function by file and source range, as Bryan suggests in (b), keeps this PR narrow. The deeper fix in the TS LSP walk, setting the enclosing function for nested functions, affects resolved calls too and deserves its own issue and PR, as Bryan proposes.

  3. Files with unresolved calls must not appear in skipped[] (Bryan's finding 3). add_skipped_summary should treat unresolved_calls like the parse-coverage phases, since those files were indexed.

  4. Real-corpus numbers. Our rule is that a change is shown working on real input before it merges. Please report:

    • index time
    • peak RSS
    • the number and total size of unresolved_calls rows
    • the share of trace_path calls that come back "unknown"

    Please measure on a Go repository and on the Linux kernel, taken after items 1–3, so they show the signal we will actually ship.

  5. A parallel-path test. The two tests use two files, which is below MIN_FILES_FOR_PARALLEL (50), so only the sequential path runs. Please add a fixture past 50 files, ideally one that forces a result spill, to cover the path most real indexes take.

  6. yyjson_mut_doc_new(NULL) in cbm_pipeline_record_unresolved_calls uses libc's allocator. Please pass the core-backed allocator the rest of the PR already uses.

Bryan has offered to open a PR against your branch with items 1 (dedupe part) and 3, plus tests. That is welcome from our side, and it is your call whether to take it. The 1-based line next to the byte offsets is a good small addition if you have room.

We will call out one behaviour in the release notes on our side: every index built before upgrading reads "unknown" on call traces until it is re-indexed. Also a heads-up: #2294 removes is_test_file() from mcp.c, which this PR calls in new places, so whichever lands second will need a small rebase.

Thank you both. This closes a real honesty gap in the tools.

@DeusData

Copy link
Copy Markdown
Owner

Thank you for keeping the branch current, @pcristin. It saves the next rebase a lot of pain. One heads-up so it doesn't surprise you: #2294 removes is_test_file() from mcp.c, and this branch adds two new calls to it in trace_path. Git merges the two cleanly, but the result won't compile. If #2294 lands first, those two calls become trace_node_is_test(node), and that's the only change needed. Our review from the 25th still describes what we need before merging: the dedupe, nested-function outbound, skipped[], corpus numbers, the parallel-path test and the allocator. Take your time, and thank you again for closing this honesty gap.

@pcristin

Copy link
Copy Markdown
Contributor Author

AI-assisted reply on behalf of Pcristin.

I’ve reworked coverage capture around the actual call sites. Per-file and cross-file duplicates are collapsed, and a diagnostic is dropped when that exact site emitted a CALLS edge. Spilled results are captured inside the resolver worker before they’re freed.

The nested procesar example now reports outbound unknown, while crearServicio stays eq. Inbound checks use the resolver’s candidate QN, so an unrelated buscar no longer affects the result. Indexed files with unresolved calls also stay out of skipped[].

Added sequential, 61-file parallel, and forced-spill regressions, including repeated calls on one line and route registrations. The JSON collector uses the core allocator. Existing indexes need reindexing for coverage version 5.

The focused sanitizer suites passed 780 tests. Final changed-line clang-tidy 20 and full CI lint passed; the serialized security fuzz run passed 32/32. The upstream C harness passed all 144 suites: 8222 passed, 0 failed, 8 platform skips.

Prometheus v3.0.0 (c5d009d57fcc) indexed in 31.052 s with 552900 KiB peak worker RSS. It stored 505 unresolved rows / 19006 sites, totaling 3758665 JSON bytes. A sample of 200 equally spaced callable QNs in lexical order, depth 1, CALLS only, both directions returned 122/400 unknown traces: 1 inbound and 121 outbound, with no query errors.

On this 4-CPU, 7941 MiB, no-swap machine, full Linux v6.12 (adc218676eef, 81924 files) could not finish. The first run was OOM-killed after 1881.23 s at 3522632 KiB worker RSS. Forced-spill runs with 2048 and 3072 MiB budgets stopped at the memory guard after 173.280 and 848.803 s, reaching 2939752 and 4539200 KiB. No partial graph was published, so kernel row sizes and trace share remain unmeasured. That remaining check needs a larger machine. The Linux timings overlapped verification jobs; the final Go run followed the test harness. No baseline overhead comparison is claimed.

The default test script remains blocked by the unchanged Scoop version-metadata contract. Existing UBSan warnings in ObjectScript scanners, project listing, and semantic-edge code also remain; the full gate is not claimed as clean.

@DeusData

Copy link
Copy Markdown
Owner

Thank you, @pcristin. This rework lands nearly everything we asked for:

  • Coverage now keys each site on the extractor's own occurrence and drops sites that emitted CALLS.
  • The caller is the innermost function.
  • Indexed files stay out of skipped[].
  • The core allocator is used throughout.
  • The 61-file parallel and forced-spill cases cover the path real indexes take.

Thanks, too, for the Prometheus numbers.

The one red is our bug, not yours. test-diag, test-lsan-macos and test-unix 2/3 show a 1 KB LeakSanitizer report from trace_path's qualified-name fallback on main. cbm_store_find_nodes_by_name allocates its array even for zero rows, and the fallback's malloc overwrites it. Your new trace test is simply the first to reach that path. We'll fix it on main in our own PR with a regression test, and once it lands, updating this branch clears the red. No change is needed from you there.

One contract decision from the maintainer. On your sample, 121 of 200 outbound traces come back "unknown", because every unresolved call site counts, including stdlib and third-party calls. We'd like outbound "unknown" to count only unresolved sites whose callee name matches an in-project symbol. Then "unknown" means "there may be a project edge we missed", not "this function calls something outside the repository". Inbound is fine as is.

We'll take the kernel measurement on our side, since the corpus is here. Also, #2294 is still open: if it lands first, your one is_test_file call in trace_path becomes trace_node_is_test(node). Thanks again for closing this honesty gap.

Signed-off-by: Pcristin <xxxokzxxx@protonmail.com>
Signed-off-by: Pcristin <xxxokzxxx@protonmail.com>
Signed-off-by: Pcristin <xxxokzxxx@protonmail.com>
@pcristin
pcristin force-pushed the fix/issue-2265-unresolved-call-coverage branch from 853192a to 8d80be7 Compare September 29, 2026 09:10
@pcristin

Copy link
Copy Markdown
Contributor Author

I’ve limited outbound unknown to unresolved names that also occur as symbols in the same project. External-only names and structural containers no longer count; classes still do. Inbound candidate matching is unchanged.

Added regressions for external names, another project's name collision, case-sensitive matching, constructors, and all six container labels. The nested-function fixture now defines a Client.buscar method, so it continues to check inner unknown versus outer eq under the new contract.

The branch is rebased onto 80eb92a7. Prometheus v3.0.0 (c5d009d57fcc) indexed in 23.396 s with 556016 KiB peak worker RSS. Coverage remains 505 rows / 19006 sites / 3758665 JSON bytes. The same 400-query sample returned 43 unknown traces: 1 inbound and 42 outbound, with no errors. The previous sample was 1 inbound / 121 outbound. These runs also differ by the upstream rebase, so I’m not claiming an isolated timing improvement.

Clang-tidy 20, full CI lint, and the full security audit passed (MCP robustness 32/32). The final MCP suite passed 324 tests with four platform skips; pipeline, incremental and the two affected store suites passed another 609 tests. The default test command still stops at the unchanged Scoop version-metadata contract, and the local sanitizer suites use detect_leaks=0.

I’ll pick up your separate fallback-leak fix when it lands. Thanks for taking the kernel measurement.

@pcristin
pcristin requested a review from DeusData October 1, 2026 16:22
timothybrush pushed a commit to timothybrush/codebase-memory-mcp that referenced this pull request Oct 1, 2026
…llback

handle_trace_call_path first looks the function up by bare name
(cbm_store_find_nodes_by_name). find_nodes_generic allocates its
initial 16-slot array before stepping the statement and hands it back
even when zero rows match. When the bare name misses and the
qualified_name fallback then resolves the node, the handler overwrote
`nodes` with a fresh one-element array, leaking the empty 1 KB array
on every trace_path call made with a qualified name - for the whole
lifetime of the daemon. The malloc-failure branch of the fallback lost
the same array.

Release the zero-row array before the fallback replaces the pointer.
The other find_nodes_* callers (get_code_snippet outline, search_code
grep classification, detect_changes seeding, Cypher label scans)
already free the container on zero rows; the not-found path in this
handler did too.

Proof (macOS leak lane, Homebrew clang 22.1.8,
`make -f Makefile.cbm test-lsan LSAN_SUITES=mcp`):
- with the new test, without the fix: 322 passed, then
  "ERROR: LeakSanitizer: detected memory leaks - Direct leak of 1024
  byte(s) in 1 object(s) allocated from find_nodes_generic store.c:2771
  <- handle_trace_call_path mcp.c:9176 <-
  test_tool_trace_call_path_qn_fallback_frees_name_miss", exit 1
  (the only leak in the suite).
- with the fix: 322 passed, 4 skipped, no leak report, exit 0.

The new test pins its own precondition (the bare-name lookup of the
qualified name returns zero rows) and that the trace resolves through
the fallback, so it keeps exercising this path. LSan (default-on under
Linux ASan, test-lsan on macOS) turns any regression red.

Found while attributing CI on DeusData#2305.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

3 participants