Repository navigation
fix(graph-buffer): log why an atomic publish failed instead of a successful dump - #2051
junk151516 wants to merge 1 commit into
Conversation
|
Thanks for opening this — it has been seen, and it is queued. This note is automated, but it is not a brush-off: it exists so you know where your PR stands instead of having to guess from silence. Current review status: working through a backlog. What that means for this PR, concretely:
Things that will genuinely speed it up whenever review does happen:
If this fixes a bug, a reproduction we can run is worth more than a description of the symptom. Thanks for contributing, and sorry in advance for the wait. |
|
Thanks a lot for this one — it is exactly the caller #1628 didn't reach, and the write-up made it easy to check. I built Blocking
Please add
Minor
One honest limit from my side: I reviewed on macOS, so the Windows |
…essful dump cbm_gbuf_dump_to_sqlite() emitted gbuf.dump regardless of the writer's result, so a run that published nothing still logged node and edge counts as if it had worked, and the failure reached the user only as the generic "Pipeline failed. Check repo_path exists and contains source files." The writer's temp -> final rename in publish_writer_output() returned a bare ERR_WRITE_FAILED. DeusData#1628 taught cbm_rename_replace() to translate the platform error into errno so callers could report why, and wired up the stage -> final rename in pipeline.c. This is the other rename on that path, and it is the one that runs first. Preserve errno across the cleanup in every publish error branch. That includes cbm_writer_open(), where the cleanup unlinks a file that was never created: without the save it leaves ENOENT behind, so a directory that denied the create is reported as a missing path — the DeusData#1620 case, stated wrongly. Read the value in graph_buffer.c before the profiling macro can clobber it, and report gbuf.dump_failed with the return code and that errno. A sticky append failure clears errno rather than attach a reason it does not have, and the field is omitted when there is none. No logging is added inside internal/cbm/, which has none today. Tests cover both halves: the dump reports a failure instead of a summary, the writer names its truncation reason, and the publish rename's errno survives a cleanup that itself fails. Refs DeusData#1620, DeusData#2001. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012FxkVF773gDGNmBmsLnzN7 Signed-off-by: juan Carlos Estrada Montoya <junk151516@users.noreply.github.com>
e2195dc to
1067e51
Compare
|
Thanks — (2) is the one that mattered, and you are right that it points the wrong way in exactly the case this PR exists to explain. All six are addressed; measurements below. 1. DCO. Signed off and force-pushed. 2. The The truncation branch above it now sets 3. Test that binds the writer half. Added two, and you were right that the suite was green without them.
One limit, stated rather than hidden: that test is 4. The 5. 6. Verification
On the Windows branch you flagged You mentioned reviewing on macOS, so the |
|
Approved on merit. You had the last word here on 2026-09-04 and got sixteen days of silence in return — sorry about that. This is a small diff fixing a failure mode that is disproportionately annoying in production: the cleanup path destroying the evidence of the failure it is cleaning up after.
The detail I want to credit is the comment:
That is the part people get wrong. It is tempting to assume Preserving the translated error from I also like that you named the truncation cases explicitly, where neither condition sets Queued to merge behind #2248, which fixes an unrelated failing test on |
|
A status note so this does not look stalled after today's approval. When I went to merge, the actual merge attempt found one conflict in Rather than ask you to rebase after we had already kept this waiting sixteen days, I carried the rebase myself as #2259 — your commit, your authorship, your sign-off, unchanged apart from that one resolution. It builds clean and passes 421 tests on the rebased result. It merges as soon as an unrelated linter regression on If you would rather rebase this branch yourself and have it merge from here instead, say so and I will close #2259 — either way is fine, and either way the fix is yours. |
|
Your change is on Closing this PR as superseded by that carry, not because anything was wrong with it — the opposite. The defect you fixed is the kind that costs hours when it bites: the cleanup path overwriting Verified before landing: Thank you for the fix, and for your patience through sixteen days of silence that you did nothing to deserve. |
|
Thank you — for carrying the rebase yourself, for keeping the authorship intact, and for the care in every reply on this thread. No need to apologize for the wait; being kept informed the whole way through was more than enough. We use codebase-memory-mcp daily across our projects and we're genuinely happy with it. Glad this small fix could give something back. |
Summary
A failed atomic publish leaves no trace in the log, and the one record it does emit says the dump
succeeded.
cbm_gbuf_dump_to_sqlite()collects the return code ofcbm_writer_finalize()and then callslog_dump_summary()unconditionally, so a run that published nothing still logsmsg=gbuf.dump nodes=… edges=…at INFO. The failure itself carries no reason: the writer'stemp → finalrename inpublish_writer_output()returns a bareERR_WRITE_FAILED. What theuser is left with is the generic
which sends them to look at their repository when the repository was never the problem.
This is the caller #1628 did not reach. That PR taught
cbm_rename_replace()to translate theplatform error into
errnoprecisely so callers could report why — its comment says "callers logerrnoafter a failed rename" — and wired up thestage → finalrename inpipeline.c, whichlogs
finalize.rename_failed. There are two renames on the publish path. The writer's own is stillsilent, and it is the one that runs first.
Refs #1620, #2001.
What this does not do
It does not fix #1620. On that host something below the DACL shape denies the rename, and no
amount of logging changes that. What it does is turn a silent failure into a stated one, which is
the obstacle both #1620 and #2001 describe when they say the log cannot tell them what went wrong.
Changes
internal/cbm/sqlite_writer.c— carry the failure reason across cleanup.publish_writer_output()callscbm_unlink()to drop the temp file after a failed rename, anddiscard_writer_output()does the same after a failed sync. A library call is free to overwriteerrnoeven when it succeeds, so the translated value could not be relied on by the time thecaller saw it. Each publish-path error branch now saves
errnobefore cleanup and restores itafter. No logging is added here: nothing in
internal/cbm/includesfoundation/log.h, and thischange does not make it the first.
src/graph_buffer/graph_buffer.c— report the failure instead of a summary.errnois captured immediately aftercbm_writer_finalize()returns, before the profiling macrothat follows it can clobber the value. When the writer failed,
gbuf.dump_failedis emitted atERROR with the return code and that errno; the success record is emitted only on success. The early
return for a writer that never opened logs the same event, so "the dump produced no database" has
exactly one signature in the log.
tests/test_graph_buffer.c— regression test.gbuf_dump_failure_logs_reasonpublishes onto an existing directory, which is howsw_publish_failure_preserves_destination_sidecarsalready forces this failure, and asserts thatgbuf.dump_failedis logged with a non-zeroerrnoand that the success record is absent.Verification
Built and run locally on Windows x64 under MSYS2/CLANG64 (clang 22.1.8), ASan+UBSan, via
scripts/test.sh --suites … CC=clang CXX=clang++:graph_bufferwithgraph_buffer.creverted to maingbuf_dump_failure_logs_reason:strstr(logs, "gbuf.dump_failed") is NULLgraph_buffer+sqlite_writerwith the changesw_publish_preserves_live_reader,SKIP_PLATFORMon Windows)clang-format --dry-run -Werroris clean on all three files.Notes for review
cbm_progress_sink_fn()dispatches on an exactstrcmpof themsgfield, sogbuf.dump_faileddoes not fall into the
gbuf.dumphandler. The visible effect on the CLI is that a failed run nolonger reports node and edge counts for a database it did not publish.
CBM_LOG_LEVEL=debugand still seeing nothing, which this addresses at any level.🤖 Generated with Claude Code
https://claude.ai/code/session_012FxkVF773gDGNmBmsLnzN7