perf(decay): write back only the rows whose decay_score actually moved - #47
Merged
Merged
Conversation
runDecayCycle recomputed decay_score for every active row and then rewrote
all of them: one full-table transaction per run. On a ~10k-row library run
every 30 minutes that is 48 full-table transactions a day, roughly 34 MB of
WAL each, nearly all of it rewriting values that had not changed.
Each computed score is now compared with the value stored in the row, and
only rows that differ by at least EPS (1e-4) are updated. The baseline is
the stored value, not the previous cycle's computation: against the stored
value the error stays bounded by EPS, against the previous computation it
would accumulate from cycle to cycle.
Two cases are always written regardless of EPS:
- a move that crosses a band edge (0.7 / 0.3 / 0.1). Readers compare against
those edges -- the cold-pool gate is decay_score >= 0.3 -- so skipping a
0.30001 -> 0.29996 move would leave the row eligible while the returned
distribution already counts it as low;
- a non-finite score, so a bad computation still fails on the NOT NULL
constraint as before instead of being silently skipped.
Measured on a copy of a 10,217-row library (before the band-edge rule,
which only adds rows that cross an edge by less than EPS):
pass 1: wrote 351, skipped 9,866 (96.6% fewer writes), 10,217 scanned
pass 2: wrote 0, skipped 10,217
1,822 of those rows (17.8%) were bit-identical to their stored value --
every one pinned at the 1.0 cap by min(1, ...), so they stay unchanged
while capped but were rewritten on every run.
Nothing relies on decay_score having just been written: the recall
multiplier, the cold-pool gate (decay_score >= 0.3), the maintenance
listings and memory-health's stale-decay check all read the value, and
trg_mem_fts_update only fires on UPDATE OF content, summary, tags, so a
decay_score write never touched FTS.
API note: `processed` now counts rows written, and the result also carries
`skipped` and `scanned`. A dry run is unchanged (processed = rows scanned; sample entries are still
{ rowid, score }).
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: 千夏 <qianxia@clawgamers.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
runDecayCyclenow writes back only the rows whose recomputeddecay_scorediffers from the stored value by at least 1e-4 (EPS). Previously every active row was rewritten on every run.Two cases are always written regardless of EPS:
decay_score >= 0.3), so skipping a 0.30001 → 0.29996 move would leave the row eligible while the returned distribution already counts it as low;Why
Each run was one full-table transaction. On a ~10k-row library run every 30 minutes that is 48 full-table transactions a day, roughly 34 MB of WAL each, nearly all of it rewriting values that had not changed.
The comparison baseline is the stored value, not the previous cycle's computation: against the stored value the error stays bounded by EPS, against the previous computation it would accumulate from cycle to cycle.
Measurements
On a copy of a 10,217-row library (before the band-edge rule, which only adds rows that cross an edge by less than EPS):
1,822 rows (17.8%) were bit-identical to their stored value; all pinned at the 1.0 cap by
min(1, ...), so they stay unchanged while capped but were rewritten every run.Why it is safe
Nothing relies on
decay_scorehaving just been written. The recall multiplier, the cold-pool gate, maintenance listings and memory-health's stale-decay check all read the value, andtrg_mem_fts_updateonly fires onUPDATE OF content, summary, tags, so adecay_scorewrite never touched FTS.API note
processednow counts rows written; the result also carriesskippedandscanned. A dry run is unchanged (processed= rows scanned, sample entries still{ rowid, score }).Testing
bBase: NaNreturns the NOT NULL constraint error instead of a silent all-skipped success; dry-run sample keys arerowid,score.embedding-timeout, which passes 18/18 assertions and then hits a libuv assertion on process exit on Windows — identical on unmodifiedmain.🤖 Generated with Claude Code