Skip to content

fix(startup): simplify tool validation and fix font loading - #5614

Merged
me2seeks merged 9 commits into
apache:mainfrom
testikun:codex/remove-tool-ledger-cache
Sep 30, 2026
Merged

me2seeks merged 9 commits into
apache:mainfrom
testikun:codex/remove-tool-ledger-cache

Conversation

@testikun

@testikun testikun commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Fixes #5613. Follow-up to #5556, based on current main.

Tool writes now reconstruct validation state from the current SQLite transaction, scoped to the invocation and explicit parent dependencies. Remove the cross-transaction reducer cache and its LRU/budgets, checkpoint/undo, invalidation versions and settlement synchronization. Scans and writes retain shared validation rules, corruption classification and exact retry semantics. Also remove the leftover MAKA_STARTUP_PROFILE timing layer from SessionManager recovery.

The trade-off is deliberate: the previous same-base ablation retained the startup improvement, while a synthetic single invocation with 200 uninterrupted tool executions was slower without caching (615→1249 ms for 512 B; 839→2056 ms for 4 KiB). Short invocations and heavily interleaved partial writes showed smaller or unstable benefits; some intermediate patterns still benefit. We accept that local cost to remove long-lived consistency obligations whose representative end-to-end value has not been established. These are prior measurements on 518fd529c0, not new performance claims for this PR's base.

Keep bundled fonts as local assets so KaTeX history rendering respects the existing CSP. The memory/model-connection startup readiness fix is already supplied by #5606 on main.

Add two real Host/SQLite regression journeys: mixed prepared/committed/unknown tool outcomes recover once; consumed quoted steering survives SIGKILL without duplicate echo or old-epoch replay. Run both with 1k and 100k background events, checking history digests and repeated recovery. npm run test:startup-correctness groups these with existing storage, tool, queue, background-task and Desktop contracts. Local benchmark drivers, databases and generated reports are excluded.

Verification

  • Current head 49ac36db0 contains review-test commit 09ec6f699 and merges current main 49dacdf13; GitHub reports the branch mergeable.
  • The startup matrix now verifies all three unresolved generic prepared effects end in interrupted_unknown and that no unsettled tool operation remains. It continues to assert that no synthetic function_response is invented.
  • The Desktop Vite contract performs a real KaTeX CSS build, rejects any data:font URL, and requires emitted font assets. It fails under Vite's default 4096-byte inlining policy.
  • Current-head dependency builds for Core, Storage, Runtime, Runtime Host, and UI pass. The strengthened startup matrix passes 2/2; the Desktop Vite contract passes 2/2; Biome, git diff --check, and the protocol epoch guard pass.
  • Ablation: restoring the default font-inline limit exposes the blocked KaTeX data URL; changing the expected recovered operation state back to prepared fails against the actual interrupted_unknown state.
  • The single-long-invocation O(k²) validation cost remains an accepted and disclosed P3 trade-off. This PR does not reintroduce a cross-transaction cache; a future optimization can narrow the SQL event selection without restoring mutable cache authority.
  • Hosted CI run 36387285595, Linux package run 36387285571, Windows package run 36387285605, owner-platform run 36387285541, and dependency audit run 36387285536 all pass on the exact head. Real provider networks and external tool side effects remain outside this local run.

AI use

  • No generative tool made a substantive contribution
  • Generative tooling made a substantive contribution

Tool(s) and scope: OpenAI Codex implemented the simplification and font fix, authored the mixed-state tests and runner, ran validation, and prepared this submission at @testikun's request. @testikun is the human contributor of record; final review and merge remain with humans.

Checklist

  • Tests cover the change and fail without it
  • Lint, format, typecheck and the affected suites pass locally

Does this PR entail a change in behavior?

  • Yes — described under Summary above
  • No

Rebuild tool validation from the current transaction's local dependency
closure, removing the retained reducer cache and its consistency protocol.
Remove recovery profiling wrappers and keep fonts as local CSP-safe assets.
Add mixed tool/steering restart checks and a reusable correctness runner.

Refs apache#5613

Generated-by: OpenAI Codex
@github-actions github-actions Bot added the effort/XL Under 2500 readable lines label Sep 23, 2026
Keep the startup correctness file inventory visible to the release contract checker. Remove dynamic path construction without changing any selected test or group.

Generated-by: OpenAI Codex
Keep the self-contained correctness runner and generated-data regression tests in the PR. Retain the handoff document only as local validation material.

Generated-by: OpenAI Codex

@hqhq1025 hqhq1025 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed the current head c106ed8d8b0e7c95403e8205eeb043ffecff7f39. I found no substantiated P0–P3 issue in the changed paths.

The change removes the cross-transaction tool-ledger reducer cache and validates each proposed append against dependency events read in the current SQLite transaction (packages/storage/src/sqlite-runtime-store.ts:4263, packages/core/src/tool-ledger-scanner.ts:228). It also removes startup profiling wrappers and keeps KaTeX font files external so the renderer CSP can load them (apps/desktop/vite.config.ts:65). I checked transaction/rollback behavior, existing-versus-candidate corruption handling, recovery flow, packaging, and the adjacent tests. There is no database schema migration in this PR.

The current-head test, audit, Linux packaging, packaging, Windows owner, and macOS owner checks pass. The branch merges cleanly with current main. I could not rerun tests locally (this checkout lacks dependencies and the host has Node 18), measure production-scale startup/write performance after cache removal, or visually verify font loading. This is not a merge approval.

Automated review notice: This comment was posted by an automated review agent operated by hqhq1025. It is not an independent human review and does not replace one.

@hqhq1025 hqhq1025 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Follow-up to my earlier review of the same head (c106ed8d8b0e7c95403e8205eeb043ffecff7f39):

[P3] Removing the reducer cache makes each tool-bearing append re-read and decode every existing RuntimeEvent in the candidate's invocation, not just the candidate's direct dependencies. packages/storage/src/sqlite-runtime-store.ts:4283-4317 selects all rows with WHERE invocation_id = ? and can expand to parent-operation invocations; packages/core/src/tool-ledger-scanner.ts:233-260 rebuilds and scans the resulting set. For a single long invocation with k accumulated events, one new append therefore requires at least O(k) validation work, and successive appends can total O(k²). This does not grow with unrelated session invocations: the operation lookup has an expression index (packages/storage/src/sqlite-runtime-schema.ts:699-706). The new 100k-history test primarily covers unrelated history, not repeated writes to one long invocation. Please add a representative long-invocation write regression or measurement and consider a bounded way to avoid repeated full-prefix decoding. I have not measured user-visible latency, so I classify this as P3 rather than a demonstrated higher-severity slowdown.

Automated review notice: This comment was posted by an automated review agent operated by hqhq1025. It is not an independent human review and does not replace one.

@me2seeks me2seeks left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated review notice: This comment was posted by an automated review agent operated by me2seeks make. It is not an independent human review and does not replace one.

Summary

The PR removes the cross-transaction tool-ledger reducer cache (LRU budgets, checkpoint/undo log, write-view invalidation, settlement synchronization) from SqliteRuntimeStore/@maka/core and replaces it with per-write validation rebuilt inside the current SQLite transaction, scoped to the candidate invocation plus its explicit parent-operation closure; it also drops the MAKA_STARTUP_PROFILE layer in SessionManager and keeps KaTeX fonts as files to satisfy the renderer CSP. The premise is real on all three fronts: the deleted machinery existed in base and carried heavy consistency obligations, and the CSP at apps/desktop/src/renderer/index.html:24-27 (default-src 'self', no font-src, no data:) does block Vite-inlined data-URL fonts. The direction is right — validation semantics (validateToolLedgerTransition, packages/core/src/tool-ledger-scanner.ts:228) are byte-for-byte the old incremental logic minus statefulness, and the dependency-closure BFS (packages/storage/src/sqlite-runtime-store.ts:4283) matches the deleted cache-miss path; SQLite read-your-writes covers in-transaction retries since inserts immediately follow validation. The disclosed per-write rescan cost is an honest, documented trade-off.

Findings

  1. [P3] scripts/test-startup-correctness.mjs:30 — the runner's default output artifacts/startup-correctness/latest is not covered by .gitignore (only *.log is ignored); report.json and working-tree.patch become untracked files after every documented run, contradicting the PR body's "generated reports are excluded". Add artifacts/startup-correctness/ to .gitignore.
  2. [P3] packages/runtime-host/src/tests/startup-state-matrix.test.ts:52 — the flagship 100k-background-event variant runs only via the manual script; no CI workflow references test:startup-correctness (repo-wide grep hits only package.json:35), so CI (test job, runtime-host test:dist) enforces only the 1k default. The headline scale journey can rot silently; consider a scheduled 100k lane or soften the claim.
  3. [P3] packages/storage/src/sqlite-runtime-store.ts:4263 — the worst case the cache absorbed (many tool writes inside one long invocation, now O(history) per write, O(N²) per invocation) is exactly what the new matrix does not exercise: historyEvents are spread across Math.min(300, …) invocations (~333 events each at 100k) and tool targets sit in their own invocations. Disclosed in the body with measurements, so acceptable — but the suite cannot detect a future blowup of that shape.

Verdict

merge-ready — premise real, simplification correct and well-tested; only three minor follow-ups (gitignore, CI wiring for 100k, unstressed worst case).

@testikun

testikun commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

@hqhq1025 @Astro-Han @me2seeks 已复核并更新到当前 head 49ac36db0:

  • 单 invocation 的重复全前缀读取确实仍是 O(k²) 累计成本。这不是漏修,而是本 PR 已披露的取舍:独立测量显示影响集中在无 streaming partial 的长连续调用,现实 streaming 路径与 main 接近,而且不在启动读路径。这里不重新引入跨事务 reducer cache;后续若优化,应优先在 SQL 层只选 tool-relevant events。
  • 启动矩阵的弱断言意见成立。三种未提交的普通 prepared 场景现在逐项读取持久化 tool operation,要求状态为 interrupted_unknown,并统一断言目标 Sessions 的 unsettled 列表为空;仍保留“不伪造 function_response”的原断言。
  • 字体缺少回归契约的意见成立。Desktop 现有 Vite 契约测试新增真实 KaTeX CSS 构建,要求 CSS 不含 data:font 且输出目录存在字体文件。临时恢复 Vite 默认 4096-byte inline limit 会稳定内联 KaTeX_Size3-Regular.woff2 并使测试失败。
  • 对行为保持型缓存删除,不再声称所有存储测试都能区分旧实现;新增字体测试能区分原始配置,恢复状态断言则钉住合并后 fix(runtime): settle unknown tool outcomes during startup recovery #5770 的 recovery contract。
  • 已合入当前 main 49dacdf13,无冲突,protocol 保持 main 的 epoch 197。

当前 head 本地验证:Core/Storage/Runtime/Runtime Host/UI 依赖构建通过;startup matrix 2/2;Vite contracts 2/2;Biome、git diff --check、protocol epoch guard 通过。exact-head CI、Linux/Windows package、owner-platform 与 dependency audit checks 均已通过。

@hqhq1025 hqhq1025 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the current head, c1ddf366, including the upstream-main integration and the two follow-up commits. The latter adjust the startup recovery matrix to the current unknown-outcome contract, exclude local review artifacts, and replace an immediate Windows sandbox-process baseline check with the existing bounded observation helper. I found no new substantiated P0–P2 issue.

[P3] The single-invocation append cost reported on the previous head remains in the current code. Each tool-bearing admission calls readToolLedgerDependencies (packages/storage/src/sqlite-runtime-store.ts:4360-4417), which loads every stored event for the candidate invocation and connected parent invocations. validateToolLedgerTransition (packages/core/src/tool-ledger-scanner.ts:228-260) then reconstructs and scans that history. For an uninterrupted invocation with k prior tool events, the next append does O(k) validation work; k sequential appends can therefore do O(k²) cumulative work. This does not imply a scan of unrelated sessions. The existing large-background tests do not measure a single long invocation's incremental append latency. A focused thousands-of-appends benchmark would make the accepted trade-off measurable; I have not measured latency on this head.

Validation: the current-head hosted test, audit, Linux/package, and Windows/macOS owner checks pass; the changed Windows script passes Node 24 syntax checking, git diff --check and merge-tree against current main are clean. I did not independently run the complete startup matrix or a packaged Windows sandbox smoke test. No schema migration is added by the follow-up commits.

Automated review notice: This comment was posted by an automated review agent operated by hqhq1025. It is not an independent human review and does not replace one.

@Astro-Han Astro-Han left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent second review of head c1ddf366 (a different model lineage from the parallel review). No P0–P2 found.

The "simplification" removes the cross-transaction reducer cache added in #5556. The reducer's interpretation rules are unchanged apart from dropping the cache's write wrappers. In detail:

  • Checks unchanged: dispatch-before-result and the orphan/order/identity/spine/duplicate checks are intact (tool-ledger-scanner.ts:314-545).
  • Parent/child linkage unchanged: its load scope matches the old cache-miss path (sqlite-runtime-store.ts:4378-4425).
  • Recovery unaffected: the unsettled-operation terminal guard is SQL over tool_operations and unaffected (:4625-4642). #5770's ensureRecoveredTerminalRuntimeEventDurable doesn't go through ledger validation (:469-556).
  • Consistency is stronger: validation now always re-reads inside the transaction instead of relying on writeViewRevision and data_version for cache invalidation.
  • No leftovers: the removed exports have no remaining users, and there is no schema change.

The font fix is correct. The renderer CSP has no font-src, so data: fonts are blocked. Size3-Regular.woff2 (3624 B) is the only KaTeX font below the 4096 B inline threshold, and a partial Vite build confirms it is now emitted as a file.

P3 (non-blocking):

  • Per-write re-read cost (performance). Each tool write re-reads the whole invocation and rebuilds the reducer (sqlite-runtime-store.ts:4359-4425). CPU-time probes, order-of-magnitude only on a shared machine:

    Scenario This PR main
    1000 back-to-back tool commits, 512 B results 54.3 s (~70 ms/tool at the tail) 12.5 s (~20 ms/tool)
    500 tools with a streaming partial before each 15.5 s 16.6 s
    1000 tools, 4 KiB results 55.7 s 43.6 s (main's cache degrades after ~600 tools)

    So the quadratic cost already exists on main in realistic streaming loops, and this PR only removes the optimisation for the narrow back-to-back case. It isn't on the startup path. A follow-up could select only tool-relevant events in SQL.

  • A test assertion is too weak. In startup-state-matrix.test.ts:129 the three prepared scenarios only assert "no function_response". They don't assert that tool_operations ended interrupted_unknown or that the unsettled list is empty, which is what #5770 guarantees.

  • The tests don't distinguish old code. The PR's storage tests (68/68) and matrix tests (2/2) also pass on the main baseline. That's expected for a behaviour-preserving removal, but it means "tests fail on the old code" doesn't hold for these. The font fix has no regression test; a build-contract check that CSS contains no data:font would pin it.

Verified:

  • core/storage touched tests 135/135;
  • startup-state-matrix (1k history) 2/2;
  • a baseline comparison run.

Not verified: Node 24 (Node 22 was used), the 100k matrix, Electron KaTeX rendering, and Windows sandbox e2e.


Automated review (Claude lineage) by the Qronos review line on behalf of @Astro-Han; please verify before acting.

@hqhq1025 hqhq1025 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact head 49ac36db04d100af82c57eca800d8507c8e8f16e. The new commit adds two contract checks: startup-state-matrix.test.ts:134-152 reads persisted tool-operation state after recovery and requires no unsettled operations, while vite-client-plugin-slots.test.mjs:60-89 builds KaTeX CSS using the renderer's asset-inline setting and requires emitted font files rather than data-font URLs. These target the startup recovery and CSP/font behaviors changed by this PR. I found no substantiated new P0–P3 issue in the follow-up or main merge; the previous findings in my area remain closed.

The current-head hosted test, audit, Linux/package, and Windows/macOS owner checks pass. Script syntax, diff-check, and a merge-tree against current main pass locally. I could not run the new Vite test locally because this checkout has no vite installation; I also did not run the full recovery matrix or packaged Windows/macOS smoke locally.

Automated review notice: This comment was posted by an automated review agent operated by hqhq1025. It is not an independent human review and does not replace one.

@testikun

Copy link
Copy Markdown
Contributor Author

@hqhq1025 @Astro-Han @me2seeks Thank you for the reviews. Follow-up on current head 49ac36db0:

  • Resolved — recovery assertions: the startup matrix now verifies that all three unresolved generic prepared effects end in interrupted_unknown, that no unsettled tool operation remains, and that recovery still does not invent a function_response.
  • Resolved — font regression coverage: the Desktop Vite contract now performs a real KaTeX CSS build, rejects data:font URLs, and requires emitted font files. Restoring Vite's default 4096-byte inline limit makes the test fail on KaTeX_Size3-Regular.woff2.
  • Accepted trade-off, not changed — long-invocation O(k²) validation: the measured regression is concentrated in back-to-back tool commits without streaming partials and is not on the startup read path. This PR intentionally does not restore cross-transaction mutable cache authority. A follow-up optimization should narrow the SQL event selection instead.
  • Clarified — behavior-preserving cache removal: the storage behavior tests are not claimed to distinguish the old cache implementation. The new font contract distinguishes the configuration fix, while the recovery-state assertions pin the merged fix(runtime): settle unknown tool outcomes during startup recovery #5770 contract.
  • Resolved — generated artifacts: artifacts/startup-correctness/ remains ignored; the 100k journey remains an explicit/manual correctness lane rather than an every-PR cost.

The branch merges cleanly with current main. Exact-head CI, Linux/Windows packaging, owner-platform checks, and dependency audit all pass. There are no unresolved review threads.

@me2seeks me2seeks left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approve after full-diff review of head 49ac36d.

The simplification removes the cross-transaction tool-ledger reducer cache introduced in #5556 and returns to per-write validation rebuilt from the current SQLite transaction, while keeping #5556's scoped-read semantics. I verified the pre-#5556 logic against this head: validation rules, idempotent dedupe, and CorruptionError/RejectionError classification are preserved; the two deliberate deltas (dependency-closure failure domain instead of workspace-wide fail-stop; reads widened to the explicit parent-operation closure) are sound — validation scope now exactly equals failure domain, and global integrity is correctly separated from local write validation.

The disclosed single-long-invocation O(k²) trade-off is accepted; it remains cheaper than the pre-#5556 workspace-scan-per-write model. CI is green on this exact head; no unresolved threads. Non-blocking transparency nits: the KaTeX font fix is a separate domain (though declared in the title), and the Windows sandbox e2e baseline-wait change is not mentioned in the PR body — worth declaring incidental fixes in the description in the future.

@me2seeks
me2seeks merged commit d00c43e into apache:main Sep 30, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

effort/XL Under 2500 readable lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(startup): resolve startup issues and simplify runtime structure

4 participants