Skip to content

Reduce native text measurement and TALA allocation overhead - #2930

Closed
alixander wants to merge 4 commits into
masterfrom
perf/native-text-routing
Closed

alixander wants to merge 4 commits into
masterfrom
perf/native-text-routing

Conversation

@alixander

@alixander alixander commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Human


AI

Native SVG compilation painted Markdown that only needed dimensions, rebuilt text measurement state, and repeatedly measured legend labels. TALA node sorting also used reflection-related temporary objects. This PR removes that work while preserving layout, output, ordering and validation behavior.

The three accepted changes were implemented and measured sequentially against the preceding accepted source:

  • Separate Markdown dimension measurement from SVG painting through the shared layout path.
  • Reuse lazy per-render text measurement state, measure legend labels once, and explicitly pass the compiler ruler from the native SVG CLI. Library callers opt in through the optional ruler field.
  • Use slices.SortFunc and cmp.Compare for node IDs, preserving the Go 1.27 sorting algorithm and comparator order.

The routing scratch-map candidate passed correctness checks but failed its predeclared performance screen, so it was removed. Its patch and all measurements remain in the evidence archive.

Performance

Negative means lower; “inconclusive” means the paired 95% interval includes no change. Full-corpus rows use the same 246 successful diagrams. Each incremental row has its own immediately preceding baseline; the final rows are a separate direct comparison with the original source.

Comparison Pipeline time CPU time Peak RSS Allocated bytes
Markdown measurement vs original −0.204% −0.257% Inconclusive Inconclusive
Ruler/legend reuse vs Markdown change Inconclusive Inconclusive −0.346% −0.349%
Typed sorting vs ruler/legend reuse −0.396% −0.437% Inconclusive −1.024%
All changes vs original: 246 diagrams −0.254% −0.472% −0.486% −1.483%
All changes vs original: 10 public diagrams −0.913% −1.542% Inconclusive −3.435%

The affected-cohort confirmations support the first two changes' targeted benefit: Markdown measurement reduced pipeline time 0.783% on 24 Markdown cases; ruler/legend reuse reduced it 0.935% on those cases plus the native legend. The latter reduced rendering time 18.736%, which is a much larger percentage than its whole-pipeline gain. Allocation counts fell 3.466% in the direct full-corpus comparison.

The overall pipeline interval is −0.541% to −0.061%; CPU is −0.747% to −0.273%. These are modest, workload-specific improvements, not evidence that every diagram becomes faster. Do not add incremental percentages. Aggregate estimates are ratios of summed per-case medians. The dated report gives every interval, absolute public-case pipeline medians, all ten public diagrams, and limitations; its evidence ZIP contains all 33,640 numeric rows across eleven campaigns, frozen runners/statistics and public reproduction instructions. Private sources, names, paths, fonts and SVGs are excluded.

Runs use Go 1.27 on an active Apple M4 desktop with unchanged TALA options, seeds and native concurrency: ten adjacent paired timing rounds after two warmups, plus three separate allocation rounds. Predeclared affected/control confirmations use twenty pairs. None of the nine initial timing regression signals across seven cases repeated in the fixed confirmation. Individual intervals are unadjusted; all unfavorable observations are retained. Pipeline and process-wall time are distinct, as are allocated bytes and peak RSS. The runner explicitly adopts native CLI ruler reuse; unchanged library callers do not automatically get that opt-in benefit.

Validation

  • Full-corpus preflight: 246 byte-identical successful SVG pairs and one matching existing error; every retained measured SVG matched its oracle.
  • Native CLI: all ten public SVGs match the original source. The actual public reproduction controller also passes its identity/output/completion smoke check; those low-round smoke timings are excluded from performance claims.
  • Text measurement, renderer, graph/target/library, CLI and full E2E tests; focused renderer/CLI race checks; exact SVG/legend fixtures; 84 exact pointer-order differential cases; independent source review.
  • Measured accepted source passed hosted CI, TALA race/fuzz checks, and Linux/macOS/Windows release smoke checks. The final evidence/comment-only commit is checked separately by the PR workflows.
  • Exact numeric recalculation, source/binary freeze validation, and public archive privacy/hash checks. Only the Markdown report and ZIP are added under ci/performance-results; frozen Go runners stay inside the ZIP.

@alixander alixander changed the title Reduce native text measurement overhead Reduce native text measurement and TALA allocation overhead Sep 15, 2026
@alixander
alixander marked this pull request as ready for review September 15, 2026 05:09
@alixander alixander closed this Sep 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant