perf(rust): restore math spans in one pass instead of rescanning per token - #948
Merged
Merged
Conversation
…token restore_math_spans searched the rest of the HTML for both token spellings on every token. The lowercase heading-anchor spelling is usually absent, so each search ran to the end of the document: quadratic in the number of math spans. Cache the next position of each spelling and search again only once it has been passed. convert_markdown on a 1.6 MB concatenation of samples/*.md (release): restore pass 2.3-3.0 s -> 2.5 ms, total 2.7-3.3 s -> 122 ms. Rendered every samples/*.md, their 2x/4x copies and the 1.6 MB file before and after: all 29 outputs byte-identical.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is
restore_math_spanssearched the rest of the HTML for both token spellings on every math token. The lowercase heading-anchor spelling is usually absent, so each search ran to the end of the document: tokens × document length. It now remembers where each spelling next occurs and searches again only once the cursor has passed it.Mechanism
convert_markdownon a 1.6 MB sample built fromsamples/, release: 3.31 s → 122 ms, of whichrestore_math_spans2.99 s → 2.5 ms. Doubling the sample took it from 3.0 s to 11.6 s before, so it was quadratic. Only the samples with math showed it.Tests
math_tokens_in_headings_and_text_restore_in_orderinterleaves heading and body math and checks both anchor ids and the order of all five spans. It passes on the old code too: the output did not change. It guards the cursor bookkeeping.Verification
samples/*.md, their 2× and 4× copies, the 1.6 MB sample and its 2×) with the old code and the new:diff -ris emptycargo fmt --check,cargo clippy --all-targets -- -D warnings: cleancargo test: 186 passedNot verified: the app itself.
encoding-gbk-largeis still superlinear inside comrak (80 ms at 2×, 311 ms at 4×), which this does not touch.