Skip to content

Parse only the last markdown block of a streaming reply #461

Description

@alex-clickhouse

After #460, the chat renders at most once per animation frame while a reply streams, and finished messages are not parsed again. But the streaming message is still parsed in full on each frame: remark-gfm, rehype-highlight, and the React render all run on the full text of the reply. For one long reply this is O(n²). It now occurs at most once per frame, not once per token.

Step 1: measure

Measure after #460 is merged. Stream a long reply, for example more than 20k characters with some code fences, and record a profile. If one frame takes less than a few milliseconds, close this issue.

Proposal

Split the reply into top-level markdown blocks: paragraphs, headings, full lists, full blockquotes, and full code fences. marked.lexer gives these blocks directly. Render each block with its own memoized MarkdownContent, keyed by its source text. All blocks except the last are finished and do not render again. Only the last block is parsed on each frame.

Nested markdown is not a problem. A nested list, or a blockquote in a list, is part of one top-level block, and that block is always parsed as one unit. There is no need to find a position inside nested markdown to go back to.

Existing work to look at:

  • The Vercel AI SDK docs describe this approach as "memoized markdown", with marked.lexer and memo.
  • Vercel's streamdown package may do this as a replacement for react-markdown. Make sure that it supports our custom code, pre, and a components in MarkdownContent.tsx before we use it.

Edge cases

  • One long block: a long code fence or a long list is still parsed in full on each frame while it grows. The O(n²) stays inside that block. For code fences, rehype-highlight is the most expensive part. Do not highlight the last, unfinished block until the fence closes, or until the stream ends.
  • The last block changes type: --- or === under a paragraph changes that paragraph into a setext heading. A list can continue after what looks like its end, for example after a blank line followed by another item. Thus the block before the last is not always finished. Treat the last two blocks as unfinished.
  • Unclosed code fence: an open fence contains all text after it until the fence closes. The lexer gives this as one last block, so it is correct. But while the fence is open, that block grows for the full length of the code.
  • Reference-style links: a definition such as [x]: https://… late in the reply changes how earlier blocks render. If each block is parsed alone, earlier blocks do not see the definition. LLMs almost never write these, so it is acceptable. If we find that it is necessary, collect the definitions first and give them to each block.
  • Lexer cost: the lexer reads the full text on each frame. This is O(n) per frame, but it is much cheaper than remark, highlighting, and React together. To make it incremental, keep the offset where the last unfinished block starts, and lex again only from that offset.
  • Different parsers: marked and remark (micromark) can split some input differently, for example HTML blocks or lazy continuation lines. A block from marked that remark then parses alone can render differently than it would in the full document. Test with real replies that have lists, tables, and HTML.
  • Block keys: key the blocks by index, not by content. Then a block that grows keeps its component and its state, for example the "Copied" state of a code block.

Where

  • web/src/components/Chat/BlockRenderer.tsx: the text case, only when streaming is true. Finished messages can keep one MarkdownContent.
  • web/src/components/Chat/MarkdownContent.tsx

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions