Skip to content

fix(session): tolerate invalid UTF-8 in saved sessions - #39

Closed
jipeng6036-del wants to merge 1 commit into
he-yufeng:mainfrom
jipeng6036-del:fix/invalid-utf8-sessions
Closed

jipeng6036-del wants to merge 1 commit into
he-yufeng:mainfrom
jipeng6036-del:fix/invalid-utf8-sessions

Conversation

@jipeng6036-del

Copy link
Copy Markdown

A saved session containing an invalid UTF-8 byte, or truncated in the middle of a multibyte character, raises UnicodeDecodeError before JSON parsing. This crashes --resume and /sessions, despite the existing handling for corrupt or truncated JSON.

Catch UnicodeDecodeError in both session readers: resume returns None, and listing skips the unreadable session while retaining valid sessions. The files are left untouched. Four regression cases cover invalid bytes and truncated UTF-8 in both paths, including a valid Chinese-language session alongside a corrupt file.

Validation on Windows, Python 3.13.12:

  • Before the fix: all four new regression cases fail with UnicodeDecodeError.
  • python -m pytest tests/ -q: 229 passed.
  • ruff check corecoder tests, python -m compileall -q corecoder tests, python -m build, and python -m twine check dist/*: passed.

AI assistance: this change and its regression tests were prepared with OpenAI Codex.

@he-yufeng

Copy link
Copy Markdown
Owner

Thanks for the catch, @jipeng6036-del. Verified locally: a session file with invalid UTF-8 raises UnicodeDecodeError before json.loads, crashing --resume, and one such file kills the whole /sessions listing. Landed the same fix as my own commit in 1f201a2 with regression coverage for both readers, so closing this one. Credit for the report stands.

@he-yufeng he-yufeng closed this Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants