Expected Behavior
PdfReader.Open should open a PDF in which an object declares a /Length larger
than the actual stream content, recovering by locating the endstream keyword
instead of trusting the declared value. Other PDF libraries open these files
without complaining.
Actual Behavior
The parser advances by the declared length, moves past the end of the file and
throws.
- 1.50.5147:
Unexpected character '0xffff' in PDF stream. The file may be corrupted. If you think this is a bug in PDFsharp, please send us your PDF file.
- 6.1.1:
Stream cannot be read. Please send us the PDF file so that we can fix this (issues (at) pdfsharp.net).
Steps to Reproduce the Behavior
var document = PdfReader.Open(path, PdfDocumentOpenMode.ReadOnly);
Two attached files, 747 bytes each:
| File |
Object 5 |
Result |
repro_invalid_stream_length.pdf |
/Length 2338 on an empty stream |
throws |
repro_corrected.pdf |
/Length 0000, everything else identical |
opens, 1 page |
The files differ by 4 characters, so the declared length is the only variable.
The offending object:
5 0 obj
<< /Length 2338 /Type /Metadata /Subtype /XML >>
stream
endstream
endobj
The stream is empty, but /Length says 2338. Starting from the stream content at
byte 457, the parser is sent to byte 2795 in a 747-byte file.
I did not use the IssueSubmissionTemplate because the repro is a single call to
PdfReader.Open on the attached files — there is no surrounding code involved.
Happy to provide the full solution if it helps.
Why this matters in practice
This is not a hand-crafted edge case. We hit it in production with scanned
documents produced by HP Scan, which writes an empty XML metadata stream with
a non-zero /Length. The original file is a customer invoice, so the attachments
are a minimal reproduction with no real data.
repro_corrected.pdf
repro_invalid_stream_length.pdf
Expected Behavior
PdfReader.Openshould open a PDF in which an object declares a/Lengthlargerthan the actual stream content, recovering by locating the
endstreamkeywordinstead of trusting the declared value. Other PDF libraries open these files
without complaining.
Actual Behavior
The parser advances by the declared length, moves past the end of the file and
throws.
Unexpected character '0xffff' in PDF stream. The file may be corrupted. If you think this is a bug in PDFsharp, please send us your PDF file.Stream cannot be read. Please send us the PDF file so that we can fix this (issues (at) pdfsharp.net).Steps to Reproduce the Behavior
Two attached files, 747 bytes each:
repro_invalid_stream_length.pdf/Length 2338on an empty streamrepro_corrected.pdf/Length 0000, everything else identicalThe files differ by 4 characters, so the declared length is the only variable.
The offending object:
The stream is empty, but
/Lengthsays 2338. Starting from the stream content atbyte 457, the parser is sent to byte 2795 in a 747-byte file.
I did not use the IssueSubmissionTemplate because the repro is a single call to
PdfReader.Openon the attached files — there is no surrounding code involved.Happy to provide the full solution if it helps.
Why this matters in practice
This is not a hand-crafted edge case. We hit it in production with scanned
documents produced by HP Scan, which writes an empty XML metadata stream with
a non-zero
/Length. The original file is a customer invoice, so the attachmentsare a minimal reproduction with no real data.
repro_corrected.pdf
repro_invalid_stream_length.pdf