Skip to content

Datamatrix GS1 fixes - #99

Merged
hschimke merged 5 commits into
rxing-core:mainfrom
JasperDeSutter:datamatrix-gs1-fixes
Sep 16, 2026
Merged

hschimke merged 5 commits into
rxing-core:mainfrom
JasperDeSutter:datamatrix-gs1-fixes

Conversation

@JasperDeSutter

@JasperDeSutter JasperDeSutter commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor
  1. This condition on minimal ECI input got lost in the port:
    https://github.com/zxing/zxing/blob/33dfdefcb35576612841e614d44ba9edc9aee2b5/core/src/main/java/com/google/zxing/common/MinimalECIInput.java#L162

  2. FNC1 characters aren't accounted for in determining datamatrix edge sizing, so some content could get cut short.

  3. Fixes a couple differences compared to the java implementation

Upstream zxing excludes the sentinel with an upper bound of 999, which was
lost in the port.
The minimal encoder prepends the FNC1 codeword of a GS1 symbol once the
solution has been assembled, so every edge computed its size one codeword
short. When the message without the FNC1 exactly filled a symbol, the
encoder concluded that the trailing C40, Text or X12 run needed no unlatch.
The FNC1 then pushed the message into the next symbol size and the padding
of that symbol was decoded as text: "01012345678901281720010110ABC123"
decoded as "01012345678901281720010110ABC123GR2u".

Upstream zxing has the same defect.
The C40 versus X12 tie break in the Data Matrix lookahead searches for an X12
terminator that comes before a character that X12 cannot encode. It iterated
the message from index 0 instead of from the position the lookahead stopped
at, so it answered the question for an unrelated part of the input, and the
position variable it computed was never read.

Upstream zxing scans from startpos + charsProcessed + 1. Over a corpus of
11260 inputs the two disagree on 570 of them: the corrected version selects a
smaller symbol for 67 and a larger one for 40, the heuristic in the
specification is not optimal.
The minimal encoder sliced the macro header and trailer off the message with
character counts used as byte offsets. A macro message holding any non ASCII
character either sliced inside a character, which panics, or lost trailing
bytes. Both affixes are ASCII, so their byte lengths are also their character
lengths.

@hschimke hschimke left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks like all tests pass and changes look good, approving for merge.

@hschimke

Copy link
Copy Markdown
Collaborator

Looking through these I see exactly how I messed them up while porting. I had a huge set of very similar issues that I found because they broke things more obviously when I first started integration testing. Thank you for this patch.

I'm going to bundle this and a bunch of other fixes up and do a release before the end of the week.

@hschimke
hschimke merged commit 3f1e73f into rxing-core:main Sep 16, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants