SubstraitLexer.g4 lets Number absorb a leading sign, so a - written immediately before a digit lexes as a single negative-number token rather than as the subtraction operator. The parser then sees two adjacent expressions and fails. Whitespace after the - is enough to avoid it; whitespace before it is not.
Measured on main at cf581f4:
| Expression |
Result |
varchar<L - 1> |
parses |
varchar<L- 1> |
parses |
varchar<L-x> |
parses |
varchar<L+1> |
parses |
varchar<L*2> |
parses |
decimal<P-S+1, 0> |
parses |
varchar<L-1> |
ParseError: extraneous input '-1' expecting '>' |
varchar<L -1> |
ParseError: extraneous input '-1' expecting '>' |
decimal<P-1, 0> |
ParseError: extraneous input '-1' expecting ',' |
varchar<2-1> |
ParseError: extraneous input '-1' expecting '>' |
So the trap is narrow but sharp: - between two identifiers is fine, and only - directly followed by a digit breaks. + and * are unaffected, which makes the asymmetry easy to hit by accident — a declaration author who writes decimal<P+1, 0> successfully has no reason to expect decimal<P-1, 0> to fail, and the error names '>' or ',' rather than the subtraction.
The same lexing also means a negative literal is accepted wherever a type parameter is expected, which feeds #1310: varchar<L - -1> parses and derives a length from a negative operand.
The grammar is owned upstream (it ships in the substrait-packaging antlr artifact, generated from grammar/SubstraitLexer.g4), so a fix is a spec change — restricting Number to unsigned digits and giving the unary minus its own production — followed by a packaging release and a catalog bump here. Filing it here so the limitation is recorded against the code that hits it; happy to move it upstream instead.
SubstraitLexer.g4letsNumberabsorb a leading sign, so a-written immediately before a digit lexes as a single negative-number token rather than as the subtraction operator. The parser then sees two adjacent expressions and fails. Whitespace after the-is enough to avoid it; whitespace before it is not.Measured on
mainat cf581f4:varchar<L - 1>varchar<L- 1>varchar<L-x>varchar<L+1>varchar<L*2>decimal<P-S+1, 0>varchar<L-1>ParseError: extraneous input '-1' expecting '>'varchar<L -1>ParseError: extraneous input '-1' expecting '>'decimal<P-1, 0>ParseError: extraneous input '-1' expecting ','varchar<2-1>ParseError: extraneous input '-1' expecting '>'So the trap is narrow but sharp:
-between two identifiers is fine, and only-directly followed by a digit breaks.+and*are unaffected, which makes the asymmetry easy to hit by accident — a declaration author who writesdecimal<P+1, 0>successfully has no reason to expectdecimal<P-1, 0>to fail, and the error names'>'or','rather than the subtraction.The same lexing also means a negative literal is accepted wherever a type parameter is expected, which feeds #1310:
varchar<L - -1>parses and derives a length from a negative operand.The grammar is owned upstream (it ships in the
substrait-packagingantlrartifact, generated fromgrammar/SubstraitLexer.g4), so a fix is a spec change — restrictingNumberto unsigned digits and giving the unary minus its own production — followed by a packaging release and a catalog bump here. Filing it here so the limitation is recorded against the code that hits it; happy to move it upstream instead.