Skip to content

fix(isthmus): preserve update targets in nested schemas - #1265

Open
bvolpato wants to merge 3 commits into
substrait-io:mainfrom
bvolpato:bvolpato/fix-nested-update-target
Open

bvolpato wants to merge 3 commits into
substrait-io:mainfrom
bvolpato:bvolpato/fix-nested-update-target

Conversation

@bvolpato

@bvolpato bvolpato commented Sep 3, 2026 •

Copy link
Copy Markdown
Member

For a table src(s ROW(x INTEGER), x INTEGER, n INTEGER), UPDATE src SET n = 99 converts back to UPDATE src SET x = 99. The transform's top-level column ordinal is incorrectly used to index the flattened depth-first name list [S, X, X, N].

Resolve target names from top-level fields in the declared schema without converting untouched types. This preserves multiple-assignment order when nested and top-level names collide. Document the library convention that target indexes address top-level struct fields, rather than the flattened name list.

Partially addresses #1175; struct-literal assignment field names are separate from target-column resolution.

Summary by CodeRabbit

  • Bug Fixes
    • UPDATE targets now resolve to the correct top-level column names in nested schemas, including when fields are reordered.
    • Invalid target indexes are rejected with an error that reports the requested index and the number of available top-level columns. Valid indexes continue to identify the intended column through plan serialization and deserialization and during Calcite conversion.

@bvolpato
bvolpato force-pushed the bvolpato/fix-nested-update-target branch 2 times, most recently from 17b5c27 to fbccef8 Compare September 10, 2026 05:51
@bvolpato
bvolpato marked this pull request as ready for review September 10, 2026 05:52

@nielspardon nielspardon left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two changes to the conversion and one to the test. The fix for #1175 itself is right, and it matches the ordinal the outbound direction already emits, so this makes the two directions agree rather than picking a reading; the adjacent WHERE-clause gap is covered by your #1259.

One point sits outside the diff, so no suggestion is attached: core/src/main/java/io/substrait/relation/AbstractUpdate.java:53. getColumnTarget()'s Javadoc says only "the index of the target column to update", while NamedStruct.names() advertises the flattened depth-first list — the list this bug came from reaching for. The spec says only "index of the column to apply the transformation to" and never says which of the two it indexes, so it is worth stating on the accessor that it is a top-level ordinal into getTableSchema().struct(), and that this is the library's reading rather than something the spec settles.

Comment thread isthmus/src/main/java/io/substrait/isthmus/SubstraitRelNodeConverter.java Outdated
For a table src(s ROW(x INTEGER), x INTEGER, n INTEGER), UPDATE src SET n = 99 converts back to UPDATE src SET x = 99. The transform's top-level column ordinal is incorrectly used to index the flattened depth-first name list [S, X, X, N].

Reconstruct the declared row type before resolving target names so nested fields do not shift top-level update columns. This also preserves multiple-assignment order when nested and top-level names collide.

Partially addresses substrait-io#1175; struct-literal assignment field names are separate from target-column resolution.
@bvolpato
bvolpato force-pushed the bvolpato/fix-nested-update-target branch from fbccef8 to 0b419ef Compare September 24, 2026 00:50
@coderabbitai

coderabbitai Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 943a4a98-29d7-4e63-8a1e-e1b3fb6f7217

📥 Commits

Reviewing files that changed from the base of the PR and between 0b419ef and 6219b29.

📒 Files selected for processing (1)
  • core/src/main/java/io/substrait/relation/AbstractUpdate.java

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 9 remain after this review.


📝 Walkthrough

Walkthrough

NamedUpdate conversion maps target indexes to top-level field names in nested schemas and rejects negative or out-of-range indexes. Tests cover flat, nested, reordered, and differing declared schemas.

Changes

UPDATE target conversion

Layer / File(s) Summary
Map and validate UPDATE targets
core/src/main/java/io/substrait/relation/AbstractUpdate.java, isthmus/src/main/java/io/substrait/isthmus/SubstraitRelNodeConverter.java, isthmus/src/test/java/io/substrait/isthmus/NestedUpdateTargetTest.java
The target documentation defines indexes as top-level field ordinals. The converter derives top-level field names from nested schema names and checks target indexes against the top-level column count. Tests verify plan round trips, converted column names and values, and conversion with a declared schema that differs from the catalog schema.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix

Suggested reviewers: nielspardon

Merge Risk: ⚪ Minimal · up to 6219b

The UPDATE conversion maps valid targets to top-level fields and rejects invalid ordinals. No actionable merge-blocking risk remains in the reviewed change.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 6219b

The change corrects which column a nested-schema UPDATE targets, and no new security bypass was identified. Its effect on actual writes still depends on authorization and execution controls outside the reviewed conversion path.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The security-relevant effect is the column selected in a logical UPDATE for a catalog-resolved table. Whether that plan can execute, and under whose permissions, is not established by the reviewed code.

Trust Boundaries and Controls

  • observed — Protocol decoding retains producer-supplied schema and target ordinals; the consumer checks ordinal bounds and requires a catalog table, but no column-authorization control is visible in this path.

Hardening Proposals

  • proposed — If untrusted plans can reach execution, ensure planned column names are bound to catalog columns and authorized for the caller before a write occurs.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 28.57% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 7 functions across 3 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title is concise, uses Conventional Commit format, and clearly identifies the fix for preserving update targets in nested schemas.
Description check ✅ Passed The description explains the defect, the implemented approach, the expected behavior, and the scope limitation. It provides the rationale required by the template.
  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@nielspardon nielspardon left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One small follow-up on the bounds check.

Comment on lines +105 to +107
assertEquals(List.of("N"), converted.getUpdateColumnList());
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pin the guard with a test — nothing currently fails if it is removed, so a later refactor can drop it silently. I checked: deleting the bounds check leaves the rest of this file green.

Suggested change
assertEquals(List.of("N"), converted.getUpdateColumnList());
}
}
assertEquals(List.of("N"), converted.getUpdateColumnList());
}
@Test
void rejectsOutOfRangeColumnTarget() throws Exception {
Prepare.CatalogReader catalog =
SubstraitCreateStatementParser.processCreateStatementsToCatalog(
"CREATE TABLE src (u INTEGER, n INTEGER)");
Plan plan = new SqlToSubstrait().convert("UPDATE src SET n = 11", catalog);
NamedUpdate update = assertInstanceOf(NamedUpdate.class, plan.getRoots().get(0).getInput());
for (int outOfRange : new int[] {-1, 2}) {
AbstractUpdate.TransformExpression transform =
AbstractUpdate.TransformExpression.builder()
.from(update.getTransformations().get(0))
.columnTarget(outOfRange)
.build();
NamedUpdate foreignUpdate =
ImmutableNamedUpdate.copyOf(update).withTransformations(List.of(transform));
IllegalArgumentException e =
org.junit.jupiter.api.Assertions.assertThrows(
IllegalArgumentException.class,
() ->
new SubstraitToCalcite(ConverterProvider.DEFAULT, catalog)
.convert(foreignUpdate));
assertEquals(
"Update column target "
+ outOfRange
+ " is outside the table schema's 2 top-level columns",
e.getMessage());
}
}
}

assertThrows is spelled out in full here only so the block applies on its own; a static import alongside the other two reads better.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants