Skip to content

Antalya 26.6 - Drop partition for Iceberg - #2361

Open
xieandrew wants to merge 5 commits into
antalya-26.6from
feature/antalya-26.6/iceberg-alter-table-drop-partition
Open

xieandrew wants to merge 5 commits into
antalya-26.6from
feature/antalya-26.6/iceberg-alter-table-drop-partition

Conversation

@xieandrew

@xieandrew xieandrew commented Sep 14, 2026 •

Copy link
Copy Markdown
Collaborator

Changelog category (leave one):

  • New Feature

Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):

Adds support for ALTER TABLE <table> DROP PARTITION <id> for Iceberg tables.

Documentation entry for user-facing changes

Adds support for ALTER TABLE <table> DROP PARTITION <id> for Iceberg tables. This resolves the correct partition to remove and writes a new Iceberg snapshot with the matching data files excluded.

Data files are physically deleted if the option iceberg_delete_data_on_drop is enabled.

CI/CD Options

Exclude tests:

  • Fast test
  • Integration Tests
  • Stateless tests
  • Stateful tests
  • Unit tests
  • Performance tests
  • Aarch64 tests
  • All with ASAN
  • All with TSAN
  • All with MSAN
  • All with UBSAN
  • All with Coverage
  • All Regression
  • Disable CI Cache

Regression jobs to run:

  • Fast suites (mostly <1h)
  • Aggregate Functions (2h)
  • Alter (1.5h)
  • Benchmark (30m)
  • CAS (content-addressed storage; Antalya only)
  • ClickHouse Keeper (1h)
  • Iceberg (2h)
  • LDAP (1h)
  • OAuth (5m)
  • Parquet (1.5h)
  • RBAC (1.5h)
  • SSL Server (1h)
  • S3 (2h)
  • S3 Export (2h)
  • Swarms (30m)
  • Tiered Storage (2h)

Closes #1046

@github-actions

github-actions Bot commented Sep 14, 2026 •

Copy link
Copy Markdown

Workflow [PR], commit [1407255]

Signed-off-by: Andrew Xie <dev@xie.is>
@xieandrew
xieandrew force-pushed the feature/antalya-26.6/iceberg-alter-table-drop-partition branch from 27e2460 to c5a11a6 Compare September 14, 2026 21:16
@xieandrew
xieandrew force-pushed the feature/antalya-26.6/iceberg-alter-table-drop-partition branch from 4591889 to 953220d Compare September 23, 2026 18:30
@xieandrew

Copy link
Copy Markdown
Collaborator Author

@blau-ai

@blau-ai

blau-ai commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

CI triage for #2361 @ 953220d

Verdict: 4 red checks, 0 PR-caused. They are 2 logical Grype image scans (keeper + server-alpine, each surfaced twice — once as a GitHub check-run, once as an S3 result). Every test/build job that has finished is green; the rest of the workflow is still running.


Not PR-related — Grype base-image CVE (infra/fleet-wide)

Check Result
GrypeScanKeeper / Grype Scan (altinityinfra/clickhouse-keeper) fail — 1 high/critical
GrypeScanServer (-alpine) / Grype Scan (altinityinfra/clickhouse-server:…-alpine) fail — 1 high/critical

Both images fail on the same single High vulnerability, CVE-2026-85091 (an nvd:cpe match; the only other finding, CVE-2025-60876, is Medium and passes the threshold). Evidence it is not caused by this PR:

  • The diff is C++ source + one test only — src/Storages/ObjectStorage/DataLakes/Iceberg/*, StorageObjectStorage.*, IDataLakeMetadata.h, and tests/integration/.../test_drop_partition.py. No Dockerfile, no dependency/package manifest, no base-image change. Grype scans OS/runtime packages baked into the image, which this PR does not touch.
  • A fresh 2026 CVE. CVE-2026-85091 was published essentially now (today is 2026-09-23), so it lights up on every image built after the grype DB picked it up — independent of source changes.
  • Reproduces on an unrelated PR. Sibling PR Iceberg: reuse the Puffin object metadata across deletion-vector reads #2419 (Iceberg Puffin, no shared code) shows the identical two failures with the same "1 high/critical" message. Notably its ubuntu-based clickhouse-server image passes with 0 high/critical — i.e. the CVE lives in the keeper + alpine base images, not in anything either PR wrote.

Suggested action: nothing to change in this PR. This is resolved at the CI/base-image level by Altinity infra — patch/rebuild the keeper and alpine base images, or add CVE-2026-85091 to the grype ignore list once triaged. Re-running the job won't clear it until the base image is updated, and it should not block review/merge of the code change.


Everything else so far: green

Finished and passing: Fast test (0 failed / 9392 passed), all Builds (amd debug/asan_ubsan/binary/release, arm release), Unit tests (asan_ubsan: 0/14839), Stateless (amd_debug parallel 0/11130; amd_asan_ubsan distributed-plan parallel 1/2 0/5550), both AST fuzzer (targeted) jobs, Integration tests (amd_asan_ubsan, targeted), Docker server/keeper images, Source upload.

⚠️ The PR workflow is still running — many jobs are PENDING/RUNNING (remaining Stateless shards, Integration db disk / old analyzer 1–8, Stress tests, Compatibility check, SQLLogic/SQLStorm, the RegressionTestsRelease / Iceberg regression suite, Finish Workflow). No functional failures have appeared yet, but the run isn't complete — worth a final glance once it settles, especially the Iceberg regression + integration jobs given what this PR changes.

— @blau-ai (analysis only; CI is the source of truth since I can't build/run ClickHouse here)

@DimensionWieldr

Copy link
Copy Markdown
Collaborator

CI Failures Analysis

Run 35903013423, commit 953220d. Four jobs failed. None of them come from the Iceberg DROP PARTITION change. Regression suites reported no new failures. Compared with base 0dbbd797, the report's "New Fails in PR" tab is empty.

Related to this PR

None.

Pre-existing Flaky Tests (Unrelated)

  • 04661_refreshable_mv_cancel_during_planning in Stateless tests (amd_asan_ubsan, cas s3 storage, parallel, 1/2). The script gives DROP TABLE 10 seconds (timeout 10) and prints drop did not finish when that expires. This run printed that instead of dropped. Upstream, the same test failed 5 times in 90 days (about 5 of 270k runs), on unrelated PRs, under asan, tsan, and debug. This PR ran it once. The diff does not touch refreshable materialized views.
  • Stress test (amd_debug): the server aborted on Logical error: 'Inconsistent KeyCondition behavior' (STID 5182-2f27) in MergeTreeDataSelectExecutor::markRangesFromPKRange. The throw is inside #ifndef NDEBUG, so only debug builds hit it. The query was a stress mutation of 04612_reverse_key_index_analysis (count() with toUInt128 on the primary key). Upstream Stress test (amd_debug) hit this STID on 6 of 20,709 runs in 60 days, including master on 2026-09-13. The same STID failed master Stress test (arm_debug) on 2026-09-19. Tracked upstream in Logical error: Inconsistent KeyCondition behavior (STID: 5182-2f27) ClickHouse/ClickHouse#112418 and the wider KeyCondition issue Multiple functions report inaccurate monotonicity information ClickHouse/ClickHouse#90461. The diff does not touch MergeTree index analysis.

Server died in that stress job is this abort (signal 6), not a later cascade.

Infrastructure Issues (Unrelated)

  • Grype on the alpine images: CVE-2026-85091 in zlib 1.3.2-r0, high, on clickhouse-keeper and clickhouse-server alpine. The non-alpine server image passed with no high or critical findings. The PR does not change image packaging.

Already-known broken tests (jobs stayed green)

  • test_s3_cache_locality/test.py::test_cache_locality[0] — INVESTIGATE, timeout or assertion
  • 00429_long_http_bufferization — known sanitizer timeout
  • 01169_old_alter_partition_isolation_stress — known sanitizer timeout

Issue/Fix References

@DimensionWieldr

Copy link
Copy Markdown
Collaborator

Regression tests for ALTER TABLE ... DROP PARTITION were run locally on the amd_release build of 953220d, against both the REST and Glue catalogs.

On REST, every scenario passed except format version 3. The suite's REST catalog rejects v3, so Spark never creates the table. On Glue that scenario is skipped for the same reason.

On both catalogs the drop left the expected rows: identity and day partitions, a tuple key, an empty partition, purge on and off, IcebergS3 with no catalog, a mixed manifest, and a file written under a finer spec. Rejections also matched: unpartitioned tables, DROP PARTITION ALL, DROP PARTITION ID, format version 1, a coarser partition spec, and allow_insert_into_iceberg left off.

Glue scenarios that then read system.iceberg_history or system.iceberg_files failed. Those tables list nothing for Glue catalog tables. That already happens on the public 26.6.4.20001.altinityantalya release and is tracked in #2457. Position-delete coverage on Glue stops at that check, before DROP PARTITION.

The new scenarios are skipped in clickhouse-regression until this is in a released build.

CI failures are unrelated. Waiting on @arthurpassos for dev review.

@DimensionWieldr

Copy link
Copy Markdown
Collaborator

AI audit note: This review comment was generated by AI.

Audit update for PR #2361 (Iceberg ALTER TABLE ... DROP PARTITION):

Confirmed defects:

High: Partial manifest rewrite turns DELETED entries back into live files

  • Impact: DROP PARTITION can bring back rows that an earlier delete or overwrite already removed, on any manifest that still holds both the dropped partition and other partitions.
  • Anchor: IcebergWrites.cpp / rewriteManifestFileExcludingFiles; scan in IcebergMetadata::scanManifestsForPartition
  • Trigger: A v2 manifest with a DELETED entry plus live files from more than one partition (typical of Spark or PyIceberg delete/overwrite). Drop one of those partitions.
  • Why defect: The scan only considers live entries (getFilesWithoutDeleted drops DELETED). The rewrite then copies every other Avro entry and forces status = EXISTING (0). A previously deleted file that shared the manifest is live in the new snapshot.
  • Fix direction (short): Copy DELETED entries unchanged, or omit them; never rewrite their status to EXISTING.
  • Regression test direction (short): Build a manifest with status=DELETED for one file and live files in two partitions, drop one partition, and assert the deleted file stays absent.

High: A failed catalog commit deletes the manifest list the new metadata already names

  • Impact: A retryable REST error (409, or 5xx after the catalog actually committed) either wedges every later DROP PARTITION or makes the current snapshot unreadable.
  • Anchor: IcebergMetadata::tryDropPartitionOnce; RestCatalog::updateMetadata
  • Trigger: writeMetadataFileAndVersionHint succeeds, then updateMetadata returns false. Cleanup removes the manifest list and rewritten manifests. The retry calls getLatestOrExplicitMetadataFileAndVersion with the explicit catalog path ignored, so it loads this higher metadata version, whose manifest list is already gone.
  • Why defect: For a 5xx after a successful commit, the published snapshot’s manifests are deleted. For a real 409, the orphan metadata file is what the retry reads, so the retry throws instead of recommitting against the catalog’s current snapshot.
  • Fix direction (short): Treat the metadata file as published once the catalog may have accepted it; on conflict, retry from the catalog location and do not delete manifests that file references.
  • Regression test direction (short): Fail updateMetadata with 409 and with a 5xx-after-commit, then assert the catalog snapshot still reads and a retry completes.

Medium: iceberg_delete_data_on_drop deletes files that older snapshots still reference

  • Impact: Time travel and any reader still on the parent snapshot lose those objects. A purge error is only logged, so the ALTER still succeeds.
  • Anchor: IcebergMetadata::tryDropPartitionOnce (purge block after the commit)
  • Trigger: iceberg_delete_data_on_drop = 1 on a table that has more than one snapshot.
  • Why defect: The new snapshot drops the partition, but ancestor snapshots still point at the same data and delete files. removeObjectsIfExist runs anyway, and exceptions are swallowed.
  • Fix direction (short): Delete a file only when no remaining snapshot references it, and fail the statement if a requested purge fails.
  • Regression test direction (short): Purge a partition, then read the parent snapshot id and expect the old rows to still be readable.

Medium: Successful drop leaves the cached “latest metadata” pointer on the pre-drop snapshot

  • Impact: With iceberg_metadata_staleness_ms greater than 0, later queries keep reading the dropped partition until the cache entry expires.
  • Anchor: IcebergMetadata::tryDropPartitionOnce (no invalidateMetadataCache); contrast IcebergStorageSink::initializeMetadata
  • Trigger: Set iceberg_metadata_staleness_ms (the storage docs suggest 120000), drop a partition, and query again inside that window. The drop itself refreshes the cache with the metadata it read before the commit.
  • Why defect: INSERT and schema ALTER clear this cache after commit. DROP PARTITION does not, so the staleness window serves the snapshot from before the drop.
  • Fix direction (short): Call persistent_components.invalidateMetadataCache after a successful commit.
  • Regression test direction (short): Drop a partition with staleness enabled and assert the next SELECT does not return the dropped rows.

Coverage summary:

  • Scope reviewed: Iceberg DROP PARTITION from StorageObjectStorage::alterPartition through manifest scan/rewrite, snapshot commit, catalog update, purge, and the new stateless/integration tests.
  • Categories failed: DELETED manifest entries on rewrite; catalog commit failure versus cleanup; purge versus snapshot reachability; metadata-cache invalidation.
  • Categories passed: partition-value evaluation and spec-evolution checks, full-manifest drop versus partial rewrite for live files, v1/v3 rejection, empty-match no-op, position/equality deletes scoped to the same partition key, filesystem metadata-version conflict retry, exception cleanup before the metadata file is written.
  • Assumptions/limits: Static review only; integration tests were not executed. No runtime check of Avro setMetadata during rewrite, or of total-files-size when delete-file bytes are included.

@arthurpassos

Copy link
Copy Markdown
Collaborator

@DimensionWieldr @xieandrew is it ready for review tho?

I see "WIP: Add support for purging data files (physically delete)"

@xieandrew

Copy link
Copy Markdown
Collaborator Author

@arthurpassos Yes it is ready for review. Sorry, forgot to update the description.

@arthurpassos arthurpassos left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is one fundamental issue I would like to discuss before proceeding with the review:

AFAIK, ClickHouse has two ways to drop partitions: by id and by value. By id is very simple, you provide the partition id and that is it. The value one is interesting, and is the one supported in this PR.

In ClickHouse MergeTree tables, the "value" for the drop partition is the value of each element of a partition expression tuple after it has gone through the "transforms".

You've implemented the opposite. On yours, you are required to provide the source values present in the columns.

For example:

CREATE TABLE xie_test
(
    `id` UInt32,
    `event_date` DateTime64
)
ENGINE = MergeTree
PARTITION BY (id, toYYYYMM(event_date))

INSERT INTO xie_test VALUES (1, now());

In the ClickHouse idiom, to drop such a partition by value you would do:

ALTER TABLE xie_test DROP PARTITION (1, toYYYYMM(now()))

On the other hand, with your implementation, you'd have to provide the raw source values.

ALTER TABLE xie_test DROP PARTITION (1, now())

I think both approaches have its pros and cons, but I would vote for keeping it consistent with MergeTree unless this is a limitation of Iceberg (tho I don't see how it could be).

Also, it would be great if you could add docs.


/// Find the partition spec object with the given spec-id inside a metadata JSON document.
/// Throws METADATA_MISMATCH if the spec is not found (indicates metadata/spec-id mismatch).
Poco::JSON::Object::Ptr lookupPartitionSpec(const Poco::JSON::Object::Ptr & meta, Int64 spec_id)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting to see this function being reused

"Schema with id {} not found in table metadata", schema_id);
}

using PartitionSpecSignature = std::vector<std::pair<Int32, String>>;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: I would say create a struct.

struct PartitionTerm
{
Int32 column_id;
String transform;
};

auto metadata_object = getMetadataJSONObject(metadata_path, object_storage, persistent_components.metadata_cache, context, log, compression_method, persistent_components.table_uuid);

if (!metadata_object->has(f_current_snapshot_id))
throw Exception(ErrorCodes::BAD_ARGUMENTS, "No snapshot exists for this Iceberg table");

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this really a bad_arguments error?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe it can just succeed with a no-op (and log info)? This doesn't mean the table is in an invalid state, there's just no data files yet. A non-empty table that doesn't match any files for drop partition already is a no-op right now.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hm.. I would say copy ClickHouse MergeTree behavior. What happens if we try to drop a partition that does not exist on MergeTree tables? If it is no-op, make this one no-op as well.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, MergeTree is a no-op as well


const Int64 current_snapshot_id = metadata_object->getValue<Int64>(f_current_snapshot_id);
if (current_snapshot_id < 0)
throw Exception(ErrorCodes::BAD_ARGUMENTS, "No snapshot exists for this Iceberg table");

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment on lines +1217 to +1224
const Int32 format_version = metadata_object->getValue<Int32>(f_format_version);
if (format_version < 2)
throw Exception(ErrorCodes::SUPPORT_IS_DISABLED, "DROP PARTITION is supported only for Iceberg format version 2 and above");

if (format_version >= 3)
throw Exception(ErrorCodes::SUPPORT_IS_DISABLED,
"DROP PARTITION is not supported for Iceberg format version {}. Dropping Puffin deletion vectors "
"is not implemented yet", format_version);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Perhaps just if (format_version != 2) { throw...

It might confuse users because you state it is only supported for version 2 and above, but then version 3 doesn't work.

const auto target_partition_key = evaluateTargetPartitionKey(partitioner, partition_source_header, source_values);

LOG_INFO(log, "Iceberg DROP PARTITION requested for partition {} of spec {}",
dumpPartitionTuple(target_partition_key), partition_spec_id);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You call dumpPartitionTuple several times, do it only once

@xieandrew

Copy link
Copy Markdown
Collaborator Author

AFAIK, ClickHouse has two ways to drop partitions: by id and by value. By id is very simple, you provide the partition id and that is it. The value one is interesting, and is the one supported in this PR.

I can add support for dropping partition by id, but is the partition id for Iceberg tables actually exposed anywhere that the user would be able to know? It looks like system.iceberg_files uses a different format with braces ({'US', 19738}) from IcebergMetadata#formatPartitionKeyValue.

In ClickHouse MergeTree tables, the "value" for the drop partition is the value of each element of a partition expression tuple after it has gone through the "transforms".

You've implemented the opposite. On yours, you are required to provide the source values present in the columns.

I agree that it should be consistent if possible. I'll change it and make sure the correct function for each transform from source -> iceberg partition format is documented somewhere.

@arthurpassos

Copy link
Copy Markdown
Collaborator

I can add support for dropping partition by id, but is the partition id for Iceberg tables actually exposed anywhere that the user would be able to know? It looks like system.iceberg_files uses a different format with braces ({'US', 19738}) from IcebergMetadata#formatPartitionKeyValue.

No need to, I was just putting context in the message. The real issue to be tackled is the below:

In ClickHouse MergeTree tables, the "value" for the drop partition is the value of each element of a partition expression tuple after it has gone through the "transforms".
You've implemented the opposite. On yours, you are required to provide the source values present in the columns.

@xieandrew

Copy link
Copy Markdown
Collaborator Author

Is there a simpler way to pass a function in a single value partition? It looks like a bare function only works in a partition tuple, so I found ALTER TABLE {ch_table} DROP PARTITION tuple(toRelativeDayNum(toDateTime64('2024-01-16', 6))) as a workaround but it looks strange.

Signed-off-by: Andrew Xie <dev@xie.is>

- DELETED entries in partial manifests are now omitted, so they can't be turned back into live files
- Failed catalog commit doesn't delete manifest list or leave orphaned files
- Fix stale cache for latest metadata
- Iceberg table with no snapshot now no-ops instead of throwing an exception to match MergeTree
- A failure when purging data files now logs all the paths that could still exist
Signed-off-by: Andrew Xie <dev@xie.is>

- Also remove a failed test that pyiceberg can't produce the right conditions for
@xieandrew
xieandrew force-pushed the feature/antalya-26.6/iceberg-alter-table-drop-partition branch from d0cf3cf to 1407255 Compare October 2, 2026 14:38

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DROP PARTITION for Iceberg tables

5 participants