Skip to content

Support aggregate function states in Parquet and Iceberg - #2301

Open
zvonand wants to merge 10 commits into
antalya-26.6from
feature/antalya-26.6/aggregate-function-states-in-parquet-iceberg
Open

zvonand wants to merge 10 commits into
antalya-26.6from
feature/antalya-26.6/aggregate-function-states-in-parquet-iceberg

Conversation

@zvonand

@zvonand zvonand commented Sep 2, 2026 •

Copy link
Copy Markdown
Member

Closes #2206

Changelog category (leave one):

  • New Feature

Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):

Support aggregate function states in Parquet and Iceberg

CI/CD Options

Exclude tests:

  • Fast test
  • Integration Tests
  • Stateless tests
  • Stateful tests
  • Unit tests
  • Performance tests
  • Aarch64 tests
  • All with ASAN
  • All with TSAN
  • All with MSAN
  • All with UBSAN
  • All with Coverage
  • All Regression
  • Disable CI Cache

Regression jobs to run:

  • Fast suites (mostly <1h)
  • Aggregate Functions (2h)
  • Alter (1.5h)
  • Benchmark (30m)
  • CAS (content-addressed storage; Antalya only)
  • ClickHouse Keeper (1h)
  • Iceberg (2h)
  • LDAP (1h)
  • OAuth (5m)
  • Parquet (1.5h)
  • RBAC (1.5h)
  • SSL Server (1h)
  • S3 (2h)
  • S3 Export (2h)
  • Swarms (30m)
  • Tiered Storage (2h)

@zvonand

zvonand commented Sep 2, 2026

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 2, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-02T07:16:26.909845Z 090c16a Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@github-actions

github-actions Bot commented Sep 2, 2026 •

Copy link
Copy Markdown

Workflow [PR], commit [e44403c]

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 090c16ac1f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/Processors/Formats/Impl/Parquet/Write.cpp Outdated
Comment thread src/Processors/Formats/Impl/Parquet/SchemaConverter.cpp Outdated
Comment thread src/Storages/ObjectStorage/DataLakes/Iceberg/DataFileStatistics.cpp
@zvonand

This comment was marked as outdated.

@blau-ai

This comment was marked as outdated.

@zvonand

This comment was marked as outdated.

@blau-ai

This comment was marked as outdated.

@zvonand
zvonand force-pushed the feature/antalya-26.6/aggregate-function-states-in-parquet-iceberg branch 2 times, most recently from d5099b8 to ea636f0 Compare September 3, 2026 10:46
@zvonand

zvonand commented Sep 3, 2026

Copy link
Copy Markdown
Member Author

@blau-ai

@blau-ai

This comment was marked as outdated.

zvonand added a commit that referenced this pull request Sep 24, 2026
Reuse `ISerialization` for aggregate state conversion, trim nonessential comments, and consolidate overlapping tests while preserving distinct format and Iceberg coverage.

PR: #2301
Parquet and Iceberg have no aggregate-state type, so neither can describe an
`AggregateFunction` or `SimpleAggregateFunction` column with its schema alone.
This adds a ClickHouse-only annotation next to the data - the
`clickhouse.column_types` key of a parquet file's footer, and a `clickhouse.type`
key on an Iceberg schema field - naming the ClickHouse type so it can be rebuilt
on read. An `AggregateFunction` state is stored as an opaque `BYTE_ARRAY` /
`binary` holding the bytes the `-State` combinator produces; a
`SimpleAggregateFunction(f, T)` is stored as an ordinary value of `T`. Other
query engines ignore the key and see a plain binary or `T`-typed column.

The recorded name always carries the state version, which pins the serialized
layout: `getName` drops a zero version, so a versioned function pinned to
version 0 would otherwise be rebuilt with the default version and its bytes
misread. `getNameForAnnotation` keeps it.

Two experimental settings gate the feature, both off by default:
`allow_experimental_aggregate_function_states_in_parquet` for writing such a
column and for reconstructing one during parquet schema inference, and
`allow_experimental_aggregate_function_states_in_iceberg` for `CREATE TABLE`
and for honouring a `clickhouse.type` key that names an `AggregateFunction`.
Honouring the annotation lets the file or the table metadata, rather than the
query, choose the deserializer the stored bytes are handed to, so a refused
annotation is rejected rather than read as `String` - reading states as strings
would be a wrong result, not an error. `SimpleAggregateFunction` needs no opt-in
on read, holding ordinary values with no state deserializer involved.

The annotation is checked against the type the schema derives on its own before
it is honoured, so a stale or crafted one cannot re-type a column to anything of
the same nesting shape. The parquet reader checks strictly, requiring a type its
own writer maps to what the file holds; Iceberg checks the nesting structure,
which is all its type system allows.

The Iceberg gate is read from the context of the query that parses the schema
rather than from one captured at `ATTACH`, so a reading query can opt in at all;
`IcebergSchemaProcessor::addIcebergTableSchema` publishes a schema-id only once
every field has parsed and drops what it wrote on failure, so a refused read can
be retried with the setting on. The metadata prefetcher asks for the schema-id
with `SchemaParsing::Skip`, since it has no query to carry the gate and must not
publish a schema that later queries would be served. The parquet schema cache is
keyed by both settings, so a schema inferred with a gate open is not served to a
query that has it closed.

`ALTER TABLE ... EXPORT PART` and `EXPORT PARTITION` from an `AggregatingMergeTree`
table into an Iceberg table carry both gates: `EXPORT PARTITION` records them in
its ZooKeeper manifest, so every replica executing the task applies the values
the `ALTER` gave instead of its own profile, and the destination is resolved
under those settings. A manifest without the fields, written before they existed,
reads them as closed.

Iceberg records no bounds for aggregate states, whose extremes cannot be
compared, and parquet min/max statistics over serialized states are never used
for pruning.

Adds `04673_parquet_aggregate_function_state`, integration tests for the export
paths and for round-tripping states through Spark, and unit tests for the
annotation matching, the name parsing, the schema processor and the metadata
generator. Documents the feature in `Parquet.md` and `iceberg.md`.

PR: #2301
@zvonand
zvonand force-pushed the feature/antalya-26.6/aggregate-function-states-in-parquet-iceberg branch from de5caea to 322a47b Compare September 24, 2026 11:23
@zvonand zvonand added the port-antalya PRs to be ported to all new Antalya releases label Sep 28, 2026
Reuse `removeLowCardinalityAndNullable`, inline the one-use schema helper,
and remove redundant footer state and comments. Share metadata-test lookup
code while preserving all regression scenarios.

Validated with a Debug build, 103 targeted unit tests, and the Parquet
aggregate-state regression matching its reference output.

Related: #2301
…et-iceberg' of github.com:Altinity/ClickHouse into feature/antalya-26.6/aggregate-function-states-in-parquet-iceberg

@arthurpassos arthurpassos left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have reviewed the docs and the parquet implementation, looks ok so far. I'll review the iceberg impl and the tests tomorrow. Few comments below

Comment thread docs/en/engines/table-engines/integrations/iceberg.md
Comment thread docs/en/interfaces/formats/Parquet/Parquet.md
Comment thread src/Processors/Formats/Impl/Parquet/PrepareForWrite.cpp Outdated
Comment thread src/Processors/Formats/Impl/Parquet/SchemaConverter.cpp Outdated
namespace parq = parquet::format;

/// Maps top-level columns to ClickHouse types not represented by the Parquet schema.
constexpr const char * clickhouse_column_types_key = "clickhouse.column_types";

@arthurpassos arthurpassos Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hm.. it would be funny if upstream one day decided to use the very same kvp for a different implementation, but I suppose it is extremely unlikely

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

well, yes, there is always a chance of something like this. but I do not see any other option actually.

@zvonand zvonand Oct 2, 2026 •

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think I just take this PR into upstream later, there is (mostly) nothing antalya-specific in it

Poco::JSON::Parser parser;
const auto object = parser.parse(kv.value).extract<Poco::JSON::Object::Ptr>();
for (const auto & name : object->getNames())
clickhouse_column_type_names[name] = object->getValue<String>(name);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A json key value pair inside a parquet key value parquet, that's funny, but probably the best way

std::unordered_map<String, GeoColumnMetadata> geo_columns;

/// Type names from the `clickhouse.column_types` footer metadata.
std::unordered_map<String, String> clickhouse_column_type_names;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would make the name a bit more descriptive by adding the custom keyword

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

agree. claude suggested annotated_column_type_names, which IMO looks even better


DataTypePtr SchemaConverter::resolveAnnotatedType(const String & column_name, const String & type_name) const
{
ASTPtr ast;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Instead of parsing text into an AST, ins't it easier to simply do the below? Ofc it assumes type_name is the "full type name" the data type factory requires

auto data_type = DataTypeFactory::instance().get(type_name));

if (WhichDataType(data_type).isAggregateFunction() && !options.format.parquet.allow_aggregate_function_states)
    throw

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, I suppose it is not that simple because it might be a nested data type. In that case you'd need something similar to what you've implemented in needsClickHouseTypeAnnotation

{
annotated_type = resolveAnnotatedType(col.name, it->second);

if (!annotatedTypeMatchesDerived(annotated_type, col.output_type, /*strict=*/ true))

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel like annotatedTypeMatchesDerived should be called from inside the resolveAnnotatedType, but it is up to you

@arthurpassos arthurpassos left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the tests, looks ok. Pending part is the iceberg implementation

Comment thread tests/integration/test_export_merge_tree_part_to_iceberg/test.py Outdated
Comment thread tests/queries/0_stateless/02902_topKGeneric_deserialization_memory.sql Outdated
`SchemaConverter::resolveAnnotatedType` now takes the type derived from the parquet schema and verifies it against the annotated type with `annotatedTypeMatchesDerived`, so it only returns a validated type.

Related: #2301 (comment)
…egate_function_states_in_open_formats`

Replace `allow_experimental_aggregate_function_states_in_parquet` and `allow_experimental_aggregate_function_states_in_iceberg` with a single setting. Both gated the same trust decision: honouring a ClickHouse type recorded in the data, which chooses the deserializer for the stored bytes. Each check keeps its current behavior. The export task and manifest now carry one flag.

Related: #2301 (comment)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

antalya-26.6 port-antalya PRs to be ported to all new Antalya releases

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Support AggregateState in Parquet/Iceberg

4 participants