Skip to content
Open
62 changes: 62 additions & 0 deletions docs/en/engines/table-engines/integrations/iceberg.md
Original file line number Diff line number Diff line change
Expand Up @@ -106,6 +106,68 @@ The following table shows how Iceberg data types are mapped to ClickHouse data t
| `map` | `Map` |
| `struct` | `Tuple` |

### Aggregate function states {#aggregate-function-states}

Iceberg has no aggregate-state type, so ClickHouse stores the two aggregate-state types as ordinary
Iceberg values and records the ClickHouse type name in a `clickhouse.type` key on the schema field:

| ClickHouse type | Iceberg type | Stored as |
|---|---|---|
| `AggregateFunction(f, T...)` | `binary` | The serialized state, the same bytes `f` writes with the `-State` combinator |
| `SimpleAggregateFunction(f, T)` | whatever `T` maps to | An ordinary value of type `T` |

Other query engines ignore the key and see a plain `binary` (or `T`-typed) column.

Creating such a table requires
[`allow_experimental_aggregate_function_states_in_open_formats`](/operations/settings/settings#allow_experimental_aggregate_function_states_in_open_formats),
and so does every query that reads a column whose `clickhouse.type` names an `AggregateFunction` - an
`INSERT` or an `ALTER TABLE ... EXPORT PART` as much as a `SELECT`. A state is an opaque blob handed
to the deserializer of the aggregate function the table's metadata names, so with the setting enabled
it is that metadata, not the query, which chooses the deserializer; keep it disabled for tables from
untrusted sources, where such a field is then rejected rather than read as `String`. A
`SimpleAggregateFunction` column is not gated on read, holding ordinary values of its storage type.

The setting is read from the query that parses the schema, so a session `SET` or a `SETTINGS` clause
on that query supplies it. The `iceberg*` table functions build a fresh storage object per query, so
for them it takes effect per query; the table engine parses the schema once and keeps it, so the
query that first touches the table after `ATTACH` decides, and that outcome holds - including for
queries that do not set it - until `DETACH TABLE` or a server restart.

Writing goes through the ordinary Parquet writer, which takes the same setting from its format
settings. An object storage table engine freezes its format settings at `CREATE TABLE` - the server
settings plus that query's `SETTINGS` clause, session settings ignored - so for `INSERT` the setting
must also be in that clause; given only on the `INSERT`, it does not reach the writer:

```sql
CREATE TABLE agg (k UInt32, u AggregateFunction(uniq, UInt64), s SimpleAggregateFunction(sum, UInt64))
ENGINE = IcebergLocal('/path/to/table/')
PARTITION BY k
SETTINGS allow_experimental_aggregate_function_states_in_open_formats = 1;

-- The states merge exactly as they do in an AggregatingMergeTree table.
SELECT k, uniqMerge(u), sum(s) FROM agg GROUP BY k
SETTINGS allow_experimental_aggregate_function_states_in_open_formats = 1;
```

`ALTER TABLE ... EXPORT PART` and `ALTER TABLE ... EXPORT PARTITION` from an `AggregatingMergeTree`
table into such an Iceberg table work as well, so a partition of pre-aggregated states can be moved
into the lake without finalizing it. Both take the setting off the `ALTER` query itself:

```sql
ALTER TABLE mt EXPORT PART 'all_1_1_0' TO TABLE agg
SETTINGS allow_experimental_aggregate_function_states_in_open_formats = 1;

-- EXPORT PARTITION, which only ReplicatedMergeTree implements, records the value in its manifest,
-- so every replica executing the task applies it in place of its own profile.
ALTER TABLE rmt EXPORT PARTITION ID '1' TO TABLE agg
Comment thread
zvonand marked this conversation as resolved.
SETTINGS allow_experimental_aggregate_function_states_in_open_formats = 1;
```

The Parquet footer of each data file carries the same information, so `DESCRIBE file('data.parquet')`
reports the aggregate types too; that path goes through Parquet schema inference, gated by the same
setting. See
[aggregate function states in Parquet](/interfaces/formats/Parquet#aggregate-function-states).

## Schema evolution {#schema-evolution}
ClickHouse supports reading Iceberg tables whose schema has evolved over time. This includes tables where columns have been added, removed, or reordered, as well as columns changed from required to nullable. Additionally, the following type casts are supported:

Expand Down
35 changes: 35 additions & 0 deletions docs/en/interfaces/formats/Parquet/Parquet.md
Original file line number Diff line number Diff line change
Expand Up @@ -95,6 +95,41 @@ On write, top-level columns of type `Point`, `LineString`, `Polygon`, `MultiLine

Geometry columns must appear at the root of the schema or nested inside `Tuple` (`struct`); nesting them inside `Array` or `Map` is not supported. `Nullable` is not supported for geo columns either.

## Aggregate function states {#aggregate-function-states}
Comment thread
zvonand marked this conversation as resolved.

Parquet has no aggregate-state type, so ClickHouse writes an [`AggregateFunction`](/sql-reference/data-types/aggregatefunction.md) column as a plain `BYTE_ARRAY` holding the serialized state - the same bytes the `-State` combinator produces - with no `STRING`/`UTF8` logical type, since a state is arbitrary binary data rather than text. A [`SimpleAggregateFunction(f, T)`](/sql-reference/data-types/simpleaggregatefunction.md) column is written as an ordinary value of type `T`.

Neither type can be recovered from the Parquet schema alone: every `AggregateFunction` state is just a binary column, and a `SimpleAggregateFunction` is indistinguishable from its storage type. ClickHouse therefore records the type names in a `clickhouse.column_types` key in the file-level Parquet metadata, as a JSON object mapping column name to ClickHouse type name. The recorded name includes the state version, which is what pins the serialized layout across server versions.

Both writing an `AggregateFunction` column and reconstructing one from that metadata on read are gated by [`allow_experimental_aggregate_function_states_in_open_formats`](/operations/settings/settings#allow_experimental_aggregate_function_states_in_open_formats), which is disabled by default. A state is an opaque blob passed to the deserializer of whichever aggregate function the file names, so with the setting enabled it is the file, not the query, choosing that deserializer; keep it disabled for files from untrusted sources, where such a file is then rejected rather than read as `String`. With it disabled, writing is refused with `UNKNOWN_TYPE`, exactly as in versions that did not support states in Parquet at all. `SimpleAggregateFunction` is gated in neither direction: it is stored as an ordinary value of its storage type and has always been written that way.

With the setting enabled the states round-trip without any hint:

```sql
SET allow_experimental_aggregate_function_states_in_open_formats = 1;

INSERT INTO FUNCTION file('states.parquet')
SELECT k, uniqState(v) AS u, sumSimpleState(v) AS s FROM source GROUP BY k;

DESCRIBE file('states.parquet');
-- k UInt64
-- u AggregateFunction(uniq, UInt64)
-- s SimpleAggregateFunction(sum, UInt64)

SELECT k, uniqMerge(u), sum(s) FROM file('states.parquet') GROUP BY k;
```

An explicit structure overrides the recorded metadata and is unaffected by the setting, so a state column can also be read as `String` to get the raw serialized bytes:

```sql
SELECT uniqMerge(CAST(u AS AggregateFunction(uniq, UInt64)))
FROM file('states.parquet', Parquet, 'u String');
```

Reading fails rather than guessing if the recorded type does not describe the data actually in the file. Min/max statistics are never used for pruning a state column, because bounds over serialized states are meaningless.

Other query engines see a plain binary (or `T`-typed) column and ignore the metadata key. The same mechanism backs [aggregate-state support in Iceberg tables](/engines/table-engines/integrations/iceberg#aggregate-function-states), whose data files are Parquet.

## Example usage {#example-usage}

### Inserting data {#inserting-data}
Expand Down
10 changes: 10 additions & 0 deletions src/Core/Settings.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -7433,6 +7433,16 @@ Query Iceberg table using the specific snapshot id.
)", 0) \
DECLARE(Bool, allow_experimental_geo_types_in_iceberg, false, R"(
Allow parsing Iceberg `geometry` and `geography` field types as ClickHouse `Geometry` (Variant) type.
)", 0) \
DECLARE(Bool, allow_experimental_aggregate_function_states_in_open_formats, false, R"(
Allow `AggregateFunction` states in Parquet files and Iceberg tables, and `SimpleAggregateFunction`
columns in Iceberg tables. A state is stored as opaque binary, with its ClickHouse type recorded in
`clickhouse.column_types` (Parquet file metadata) or `clickhouse.type` (Iceberg schema field).

Gates both writing such columns and reconstructing their types on read. When enabled, the data, not
the query, chooses the deserializer for the stored bytes, so keep it disabled for untrusted sources.
More [in Parquet](/interfaces/formats/Parquet#aggregate-function-states) and
[in Iceberg](/engines/table-engines/integrations/iceberg#aggregate-function-states).
)", 0) \
DECLARE(Bool, show_data_lake_catalogs_in_system_tables, false, R"(
Enables showing data lake catalogs in system tables.
Expand Down
1 change: 1 addition & 0 deletions src/Core/SettingsChangesHistory.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,7 @@ const VersionToSettingsChangesMap & getSettingsChangesHistory()
addSettingsChanges(settings_changes_history, "26.6.2.20001.altinityantalya",
{
{"use_puffin_files_cache", false, true, "Enables cache of parsed Puffin file content such as deletion vectors."},
{"allow_experimental_aggregate_function_states_in_open_formats", false, false, "New setting gating aggregate function states in Parquet files and Iceberg tables. Disabled by default, so writing such a Parquet column keeps throwing `UNKNOWN_TYPE` and such an Iceberg column keeps being refused with `SUPPORT_IS_DISABLED` as in versions without the feature."},
});

addSettingsChanges(settings_changes_history, "26.6",
Expand Down
211 changes: 206 additions & 5 deletions src/DataTypes/DataTypeAggregateFunction.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -5,10 +5,20 @@

#include <Common/SipHash.h>
#include <Common/AlignedBuffer.h>
#include <Common/quoteString.h>
#include <Common/FieldVisitorToString.h>

#include <Formats/FormatSettings.h>
#include <DataTypes/DataTypeAggregateFunction.h>
#include <DataTypes/DataTypeCustomSimpleAggregateFunction.h>
#include <DataTypes/DataTypeArray.h>
#include <DataTypes/DataTypeMap.h>
#include <DataTypes/DataTypeDateTime64.h>
#include <DataTypes/DataTypeFixedString.h>
#include <DataTypes/DataTypeLowCardinality.h>
#include <DataTypes/DataTypeNullable.h>
#include <DataTypes/DataTypeTime64.h>
#include <DataTypes/DataTypeTuple.h>
#include <DataTypes/Serializations/SerializationAggregateFunction.h>
#include <DataTypes/DataTypeFactory.h>
#include <DataTypes/transformTypesRecursively.h>
Expand All @@ -20,7 +30,9 @@

#include <AggregateFunctions/AggregateFunctionFactory.h>
#include <AggregateFunctions/IAggregateFunction.h>
#include <Parsers/ASTDataType.h>
#include <Parsers/ASTFunction.h>
#include <Parsers/ASTIdentifier.h>
#include <Parsers/ASTIdentifier_fwd.h>
#include <Parsers/ASTLiteral.h>

Expand Down Expand Up @@ -54,13 +66,19 @@ String DataTypeAggregateFunction::getFunctionName() const

String DataTypeAggregateFunction::doGetName() const
{
return getNameImpl(true);
return getNameImpl(true, false);
}


String DataTypeAggregateFunction::getNameWithoutVersion() const
{
return getNameImpl(false);
return getNameImpl(false, false);
}


String DataTypeAggregateFunction::getNameForAnnotation() const
{
return getNameImpl(true, true);
}


Expand Down Expand Up @@ -91,14 +109,14 @@ void DataTypeAggregateFunction::updateVersionFromRevision(size_t revision, bool
setVersion(function->getVersionFromRevision(revision), if_empty);
}

String DataTypeAggregateFunction::getNameImpl(bool with_version) const
String DataTypeAggregateFunction::getNameImpl(bool with_version, bool always_emit_version) const
{
WriteBufferFromOwnString stream;
stream << "AggregateFunction(";

/// If aggregate function does not support versioning its version is 0 and is not printed.
/// Version 0 is normally omitted, but annotations must distinguish it from the default version.
auto data_type_version = getVersion();
if (with_version && data_type_version)
if (with_version && (data_type_version || (always_emit_version && isVersioned())))
stream << data_type_version << ", ";
stream << function->getName();

Expand Down Expand Up @@ -478,4 +496,187 @@ bool hasAggregateFunctionType(const DataTypePtr & type)
return result;
}

bool astHasAggregateFunctionType(const ASTPtr & ast)
{
std::string_view name;
if (const auto * data_type = ast->as<ASTDataType>())
name = data_type->name;
else if (const auto * identifier = ast->as<ASTIdentifier>())
name = identifier->name();

if (name == "AggregateFunction")
return true;

for (const auto & child : ast->children)
if (astHasAggregateFunctionType(child))
return true;

return false;
}

bool needsClickHouseTypeAnnotation(const DataTypePtr & type)
{
auto result = false;
auto check = [&](const IDataType & t)
{
result |= WhichDataType(t).isAggregateFunction()
|| typeid_cast<const DataTypeCustomSimpleAggregateFunction *>(t.getCustomName()) != nullptr;
};

check(*type);
type->forEachChild(check);
return result;
}

String getClickHouseTypeAnnotationName(const DataTypePtr & type)
{
if (type->getCustomName())
return type->getName();

switch (type->getTypeId())
{
case TypeIndex::AggregateFunction:
return assert_cast<const DataTypeAggregateFunction &>(*type).getNameForAnnotation();
case TypeIndex::Array:
return "Array(" + getClickHouseTypeAnnotationName(assert_cast<const DataTypeArray &>(*type).getNestedType()) + ")";
case TypeIndex::Map:
{
const auto & map_type = assert_cast<const DataTypeMap &>(*type);
return "Map(" + getClickHouseTypeAnnotationName(map_type.getKeyType()) + ", "
+ getClickHouseTypeAnnotationName(map_type.getValueType()) + ")";
}
case TypeIndex::Tuple:
{
const auto & tuple_type = assert_cast<const DataTypeTuple &>(*type);
const auto & elements = tuple_type.getElements();
const auto & names = tuple_type.getElementNames();
WriteBufferFromOwnString stream;
stream << "Tuple(";
for (size_t i = 0; i < elements.size(); ++i)
{
if (i)
stream << ", ";
if (tuple_type.hasExplicitNames())
stream << backQuoteIfNeed(names[i]) << ' ';
stream << getClickHouseTypeAnnotationName(elements[i]);
}
stream << ")";
return stream.str();
}
default:
return type->getName();
}
}

namespace
{

bool isFixedStringOfSize(const DataTypePtr & type, size_t size)
{
const auto * fixed_string = typeid_cast<const DataTypeFixedString *>(type.get());
return fixed_string && fixed_string->getN() == size;
}

bool isDateTime64WithScale(const DataTypePtr & type, UInt32 scale)
{
const auto * date_time64 = typeid_cast<const DataTypeDateTime64 *>(type.get());
return date_time64 && date_time64->getScale() == scale;
}

/// Mirrors the non-injective mappings in Parquet's `preparePrimitiveColumn`.
bool parquetWriterCouldProduce(const DataTypePtr & annotated, const DataTypePtr & derived)
{
if (annotated->equals(*derived))
return true;

auto next_timestamp_unit = [](UInt32 scale) -> UInt32
{
if (scale <= 3)
return 3;
if (scale <= 6)
return 6;
return 9;
};

switch (annotated->getTypeId())
{
case TypeIndex::Date:
return WhichDataType(derived).isDate32() || WhichDataType(derived).isUInt16();
case TypeIndex::DateTime:
return isDateTime64WithScale(derived, 3) || WhichDataType(derived).isUInt32();
case TypeIndex::DateTime64:
return isDateTime64WithScale(
derived, next_timestamp_unit(assert_cast<const DataTypeDateTime64 &>(*annotated).getScale()));
case TypeIndex::Time:
return isDateTime64WithScale(derived, 6);
case TypeIndex::Time64:
return isDateTime64WithScale(
derived, assert_cast<const DataTypeTime64 &>(*annotated).getScale() <= 6 ? 6 : 9);
case TypeIndex::Enum8:
return isString(derived) || WhichDataType(derived).isInt8();
case TypeIndex::Enum16:
return isString(derived) || WhichDataType(derived).isInt16();
case TypeIndex::IPv4:
return WhichDataType(derived).isUInt32();
case TypeIndex::IPv6:
case TypeIndex::UInt128:
case TypeIndex::Int128:
return isFixedStringOfSize(derived, 16);
case TypeIndex::UInt256:
case TypeIndex::Int256:
return isFixedStringOfSize(derived, 32);
case TypeIndex::FixedString:
return isString(derived);
case TypeIndex::Object:
return isString(derived);
default:
return false;
}
}

}

bool annotatedTypeMatchesDerived(const DataTypePtr & annotated_type, const DataTypePtr & derived_type, bool strict)
{
auto annotated = removeLowCardinalityAndNullable(annotated_type);
auto derived = removeLowCardinalityAndNullable(derived_type);

if (annotated->getTypeId() == TypeIndex::AggregateFunction)
return isString(derived);

switch (annotated->getTypeId())
{
case TypeIndex::Array:
{
const auto * derived_array = typeid_cast<const DataTypeArray *>(derived.get());
if (!derived_array)
return false;
return annotatedTypeMatchesDerived(
assert_cast<const DataTypeArray &>(*annotated).getNestedType(), derived_array->getNestedType(), strict);
}
case TypeIndex::Tuple:
{
const auto & annotated_tuple = assert_cast<const DataTypeTuple &>(*annotated);
const auto * derived_tuple = typeid_cast<const DataTypeTuple *>(derived.get());
if (!derived_tuple || derived_tuple->getElements().size() != annotated_tuple.getElements().size())
return false;
for (size_t i = 0; i < annotated_tuple.getElements().size(); ++i)
if (!annotatedTypeMatchesDerived(annotated_tuple.getElement(i), derived_tuple->getElement(i), strict))
return false;
return true;
}
case TypeIndex::Map:
{
const auto & annotated_map = assert_cast<const DataTypeMap &>(*annotated);
const auto * derived_map = typeid_cast<const DataTypeMap *>(derived.get());
if (!derived_map)
return false;
return annotatedTypeMatchesDerived(annotated_map.getKeyType(), derived_map->getKeyType(), strict)
&& annotatedTypeMatchesDerived(annotated_map.getValueType(), derived_map->getValueType(), strict);
}
default:
return !strict || parquetWriterCouldProduce(annotated, derived);
}
}

}
Loading
Loading