Skip to content

Antalya 26.8: Allow empty object storage cluster - #2469

Merged
zvonand merged 2 commits into
antalya-26.8from
feature/antalya-26.8/pr-2221
Oct 6, 2026
Merged

zvonand merged 2 commits into
antalya-26.8from
feature/antalya-26.8/pr-2221

Conversation

@zvonand

@zvonand zvonand commented Oct 1, 2026

Copy link
Copy Markdown
Member

Changelog category (leave one):

  • Improvement

Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):

Allow empty object storage cluster (#2221 by @ianton-ru).

CI/CD Options

Exclude tests:

  • Fast test
  • Integration Tests
  • Stateless tests
  • Stateful tests
  • Performance tests
  • All with ASAN
  • All with TSAN
  • All with MSAN
  • All with UBSAN
  • All with Coverage
  • All with Aarch64
  • All Regression
  • Disable CI Cache

Regression jobs to run:

  • Fast suites (mostly <1h)
  • Aggregate Functions (2h)
  • Alter (1.5h)
  • Benchmark (30m)
  • ClickHouse Keeper (1h)
  • Iceberg (2h)
  • LDAP (1h)
  • Parquet (1.5h)
  • RBAC (1.5h)
  • SSL Server (1h)
  • S3 (2h)
  • S3 Export (2h)
  • Swarms (30m)
  • Tiered Storage (2h)

Cherry-picked from #2221.


Rebase of #2028

Documentation entry for user-facing changes

With 'object_storage_cluster' setting query to s3,iceberg and some other sources are executed as cluster request.
But with swarm cluster, when initiator is not a cluster member, may be situation when no one swarm node is alive at the moment. In this case query is failed with CLUSTER_DOESNT_EXIST error.

New setting object_storage_cluster_fallback_if_empty allow to execute read query on local node in this case.

Write query is not executed on cluster right now, so attempt to write is still failed in this case to avoid situation when query is success when swarm is empty and failed when has some nodes alive.

PR is a little bit complex because:
s3(...) - can fall back if object_storage_cluster is empty (cluster does not have active nodes, not 'empty setting value')
s3(...) SETTINGS object_storage_remote_initiator=1 - failed on local node if object_storage_cluster is empty
s3(...) SETTINGS object_storage_remote_initiator=1, object_storage_remote_initiator_cluster='...' - decision about falling back must be made on remote initiator, on local node object_storage_cluster can be unknown.

But behavior is not changed for Cluster functions:

s3Cluster(...) - can't fall back
s3Cluster(...) SETTINGS object_storage_remote_initiator=1 - must failed on remote initiator if object_storage_cluster is empty.

…next commit)

---
Original cherry-pick message follows:

Merge pull request #2221 from Altinity/feature/antalya-26.6/object_storage_cluster_allow_empty

Antalya 26.6: Allow empty object storage cluster
# Conflicts:
#	src/Storages/ObjectStorage/StorageObjectStorageCluster.cpp
Took the source PR's `getQueryProcessingStage` body (`resolveClusterRead` + `shouldReadLocallyOnFallbackToPure`). Kept `readsFromPureStorage` from antalya-26.8 since `appliesRowLevelFilterInRead` still uses it.
@zvonand zvonand added releasy Created/managed by RelEasy antalya-26.8 Session label (releasy session config) forwardport This is a frontport of code that existed in previous Antalya versions ai-resolved Port conflict auto-resolved by Claude labels Oct 1, 2026
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

Workflow [PR], commit [a4dc537]

@zvonand zvonand mentioned this pull request Oct 3, 2026
37 tasks done
@svb-alt svb-alt added the antalya label Oct 4, 2026
@zvonand

zvonand commented Oct 6, 2026

Copy link
Copy Markdown
Member Author

@blau-ai

@blau-ai

blau-ai commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

CI triage — @zvonand

Verdict: 0 failures attributable to this PR. The red checks are 1 pre-existing/flaky crash in an unrelated subsystem and 2 base-image vulnerability scans. Every functional check that exercises this PR's code is green (builds on all targets, stateless incl. the new 04303_object_storage_cluster_fallback_to_local_if_empty, integration incl. the modified test_s3_cluster, AST fuzzer, BuzzHouse, compatibility, docker images).

Failure-by-failure

1. Stress test (amd_debug) — 1/5 — NOT PR-caused (pre-existing/flaky, unrelated subsystem)

Logical error: 'Cannot write to canceled buffer'.
...
src/IO/WriteBuffer.cpp:67: DB::WriteBuffer::write(char)
src/Core/PostgreSQLProtocol.h:362: ...ErrorOrNoticeResponse::serialize(WriteBuffer&)
src/Server/PostgreSQLHandler.cpp:1063: DB::PostgreSQLHandler::processQuery()
src/Server/PostgreSQLHandler.cpp:441: DB::PostgreSQLHandler::run()

This is the PostgreSQL wire-protocol handler trying to serialize an error response onto a connection whose buffer was already canceled (client disconnected mid-query) — a known race that surfaces under stress. It has nothing to do with this PR, which touches only Settings, Cluster, and IStorageCluster/StorageObjectStorageCluster (S3/object-storage cluster reads). It matches the open upstream bug ClickHouse/ClickHouse#123785 ("PostgreSQL auth path throws LOGICAL_ERROR when client resets mid-authentication"). Only 1 of 5 stress shards hit it; arm_debug and arm_release stress passed.
→ Next step: safe to re-run. Not a blocker for this PR; the underlying PG-handler bug is an upstream matter, not for this backport.

2 & 3. Grype Scan (keeper: 4 high/critical, server-alpine: 1 high/critical) — NOT PR-caused (base-image CVEs / infra)
These scan the built Docker images for OS-package CVEs in the base image (Alpine / keeper base), not C++ source. This PR changes no Dockerfile, dependency, or submodule, so it cannot add or remove these findings — they are present independent of the diff and are the same class of finding that lands on essentially every PR. They're addressed by the infra team via base-image bumps, not inside a feature backport.
→ Next step: no action for this PR. Track/clear via the usual base-image update process.

4. PR (aggregate gate) — NOT independent. This is the roll-up status; it is red only because of the three checks above. It will clear once they do.

Health check

This is a clean backport of #2221 ("Allow empty object storage cluster"). The functionally meaningful signal is all-green: it builds on every target, and the new behavior is directly exercised by both the new stateless test 04303_object_storage_cluster_fallback_to_local_if_empty and the updated test_s3_cluster integration test, both passing. No real test regression is visible — the PR looks ready from a CI standpoint; a re-run of the amd_debug stress shard should clear the only non-infra red.

@zvonand
zvonand merged commit a52c43f into antalya-26.8 Oct 6, 2026
314 of 321 checks passed
@zvonand zvonand added verified Approved for release port-antalya PRs to be ported to all new Antalya releases labels Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ai-resolved Port conflict auto-resolved by Claude antalya antalya-26.8 Session label (releasy session config) forwardport This is a frontport of code that existed in previous Antalya versions port-antalya PRs to be ported to all new Antalya releases releasy Created/managed by RelEasy verified Approved for release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants