Skip to content

feat(physical): compile summaries over raw-sample precompute boundaries - #488

Open
zzylol wants to merge 3 commits into
feat/physical-compile-coverage-3from
feat/precompute-raw-sample-input
Open

zzylol wants to merge 3 commits into
feat/physical-compile-coverage-3from
feat/precompute-raw-sample-input

Conversation

@zzylol

@zzylol zzylol commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Stacked on #487.

Why

The backend still evaluates raw-ingest summary updates itself (value/item expressions, per-series and grouped update), because Planner's precompute compiler only accepts stored summary states as boundaries. For a raw time-series scan it failed with stored population requires one typed summary state; the general compiler failed per-series summaries with per-entity summary requires complete source identity. Planner should own this computation; a deployment should only supply rows and panes.

What

  • precompute::compile accepts a raw time-series scan (Fallback over Scan{TimeSeries}, optionally under TimeRange) as a boundary, bound as raw sample rows [$population: label map, $timestamp, value] (precompute::raw_sample_schema, rows via raw_sample_row). The canonical label map is the complete source identity, so per-series summaries need no identity-typed plan.
  • SummaryAgg over that boundary lowers per-series and grouped (by/without), with constant or column weights; HLL unit-frequency updates observe the sample value; CMS/CountSketch heaps resolve items from labels, the sample value, or the canonical label identity (optionally excluding labels, as topk by emits) via new Expression::Label / Expression::LabelIdentity.
  • Rejected: counter-derivative weights over raw cumulative samples, signed CMS weights, non-label scan columns as items, the series identity as a grouping key. Pane geometry and storage formats stay outside Planner.

Before this PR

Backend probe over 16 PromQL queries (exact, Epsilon, EpsilonDelta; 75 raw stored outputs): 11 compile, 64 fail (per-entity summary requires complete source identity); no raw boundary is accepted by precompute::compile.

After this PR

The same probe: 75/75 raw outputs compile through precompute::compile(dag, &[raw_source], &[summary]). New acceptance test precompute_raw_samples runs every Planner raw-input candidate for 18 queries (Sum, Count, Min, Max, Rate, Increase, KLL, DDSketch, HLL, CountSketchWithHeap from Planner; CmsWithHeap hand-built) and matches each population's estimates against its kernel fed sample by sample. Families without a native state (plain CMS/CountSketch, Kmv, Theta, UnivMon) still fail to compile, as before.

Behaviour differences

  • The existing finalized-readout SummaryAgg fragment is unchanged (items over finalized readouts remain rejected).
  • Raw values must be finite (FiniteFloat64); stale markers are not samples.
  • Raw label maps must be canonical (sorted, unique, no empty values); raw_sample_row guarantees it.

Remaining

  • Heap items with excluded identity labels (topk by) have no readout-side decoder in this crate yet; a reader must add group labels back from $population.
  • sum without (instance) (sum_over_time(m[5m])) candidates hit a pre-existing compile_post_asap_dag SummaryFamilySchemaMismatch; without is covered by an edited DAG.

Validation

cargo fmt --all -- --check, cargo clippy --workspace --all-targets --all-features --locked -- -D warnings, cargo test --workspace --no-fail-fast (pass). The acceptance test fails on the base (stored population requires one typed summary state) and passes here. Reviewed by a separate reviewer agent; its findings are addressed in the follow-up commits.

🤖 Generated with Claude Code

zzylol and others added 3 commits September 30, 2026 06:29
A precompute boundary at a raw time-series scan binds as raw sample rows
($population label map, $timestamp, value). The label map is the complete
source identity, so per-series and grouped SummaryAgg lower over it with
constant, column or unit-frequency (HLL) updates, and keyed heaps resolve
their items from labels, the sample value or the canonical label identity.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Raw sample rows drop empty label values so one series has one population.
Heap items resolve names against the scan (the value column, the series
identity, labels; other scan columns are rejected), and identity items may
exclude labels, as `topk by` emits. Unit-frequency HLL updates are raw-only
and never apply to keyed families. Tests cover Planner-generated heaps,
missing and empty labels, and `without` grouping.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
State the canonical label-set obligation and assert per-query coverage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant