Conversation
A precompute boundary at a raw time-series scan binds as raw sample rows ($population label map, $timestamp, value). The label map is the complete source identity, so per-series and grouped SummaryAgg lower over it with constant, column or unit-frequency (HLL) updates, and keyed heaps resolve their items from labels, the sample value or the canonical label identity. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Raw sample rows drop empty label values so one series has one population. Heap items resolve names against the scan (the value column, the series identity, labels; other scan columns are rejected), and identity items may exclude labels, as `topk by` emits. Unit-frequency HLL updates are raw-only and never apply to keyed families. Tests cover Planner-generated heaps, missing and empty labels, and `without` grouping. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
State the canonical label-set obligation and assert per-query coverage. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #487.
Why
The backend still evaluates raw-ingest summary updates itself (value/item expressions, per-series and grouped update), because Planner's precompute compiler only accepts stored summary states as boundaries. For a raw time-series scan it failed with
stored population requires one typed summary state; the general compiler failed per-series summaries withper-entity summary requires complete source identity. Planner should own this computation; a deployment should only supply rows and panes.What
precompute::compileaccepts a raw time-series scan (FallbackoverScan{TimeSeries}, optionally underTimeRange) as a boundary, bound as raw sample rows[$population: label map, $timestamp, value](precompute::raw_sample_schema, rows viaraw_sample_row). The canonical label map is the complete source identity, so per-series summaries need no identity-typed plan.SummaryAggover that boundary lowers per-series and grouped (by/without), with constant or column weights; HLL unit-frequency updates observe the sample value; CMS/CountSketch heaps resolve items from labels, the sample value, or the canonical label identity (optionally excluding labels, astopk byemits) via newExpression::Label/Expression::LabelIdentity.Before this PR
Backend probe over 16 PromQL queries (exact, Epsilon, EpsilonDelta; 75 raw stored outputs): 11 compile, 64 fail (
per-entity summary requires complete source identity); no raw boundary is accepted byprecompute::compile.After this PR
The same probe: 75/75 raw outputs compile through
precompute::compile(dag, &[raw_source], &[summary]). New acceptance testprecompute_raw_samplesruns every Planner raw-input candidate for 18 queries (Sum, Count, Min, Max, Rate, Increase, KLL, DDSketch, HLL, CountSketchWithHeap from Planner; CmsWithHeap hand-built) and matches each population's estimates against its kernel fed sample by sample. Families without a native state (plain CMS/CountSketch, Kmv, Theta, UnivMon) still fail to compile, as before.Behaviour differences
SummaryAggfragment is unchanged (items over finalized readouts remain rejected).FiniteFloat64); stale markers are not samples.raw_sample_rowguarantees it.Remaining
topk by) have no readout-side decoder in this crate yet; a reader must add group labels back from$population.sum without (instance) (sum_over_time(m[5m]))candidates hit a pre-existingcompile_post_asap_dagSummaryFamilySchemaMismatch;withoutis covered by an edited DAG.Validation
cargo fmt --all -- --check,cargo clippy --workspace --all-targets --all-features --locked -- -D warnings,cargo test --workspace --no-fail-fast(pass). The acceptance test fails on the base (stored population requires one typed summary state) and passes here. Reviewed by a separate reviewer agent; its findings are addressed in the follow-up commits.🤖 Generated with Claude Code