Repository navigation
feat(ir): add summary coverage metadata for time and population - #567
Conversation
|
TODO: should write a design doc for it, not developer doc, human written. |
|
According to this PR doc, I still don't think we are solving the right problem here. This is indeed annoying but it is not schema problem. Since the responsibility of schema is mostly to encode data structure, not data semantic. To determine if the above two sums need to be merged, we can either
|
Yes, this is not a schema problem, I will change the PR problem description title. Actually in the code implementation, the "what is summarized" is a field in the Node, in parallel with the field "schema" in the node struct. |
4a91371 to
02833d1
Compare
What does this sentence in the PR doc mean |
|
Random thinking: I feel the semantic Say, a scan node can report "I am scanning columns A, B on data D" This is a generalization of the coverage semantic, and if exist, can be useful for other stuffs like CSE reducing as well. That said, this is just a rough, in-mature idea. We can save that for the future and only focus on summary coverage on this PR. |
e5e4f67 to
fc8eec9
Compare
…unified IR (#645) #536 defined the unified OperatorNode DAG, but canonicalization, common sub-DAG sharing and a flat, serializable form existed only for the old split IRs. Add them on OperatorNode: - ir::canonicalize: heavy-hitter promotion and EXISTS / NOT EXISTS / IN subquery lowering, bottom-up and memoized so shared sub-DAGs stay shared. - ir::cse: hash-consing over a workload batch (the identical-expression rule of #509 Pass 2), following scalar-referenced operator nodes too. - Generic child references: Operator, NonASAPOp, ASAPOp, ScalarExpr, Predicate, ProjectItem, SortKey and QueryRoot take the child reference as a type parameter (default Rc<OperatorNode>). ir::flat::flatten writes a DAG as nodes whose operators are Operator<NodeId>. Summary coverage (#567) and SummaryMerge (#560) build on this. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…ts fields Name the metadata coverage to match #560 and the docs; drop the single-variant multiplicity and deployment-specific revision; rename grouping to reduction to match SummaryAgg; report failures through SchemaDerivationError::Coverage; revert the unrelated PaneCoverageError rename. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…me bounds validate_structure rejects a SummaryAgg without coverage (CoverageError::Missing). CoverageRegion time bounds become optional so tabular sources without a time column can declare coverage. Population stays trusted; #570 tracks checking it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Given required coverage on summary nodes, input and reduction duplicated the producing SummaryAgg fields; drop them along with ProducerMismatch. Type source as Source, matching Scan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SourceCoverage names the rows a physical scan reads for cost comparison, not which observations a summary state holds; rename it so it is not confused with SummaryCoverage. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move it to docs/design_docs/proposals with problem and motivation, requirements, design, alternatives and key code interfaces. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tive schema design Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The design document and the docs it consolidates are reviewed separately on main. This PR keeps code, tests and the ScanSelection rename in docs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fc8eec9 to
bb3bf69
Compare
Coverage is part of a summary state's identity: two states over different observations are never shared, so CSE hashes and compares it, and a flat node keeps it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bb3bf69 to
4750e71
Compare
|
According to my (and my agents', mostly my agents') understanding, this PR only defines the coverage types but did not answer how to calculate such a coverage. I think this is an important information that should be mentioned in the PR doc clearly but idk why the agent writing the PR didn't mention it. |
Yes, it is interesting. It's saying how to represent the semantic of a sub-DAG. |
|
You're right, the PR doc should say this. #567 only defines the coverage type and the merge rule; computing it is split across later PRs:
Producers don't write coverage by hand: they call |
Short answer: two states worth merging always cover different data. If coverage were part of Why. Example. Merge two KLL states for "latency by job", one built from minute 0–1 and one from minute 1–2:
If coverage were a field of So |


Rebased on main d4869a7 (DF 54). At this head
cargo fmt --all --check,cargo clippy --workspace --all-targets -- -D warningsandcargo test --workspacepass (1575 passed, 0 failed, 0 ignored).Closes #571.
Problem: a summary state does not record which observations it covers
Once the planner combines existing summary states (#560
SummaryMerge, reuse of ingested panes, sub-DAG sharing from #645), it must know which observations each state was built from: which time range and which population (label values). Nothing on anOperatorNoderecords this today. The producer that built the state is no longer visible after composition, and the edge'sSchema(#535) is the layout contract: field types and the committed state type, nothing more:That is correct for a schema: #535 deliberately keeps filters, grouping keys, timing and windows out of
SchemaandField("filters, reduction/group keys, … execution timing, window framework … are not additionalSchemaorFieldmembers"). So the information has to live somewhere else. The examples below show why it is needed: each uses two states with byte-for-byte equal schemas that must not be combined:Example 1: time. Equal schemas, different answers
[00:00, 00:01)[00:01, 00:02)[00:00, 00:02)[00:00, 00:02)[00:01, 00:03)[00:01, 00:02)is counted twice, which skews the quantile and doubles counts or frequencies[00:00, 00:01)[00:02, 00:03)[0,1) ∪ [2,3); wrong if the result is used for the continuous window[00:00, 00:03)time_indexis a column position. A KLL state has no timestamp column at all, sotime_indexisNonein all three rows, and the schema cannot tell these cases apart.Example 2: population (label values). Equal schemas, different answers
region='us'region='eu'us ∪ euwithin eachjobregion='us'tier='premium'region='us'region='us'regionis a filter label, not an output column, so it never appears in the schema. Thejobfield only says the state is grouped by job. It does not say which jobs or which rows contributed.Example 3: time and population together
A =
us × [0,1)and B =eu × [1,2). The merged state covers exactly those two blocks. Describing it as{us,eu} × [0,2)(the result of storing a time range and a label set separately) would claim EU data for[0,1)and US data for[1,2)that was never read. The metadata has to keep time and population paired per region.Example 4: answering a query from a stored state
Query:
p99(latency) WHERE region='us' AND ts IN [10:00, 10:05) GROUP BY job. A stored state with the matching schema could hold US data for 10:00–10:05, EU data, or US data for only 10:00–10:03. All three have the same schema. Today the planner can confirm that the state type fits, but not that the contents fit.Conclusion. Equal schemas are necessary but not sufficient for composing or reusing summaries; the planner also needs each state's coverage, which this PR adds as a node property next to
Schema, not inside it. Without time/population metadata, the planner must either refuse every composition or accept silent double counting and missing data.What this PR adds
Schemastays the layout contract from #535 and does not describe coverage. Coverage is a sibling field on the node, next toschema:coverageis required on summary nodes:validate_structurerejects aSummaryAgg(and, in #560, aSummaryMerge) whose coverage isNonewithCoverageError::Missing. Plain relational nodes leave itNone; the field is anOptiononly because all operators shareOperatorNode.Why coverage is not part of
Schema. Two states worth merging always cover different data.SummaryMerge(#560) requires all inputs to have the sameSchema; that check is how it knows they are the same kind of state (same sketch, parameters and grouping). For example, two KLL states for "latency by job", built from minute 0–1 and minute 1–2:schema(job: Utf8, state: KLL{k=200})(job: Utf8, state: KLL{k=200})coverage[0,1)[1,2)If coverage were a field of
Schema, these schemas would differ and the merge would be rejected; the only merge left would be a state with an exact copy of itself, which counts every observation twice. Soschemasays what kind of state this is, andcoveragesays which data it was built from.CSE (#645) hashes and compares
coverage, so two summary states over different observations are never shared, andir::flatkeeps it on each flat node.Every observation in a region is assumed to contribute once to the state.
Removing duplication
Given the new definition, coverage holds only what no other type records: time × population. This PR also removes the duplication that the first draft introduced or exposed:
inputandreductiondropped from coverage. They copiedSummaryAgg.input/SummaryAgg.reductionon the same node, andwith_coverageneeded aProducerMismatchcheck to keep the copies in sync. feat(ir): define compatible logical summary merges #560'sSummaryMergenow compares them on its producers throughOperatorNode::summary_update().sourceis aSource, not aString. It uses the same type asScan.source, so one table cannot have two spellings.SourceCoveragerenamed toScanSelection(asap-aware-mapping, about 100 call sites plus docs). It names the rows a physical scan reads for cost comparison, a different concept fromSummaryCoverage.revisionandmultiplicitywere removed earlier, as a deployment concern and a single-variant enum respectively.How the examples come out under
merge_disjoint:[0,1)+[1,2), same population[0,2)[0,1)+[2,3)[0,2)+[1,3)PossibleOverlapregion=us+region=eu, same timeregion=us+region=usPossibleOverlapregion=us+tier=premiumPossibleOverlap. Different label names prove nothing.us×[0,1)+eu×[1,2){us,eu}×[0,2)SourceMismatchSummaryMerge(#560)Rules:
validate_structure.Somewithregions = []means known empty.time_ms: Noneis for sources without a time column (plain tabular data). Such a region overlaps every region it is not population-disjoint from.with_coveragevalidates the declaration and requiresStateoutput (CoverageError::NotState).validate_structurere-checks it.with_coverage.How coverage is computed
This PR defines the type and the merge rule. Each part is filled in by a later PR, not written by hand:
sourceScanunder theSummaryAggpopulationfield = 'text'equality filters on the path from theSummaryAggto itsScan(SummaryAgg.filter,Filter,Scan.predicates). A declaration that does not match is rejected. If the filters cannot be read (another operator on the path, or>, regex,IN, …), the population is unrestricted and such states merge only as time panes of the same computation.time_ms[0, 1min).Producers call
SummaryCoverage::for_summary(&agg, time_ms)(#646), which builds a declaration that passes the check. Until #646 lands, a declaration is not checked against the subtree, so a wrong one could pass:#646 rejects A with
PopulationMismatch.Out of scope
sourceandpopulation: feat(ir): derive summary coverage as definition + selection #646 (closes Check declared summary coverage population against subtree filters #570). Setting pane time bounds: Pass 2 window composition: tumbling panes #601.Stack and validation
Order: #645 (merged) → #567 → #560 (
SummaryMergerequires and derives coverage) → #646 (coverage computed from the subtree) → #539 → #540 → #541 → #542 → #543.Tests in
crates/types/tests/summary_coverage.rscover adjacency, gaps, population disjointness, joint regions, overlap, regions without time bounds, source mismatch, serde round-trip, invalid declarations, required coverage, and clearing after rewrites. Each example in this body and the doc is also built as a realSummaryAgg→SummaryMergeplan in #560'ssummary_coverage_examples.rs. The design document is reviewed separately in #573.🤖 Generated with Claude Code