Skip to content

feat(ir): add summary coverage metadata for time and population - #567

Merged
zzylol merged 18 commits into
mainfrom
feat/summary-coverage-contract
Oct 6, 2026
Merged

zzylol merged 18 commits into
mainfrom
feat/summary-coverage-contract

Conversation

@zzylol

@zzylol zzylol commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Rebased on main d4869a7 (DF 54). At this head cargo fmt --all --check, cargo clippy --workspace --all-targets -- -D warnings and cargo test --workspace pass (1575 passed, 0 failed, 0 ignored).

Design doc: the schema and summary-semantics design for this PR (problem, design considerations, per-operator examples, key interfaces) is in #573, docs/design_docs/proposals/asap-primitive-schema.md, reviewed separately against main. This PR contains the code, tests and the ScanSelection doc renames only.

Closes #571.

Problem: a summary state does not record which observations it covers

Once the planner combines existing summary states (#560 SummaryMerge, reuse of ingested panes, sub-DAG sharing from #645), it must know which observations each state was built from: which time range and which population (label values). Nothing on an OperatorNode records this today. The producer that built the state is no longer visible after composition, and the edge's Schema (#535) is the layout contract: field types and the committed state type, nothing more:

pub struct Schema {
    pub fields: Vec<Field>,            // e.g. job: Plain(Utf8), state: Sketch(KLL{k=200}, PerSubpopulationInstance)
    pub time_index: Option<ColumnId>,  // position of a timestamp column, not a time range
    pub unique_keys: Vec<Vec<ColumnId>>,
    pub closed: bool,
}

That is correct for a schema: #535 deliberately keeps filters, grouping keys, timing and windows out of Schema and Field ("filters, reduction/group keys, … execution timing, window framework … are not additional Schema or Field members"). So the information has to live somewhere else. The examples below show why it is needed: each uses two states with byte-for-byte equal schemas that must not be combined:

Schema(job: Plain(Utf8), state: Sketch(KLL{k=200}, PerSubpopulationInstance)), result_kind = State

Example 1: time. Equal schemas, different answers

Input A Input B Merging A and B is…
latency, [00:00, 00:01) latency, [00:01, 00:02) correct: p99 over [00:00, 00:02)
latency, [00:00, 00:02) latency, [00:01, 00:03) wrong: every observation in [00:01, 00:02) is counted twice, which skews the quantile and doubles counts or frequencies
latency, [00:00, 00:01) latency, [00:02, 00:03) correct only for [0,1) ∪ [2,3); wrong if the result is used for the continuous window [00:00, 00:03)

time_index is a column position. A KLL state has no timestamp column at all, so time_index is None in all three rows, and the schema cannot tell these cases apart.

Example 2: population (label values). Equal schemas, different answers

Input A Input B Merging A and B is…
region='us' region='eu' correct: p99 for us ∪ eu within each job
region='us' tier='premium' wrong: premium US requests are counted in both inputs
region='us' region='us' wrong: everything is counted twice

region is a filter label, not an output column, so it never appears in the schema. The job field only says the state is grouped by job. It does not say which jobs or which rows contributed.

Example 3: time and population together

A = us × [0,1) and B = eu × [1,2). The merged state covers exactly those two blocks. Describing it as {us,eu} × [0,2) (the result of storing a time range and a label set separately) would claim EU data for [0,1) and US data for [1,2) that was never read. The metadata has to keep time and population paired per region.

Example 4: answering a query from a stored state

Query: p99(latency) WHERE region='us' AND ts IN [10:00, 10:05) GROUP BY job. A stored state with the matching schema could hold US data for 10:00–10:05, EU data, or US data for only 10:00–10:03. All three have the same schema. Today the planner can confirm that the state type fits, but not that the contents fit.

Conclusion. Equal schemas are necessary but not sufficient for composing or reusing summaries; the planner also needs each state's coverage, which this PR adds as a node property next to Schema, not inside it. Without time/population metadata, the planner must either refuse every composition or accept silent double counting and missing data.

What this PR adds

Schema stays the layout contract from #535 and does not describe coverage. Coverage is a sibling field on the node, next to schema:

OperatorNode
├── schema: Schema                    what each output row looks like   (#535)
└── coverage: Option<SummaryCoverage> which observations the state holds (this PR)

coverage is required on summary nodes: validate_structure rejects a SummaryAgg (and, in #560, a SummaryMerge) whose coverage is None with CoverageError::Missing. Plain relational nodes leave it None; the field is an Option only because all operators share OperatorNode.

Why coverage is not part of Schema. Two states worth merging always cover different data. SummaryMerge (#560) requires all inputs to have the same Schema; that check is how it knows they are the same kind of state (same sketch, parameters and grouping). For example, two KLL states for "latency by job", built from minute 0–1 and minute 1–2:

State A State B Equal?
schema (job: Utf8, state: KLL{k=200}) (job: Utf8, state: KLL{k=200}) yes, so the merge is allowed
coverage time [0,1) time [1,2) no, which is why merging them is useful

If coverage were a field of Schema, these schemas would differ and the merge would be rejected; the only merge left would be a state with an exact copy of itself, which counts every observation twice. So schema says what kind of state this is, and coverage says which data it was built from.

CSE (#645) hashes and compares coverage, so two summary states over different observations are never shared, and ir::flat keeps it on each flat node.

pub struct SummaryCoverage {
    pub source: Source,               // as named by Scan: Table { table_ref } or TimeSeries { metric }
    pub regions: Vec<CoverageRegion>, // union of time × population blocks
}
pub struct CoverageRegion {
    pub time_ms: Option<Range<i64>>,            // half-open; None = no time restriction
    pub population: BTreeMap<String, String>,   // label = value AND …; empty = all
}

// OperatorNode
pub coverage: Option<SummaryCoverage>;   // required on summary nodes
pub fn requires_coverage(&self) -> bool;
pub fn with_coverage(self, c: SummaryCoverage) -> Result<Self, SchemaDerivationError>;
// SummaryCoverage
pub fn merge_disjoint(inputs: &[Self]) -> Result<Self, CoverageError>;
// SchemaDerivationError
Coverage(CoverageError)

Every observation in a region is assumed to contribute once to the state.

Removing duplication

Given the new definition, coverage holds only what no other type records: time × population. This PR also removes the duplication that the first draft introduced or exposed:

  • input and reduction dropped from coverage. They copied SummaryAgg.input / SummaryAgg.reduction on the same node, and with_coverage needed a ProducerMismatch check to keep the copies in sync. feat(ir): define compatible logical summary merges #560's SummaryMerge now compares them on its producers through OperatorNode::summary_update().
  • source is a Source, not a String. It uses the same type as Scan.source, so one table cannot have two spellings.
  • SourceCoverage renamed to ScanSelection (asap-aware-mapping, about 100 call sites plus docs). It names the rows a physical scan reads for cost comparison, a different concept from SummaryCoverage.
  • revision and multiplicity were removed earlier, as a deployment concern and a single-variant enum respectively.

How the examples come out under merge_disjoint:

Case Result
[0,1) + [1,2), same population accepted, coalesced to one region [0,2)
[0,1) + [2,3) accepted, two regions (gap kept)
[0,2) + [1,3) PossibleOverlap
region=us + region=eu, same time accepted, two regions
region=us + region=us PossibleOverlap
region=us + tier=premium PossibleOverlap. Different label names prove nothing.
us×[0,1) + eu×[1,2) accepted, two regions, never widened to {us,eu}×[0,2)
different source SourceMismatch
different update expression or reduction rejected by SummaryMerge (#560)

Rules:

  • A summary node without coverage fails validate_structure. Some with regions = [] means known empty.
  • time_ms: None is for sources without a time column (plain tabular data). Such a region overlaps every region it is not population-disjoint from.
  • with_coverage validates the declaration and requires State output (CoverageError::NotState). validate_structure re-checks it.
  • Rewriting a node's inputs clears the coverage, like other assessed metadata, so the rewriter must declare it again with with_coverage.
  • Declarations come from trusted composition rules or catalogs. Nothing is inferred from arbitrary SQL predicates.

How coverage is computed

This PR defines the type and the merge rule. Each part is filled in by a later PR, not written by hand:

Part How it is computed PR
source the one Scan under the SummaryAgg #646
population field = 'text' equality filters on the path from the SummaryAgg to its Scan (SummaryAgg.filter, Filter, Scan.predicates). A declaration that does not match is rejected. If the filters cannot be read (another operator on the path, or >, regex, IN, …), the population is unrestricted and such states merge only as time panes of the same computation. #646 (closes #570)
time_ms not in the logical plan: a pane's bounds are a materialization fact. Today's planner leaves it unset and builds no merges; Pass 2 window composition sets pane bounds such as [0, 1min). #601
merged coverage disjoint union of the inputs' coverage #560

Producers call SummaryCoverage::for_summary(&agg, time_ms) (#646), which builds a declaration that passes the check. Until #646 lands, a declaration is not checked against the subtree, so a wrong one could pass:

A = SummaryAgg(filter: region='us')   declared {region: eu} × [0,1)   ← wrong; A holds US data
B = SummaryAgg(filter: region='us')   declared {region: us} × [0,1)
merge → accepted ("eu" ≠ "us"), every US observation in [0,1) counted twice

#646 rejects A with PopulationMismatch.

Out of scope

Stack and validation

Order: #645 (merged) → #567 → #560 (SummaryMerge requires and derives coverage) → #646 (coverage computed from the subtree) → #539 → #540 → #541 → #542 → #543.

Tests in crates/types/tests/summary_coverage.rs cover adjacency, gaps, population disjointness, joint regions, overlap, regions without time bounds, source mismatch, serde round-trip, invalid declarations, required coverage, and clearing after rewrites. Each example in this body and the doc is also built as a real SummaryAgg → SummaryMerge plan in #560's summary_coverage_examples.rs. The design document is reviewed separately in #573.

🤖 Generated with Claude Code

@zzylol

zzylol commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

TODO: should write a design doc for it, not developer doc, human written.

@Selvomega

Copy link
Copy Markdown
Collaborator
image I suggest you change this problem headline: This is misleading. Encoding "what is summarized" is not schema's job

@Selvomega

Selvomega commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

According to this PR doc, I still don't think we are solving the right problem here.
You cannot ask PR to encode that much semantic information of data.
For example, consider two SQL queries:
SELECT sum_1 AS SUM(val) FROM table WHERE time in [Jan, Feb)
SELECT sum_2 AS SUM(val) FROM table WHERE time in [Feb, Mar)
Each SQL gives you a sum and now, say, you want to get SUM(val) from Jan to Mar. You naturally want to merge sum_1 and sum_2 but find the schema of the aggregation node cannot encode the sum range. (since schema only tells you that after this aggregate operator there is only one column with type, say, float)

This is indeed annoying but it is not schema problem. Since the responsibility of schema is mostly to encode data structure, not data semantic. To determine if the above two sums need to be merged, we can either

  • Go deep into the tree and analyze the SQL semantic
  • Maintain some side-cart data structure holding the semantic of the subtree / subDAG below some certain operator.
    But we should not do that in schema. For example, when we implemented CSE analysis, we also did not resort to schema, cuz that is not the correct place to go to. :)

@zzylol

zzylol commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

This is indeed annoying but it is not schema problem. Since the responsibility of schema is mostly to encode data structure, not data semantic. To determine if the above two sums need to be merged, we can either

  • Go d

Yes, this is not a schema problem, I will change the PR problem description title. Actually in the code implementation, the "what is summarized" is a field in the Node, in parallel with the field "schema" in the node struct.

@zzylol
zzylol force-pushed the feat/summary-coverage-contract branch from 4a91371 to 02833d1 Compare October 5, 2026 06:21
@zzylol zzylol mentioned this pull request Oct 5, 2026
2 tasks done
@zzylol
zzylol marked this pull request as ready for review October 6, 2026 18:33
@zzylol
zzylol requested a review from Selvomega October 6, 2026 18:34
@zzylol

zzylol commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

image I suggest you change this problem headline: This is misleading. Encoding "what is summarized" is not schema's job

I have changed the problem title. thanks

@Selvomega

Copy link
Copy Markdown
Collaborator

Coverage cannot go inside Schema. #560 only allows a merge when the input schemas are equal, and the inputs of a useful merge ([0,1) + [1,2)) always have different coverage.

What does this sentence in the PR doc mean

@Selvomega

Copy link
Copy Markdown
Collaborator

Random thinking: I feel the semantic coverage is trying to capture here is very interesting. Maybe such concept can be more generally useful: Each operator can have one such field describing "What current data is".

Say, a scan node can report "I am scanning columns A, B on data D"
Then after a filter node, the filter can also report "I am scanning columns A, B with 0<A<10 on data D"
Then after an aggregation node, the aggregation can report "I averaged B per A-value, I only have 0<A<10, I calculate on data D"

This is a generalization of the coverage semantic, and if exist, can be useful for other stuffs like CSE reducing as well. That said, this is just a rough, in-mature idea. We can save that for the future and only focus on summary coverage on this PR.

@zzylol
zzylol force-pushed the feat/summary-coverage-contract branch from e5e4f67 to fc8eec9 Compare October 6, 2026 20:17
@zzylol
zzylol changed the base branch from main to stack/528-03-graph October 6, 2026 20:17
zzylol added a commit that referenced this pull request Oct 6, 2026
…unified IR (#645)

#536 defined the unified OperatorNode DAG, but canonicalization, common
sub-DAG sharing and a flat, serializable form existed only for the old
split IRs. Add them on OperatorNode:

- ir::canonicalize: heavy-hitter promotion and EXISTS / NOT EXISTS / IN
  subquery lowering, bottom-up and memoized so shared sub-DAGs stay shared.
- ir::cse: hash-consing over a workload batch (the identical-expression
  rule of #509 Pass 2), following scalar-referenced operator nodes too.
- Generic child references: Operator, NonASAPOp, ASAPOp, ScalarExpr,
  Predicate, ProjectItem, SortKey and QueryRoot take the child reference
  as a type parameter (default Rc<OperatorNode>). ir::flat::flatten writes
  a DAG as nodes whose operators are Operator<NodeId>.

Summary coverage (#567) and SummaryMerge (#560) build on this.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
zzylol and others added 17 commits October 6, 2026 20:29
…ts fields

Name the metadata coverage to match #560 and the docs; drop the single-variant
multiplicity and deployment-specific revision; rename grouping to reduction to
match SummaryAgg; report failures through SchemaDerivationError::Coverage; revert
the unrelated PaneCoverageError rename.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…me bounds

validate_structure rejects a SummaryAgg without coverage (CoverageError::Missing).
CoverageRegion time bounds become optional so tabular sources without a time
column can declare coverage. Population stays trusted; #570 tracks checking it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Given required coverage on summary nodes, input and reduction duplicated the
producing SummaryAgg fields; drop them along with ProducerMismatch. Type source
as Source, matching Scan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SourceCoverage names the rows a physical scan reads for cost comparison, not
which observations a summary state holds; rename it so it is not confused with
SummaryCoverage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move it to docs/design_docs/proposals with problem and motivation,
requirements, design, alternatives and key code interfaces.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tive schema design

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The design document and the docs it consolidates are reviewed separately on
main. This PR keeps code, tests and the ScanSelection rename in docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the feat/summary-coverage-contract branch from fc8eec9 to bb3bf69 Compare October 6, 2026 20:29
@zzylol
zzylol changed the base branch from stack/528-03-graph to main October 6, 2026 20:29
Coverage is part of a summary state's identity: two states over different
observations are never shared, so CSE hashes and compares it, and a flat
node keeps it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the feat/summary-coverage-contract branch from bb3bf69 to 4750e71 Compare October 6, 2026 20:34
@Selvomega

Copy link
Copy Markdown
Collaborator

According to my (and my agents', mostly my agents') understanding, this PR only defines the coverage types but did not answer how to calculate such a coverage. I think this is an important information that should be mentioned in the PR doc clearly but idk why the agent writing the PR didn't mention it.

@zzylol

zzylol commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Random thinking: I feel the semantic coverage is trying to capture here is very interesting. Maybe such concept can be more generally useful: Each operator can have one such field describing "What current data is".

Say, a scan node can report "I am scanning columns A, B on data D" Then after a filter node, the filter can also report "I am scanning columns A, B with 0<A<10 on data D" Then after an aggregation node, the aggregation can report "I averaged B per A-value, I only have 0<A<10, I calculate on data D"

This is a generalization of the coverage semantic, and if exist, can be useful for other stuffs like CSE reducing as well. That said, this is just a rough, in-mature idea. We can save that for the future and only focus on summary coverage on this PR.

Yes, it is interesting. It's saying how to represent the semantic of a sub-DAG.

Comment thread crates/types/src/ir/summary_coverage.rs
Comment thread docs/design_docs/architecture/physical-plan-integration.md
@zzylol

zzylol commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

You're right, the PR doc should say this. #567 only defines the coverage type and the merge rule; computing it is split across later PRs:

Part How it is computed PR
source the one Scan under the SummaryAgg #646
population field = 'text' equality filters on the path from the SummaryAgg to its Scan (SummaryAgg.filter, Filter, Scan.predicates); a declaration that doesn't match is rejected. If the filters can't be read (other operators, >, regex, …), the population is unrestricted and such states only merge as time panes of the same computation. #646 (closes #570)
time_ms not in the logical plan: it is a materialization fact. Today's planner leaves it unset; Pass 2 window composition sets pane bounds such as [0, 1min). #601
merged coverage disjoint union of the inputs' coverage #560

Producers don't write coverage by hand: they call SummaryCoverage::for_summary(&agg, time_ms) (#646), which builds a declaration that passes the check. I'll add this to the PR description too.

@zzylol

zzylol commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Coverage cannot go inside Schema. #560 only allows a merge when the input schemas are equal, and the inputs of a useful merge ([0,1) + [1,2)) always have different coverage.

Short answer: two states worth merging always cover different data. If coverage were part of Schema, they would never have the same schema, and #560 would refuse to merge them.

Why. SummaryMerge (#560) requires all inputs to have the same Schema. That check is how it knows the inputs are the same kind of state: same sketch, same parameters, same grouping.

Example. Merge two KLL states for "latency by job", one built from minute 0–1 and one from minute 1–2:

State A State B Equal?
schema (job: Utf8, state: KLL{k=200}) (job: Utf8, state: KLL{k=200}) yes, so the merge is allowed
coverage time [0,1) time [1,2) no, which is why merging them is useful

If coverage were a field of Schema, the two schemas would differ ([0,1) vs [1,2)) and the merge would be rejected. The only merge left would be a state with an exact copy of itself, which counts every observation twice.

So schema says what kind of state this is, and coverage, a separate field on OperatorNode, says which data it was built from. I updated the PR description with the same explanation.

@zzylol
zzylol merged commit c3471c0 into main Oct 6, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants