Skip to content

test: #509 Examples 2–4 acceptance specs and end-to-end tests - #590

Draft
zzylol wants to merge 76 commits into
stack/509-e2-keyed-additive-topkfrom
stack/509-acceptance-examples-2-4
Draft

zzylol wants to merge 76 commits into
stack/509-e2-keyed-additive-topkfrom
stack/509-acceptance-examples-2-4

Conversation

@zzylol

@zzylol zzylol commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Why

#509 defines four end-to-end examples. Only Example 1 had an acceptance spec and tests. AGENTS.md asks for acceptance behavior to be defined before implementation, and by someone other than the implementer. This PR does that for Examples 2–4, ahead of the Stage 2 materialization and #580 work (window composition, the summary-capability rule, Hydra).

What

  • Acceptance specs: docs/design_docs/proposals/planner-layering-example{2,3,4}-acceptance.md. Each lists the per-stage candidate expectations taken from docs: propose workload-wide planning, summary sharing, and materialization #509, the invariants under test, and the places where the doc is ambiguous or conflicts with today's planner.
  • Tests: crates/integration-tests/tests/planner_layering_example{2,3,4}.rs, using the real plan_stages.
  • Shared adapters in tests/planner_layering_common/mod.rs (window_form, materialization, retention_ms) return today's only possible answer. The implementer of each feature fills them in.
  • Devtool examples --example planner-layering-3a / 3b. Example 1's output is byte-identical.

Before / After

Findings for the implementers

These are also listed in the specs.

Written as an independent test design; the agent implemented no planner features. Base: #589. Links: #509, #580, #572.

🤖 Generated with Claude Code

zzylol and others added 30 commits October 3, 2026 15:59
…ts fields

Name the metadata coverage to match #560 and the docs; drop the single-variant
multiplicity and deployment-specific revision; rename grouping to reduction to
match SummaryAgg; report failures through SchemaDerivationError::Coverage; revert
the unrelated PaneCoverageError rename.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…me bounds

validate_structure rejects a SummaryAgg without coverage (CoverageError::Missing).
CoverageRegion time bounds become optional so tabular sources without a time
column can declare coverage. Population stays trusted; #570 tracks checking it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Given required coverage on summary nodes, input and reduction duplicated the
producing SummaryAgg fields; drop them along with ProducerMismatch. Type source
as Source, matching Scan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SourceCoverage names the rows a physical scan reads for cost comparison, not
which observations a summary state holds; rename it so it is not confused with
SummaryCoverage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move it to docs/design_docs/proposals with problem and motivation,
requirements, design, alternatives and key code interfaces.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tive schema design

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The design document and the docs it consolidates are reviewed separately on
main. This PR keeps code, tests and the ScanSelection rename in docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…coverage examples

SummaryCoverage no longer repeats input/reduction, so SummaryMerge compares
them through OperatorNode::summary_update. summary_coverage_examples.rs builds
each example in docs/develop_docs/summary-coverage.md as a SummaryAgg ->
SummaryMerge plan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… doc

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A workload DAG has one root per batch query; queries that share a sub-DAG
reference the same exported nodes. Single-query export is a batch of one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The logical export is phase-free. Physical planning still needs the lifecycle
timing expansion; bring it back unchanged except for carrying coverage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol and others added 25 commits October 3, 2026 21:40
…r naming fix

The finalized exact accumulator now names its column after the aggregate
it realizes (value/sum instead of state). Point the stage_pipeline test at
the renamed fixture.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ction (MVP)

Stage 2 implements one logical candidate physically: exact Aggregate{[TopK]}
becomes Limit(Sort) per group, summaries stay build -> estimate, every node
runs at query time, and the timed roots export as one PhysicalASAPDAG.

Stage 3 is the only stage that computes cost. It rejects candidates whose
summary misses its query's accuracy target (accuracy model, analytical
guarantee) or uses Count-Min over weights not proven non-negative, prices
every valid candidate node by node with analytical_cost (illustrative
statistics), selects the cheapest and reports the rest as costlier.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add stage2_physical_asap (no cost) and stage3_selection (costs, selected,
every other candidate rejected as invalid or costlier) per the v2 contract.
PromQL rows carry the series identity from Stage 0, as runtime state needs.
Regenerate tools/dag-viewer/examples/planner-layering-example1.json.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Define MVP acceptance for #509 Example 1 before the Phase C stage APIs
exist: 1 -> 6 -> 6 -> 1 candidates (Pass 1 x identical-expression
sharing; physical operator implementation only, no materialization).

- Spec with per-stage candidate tables, invariants and doc ambiguities.
- Integration tests against todo!() stage stubs, ignored until the
  stages land; workload and Stage 0 frontend-shape tests run now.
- Expected-only asap-stage-pipeline/v1 fixture for the DAG viewer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace the todo!() stubs with adapters over Stage 0-3, un-ignore the 11
tests that pass, and give each still-ignored test the precise difference
between the implementation and the spec. Add runtime capability checks:
the selected plan compiles in the physical planner; compiling every
candidate stays ignored with the two runtime gaps it finds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace the pre_asap/post_asap split with modules named for what they hold:

- ir::operator: node, non_asap, asap, operator_properties, agg_intent,
  maintained_population
- ir::scalar: scalar expressions, expr_ir items, scalar_type_rules,
  column_resolution
- ir::schema: Schema/Field/DataType, aggregate_schema, error, and the
  summary state types (state_type, formerly post_asap::sketch)
- ir::properties: guarantee, timing, summary_coverage, and the execution
  timing/data-state types split out of execution_data_state
- workload: workload, parsed_workload, resources
- physical: the rest of execution_data_state, and summary_window

post_asap::query_time had no callers and is deleted. Public type names are
unchanged; every import across the workspace is rewritten, with no
compatibility re-exports from the old paths.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SummaryEstimate{TopK} derived `partition keys + topk Utf8`, an encoding
no runtime operator produces. The runtime's keyed evaluation, the exact
Sort -> Limit path and EvaluatePopulation{TopK} all return the selected
rows, so every CountSketch+heap Example 1 candidate failed to compile.
Derive partition keys + item identity columns + Float64 `value`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The 6-candidate spec did not account for Pass 1's exact-accumulator
options (user decision). Un-ignore the count-only tests, check that the
runtime compiles exactly the candidates Stage 3 finds valid and rejects
Count-Min + heap for the same reason, and list the 24 candidates in the
acceptance spec for manual review. Hydra and shared-input tests stay
ignored, naming the missing feature.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eadout

Aggregate{TopK} now derives partition keys + item identity + value, the
shape #579 gave sketch readouts, so every top-k realization of one query
has one schema and an aggregate over a pass-through top-k composes.
The keyed-heap state column keeps its topk_<k> name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sult

A top-k readout now reports min(input rows, k x groups) rows, as exact
Sort -> Limit does, so its consumers are priced alike whichever
realization is chosen.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tire MajorPass

The facade's default pass is now StagePipeline: Stage 1 inventory,
selection by a dynamic program over target nesting priced through
Stage 2 + Stage 3, then the winner built with identical producers
merged and checked in full. The program checks that every target and
the target beneath it combine additively and admissibly; otherwise it
enumerates (at most 64 combinations) or flags the result as not
guaranteed optimal.

- PlanningModels moves into plan_selection; Stage 3 does not read cost.
- Choice enumeration moves into logical_candidates; the devtool reuses
  select_exhaustive.
- Build, check and pricing failures reject one candidate, not the plan.
- PlanOutput carries the Selection.
- Sharing tests that depend on Pass 2 or on summaries being selected
  are ignored with #580.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move the cost- and recurrence-dependent selection over
`CandidateLogicalASAPDAGs` (cost ranking, recurrence profiles, global
selection, selected-DAG assembly, runtime support evidence) and its tests
from `replacement` into `plan_selection::candidate_selection`. The code
is unchanged; only visibility and imports were adjusted.
`plan_selection.rs` becomes `plan_selection/mod.rs`.

`recurrence_profiles_from_workload` is deleted rather than moved: it has
no caller anywhere in the workspace.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t model

`realizations_for_intent` lists sketch candidates in `summary_candidates`
order and sizes them with the analytical `accuracy::estimators::size_params`,
which is what `CostModel::size_params` defaulted to. Extension intents stay
PassThrough, the previous default.

- `CandidatePlanningInputs` drops its `cost` field; `realize_child` drops its
  cost argument.
- `ASAPStrategies` and `HydraGroupingStrategy` are built with `default()`
  or `new_with_planning_inputs*` without a cost model; `ASAPStrategies::new`,
  the `default_cost_model` constructors and `default_strategies_with` are
  removed, and `default_strategies_with_evidence` takes only the evidence.
- `ExactCompositionStrategy` is a unit struct and no longer checks runtime
  support; selection already requires positive support evidence.
- `CostModel` loses `size_params`, `realize_extension` and
  `evaluation_extension`.

Tests that exercised only the removed hooks are deleted; call sites across
the workspace use the new constructors.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scans the Stage 1 modules' non-test code and fails on any import of
`cost_model` or `recurrence`, including names `lib.rs` re-exports from them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Update the developer docs and the `CostModel` trait doc: strategies take no
cost model, sketch sizing is analytical, extension intents stay
pass-through, and the cost model is consulted only at selection.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…crate split

Stage 1 no longer imports `plan_selection::candidate_selection`, and the
legacy selection code no longer adds inherent methods to Stage 1 types, so
Stage 1 can move into its own crate (#572 B2, decision Q42 (a)).

- `GlobalSelection` and `TargetSubDAGSelection` move back into
  `replacement` with the DAG assembly (`assemble_*`): building a DAG from
  given choices is Stage 1 composition and needs no cost. Exact-composition
  plans are passed in by `GlobalSelection::new`.
- `global_selection*` returns `CostedGlobalSelection`, which derefs to
  `GlobalSelection` and keeps the cost comparison behind each exact
  composition (`composition(target)`), formerly the
  `TargetSubDAGSelection::composition` field.
- The `impl ReplacementSubDAG` / `impl CandidateLogicalASAPDAGs` blocks in
  `candidate_selection` become free functions (`cost_sorted`,
  `cost_sorted_with_recurrence`, `recurrence_profiles`, `global_selection`,
  `global_selection_with_recurrence`, `runtime_support_evidence`); callers
  across the workspace and docs are updated.
- Stage 1 exposes what they need: read accessors `groups()`, `order()` and
  `composition_plans()`, and public `PreparedComposition`,
  `cse_candidate_pair`, `direct_child_counts`, `realize_child` and
  `ExactComposition::same_as`.
- Selection assertions in Stage 1 test modules move to
  `candidate_selection` tests: four reconciliation tests and
  `global_selection_can_choose_nested_summaries` move whole; five mixed
  tests are split, adding three selection tests. Assembly tests move to
  `replacement`.

Behavior is unchanged; the Example 1 fixture regenerates byte for byte.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Stage 1 of #509 (logical candidate generation) moves out of
asap-aware-mapping into the new crate `asap-logical-optimizer`
(`crates/logical-optimizer`), so Cargo enforces that Stage 1 depends on no
later stage (#572, step B2). This is a pure move plus import updates.

Layout:
- `pass1/`: replacement, rewrite, rollup, grouping, exact_composition,
  function_rules, maintained_population, explanation, logical_candidates.
- `pass2/`: topk_reuse, reconciliation (was `accuracy::reconciliation`).
- `accuracy/`: the analytical model (mod, composition, allocation, evidence,
  estimators). `AccuracyModel` stays here for now; moving its pluggability
  to plan selection (Q22) is a later step.

The new crate depends only on asap-types (plus asap_sketchlib, thiserror and
serde_json); asap-aware-mapping, which keeps Stage 2, Stage 3 and the facade,
depends on it. asap-aware-mapping no longer re-exports any Stage 1 item:
every import across the workspace and docs now names
`asap_logical_optimizer`. Files move with `git mv`, together with their tests
(`tests/logical_candidates.rs`, the Stage 1 guard) and `test_support`;
asap-aware-mapping keeps a trimmed copy of `test_support` for its own tests.

The Stage 1 guard now scans every production file of the new crate and also
checks that its manifest names neither asap-aware-mapping nor
asap-physical-operators. Two candidate_selection tests build their fixture
through public Stage 1 API instead of crate-private constructors.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move asap-aware-mapping's physical_candidates into the new
asap-physical-optimizer crate (implementation/physical_candidates.rs), with
a trimmed test_support copy. Production code depends only on asap-types; a
new manifest guard rejects plan-selection, aware-mapping and
physical-operators. Imports are rewritten without compatibility
re-exports. Pure move (#572, B3).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move asap-aware-mapping's Stage 3 modules into the new asap-plan-selection
crate: the cost model and its inputs under cost/, plan_selection/mod.rs as
the crate root (select_plan, PlanningModels) and candidate_selection, with
their tests and the test_support copy. Delete pane_sharing, erp and
empirical_comparison, which have no callers (#572), and the
snapshot_prepare_cpu helper only empirical_comparison used. A manifest
guard rejects the facade and executor. Imports are rewritten without
compatibility re-exports (#572, B4).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…asap-aware-mapping

Move asap-aware-mapping's pass/ (OptimizationPass, StagePipeline,
PassRegistry, optimize) into crates/planner, the facade per #572, and
remove the now-empty asap-aware-mapping crate from the workspace.
PlanningModels is re-exported by the planner from asap-plan-selection.
The stage manifest guards now name asap-planner instead of the deleted
crate. Docs and comments that cited asap-aware-mapping as a crate now
name the stage crate that owns the code (#572, B4b).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-executor crate

Rename crates/asap-physical-operators to crates/executor (package
asap-executor), keeping plan/ and physical_planner/ in it as #572's
2026-10-03 update decides. Imports are rewritten without compatibility
re-exports. The three stage manifest guards now reject asap-executor;
only integration-tests depends on it (#572, B5).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…variant

Stage 1 now keeps two forms of the workload (#509 Pass 2): the queries as
written, and, when CSE merges something, the queries with identical
sub-DAGs shared. Stage 3 treats the variant as an outer choice: the
per-target dynamic program runs once per variant (the coupling guard also
checks targets reading a common input in the shared variant) and the
cheapest result wins.

plan_selection::plan_stages runs Stage 1 -> 2 -> 3; the facade and the
stage_pipeline devtool both call it, so the facade no longer pre-merges
and the devtool shows the shared variants.

A PromQL series has one sample per timestamp, so the identified scan now
declares (series identity, ts) as a key, which CSE's legality rule needs
to share the range selector.

Example 1: 1 -> 48 -> 48 -> 1; P44 (P20's choices over a shared input)
is selected at 52.201 vs 79.401.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Pass 1 now offers, for a top-k over a per-item sum or count whose inner
aggregate has no other consumer, Count-Min and CountSketch + heap
alternatives that read the inner aggregate's input and absorb it (the
legacy keyed-additive rule decides applicability and the update). For
#509 Example 1's Q2 that is one heap per job over the raw samples, keyed
by series identity and weighted by the sample value.

An absorbed target has no choice of its own: enumeration skips its
choices, choice_index ranks in that order, and the tree DP sums only the
targets an alternative still reads (read_targets) and sets the absorbed
target to its pass-through. Top-k readout rows are bounded by the series
count, so a sketch reading several samples per item is sized like the
other realizations (the coupling guard caught the difference).

Example 1: 1 -> 64 -> 64 -> 1; P58 (all exact, shared input) at 52.201;
the best whole-expression CountSketch + heap plan, P64, costs 537.201.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Independent test design for the summary-capability rule (Example 2),
the window-composition rule (Example 3) and Stage 2 materialization
(Example 4). Each spec extracts the doc's per-stage candidates, lists the
invariants and their tests, and records ambiguities and conflicts with
today's planner.

Tests run plan_selection::plan_stages. Behavior that works today passes
(33 tests); tests that need a missing feature are ignored with the
feature's name (25). Window summaries, materialization and retention are
read through pending adapters in tests/planner_layering_common, for the
implementers to fill in.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Adds --example planner-layering-3a (Pattern A, sub-interval batch) and
planner-layering-3b (Pattern B, repeating 5-min p99), matching the
Example 3 acceptance tests' workloads. Example 1's output is unchanged.
Examples 2 and 4 need planner changes first; their specs list them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant