Skip to content

refactor!: keep only deployment fields on precompute materializations - #806

Open
zzylol wants to merge 3 commits into
refactor/codecs-over-planner-kernelsfrom
refactor/materialization-deployment-only
Open

zzylol wants to merge 3 commits into
refactor/codecs-over-planner-kernelsfrom
refactor/materialization-deployment-only

Conversation

@zzylol

@zzylol zzylol commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #805.

Why

Installed precompute plans still described each stored output's computation
twice: once as the Planner physical DAG that now runs ingest (#796/#797/#800),
and again as backend fields on PrecomputeMaterialization (aggregation_type,
parameters, spatial_filter, aggregated_labels, …) that were cross-checked
against the DAG. The Planner DAG is the only computation definition; the backend
keeps deployment facts. The owner accepted that all old plans become invalid
(development stage, no migration).

What

  • PrecomputeMaterialization keeps only deployment fields: stored output id,
    source binding, grouping, cadence and window layout, retention, semantic
    fragment.
  • Removed: aggregation_type, aggregation_sub_type, parameters,
    spatial_filter(_normalized), aggregated_labels, rollup_labels,
    original_yaml, accumulator_spec(), sample_update_rule(), plus
    PolicyRegistry and RoutingIndex.
  • Computation facts now come from the Planner DAG: the stored state family is the
    full Planner SummaryFamilyType in StateSchemaContract (validated against
    the producing DAG node); input filter, item label and the PromQL right-closed
    pane rule are read from the producer via PrecomputePlan::summary_producer /
    population_filter.
  • stored_output_id is assigned explicitly from deployment fields plus the
    producer's computation.
  • Install-time validation rejects a stored output whose producers disagree on
    population or update, and a raw output with no Planner producer (previously
    treated as unfiltered).

Breaking change

BACKEND_COMPAT is asap-query-backend.v2 and the SummaryCatalog schema is 7.
Plans from earlier builds fail to decode with a version error (tested).

Before this PR

PrecomputeMaterialization { aggregation_type, parameters, spatial_filter, … }  ⇄  Planner DAG (cross-checked)

After this PR

PrecomputeMaterialization { stored output id, source, grouping, windows, retention, semantic fragment }
Planner DAG: the only computation definition

Known follow-ups (from independent review)

  • Plans decode through a JSON value to check the version first; YAML round-trips
    may not decode (plans are installed as JSON).
  • Compiler and runtime derive item labels differently; OTLP series may get an item
    label for keyed non-top-k sketches.
  • DAG documents are decoded on each lookup; the state-family lookup is linear.
  • A few design docs still mention removed fields.

Validation

cargo fmt --check, cargo clippy --workspace --all-targets -- -D warnings,
cargo test --workspace (lib, integration and process e2e targets). An
independent reviewer found two install-time gaps (conflicting producers; missing
producer), both fixed with regression tests that fail without the fix.

🤖 Generated with Claude Code

zzylol and others added 3 commits September 30, 2026 14:20
An installed PrecomputeMaterialization now carries its stored output id,
source binding, grouping layout, cadence and window layout, retention and
semantic fragment. What an output computes is defined only by the Planner
DAG node bound to it.

Removed: aggregation_type, aggregation_sub_type, parameters,
spatial_filter, spatial_filter_normalized, aggregated_labels,
rollup_labels, original_yaml, and the derived accumulator_spec() and
sample_update_rule() projections.

- The stored state family lives in the existing StateSchemaContract
  (now the lossless Planner SummaryFamilyType); validation requires it
  to equal the family of any SummaryAgg node that produces the output.
- Input predicates, item labels and the PromQL right-closed pane rule are
  read from the bound DAG producer; the catalog's data descriptors use
  the same predicate.
- The compiler allocates stored_output_id explicitly from the deployment
  fields and the producer's computation identity.
- PolicyRegistry and RoutingIndex are removed; OTLP content matching and
  the sketch sink read the installed plan directly.
- BACKEND_COMPAT is asap-query-backend.v2 and SummaryCatalog schema 7;
  older plans are rejected on decode by version.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Materializations no longer carry their input filter, so an output's
population comes only from the DAG node that produces it. Plan
validation now rejects:

- an output produced by several DAGs that disagree on the update or on
  the population (metric and canonical filter) they read;
- a raw time-series output that has no Planner producer in a plan that
  carries DAGs, which would otherwise be read as unfiltered;
- a bound raw producer whose input is not a source scan.

Also keeps plain OTLP sketch envelopes out of shared (Hydra) states,
restores the single-catalog descriptor test, and drops comments that
still named removed fields.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol changed the base branch from fix/topk-by-heap-groups to refactor/codecs-over-planner-kernels September 30, 2026 14:38
@zzylol
zzylol force-pushed the refactor/materialization-deployment-only branch from 4f367fe to 57baf5d Compare September 30, 2026 14:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant