Skip to content

feat(planner): optimizer accepts label-set facts and costs by cardinality - #780

Merged
milindsrivastava1997 merged 5 commits into
mainfrom
756-optimizer-accept-externally-provided-label-set-facts-cardinality-arrival-rate
Oct 4, 2026
Merged

milindsrivastava1997 merged 5 commits into
mainfrom
756-optimizer-accept-externally-provided-label-set-facts-cardinality-arrival-rate

Conversation

@milindsrivastava1997

@milindsrivastava1997 milindsrivastava1997 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Closes #756. Part of #753.

Summary

  • Label-set facts replace --dataset / --rho. asap-optimizer-cli --label-set-facts <yaml> takes series_count per (metric, spatial filter) and cardinality per (metric, spatial filter, grouping labels). Arrival rate is derived as series_count / scrape interval. The label schema now comes from the workload's metrics: hints, which are required. Missing facts produce an error listing every expected key; unused facts only warn. dataset.rs and SUBPOPULATION_COUNT are removed.
  • Cardinality-aware costing. Memory and merge/subtract work scale with output_group_count for keyed maps (Multiple*) and with instance_count otherwise. Query reads scale with output groups, except top-k (one heap read per bucket). Trivial accumulators get an analytical per-key memory estimate (dictionary-encoded label codes + value, hash-table slack). Single-group Sum/MinMax/Increase are no longer proposed, since they tie with their Multiple* forms.
  • Optimizer configs now match what the legacy planner deploys (pre-existing bug):
    • build_config now uses set_subpopulation_labels, so keyed sketches are one instance rather than one per group (or one top-k heap per series).
    • CMS/HydraKLL now deploy their paired DeltaSet key aggregation, referenced alongside the value in the query config.
  • does_precompute_operator_support_subpopulations handles Set/DeltaSet and non-top-k CMS+heap instead of panicking. Legacy planner calls are unchanged.
  • Docs: .design_docs/optimizer-v1-implementation-plan.md updated, plus runnable example workload and facts files.

Notes

Test plan

  • cargo test -p asap_planner -p promql_utilities, clippy, cargo check --workspace
  • New tests: facts parsing/validation/missing-key reporting (incl. topk by), metric-hint errors, label split per sketch type, cost scaling per class, top-k reads, end-to-end CMS + key tracker deployment and query-config references
  • asap-optimizer-cli on the example files: HydraKLL + DeltaSet for the quantile query, MultipleSum for sum by; without --atomic-costs the run reports the top-k query as unservable (EXACT was removed in feat(planner): model assignments as SLA-aware items #776)

🤖 Generated with Claude Code

milindsrivastava1997 and others added 3 commits October 4, 2026 16:17
Replace --dataset/--rho with --label-set-facts: series_count per
(metric, spatial filter) and cardinality per (metric, spatial filter,
grouping labels). Arrival rate is derived as series_count / scrape
interval. The label schema now comes from the workload's metrics: hints.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Cardinality now scales cost by how each sketch holds its groups: keyed maps
grow per key, fixed-size sketches don't, and grouped queries read one value
per output group (one heap read per top-k bucket). Trivial accumulators get
an analytical per-key memory estimate, and only their Multiple* forms are
proposed.

Optimizer configs now use the legacy planner's grouping/aggregated label
split, so keyed sketches deploy as one instance rather than one per group,
and CMS/HydraKLL deploy their paired DeltaSet key aggregation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
milindsrivastava1997 and others added 2 commits October 4, 2026 17:40
Resolutions:
- Optimizer items (#776) replace AQEs throughout; label-set facts,
  group counts, and the paired key aggregation carry over onto items.
- EXACT is removed: greedy reports unservable items and the facts
  pipeline returns OptimizerPipelineError (LabelSetFacts | Optimizer).
- dataset.rs stays deleted (replaced by label_set_facts.rs).
- Example workload drops its top-k query, which is unservable without
  atomic costs now that EXACT is gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@milindsrivastava1997
milindsrivastava1997 merged commit 8132a74 into main Oct 4, 2026
11 checks passed
@milindsrivastava1997
milindsrivastava1997 deleted the 756-optimizer-accept-externally-provided-label-set-facts-cardinality-arrival-rate branch October 4, 2026 22:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Optimizer: accept externally provided label-set facts (series count, cardinality)

1 participant