Problem
The optimizer derives N_g from --dataset (CSV series inventory) and uses one global placeholder arrival rate --rho (#563). Subpopulation-aware sketches are costed with a constant SUBPOPULATION_COUNT. Memory for Multiple* accumulators (unbounded hashmaps) is costed as one flat-stub instance regardless of group count, so they always beat their per-group twins.
Change
Accept externally provided label-set facts from a file (--label-set-facts <path>, required). The optimizer does not estimate them. Removes --dataset, --rho, SeriesDataset (dataset.rs), and SUBPOPULATION_COUNT.
Inputs
- Metric schema comes from the workload YAML's
metrics: hints (ControllerConfig::schema_from_hints), same as asap-planner. Error if metrics: is absent or any workload metric is missing from it.
- Facts file (YAML), two tables:
series: # per (metric, spatial_filter)
- metric: http_requests_total
spatial_filter: 'job="api"' # required; "" = unfiltered
series_count: 1000
groups: # per (metric, spatial_filter, grouping_labels)
- metric: http_requests_total
spatial_filter: 'job="api"'
grouping_labels: [service] # required; resolved `by` labels (also for `without` queries)
cardinality: 5
- Arrival rate is derived, not provided:
λ(metric, filter) = series_count × 1000 / scrape_interval_ms (scrape interval = --data-ingestion-interval-ms). Assumes one sample per series per scrape; overestimates sparse/irregular series.
Matching and validation
- Exact-key match only; no inference across label sets.
spatial_filter is normalized with asap_types::normalize_spatial_filter (same as the workload side). Missing-key errors print the exact normalized key expected.
- Missing
series or groups facts for any item → error listing all missing keys.
1 ≤ cardinality ≤ series_count; grouping_labels: [] requires cardinality == 1; a groups row without a series row → error; duplicate keys → error.
- Facts matching no item → warn (one file may serve several workloads).
Cost model
Cardinality card enters per sketch class (one label set per candidate; no inner/outer key split):
| Class |
Types |
Memory |
Query CPU |
Insert CPU |
| Per-group |
Sum, MinMax, Increase, KLL, HLL, Set/DeltaSet |
card × mem_per_instance |
card × query |
λ × insert |
| Keyed, unbounded |
MultipleSum, MultipleMinMax, MultipleIncrease |
card × mem_per_key |
card × query |
λ × insert |
| Keyed, bounded |
CMS, CMS+heap, HydraKLL |
1 × mem_per_instance |
card × query (one point query per group) |
λ × insert |
Merge/subtract CPU follow the memory multiplier.
Analytical per-key memory for the trivial accumulators, applied to both per-group (Sum/MinMax/Increase) and Multiple* twins so each pair ties on memory:
mem_per_key = (n_grouping_labels × 4 + value_bytes) × 8/7
value_bytes: Sum/MinMax = 8, Increase = 72
Label values assumed dictionary-encoded (u32 codes, dictionary amortized and not charged). Ignores Increase counter-reset events. Underestimates Multiple* in today's engine, whose inner maps hold Vec<String> keys. KLL/HLL/CMS keep measured/stub costs (#524); Set/DeltaSet stay on the stub (#763).
Out of scope
Deliverables
- Inline YAML fixtures in tests; update
.design_docs/optimizer-v1-implementation-plan.md CLI examples; one example facts file next to the doc.
Refs #693, #563, #525, #524.
Part of #753.
🤖 Generated with Claude Code
Problem
The optimizer derives N_g from
--dataset(CSV series inventory) and uses one global placeholder arrival rate--rho(#563). Subpopulation-aware sketches are costed with a constantSUBPOPULATION_COUNT. Memory forMultiple*accumulators (unbounded hashmaps) is costed as one flat-stub instance regardless of group count, so they always beat their per-group twins.Change
Accept externally provided label-set facts from a file (
--label-set-facts <path>, required). The optimizer does not estimate them. Removes--dataset,--rho,SeriesDataset(dataset.rs), andSUBPOPULATION_COUNT.Inputs
metrics:hints (ControllerConfig::schema_from_hints), same asasap-planner. Error ifmetrics:is absent or any workload metric is missing from it.λ(metric, filter) = series_count × 1000 / scrape_interval_ms(scrape interval =--data-ingestion-interval-ms). Assumes one sample per series per scrape; overestimates sparse/irregular series.Matching and validation
spatial_filteris normalized withasap_types::normalize_spatial_filter(same as the workload side). Missing-key errors print the exact normalized key expected.seriesorgroupsfacts for any item → error listing all missing keys.1 ≤ cardinality ≤ series_count;grouping_labels: []requirescardinality == 1; agroupsrow without aseriesrow → error; duplicate keys → error.Cost model
Cardinality
cardenters per sketch class (one label set per candidate; no inner/outer key split):card × mem_per_instancecard × queryλ × insertcard × mem_per_keycard × queryλ × insert1 × mem_per_instancecard × query(one point query per group)λ × insertMerge/subtract CPU follow the memory multiplier.
Analytical per-key memory for the trivial accumulators, applied to both per-group (
Sum/MinMax/Increase) andMultiple*twins so each pair ties on memory:Label values assumed dictionary-encoded (
u32codes, dictionary amortized and not charged). Ignores Increase counter-reset events. UnderestimatesMultiple*in today's engine, whose inner maps holdVec<String>keys. KLL/HLL/CMS keep measured/stub costs (#524); Set/DeltaSet stay on the stub (#763).Out of scope
normalize_spatial_filterTODO).Deliverables
.design_docs/optimizer-v1-implementation-plan.mdCLI examples; one example facts file next to the doc.Refs #693, #563, #525, #524.
Part of #753.
🤖 Generated with Claude Code