From e161468de1417eb0020c22a09aa59e1128c506d6 Mon Sep 17 00:00:00 2001 From: zz_y Date: Sun, 4 Oct 2026 21:14:42 +0000 Subject: [PATCH 01/39] docs(planner): plan the AutoSketch vs. planner evaluation Co-Authored-By: Claude Opus 5.5 --- docs/README.md | 4 + docs/evaluation/autosketch-vs-planner.md | 234 +++++++++++++++++++++++ 2 files changed, 238 insertions(+) create mode 100644 docs/evaluation/autosketch-vs-planner.md diff --git a/docs/README.md b/docs/README.md index 64fa2103..14f35ee0 100644 --- a/docs/README.md +++ b/docs/README.md @@ -43,6 +43,10 @@ Task-oriented guides for common operations: - [Deploy to CloudLab](03-how-to-guides/operations/deploy-cloudlab.md) - Deployment guide - [Troubleshooting](03-how-to-guides/operations/troubleshooting.md) - Common issues & solutions +## Evaluation plans + +- [AutoSketch vs. the ASAPQuery planner](evaluation/autosketch-vs-planner.md) - Paper §6.3 plan: methods, cost model, PRs + ## 04. Development Developer practices and infrastructure: diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md new file mode 100644 index 00000000..3ec17170 --- /dev/null +++ b/docs/evaluation/autosketch-vs-planner.md @@ -0,0 +1,234 @@ +# Evaluation plan: AutoSketch vs. the ASAPQuery planner (paper §6.3) + +Status: plan, no results yet. Open decisions are listed in §9 with a +recommended answer each; nothing in sketch-bench is implemented until they are +settled. + +## 1. Question + +Given the same measured sketch costs and the same workload of repeating query +expressions (RQEs), how much cheaper is the plan from ASAPQuery's planner than +the plans from AutoSketch, and how long does each take to plan? + +AutoSketch ([NSDI '24](https://www.usenix.org/system/files/nsdi24-sun.pdf), +Algorithm 4) differs from ASAPQuery's planner in four ways that matter here: + +| | AutoSketch | ASAPQuery planner | +| --- | --- | --- | +| Unit of optimization | One query at a time | The whole batch of RQEs, jointly | +| Repetition over time | Not modeled | Lookback `S` and repeat interval `T` drive window/slide choice | +| Sharing | None: each query gets its own sketch | One deployment may serve several compatible RQEs | +| Constraints | Accuracy only | Accuracy and per-RQE query latency | +| Objective | Resource use (memory) | Weighted CPU + memory cost from EC2 prices | + +So the comparison invokes AutoSketch **once per RQE** (one QE at one repeat +interval) and sums the results; the planner is invoked once for the batch. + +## 2. What already exists + +| Piece | Where | Status | +| --- | --- | --- | +| AutoSketch Algorithm 4 adaptation (LHS seeds, feasibility-directed width/depth neighbor search, pruning) | ASAPQuery-backend `data_plane/examples/autosketch_comparison.rs` ([#547](https://github.com/ProjectASAP/ASAPQuery-backend/pull/547)) | Merged. CMS/Count Sketch/Bloom only; hardcoded CMS grid; executes sketches to measure accuracy. | +| Earlier protocol (E1–E3, end-to-end execution) | ASAPQuery-backend `docs/evaluation/autosketch-comparison.md` ([#545](https://github.com/ProjectASAP/ASAPQuery-backend/pull/545)) | Merged. Execution-based; this plan is planner-level and uses estimated costs instead. | +| Top-K dashboard comparison | ASAPQuery-backend [#602](https://github.com/ProjectASAP/ASAPQuery-backend/pull/602) | Closed, not merged. | +| RQE deployment MILP (HiGHS): candidates `(capability, config, labels, x, y)`, sharing, latency bounds, minimum-CPU objective | sketch-bench `rqe-optimizer/` ([#129](https://github.com/ProjectASAP/sketch-bench/pull/129)) | Merged 2026-10-04. **This is the planner we evaluate for now.** | +| Measured per-operation costs (`AtomicCostEntry`: memory/instance, insert/merge/query CPU, accuracy) | sketch-bench `scripts/export_rqe_optimizer_costs.sh` | Merged; 18 rows, 2 configs per sketch variant. | +| Moving the MILP into ASAPQuery's planner | ASAPQuery `asap-planner-rs/src/optimizer/` (Milind; related: [#776](https://github.com/ProjectASAP/ASAPQuery/pull/776), [#725](https://github.com/ProjectASAP/ASAPQuery/pull/725)) | In progress. The evaluation does not wait for it (§8, PR 5). | + +No AutoSketch implementation exists in sketch-bench or in this repository. + +## 3. Methods compared + +All methods read the same `AtomicCostTable`, the same RQEs and the same +label-set cardinalities and arrival rates, and are scored by the same cost +function (§4). + +1. **ASAP** — sketch-bench `rqe-optimizer` MILP over the whole batch, with + accuracy and latency constraints, minimizing the §4 cost for one machine + family. +2. **AutoSketch-Adapted** — Algorithm 4 run independently per RQE: + - search space: the measured configs of the RQE's capability families + (`Capability::families()`); + - constraint: the RQE's accuracy tolerance, read from the measured table; + - objective: per-instance memory (AutoSketch's register-memory objective), + tie-break on insert CPU; + - window adapter: one sliding sketch per query, `x = S`, `y = T` (if + `S % T != 0`, `y = gcd(S, T)`), so each evaluation reads one instance and + merges nothing; + - no sharing: every RQE gets its own deployment, even when two RQEs pick an + identical one, so ingest and memory are paid per RQE; + - latency is ignored during search, then checked after. +3. **PerQuery-CostAware** (ablation) — the ASAP MILP solved on each RQE alone, + without latency bounds, and the results summed. It uses the same objective + and window choices as ASAP but no batching or sharing. ASAP vs. this ablation + isolates the batch/sharing benefit; this ablation vs. AutoSketch-Adapted + isolates the objective/window benefit. + +AutoSketch-Adapted is a planner baseline, not a reproduction of the P4 +compiler; stage/page/ALU constraints are dropped, and its accuracy oracle is +the sketch-bench table instead of online sketch execution (§9 Q3). + +## 4. Cost model + +Disk is excluded (agreed with Milind). Units: CPU in vCPU (CPU-seconds per +second), memory in GiB. + +**CPU** — already in `rqe_optimizer::objectives::score`: + +```text +CPU = Σ_active D λ(ℓ_D) · (x_D / y_D) · insert_cpu_D (ingest) + + Σ_r card(ℓ_r) · (query_cpu_D(r) + (S_r / x_D(r) − 1) · merge_cpu_D(r)) / T_r (query + merge) +``` + +**Memory** — new. Today's objective tracks peak per-query memory, not retained +state. Retained state of an active deployment `D` holds `x/y` open instances +plus the closed instances needed by the longest lookback it serves: + +```text +Mem_D = card(ℓ_D) · mem_bytes_per_instance_D · (x_D + max_{r→D} S_r) / y_D +Mem = Σ_active D Mem_D +``` + +In the MILP the `max` is linear: `Mem_D ≥ coef(r, D) · z_{r,D}` for each +eligible `r`. + +**Price per machine family** — for family `f` with `vCPU_f`, `GiB_f` and +on-demand `price_f` ($/hour), the plan needs a fractional instance count +`n_f ≥ CPU / vCPU_f` and `n_f ≥ Mem / GiB_f`; cost is `price_f · n_f`. This +adds one continuous variable and two constraints, and needs no arbitrary split +of an instance's price between CPU and memory. Families: + +| Family | Example instance | Role | +| --- | --- | --- | +| Compute-optimized | c7i.xlarge | cheap CPU, scarce memory | +| General purpose | m7i.xlarge | balanced | +| Memory-optimized | r7i.xlarge | cheap memory, scarce CPU | + +Storage-optimized families are dropped with disk. Prices are fetched once from +the AWS Pricing API (us-east-1, Linux, on-demand), committed as +`rqe-optimizer/data/ec2-pricing-.json` with the query used, and never +edited by hand. Since `n_f` is fractional, instance size within a family does +not change the result. + +ASAP is solved once per family. AutoSketch-Adapted's plan does not depend on +the family; it is scored under each family's cost. + +**Latency** — per-RQE estimate already in sketch-bench: +`card(ℓ) · (query_cpu + (S/x − 1) · merge_cpu)`. + +## 5. Constraints + +**Accuracy target 95%**, mapped per capability to the metrics the cost table +already records: + +| Capability | Metric | Constraint | +| --- | --- | --- | +| Freq | relative error | ≤ 0.05 | +| Quantile | rank error | ≤ 0.05 | +| Cardinality | relative error | ≤ 0.05 | +| TopK | precision@k | ≥ 0.95 | + +**Latency** — per-RQE limit `L_r = α · min latency over r's eligible +deployments`, swept over `α ∈ {1.5, 2, 5, ∞}`. Using a multiple of the +fastest option keeps every point feasible for ASAP and makes the bound bind. +AutoSketch-Adapted ignores it; its violations are counted and reported, and its +cost is shown for those points but marked as infeasible. + +## 6. Workloads + +| ID | Description | Purpose | +| --- | --- | --- | +| W0 | `small_problem`'s 8 RQEs (freq, quantile, cardinality, top-k; 1h–1d lookbacks; 60s/300s intervals), tolerances moved to §5 | Readable worked example; one table in the paper | +| W1 | Seeded synthetic batches, `N ∈ {8, 32, 128, 512, 2048}` RQEs. Lookbacks {5m, 15m, 1h, 6h, 1d}, intervals {10s, 60s, 300s}, 4 capabilities, 3 label sets with fixed cardinality/rate. Knob: fraction of RQEs drawn from shared (capability, labels) cohorts, {0, 0.5, 1}. | Planning-time scaling; cost vs. shareability | +| W2 (optional) | RQEs from the Google cluster-trace query sets ([#746](https://github.com/ProjectASAP/ASAPQuery/pull/746)), with label cardinalities and rates measured from the trace | Realistic mix | + +The cost table used for the evaluation needs a wider grid than today's two +configs per variant, otherwise Algorithm 4's neighbor search has nothing to +search: CMS/Count Sketch depth {2..8} × width {256..8192}, KLL k +{50..800}, DD α {0.005..0.05}, HLL precision {10..16}. Data: Zipf, as in the +existing export script. + +## 7. Metrics and figures + +Reported per (workload, method, machine family, α), median of 10 runs for +timings: + +- **Planning time.** AutoSketch: sum of per-RQE search wall time, plus the + number of accuracy probes. ASAP: candidate generation + dominance pruning + + MILP solve. Offline sketch-bench profiling is shared by both and reported + separately, once. +- **Total cost** ($/hour) and its breakdown: ingest CPU, query + merge CPU, + memory, and which resource binds `n_f`. +- **Latency SLA violations** per method. +- **Estimated accuracy** per RQE (all methods meet it on single-instance + measurements by construction). +- Active deployments and total sketch instances. + +Figures: + +1. Total cost by method, grouped by machine family (W0 and W1 at N = 128). +2. Planning time vs. N, log–log (W1). +3. Cost vs. shareability (W1). +4. Cost vs. latency limit α (W0, W1). + +## 8. Who implements what, in which PR + +| # | Repo | Change | Owner | +| --- | --- | --- | --- | +| this | ASAPQuery | This plan (`docs/evaluation/autosketch-vs-planner.md`) | Zeying | +| 1 | sketch-bench | `rqe-optimizer`: retained-memory term in `objectives.rs`; `milp::minimize_cost` with per-family price (§4); committed EC2 pricing JSON; dominance pruning also compares retained memory, so it cannot drop a candidate that is cheaper under the new objective. Tests: brute-force agreement on the tiny workload, as `#129` already does for CPU. | Zeying, coordinated with Milind since he is porting `milp.rs` | +| 2 | sketch-bench | Wider evaluation grid in `scripts/export_rqe_optimizer_costs.sh` (or a sibling script) and the resulting committed table | Zeying | +| 3 | sketch-bench | `rqe-optimizer/src/autosketch.rs`: Algorithm 4 ported from ASAPQuery-backend `autosketch_comparison.rs`, generalized from the CMS width/depth grid to each variant's measured parameter axes; one dedicated `Deployment` per RQE. Tests: picks the smallest feasible config on a grid; never shares; its window adapter output is eligible under `candidates::is_eligible`. | Zeying | +| 4 | sketch-bench | `rqe-optimizer/examples/autosketch_vs_asap.rs` (W0/W1 generators, all three methods, JSON output) and `scripts/plot_autosketch_vs_asap.py`; committed results and figures | Zeying | +| 5 | ASAPQuery | After the MILP lands in `asap-planner-rs`: port PR 1's objective there and rerun PR 4 against it, so the paper reports the planner that ships | Zeying + Milind | + +PRs 1 and 2 are independent; 3 depends on 2 for a meaningful grid only; 4 +depends on 1–3. + +## 9. Open decisions + +**Q1. AutoSketch's objective.** Recommended: per-instance memory (faithful to +the paper), plus PerQuery-CostAware as the strong ablation. Alternative: give +AutoSketch the §4 cost per query directly, which removes the objective +difference but no longer resembles AutoSketch. + +**Q2. "Once per repeating time".** Recommended: one AutoSketch call per RQE +(QE × repeat interval), reused for every evaluation, since the chosen config +would not change between evaluations. Alternative: one call per evaluation over +a horizon `H`, i.e. planning time `Σ_r (H / T_r) · t_r`, which inflates +AutoSketch's planning time without changing its plan. + +**Q3. AutoSketch accuracy probes.** Recommended: table lookup, the same +evidence ASAP uses, so the plans differ only by algorithm. AutoSketch's +planning time then excludes sketch execution; report its probe count so the +cost of online probing can be stated. Alternative: execute each probe, which +needs sketch-bench in the loop and makes planning-time comparisons dominated by +profiling. + +**Q4. AutoSketch window adapter.** Recommended: `x = S, y = T`, one sliding +sketch per query. Alternative: tumbling `x = y = gcd(S, T)` with merges at +query time, as a sensitivity run. + +**Q5. Memory model.** Recommended: retained state as in §4. Alternative: keep +today's peak per-query memory, which undercounts state for long lookbacks. + +**Q6. Machine-family cost.** Recommended: fractional-instance `max` model (§4). +Alternative: fixed linear weights per family, which need an arbitrary +CPU/memory split of each instance's price. + +## 10. Known limitations + +- **Merged accuracy is not validated.** The cost table measures single + instances. ASAP plans often merge `S/x` instances, while AutoSketch-Adapted + merges none, so this gap affects ASAP only. Before the paper claims accuracy + parity, replay at least W0's chosen plans in sketch-bench and report + post-merge error (`docs/rqe_optimizer_TODO.md`, "Next"). +- Costs and latencies are estimates from per-operation measurements, not + end-to-end executions. The execution-based comparison is ASAPQuery-backend + #545/#547. +- AutoSketch-Adapted's deployments (`x = S`, `y = gcd(S, T)`) are in the ASAP + candidate set (`candidates.rs` generates every divisor of `S` as a window and + `gcd(x, T)` as a slide). So when its plan meets the latency bounds, the ASAP + MILP can choose the same deployments and pay for shared ones once; ASAP's + cost is never higher. The result to report is the size of the gap and where it comes + from, not that a gap exists. From f006b130ecba454a3d6546f34adee42f35ef0ea9 Mon Sep 17 00:00:00 2001 From: zz_y Date: Sun, 4 Oct 2026 21:52:02 +0000 Subject: [PATCH 02/39] docs(planner): settle AutoSketch planning-time and benchmark-input decisions Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 116 +++++++++++++++-------- 1 file changed, 78 insertions(+), 38 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 3ec17170..552e48c6 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -1,8 +1,6 @@ # Evaluation plan: AutoSketch vs. the ASAPQuery planner (paper §6.3) -Status: plan, no results yet. Open decisions are listed in §9 with a -recommended answer each; nothing in sketch-bench is implemented until they are -settled. +Status: plan, no results yet. The design decisions are settled in §9. ## 1. Question @@ -49,7 +47,8 @@ function (§4). 2. **AutoSketch-Adapted** — Algorithm 4 run independently per RQE: - search space: the measured configs of the RQE's capability families (`Capability::families()`); - - constraint: the RQE's accuracy tolerance, read from the measured table; + - constraint: the RQE's accuracy tolerance must hold on **every** benchmark + input (§6, "Benchmark input"), as in the paper's §5.2; - objective: per-instance memory (AutoSketch's register-memory objective), tie-break on insert CPU; - window adapter: one sliding sketch per query, `x = S`, `y = T` (if @@ -65,8 +64,9 @@ function (§4). isolates the objective/window benefit. AutoSketch-Adapted is a planner baseline, not a reproduction of the P4 -compiler; stage/page/ALU constraints are dropped, and its accuracy oracle is -the sketch-bench table instead of online sketch execution (§9 Q3). +compiler: stage/page/ALU constraints are dropped. Its accuracy probes read +sketch-bench measurements instead of running a benchmark inside the search, but +the cost of running those benchmarks is charged to its planning time (§9 Q3). ## 4. Cost model @@ -142,21 +142,60 @@ cost is shown for those points but marked as infeasible. | W1 | Seeded synthetic batches, `N ∈ {8, 32, 128, 512, 2048}` RQEs. Lookbacks {5m, 15m, 1h, 6h, 1d}, intervals {10s, 60s, 300s}, 4 capabilities, 3 label sets with fixed cardinality/rate. Knob: fraction of RQEs drawn from shared (capability, labels) cohorts, {0, 0.5, 1}. | Planning-time scaling; cost vs. shareability | | W2 (optional) | RQEs from the Google cluster-trace query sets ([#746](https://github.com/ProjectASAP/ASAPQuery/pull/746)), with label cardinalities and rates measured from the trace | Realistic mix | +### Benchmark input + The cost table used for the evaluation needs a wider grid than today's two configs per variant, otherwise Algorithm 4's neighbor search has nothing to search: CMS/Count Sketch depth {2..8} × width {256..8192}, KLL k -{50..800}, DD α {0.005..0.05}, HLL precision {10..16}. Data: Zipf, as in the -existing export script. +{50..800}, DD α {0.005..0.05}, HLL precision {10..16}. + +Accuracy depends on how many events one sketch instance absorbs. Today's table +measures every config at one fixed size (`--size 1000000`), whatever the window. +An instance with window `x` on label set `ℓ` absorbs about + +```text +n(x, ℓ) = λ(ℓ) · x / card(ℓ) events +``` + +so a 1-day sketch sees 1440× the events of a 1-minute sketch for the same +query, and needs a larger config to meet the same error. The table is therefore +measured over a size axis, `--size ∈ {10^3, …, 10^8}` in powers of ten. Both +methods look up accuracy at the smallest measured size `≥ n(x, ℓ)`, which is +conservative. AutoSketch-Adapted uses `n(S, ℓ)`, since it keeps one sketch per +query window; ASAP uses `n(x, ℓ)` for the window it picks. + +Following AutoSketch §5.2, each (config, size) point is benchmarked on several +inputs, and a config passes only if it meets the target on all of them: + +- Zipf skew `s ∈ {0.8, 1.1, 1.4}`; +- each with and without traffic bursts, using sketch-bench's AutoSketch-style + burst injection (`--burst-intervals 2 --burst-extra-fraction 0.5`, from + sketch-bench #126). + +The windows of a repeating query see different data each time. AutoSketch does +not re-tune for this: it configures once, before deployment, against these +varied inputs (§9 Q2). ASAP reads the same worst-case accuracy, so both methods +share the same evidence. ## 7. Metrics and figures Reported per (workload, method, machine family, α), median of 10 runs for timings: -- **Planning time.** AutoSketch: sum of per-RQE search wall time, plus the - number of accuracy probes. ASAP: candidate generation + dominance pruning + - MILP solve. Offline sketch-bench profiling is shared by both and reported - separately, once. +- **Planning time.** Reported in two parts, because the two planners spend + their time differently: + - *Search time:* AutoSketch is the sum over RQEs of Algorithm 4 wall time, + using table lookups. ASAP is candidate generation, dominance pruning and + MILP solve. + - *Benchmark time:* AutoSketch benchmarks every probed (config, input size) + per RQE, as in the paper (§5.2, Exp#9: 1–2 minutes per config, about + 6.5 minutes per application). We charge `Σ_r Σ_probes t_bench(config, + n(S_r, ℓ_r))`, where `t_bench` is sketch-bench's measured wall time for + that point over all benchmark inputs; probes already charged for the same + (config, size) are not charged again. ASAP's benchmark time is one profiling + pass over the grid, shared by all RQEs and reusable across workloads. It is + reported once, next to how many RQEs it served. + - The paper's figure shows search + benchmark per method, stacked. - **Total cost** ($/hour) and its breakdown: ingest CPU, query + merge CPU, memory, and which resource binds `n_f`. - **Latency SLA violations** per method. @@ -177,7 +216,7 @@ Figures: | --- | --- | --- | --- | | this | ASAPQuery | This plan (`docs/evaluation/autosketch-vs-planner.md`) | Zeying | | 1 | sketch-bench | `rqe-optimizer`: retained-memory term in `objectives.rs`; `milp::minimize_cost` with per-family price (§4); committed EC2 pricing JSON; dominance pruning also compares retained memory, so it cannot drop a candidate that is cheaper under the new objective. Tests: brute-force agreement on the tiny workload, as `#129` already does for CPU. | Zeying, coordinated with Milind since he is porting `milp.rs` | -| 2 | sketch-bench | Wider evaluation grid in `scripts/export_rqe_optimizer_costs.sh` (or a sibling script) and the resulting committed table | Zeying | +| 2 | sketch-bench | Evaluation table (§6, "Benchmark input"): wider config grid × size axis × skews × bursts, built with a sibling of `scripts/export_rqe_optimizer_costs.sh`. Records benchmark wall time per point (needed for §7). `AtomicCostEntry` gains the measured size and keeps the worst accuracy across inputs; lookup takes the smallest size `≥ n`. Committed table. | Zeying | | 3 | sketch-bench | `rqe-optimizer/src/autosketch.rs`: Algorithm 4 ported from ASAPQuery-backend `autosketch_comparison.rs`, generalized from the CMS width/depth grid to each variant's measured parameter axes; one dedicated `Deployment` per RQE. Tests: picks the smallest feasible config on a grid; never shares; its window adapter output is eligible under `candidates::is_eligible`. | Zeying | | 4 | sketch-bench | `rqe-optimizer/examples/autosketch_vs_asap.rs` (W0/W1 generators, all three methods, JSON output) and `scripts/plot_autosketch_vs_asap.py`; committed results and figures | Zeying | | 5 | ASAPQuery | After the MILP lands in `asap-planner-rs`: port PR 1's objective there and rerun PR 4 against it, so the paper reports the planner that ships | Zeying + Milind | @@ -185,41 +224,42 @@ Figures: PRs 1 and 2 are independent; 3 depends on 2 for a meaningful grid only; 4 depends on 1–3. -## 9. Open decisions +## 9. Decisions -**Q1. AutoSketch's objective.** Recommended: per-instance memory (faithful to -the paper), plus PerQuery-CostAware as the strong ablation. Alternative: give -AutoSketch the §4 cost per query directly, which removes the objective -difference but no longer resembles AutoSketch. +**Q1. AutoSketch's objective.** Per-instance memory, as in the paper +(`SC(c) = α·n_ALU + β·n_mem` with no ALUs in software). PerQuery-CostAware is +the strong ablation. -**Q2. "Once per repeating time".** Recommended: one AutoSketch call per RQE -(QE × repeat interval), reused for every evaluation, since the chosen config -would not change between evaluations. Alternative: one call per evaluation over -a horizon `H`, i.e. planning time `Σ_r (H / T_r) · t_r`, which inflates -AutoSketch's planning time without changing its plan. +**Q2. One AutoSketch call per RQE, reused for every evaluation.** The paper +configures statically: "AutoSketch adopts static configuration instead of +dynamic adjusting" (§3.2), and "the searching is performed once before an +application is deployed" (§7, Exp#9). Data varies between windows of a repeating +query, but AutoSketch handles that through its benchmark inputs, not by +re-planning: a config must meet the target on every benchmark workload, +including random burst intervals (§5.2). So the benchmark input changes per RQE +in one way only: its size follows the RQE's window, `n(S, ℓ)` (§6). It does not +change per evaluation. -**Q3. AutoSketch accuracy probes.** Recommended: table lookup, the same -evidence ASAP uses, so the plans differ only by algorithm. AutoSketch's -planning time then excludes sketch execution; report its probe count so the -cost of online probing can be stated. Alternative: execute each probe, which -needs sketch-bench in the loop and makes planning-time comparisons dominated by -profiling. +**Q3. Probes read sketch-bench measurements, and their benchmark cost is +charged.** In the paper every probe is a benchmark run (§5.2, Algorithm 4 +line 5: "Evaluate T by c"), and benchmarking dominates search time (Exp#9). +Running sketch-bench inside the search would give the same accuracy answers as +reading the same measurements, so the plan is unchanged. Planning time adds the +measured benchmark time of each distinct probed (config, size) point (§7). A +lookup-only time would understate AutoSketch's planning cost. -**Q4. AutoSketch window adapter.** Recommended: `x = S, y = T`, one sliding -sketch per query. Alternative: tumbling `x = y = gcd(S, T)` with merges at -query time, as a sensitivity run. +**Q4. AutoSketch window adapter.** `x = S`, `y = T` (or `gcd(S, T)`): one +sliding sketch per query. -**Q5. Memory model.** Recommended: retained state as in §4. Alternative: keep -today's peak per-query memory, which undercounts state for long lookbacks. +**Q5. Memory model.** Retained state, as in §4. -**Q6. Machine-family cost.** Recommended: fractional-instance `max` model (§4). -Alternative: fixed linear weights per family, which need an arbitrary -CPU/memory split of each instance's price. +**Q6. Machine-family cost.** The fractional-instance `max` model in §4. ## 10. Known limitations - **Merged accuracy is not validated.** The cost table measures single - instances. ASAP plans often merge `S/x` instances, while AutoSketch-Adapted + instances at size `n(x, ℓ)`; merging `S/x` instances is not the same as one + instance of size `n(S, ℓ)`. ASAP plans often merge `S/x` instances, while AutoSketch-Adapted merges none, so this gap affects ASAP only. Before the paper claims accuracy parity, replay at least W0's chosen plans in sketch-bench and report post-merge error (`docs/rqe_optimizer_TODO.md`, "Next"). From 7fadc632e84302445652b410f6b4c955c0c9fbb1 Mon Sep 17 00:00:00 2001 From: zz_y Date: Sun, 4 Oct 2026 22:01:16 +0000 Subject: [PATCH 03/39] docs(planner): look up saturated accuracy and cost, with merged curves for KLL/top-k Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 66 +++++++++++++++++------- 1 file changed, 47 insertions(+), 19 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 552e48c6..abee8172 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -31,6 +31,8 @@ interval) and sums the results; the planner is invoked once for the batch. | Top-K dashboard comparison | ASAPQuery-backend [#602](https://github.com/ProjectASAP/ASAPQuery-backend/pull/602) | Closed, not merged. | | RQE deployment MILP (HiGHS): candidates `(capability, config, labels, x, y)`, sharing, latency bounds, minimum-CPU objective | sketch-bench `rqe-optimizer/` ([#129](https://github.com/ProjectASAP/sketch-bench/pull/129)) | Merged 2026-10-04. **This is the planner we evaluate for now.** | | Measured per-operation costs (`AtomicCostEntry`: memory/instance, insert/merge/query CPU, accuracy) | sketch-bench `scripts/export_rqe_optimizer_costs.sh` | Merged; 18 rows, 2 configs per sketch variant. | +| Saturation study: error vs. `N`, `N_sat`, cost per config and shape | sketch-bench [#130](https://github.com/ProjectASAP/sketch-bench/pull/130) | Merged. Source of the saturated lookup (§6). | +| Accuracy after merging `m` shards (KLL, top-k) | sketch-bench [#131](https://github.com/ProjectASAP/sketch-bench/pull/131) | Open. Needed for KLL/top-k lookups when `m > 1`. | | Moving the MILP into ASAPQuery's planner | ASAPQuery `asap-planner-rs/src/optimizer/` (Milind; related: [#776](https://github.com/ProjectASAP/ASAPQuery/pull/776), [#725](https://github.com/ProjectASAP/ASAPQuery/pull/725)) | In progress. The evaluation does not wait for it (§8, PR 5). | No AutoSketch implementation exists in sketch-bench or in this repository. @@ -149,22 +151,47 @@ configs per variant, otherwise Algorithm 4's neighbor search has nothing to search: CMS/Count Sketch depth {2..8} × width {256..8192}, KLL k {50..800}, DD α {0.005..0.05}, HLL precision {10..16}. -Accuracy depends on how many events one sketch instance absorbs. Today's table -measures every config at one fixed size (`--size 1000000`), whatever the window. -An instance with window `x` on label set `ℓ` absorbs about +Accuracy and cost are read at saturation, using sketch-bench's saturation +study (#130, `docs/saturation_conclusions.md`). Past `N_sat`, a sketch's error +depends on its config alone, not on the stream length `N`. Per-item insert, +merge and query CPU are flat in `N`, and memory is fixed by the config (DDSketch +and KLL grow only with `ln N`). So each (config, data shape) point contributes +its error plateau and its costs at `N_sat`. + +A query with lookback `S` on label set `ℓ` reads about ```text -n(x, ℓ) = λ(ℓ) · x / card(ℓ) events +n(S, ℓ) = λ(ℓ) · S / card(ℓ) events per group ``` -so a 1-day sketch sees 1440× the events of a 1-minute sketch for the same -query, and needs a larger config to meet the same error. The table is therefore -measured over a size axis, `--size ∈ {10^3, …, 10^8}` in powers of ten. Both -methods look up accuracy at the smallest measured size `≥ n(x, ℓ)`, which is -conservative. AutoSketch-Adapted uses `n(S, ℓ)`, since it keeps one sketch per -query window; ASAP uses `n(x, ℓ)` for the window it picks. - -Following AutoSketch §5.2, each (config, size) point is benchmarked on several +The saturated value applies when `n(S, ℓ) ≥ N_sat`. This is what matters when +ASAP merges: a deployment with window `x < S` answers the query by merging +`m = S/x` instances, each holding only `n(x, ℓ)` events, which may be below +`N_sat`. + +- **CMS, Count Sketch, HLL and DDSketch merge exactly.** The merged sketch + equals one sketch over all `n(S, ℓ)` events, so the saturated single-sketch + value is the right lookup however small each pane is. +- **KLL and top-k do not.** Merging raises KLL's error, by 1.0–1.1× at + k = 50/200 and 1.12–1.32× at k = 800, and up to 3–4× at small `N`. Top-k + loses up to 40% precision at large `K`. For these sketches, a deployment + with `m > 1` uses the saturated value from the merged curves at `m` shards + (sketch-bench #131). + +Two cases where the saturated value is not available: + +- If `n(S, ℓ) < N_sat`, use the curve's value at the smallest measured + checkpoint `≥ n(S, ℓ)`. +- With uniform keys and large `K`, CMS, Count Sketch and top-k do not saturate + by 1e9 events. Their `N_sat` is a lower bound; flag these points in the + results instead of treating them as saturated. + +AutoSketch-Adapted keeps one sketch per query window (`m = 1`), so it reads the +single-sketch value at `n(S, ℓ)`. For a given config, both methods therefore +read identical accuracy for CMS, Count Sketch, HLL and DDSketch. For KLL and +top-k, only ASAP pays the merge penalty, which works against ASAP. + +Following AutoSketch §5.2, each config's saturation curve is measured on several inputs, and a config passes only if it meets the target on all of them: - Zipf skew `s ∈ {0.8, 1.1, 1.4}`; @@ -216,7 +243,7 @@ Figures: | --- | --- | --- | --- | | this | ASAPQuery | This plan (`docs/evaluation/autosketch-vs-planner.md`) | Zeying | | 1 | sketch-bench | `rqe-optimizer`: retained-memory term in `objectives.rs`; `milp::minimize_cost` with per-family price (§4); committed EC2 pricing JSON; dominance pruning also compares retained memory, so it cannot drop a candidate that is cheaper under the new objective. Tests: brute-force agreement on the tiny workload, as `#129` already does for CPU. | Zeying, coordinated with Milind since he is porting `milp.rs` | -| 2 | sketch-bench | Evaluation table (§6, "Benchmark input"): wider config grid × size axis × skews × bursts, built with a sibling of `scripts/export_rqe_optimizer_costs.sh`. Records benchmark wall time per point (needed for §7). `AtomicCostEntry` gains the measured size and keeps the worst accuracy across inputs; lookup takes the smallest size `≥ n`. Committed table. | Zeying | +| 2 | sketch-bench | Evaluation table (§6, "Benchmark input"): for the wider config grid × skews × bursts, export each point's saturated error, `N_sat`, its saturation curve, and costs at `N_sat`, from #130's `study_saturation.py` outputs. For KLL and top-k, add the merged-curve values per `m` from #131. Keep the worst accuracy across inputs. Record benchmark wall time per point (needed for §7). `rqe-optimizer`'s lookup uses `m = S/x` and `n(S, ℓ)` as in §6. Committed table. Depends on #131 for KLL/top-k. | Zeying | | 3 | sketch-bench | `rqe-optimizer/src/autosketch.rs`: Algorithm 4 ported from ASAPQuery-backend `autosketch_comparison.rs`, generalized from the CMS width/depth grid to each variant's measured parameter axes; one dedicated `Deployment` per RQE. Tests: picks the smallest feasible config on a grid; never shares; its window adapter output is eligible under `candidates::is_eligible`. | Zeying | | 4 | sketch-bench | `rqe-optimizer/examples/autosketch_vs_asap.rs` (W0/W1 generators, all three methods, JSON output) and `scripts/plot_autosketch_vs_asap.py`; committed results and figures | Zeying | | 5 | ASAPQuery | After the MILP lands in `asap-planner-rs`: port PR 1's objective there and rerun PR 4 against it, so the paper reports the planner that ships | Zeying + Milind | @@ -257,12 +284,13 @@ sliding sketch per query. ## 10. Known limitations -- **Merged accuracy is not validated.** The cost table measures single - instances at size `n(x, ℓ)`; merging `S/x` instances is not the same as one - instance of size `n(S, ℓ)`. ASAP plans often merge `S/x` instances, while AutoSketch-Adapted - merges none, so this gap affects ASAP only. Before the paper claims accuracy - parity, replay at least W0's chosen plans in sketch-bench and report - post-merge error (`docs/rqe_optimizer_TODO.md`, "Next"). +- **Merged accuracy comes from shard-merge measurements, not from replaying + the plans.** It is exact by construction for CMS, Count Sketch, HLL and + DDSketch, and taken from #131 for KLL and top-k. #131 covers `N ≤ 1e7` and + `m ≤ 64`. A deployment needing `m > 64` (e.g. a 1-day lookback over 1-minute + windows, `m = 1440`) is outside the measured range. Mark it as extrapolated, + or exclude it for KLL/top-k. Replay W0's chosen plans in sketch-bench once to + confirm the lookups. - Costs and latencies are estimates from per-operation measurements, not end-to-end executions. The execution-based comparison is ASAPQuery-backend #545/#547. From 5b2510bcbf1cf50a9be2715fdf9d9d2e813c256d Mon Sep 17 00:00:00 2001 From: zz_y Date: Sun, 4 Oct 2026 22:15:28 +0000 Subject: [PATCH 04/39] docs(planner): saturation applies to ASAP only; fit data parameters on the whole dataset Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 80 ++++++++++++++---------- 1 file changed, 48 insertions(+), 32 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index abee8172..e49d73f1 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -142,7 +142,7 @@ cost is shown for those points but marked as infeasible. | --- | --- | --- | | W0 | `small_problem`'s 8 RQEs (freq, quantile, cardinality, top-k; 1h–1d lookbacks; 60s/300s intervals), tolerances moved to §5 | Readable worked example; one table in the paper | | W1 | Seeded synthetic batches, `N ∈ {8, 32, 128, 512, 2048}` RQEs. Lookbacks {5m, 15m, 1h, 6h, 1d}, intervals {10s, 60s, 300s}, 4 capabilities, 3 label sets with fixed cardinality/rate. Knob: fraction of RQEs drawn from shared (capability, labels) cohorts, {0, 0.5, 1}. | Planning-time scaling; cost vs. shareability | -| W2 (optional) | RQEs from the Google cluster-trace query sets ([#746](https://github.com/ProjectASAP/ASAPQuery/pull/746)), with label cardinalities and rates measured from the trace | Realistic mix | +| W2 | RQEs from the Google cluster-trace query sets ([#746](https://github.com/ProjectASAP/ASAPQuery/pull/746)). Label cardinalities, rates and data parameters are fit over the whole trace. | Reported results | ### Benchmark input @@ -151,12 +151,26 @@ configs per variant, otherwise Algorithm 4's neighbor search has nothing to search: CMS/Count Sketch depth {2..8} × width {256..8192}, KLL k {50..800}, DD α {0.005..0.05}, HLL precision {10..16}. -Accuracy and cost are read at saturation, using sketch-bench's saturation -study (#130, `docs/saturation_conclusions.md`). Past `N_sat`, a sketch's error -depends on its config alone, not on the stream length `N`. Per-item insert, -merge and query CPU are flat in `N`, and memory is fixed by the config (DDSketch -and KLL grow only with `ln N`). So each (config, data shape) point contributes -its error plateau and its costs at `N_sat`. +#### Data parameters, shared by both methods + +The benchmark inputs are generated from data parameters fit over the **whole +measured dataset**, not from a short sample or the average. These are key skew +`θ`, distinct keys `K` per window, value tail index `a`, and per-label-set +`λ` and `card`. Use the worst case across the dataset; for example, size CMS +and top-k from the lower `θ` bound. Longer samples expose worse cases (sketch-bench +`docs/saturation_conclusions.md`, conclusions 1–5). Fitting follows ASAPQuery +#746 and sketch-bench `scripts/recommend_config.py`. + +- **W2 (trace) gives the reported results.** Its parameters are fit on the + full trace. +- **W0/W1 (synthetic)** use their generators' parameters. + +Following AutoSketch §5.2, each config is benchmarked at these parameters with +and without traffic bursts. The bursts use sketch-bench's AutoSketch-style +injection (`--burst-intervals 2 --burst-extra-fraction 0.5`, sketch-bench +#126). A config passes only if it meets the target on all of these inputs. +AutoSketch's benchmark uses the same inputs, which matches the paper: it lets +users "use their own trace". A query with lookback `S` on label set `ℓ` reads about @@ -164,45 +178,47 @@ A query with lookback `S` on label set `ℓ` reads about n(S, ℓ) = λ(ℓ) · S / card(ℓ) events per group ``` -The saturated value applies when `n(S, ℓ) ≥ N_sat`. This is what matters when -ASAP merges: a deployment with window `x < S` answers the query by merging -`m = S/x` instances, each holding only `n(x, ℓ)` events, which may be below -`N_sat`. +#### ASAPQuery: saturated values, because it merges + +ASAP may answer a query by merging `m = S/x` smaller-window sketches into one. +Each of them holds only `n(x, ℓ)` events, which may be below the length at +which its error has settled. So ASAP reads accuracy and cost at saturation, as +measured in sketch-bench's saturation study (#130): + +- Past `N_sat`, error depends on the config alone. +- Per-item insert, merge and query CPU are flat in `N`. +- Memory is fixed by the config; DDSketch and KLL grow only with `ln N`. + +Each (config, dataset parameters) point therefore contributes its error +plateau and its costs at `N_sat`. `N_sat` is taken over the whole measured +dataset: it is the saturation length at the dataset's worst-case parameters +above. A lookup is valid when `n(S, ℓ) ≥ N_sat`. - **CMS, Count Sketch, HLL and DDSketch merge exactly.** The merged sketch equals one sketch over all `n(S, ℓ)` events, so the saturated single-sketch - value is the right lookup however small each pane is. + value applies however small each pane is. - **KLL and top-k do not.** Merging raises KLL's error, by 1.0–1.1× at k = 50/200 and 1.12–1.32× at k = 800, and up to 3–4× at small `N`. Top-k loses up to 40% precision at large `K`. For these sketches, a deployment with `m > 1` uses the saturated value from the merged curves at `m` shards (sketch-bench #131). - -Two cases where the saturated value is not available: - -- If `n(S, ℓ) < N_sat`, use the curve's value at the smallest measured - checkpoint `≥ n(S, ℓ)`. +- If `n(S, ℓ) < N_sat`, the query's window never saturates. ASAP uses the curve's + value at the smallest measured checkpoint `≥ n(S, ℓ)`. - With uniform keys and large `K`, CMS, Count Sketch and top-k do not saturate by 1e9 events. Their `N_sat` is a lower bound; flag these points in the - results instead of treating them as saturated. - -AutoSketch-Adapted keeps one sketch per query window (`m = 1`), so it reads the -single-sketch value at `n(S, ℓ)`. For a given config, both methods therefore -read identical accuracy for CMS, Count Sketch, HLL and DDSketch. For KLL and -top-k, only ASAP pays the merge penalty, which works against ASAP. + results. -Following AutoSketch §5.2, each config's saturation curve is measured on several -inputs, and a config passes only if it meets the target on all of them: +#### AutoSketch: no saturation requirement -- Zipf skew `s ∈ {0.8, 1.1, 1.4}`; -- each with and without traffic bursts, using sketch-bench's AutoSketch-style - burst injection (`--burst-intervals 2 --burst-extra-fraction 0.5`, from - sketch-bench #126). +AutoSketch never merges: it keeps one sketch per query window. In the paper it +benchmarks a config on its workloads and accepts it if the accuracy intent +holds, with no notion of `N_sat`. AutoSketch-Adapted therefore reads the +measured value at its own input size `n(S, ℓ)`, using the dataset-derived +inputs above. The windows of a repeating query see different data each time. AutoSketch does not re-tune for this: it configures once, before deployment, against these -varied inputs (§9 Q2). ASAP reads the same worst-case accuracy, so both methods -share the same evidence. +inputs (§9 Q2). ## 7. Metrics and figures @@ -243,7 +259,7 @@ Figures: | --- | --- | --- | --- | | this | ASAPQuery | This plan (`docs/evaluation/autosketch-vs-planner.md`) | Zeying | | 1 | sketch-bench | `rqe-optimizer`: retained-memory term in `objectives.rs`; `milp::minimize_cost` with per-family price (§4); committed EC2 pricing JSON; dominance pruning also compares retained memory, so it cannot drop a candidate that is cheaper under the new objective. Tests: brute-force agreement on the tiny workload, as `#129` already does for CPU. | Zeying, coordinated with Milind since he is porting `milp.rs` | -| 2 | sketch-bench | Evaluation table (§6, "Benchmark input"): for the wider config grid × skews × bursts, export each point's saturated error, `N_sat`, its saturation curve, and costs at `N_sat`, from #130's `study_saturation.py` outputs. For KLL and top-k, add the merged-curve values per `m` from #131. Keep the worst accuracy across inputs. Record benchmark wall time per point (needed for §7). `rqe-optimizer`'s lookup uses `m = S/x` and `n(S, ℓ)` as in §6. Committed table. Depends on #131 for KLL/top-k. | Zeying | +| 2 | sketch-bench | Evaluation table (§6, "Benchmark input"): for the wider config grid at the dataset-fit parameters × bursts, export each point's saturated error, `N_sat`, its saturation curve, and costs at `N_sat`, from #130's `study_saturation.py` outputs. For KLL and top-k, add the merged-curve values per `m` from #131. Keep the worst accuracy across inputs. Also keep each curve's value at every checkpoint, which AutoSketch's lookup at `n(S, ℓ)` needs. Record benchmark wall time per point (needed for §7). Lookups follow §6: ASAP uses saturated values with `m = S/x`; AutoSketch uses the curve value at `n(S, ℓ)`. Committed table. Depends on #131 for KLL/top-k. | Zeying | | 3 | sketch-bench | `rqe-optimizer/src/autosketch.rs`: Algorithm 4 ported from ASAPQuery-backend `autosketch_comparison.rs`, generalized from the CMS width/depth grid to each variant's measured parameter axes; one dedicated `Deployment` per RQE. Tests: picks the smallest feasible config on a grid; never shares; its window adapter output is eligible under `candidates::is_eligible`. | Zeying | | 4 | sketch-bench | `rqe-optimizer/examples/autosketch_vs_asap.rs` (W0/W1 generators, all three methods, JSON output) and `scripts/plot_autosketch_vs_asap.py`; committed results and figures | Zeying | | 5 | ASAPQuery | After the MILP lands in `asap-planner-rs`: port PR 1's objective there and rerun PR 4 against it, so the paper reports the planner that ships | Zeying + Milind | From 57d6498fc8eb8c177d3a67d334036ad840738322 Mon Sep 17 00:00:00 2001 From: zz_y Date: Sun, 4 Oct 2026 22:23:31 +0000 Subject: [PATCH 05/39] docs(planner): drop bursts; W2 covers all three traces Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 19 +++++++++---------- 1 file changed, 9 insertions(+), 10 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index e49d73f1..d25eeb3c 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -142,7 +142,7 @@ cost is shown for those points but marked as infeasible. | --- | --- | --- | | W0 | `small_problem`'s 8 RQEs (freq, quantile, cardinality, top-k; 1h–1d lookbacks; 60s/300s intervals), tolerances moved to §5 | Readable worked example; one table in the paper | | W1 | Seeded synthetic batches, `N ∈ {8, 32, 128, 512, 2048}` RQEs. Lookbacks {5m, 15m, 1h, 6h, 1d}, intervals {10s, 60s, 300s}, 4 capabilities, 3 label sets with fixed cardinality/rate. Knob: fraction of RQEs drawn from shared (capability, labels) cohorts, {0, 0.5, 1}. | Planning-time scaling; cost vs. shareability | -| W2 | RQEs from the Google cluster-trace query sets ([#746](https://github.com/ProjectASAP/ASAPQuery/pull/746)). Label cardinalities, rates and data parameters are fit over the whole trace. | Reported results | +| W2 | Real-trace RQEs, one workload per dataset: Alibaba 2022, BOOM and Google 2011. Taken from `asap-tools/dataset-analysis/results/skew_summary.csv` ([#746](https://github.com/ProjectASAP/ASAPQuery/pull/746)): each row's `range_s` is `S` and its `step_s` is `T`. Data parameters and accuracy targets are fit over each whole trace. | Reported results | ### Benchmark input @@ -165,12 +165,11 @@ and top-k from the lower `θ` bound. Longer samples expose worse cases (sketch-b full trace. - **W0/W1 (synthetic)** use their generators' parameters. -Following AutoSketch §5.2, each config is benchmarked at these parameters with -and without traffic bursts. The bursts use sketch-bench's AutoSketch-style -injection (`--burst-intervals 2 --burst-extra-fraction 0.5`, sketch-bench -#126). A config passes only if it meets the target on all of these inputs. -AutoSketch's benchmark uses the same inputs, which matches the paper: it lets -users "use their own trace". +Each config is benchmarked at these worst-case parameters. AutoSketch §5.2 +injects random traffic bursts into synthetic workloads to cover variation over +time. We don't need them: worst-case fits over every window of the whole dataset +already cover that variation. AutoSketch's benchmark uses the same inputs, +which matches the paper: it lets users "use their own trace". A query with lookback `S` on label set `ℓ` reads about @@ -259,7 +258,7 @@ Figures: | --- | --- | --- | --- | | this | ASAPQuery | This plan (`docs/evaluation/autosketch-vs-planner.md`) | Zeying | | 1 | sketch-bench | `rqe-optimizer`: retained-memory term in `objectives.rs`; `milp::minimize_cost` with per-family price (§4); committed EC2 pricing JSON; dominance pruning also compares retained memory, so it cannot drop a candidate that is cheaper under the new objective. Tests: brute-force agreement on the tiny workload, as `#129` already does for CPU. | Zeying, coordinated with Milind since he is porting `milp.rs` | -| 2 | sketch-bench | Evaluation table (§6, "Benchmark input"): for the wider config grid at the dataset-fit parameters × bursts, export each point's saturated error, `N_sat`, its saturation curve, and costs at `N_sat`, from #130's `study_saturation.py` outputs. For KLL and top-k, add the merged-curve values per `m` from #131. Keep the worst accuracy across inputs. Also keep each curve's value at every checkpoint, which AutoSketch's lookup at `n(S, ℓ)` needs. Record benchmark wall time per point (needed for §7). Lookups follow §6: ASAP uses saturated values with `m = S/x`; AutoSketch uses the curve value at `n(S, ℓ)`. Committed table. Depends on #131 for KLL/top-k. | Zeying | +| 2 | sketch-bench | Evaluation table (§6, "Benchmark input"): for the wider config grid at the worst-case parameters of each dataset, export each point's saturated error, `N_sat`, its saturation curve, and costs at `N_sat`, from #130's `study_saturation.py` outputs. For KLL and top-k, add the merged-curve values per `m` from #131. Keep the worst accuracy across inputs. Also keep each curve's value at every checkpoint, which AutoSketch's lookup at `n(S, ℓ)` needs. Record benchmark wall time per point (needed for §7). Lookups follow §6: ASAP uses saturated values with `m = S/x`; AutoSketch uses the curve value at `n(S, ℓ)`. Committed table. Depends on #131 for KLL/top-k. | Zeying | | 3 | sketch-bench | `rqe-optimizer/src/autosketch.rs`: Algorithm 4 ported from ASAPQuery-backend `autosketch_comparison.rs`, generalized from the CMS width/depth grid to each variant's measured parameter axes; one dedicated `Deployment` per RQE. Tests: picks the smallest feasible config on a grid; never shares; its window adapter output is eligible under `candidates::is_eligible`. | Zeying | | 4 | sketch-bench | `rqe-optimizer/examples/autosketch_vs_asap.rs` (W0/W1 generators, all three methods, JSON output) and `scripts/plot_autosketch_vs_asap.py`; committed results and figures | Zeying | | 5 | ASAPQuery | After the MILP lands in `asap-planner-rs`: port PR 1's objective there and rerun PR 4 against it, so the paper reports the planner that ships | Zeying + Milind | @@ -278,8 +277,8 @@ configures statically: "AutoSketch adopts static configuration instead of dynamic adjusting" (§3.2), and "the searching is performed once before an application is deployed" (§7, Exp#9). Data varies between windows of a repeating query, but AutoSketch handles that through its benchmark inputs, not by -re-planning: a config must meet the target on every benchmark workload, -including random burst intervals (§5.2). So the benchmark input changes per RQE +re-planning: a config must meet the target on every benchmark workload (§5.2). +Here that is the worst case over the whole dataset (§6). So the benchmark input changes per RQE in one way only: its size follows the RQE's window, `n(S, ℓ)` (§6). It does not change per evaluation. From 9a74b9cb658b37a015accd1ee3090f1cf2f6caf1 Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 02:16:57 +0000 Subject: [PATCH 06/39] =?UTF-8?q?docs(planner):=20define=20the=20synthetic?= =?UTF-8?q?=20PromQL=20workload=20for=20=C2=A76.3?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 43 ++++++++++++++++++++++++ 1 file changed, 43 insertions(+) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index d25eeb3c..e10ed618 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -141,9 +141,52 @@ cost is shown for those points but marked as infeasible. | ID | Description | Purpose | | --- | --- | --- | | W0 | `small_problem`'s 8 RQEs (freq, quantile, cardinality, top-k; 1h–1d lookbacks; 60s/300s intervals), tolerances moved to §5 | Readable worked example; one table in the paper | +| WS | Synthetic PromQL workload: the 10 queries in "Synthetic workload" below, over Zipf/Pareto data. Main figure. | Cost–latency trade-off across data and requirements | | W1 | Seeded synthetic batches, `N ∈ {8, 32, 128, 512, 2048}` RQEs. Lookbacks {5m, 15m, 1h, 6h, 1d}, intervals {10s, 60s, 300s}, 4 capabilities, 3 label sets with fixed cardinality/rate. Knob: fraction of RQEs drawn from shared (capability, labels) cohorts, {0, 0.5, 1}. | Planning-time scaling; cost vs. shareability | | W2 | Real-trace RQEs, one workload per dataset: Alibaba 2022, BOOM and Google 2011. Taken from `asap-tools/dataset-analysis/results/skew_summary.csv` ([#746](https://github.com/ProjectASAP/ASAPQuery/pull/746)): each row's `range_s` is `S` and its `step_s` is `T`. Data parameters and accuracy targets are fit over each whole trace. | Reported results | +### Synthetic workload (WS) + +**Data.** +- Series carry `label_0`, with cardinality in {10^1, …, 10^6}, and an + `instance` label with 100 values per `label_0` value. +- Each series is scraped every 10 ms, i.e. 100 samples per second, so + `λ = 100 · 100 · card(label_0)` samples/s. +- The data volume is chosen so that even the smallest windows hold enough + samples for a sketch (1e4 per group for 1 s spatial queries, 6e3 per series + for 1 m temporal ones). Points that still fall below `N_sat` are flagged. +- Key weights for frequency and top-k follow Zipf θ ∈ {0, 0.5, 1.0, 1.5, 2.0}. + Values for quantiles follow Pareto a ∈ {1.1, 2, 3}. +- The saturation curves are extended to K ∈ {1e1, 1e2, 1e4, 1e6} by + measuring, not by interpolation. + +**Queries.** Each spatial query repeats every 1 s, with `S = T = 1 s`, so each +evaluation reads the last second. Each temporal query repeats every 1 m, with +`S = T_range ∈ {1m, 10m, 1h, 6h, 24h}`. + +| # | Query | Capability | Grouping | +|---|---|---|---| +| 1 | `sum by (label_0) (data)` | Freq | `label_0` | +| 2 | `topk by (3, label_0) (data)` | TopK | `label_0` | +| 3 | `quantile by (q, label_0) (data)`, q ∈ {.5, .75, .9, .95, .99} | Quantile | per `label_0` group | +| 4 | `sum_over_time(data[T])` | Freq | per series | +| 5 | `quantile_over_time(q, data[T])`, same five q | Quantile | per series | +| 6 | `rate(data[T])` | Freq over per-series increments | per series | +| 7 | `sum by (label_0) (rate(data[T]))` | Freq over increments | `label_0` | +| 8 | `sum by (label_0) (sum_over_time(data[T]))` | Freq | `label_0` | +| 9 | `topk by (3, label_0) (rate(data[T]))` | TopK over increments | `label_0` | +| 10 | `quantile_over_time(0.9, data[T]) / quantile_over_time(0.5, data[T])` | Two Quantile RQEs | per series | + +Notes on the mapping: +- `rate`/`increase` are modeled as a frequency sum of per-series increments, + equivalent to `sum_over_time` over deltas. +- The quantiles of one query, and the two operands of query 10, read the same + stream. ASAP can serve them from one deployment; AutoSketch gets one per RQE. +- One workload instance is the 42 RQEs above, for one (cardinality, θ or a, + accuracy target, SLA) combination. The figure sweeps the accuracy target over + {90%, 95%, 99%} and an absolute latency SLA grid. For the scalability study, + the RQE set is replicated with distinct `label_0` filters. + ### Benchmark input The cost table used for the evaluation needs a wider grid than today's two From 9a731e248ad7e59b49d8c89d71f1cbf92ed74c80 Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 12:46:53 +0000 Subject: [PATCH 07/39] docs(planner): strawmen meet the latency SLA; add FewestPlans Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 11 +++++++++-- 1 file changed, 9 insertions(+), 2 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index e10ed618..7732d97f 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -59,11 +59,18 @@ function (§4). - no sharing: every RQE gets its own deployment, even when two RQEs pick an identical one, so ingest and memory are paid per RQE; - latency is ignored during search, then checked after. -3. **PerQuery-CostAware** (ablation) — the ASAP MILP solved on each RQE alone, - without latency bounds, and the results summed. It uses the same objective +3. **PerQuery-CostAware** (strawman) — the ASAP MILP solved on each RQE alone, + with the same accuracy and latency requirements, and the results summed. It uses the same objective and window choices as ASAP but no batching or sharing. ASAP vs. this ablation isolates the batch/sharing benefit; this ablation vs. AutoSketch-Adapted isolates the objective/window benefit. +4. **FewestPlans** (strawman) — the ASAP MILP minimizing the number of active + deployments first, then cost among plans with that minimum count, under + the same accuracy and latency requirements. + +Only AutoSketch-Adapted ignores latency; the two strawmen must meet the same +requirements as ASAP, so every point in the cost–latency figure except +AutoSketch's is a feasible plan. AutoSketch-Adapted is a planner baseline, not a reproduction of the P4 compiler: stage/page/ALU constraints are dropped. Its accuracy probes read From da5f69abb6a906c8e4aa4dc012dfda62f4e485fb Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 13:11:57 +0000 Subject: [PATCH 08/39] docs(planner): WS has 67 RQEs, not 42 Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 7732d97f..f4f39b22 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -189,7 +189,7 @@ Notes on the mapping: equivalent to `sum_over_time` over deltas. - The quantiles of one query, and the two operands of query 10, read the same stream. ASAP can serve them from one deployment; AutoSketch gets one per RQE. -- One workload instance is the 42 RQEs above, for one (cardinality, θ or a, +- One workload instance is the 67 RQEs above (2 spatial + 5 spatial quantiles + 5×5 temporal for queries 4, 6–9 + 5×5 for query 5 + 2×5 for query 10), for one (cardinality, θ or a, accuracy target, SLA) combination. The figure sweeps the accuracy target over {90%, 95%, 99%} and an absolute latency SLA grid. For the scalability study, the RQE set is replicated with distinct `label_0` filters. From 0c6521a976b9498ab5d7de89d5de013334194d85 Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 13:29:49 +0000 Subject: [PATCH 09/39] docs(planner): name the workloads example, scaling, traces and synthetic Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 26 ++++++++++++------------ 1 file changed, 13 insertions(+), 13 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index f4f39b22..713bbe2d 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -147,12 +147,12 @@ cost is shown for those points but marked as infeasible. | ID | Description | Purpose | | --- | --- | --- | -| W0 | `small_problem`'s 8 RQEs (freq, quantile, cardinality, top-k; 1h–1d lookbacks; 60s/300s intervals), tolerances moved to §5 | Readable worked example; one table in the paper | -| WS | Synthetic PromQL workload: the 10 queries in "Synthetic workload" below, over Zipf/Pareto data. Main figure. | Cost–latency trade-off across data and requirements | -| W1 | Seeded synthetic batches, `N ∈ {8, 32, 128, 512, 2048}` RQEs. Lookbacks {5m, 15m, 1h, 6h, 1d}, intervals {10s, 60s, 300s}, 4 capabilities, 3 label sets with fixed cardinality/rate. Knob: fraction of RQEs drawn from shared (capability, labels) cohorts, {0, 0.5, 1}. | Planning-time scaling; cost vs. shareability | -| W2 | Real-trace RQEs, one workload per dataset: Alibaba 2022, BOOM and Google 2011. Taken from `asap-tools/dataset-analysis/results/skew_summary.csv` ([#746](https://github.com/ProjectASAP/ASAPQuery/pull/746)): each row's `range_s` is `S` and its `step_s` is `T`. Data parameters and accuracy targets are fit over each whole trace. | Reported results | +| `example` | `small_problem`'s 8 RQEs (freq, quantile, cardinality, top-k; 1h–1d lookbacks; 60s/300s intervals), tolerances moved to §5 | Readable worked example; one table in the paper | +| `synthetic` | Synthetic PromQL workload: the 10 queries in "Synthetic workload" below, over Zipf/Pareto data. Main figure. | Cost–latency trade-off across data and requirements | +| `scaling` | Seeded random batches, `N ∈ {8, 32, 128, 512, 2048}` RQEs. Lookbacks {5m, 15m, 1h, 6h, 1d}, intervals {10s, 60s, 300s}, 4 capabilities, 3 label sets with fixed cardinality/rate. Knob: fraction of RQEs drawn from shared (capability, labels) cohorts, {0, 0.5, 1}. | Planning-time scaling; cost vs. shareability | +| `traces` | Real-trace RQEs, one workload per dataset: Alibaba 2022, BOOM and Google 2011. Taken from `asap-tools/dataset-analysis/results/skew_summary.csv` ([#746](https://github.com/ProjectASAP/ASAPQuery/pull/746)): each row's `range_s` is `S` and its `step_s` is `T`. Data parameters and accuracy targets are fit over each whole trace. | Appendix: real-trace results | -### Synthetic workload (WS) +### Synthetic workload **Data.** - Series carry `label_0`, with cardinality in {10^1, …, 10^6}, and an @@ -211,9 +211,9 @@ and top-k from the lower `θ` bound. Longer samples expose worse cases (sketch-b `docs/saturation_conclusions.md`, conclusions 1–5). Fitting follows ASAPQuery #746 and sketch-bench `scripts/recommend_config.py`. -- **W2 (trace) gives the reported results.** Its parameters are fit on the +- **`traces` gives the appendix results.** Its parameters are fit on the full trace. -- **W0/W1 (synthetic)** use their generators' parameters. +- **`example`, `scaling` and `synthetic`** use their generators' parameters. Each config is benchmarked at these worst-case parameters. AutoSketch §5.2 injects random traffic bursts into synthetic workloads to cover variation over @@ -297,10 +297,10 @@ timings: Figures: -1. Total cost by method, grouped by machine family (W0 and W1 at N = 128). -2. Planning time vs. N, log–log (W1). -3. Cost vs. shareability (W1). -4. Cost vs. latency limit α (W0, W1). +1. Total cost by method, grouped by machine family (`example` and `scaling` at N = 128). +2. Planning time vs. N, log–log (`scaling`). +3. Cost vs. shareability (`scaling`). +4. Cost vs. absolute latency SLA (`example`, `scaling`). ## 8. Who implements what, in which PR @@ -310,7 +310,7 @@ Figures: | 1 | sketch-bench | `rqe-optimizer`: retained-memory term in `objectives.rs`; `milp::minimize_cost` with per-family price (§4); committed EC2 pricing JSON; dominance pruning also compares retained memory, so it cannot drop a candidate that is cheaper under the new objective. Tests: brute-force agreement on the tiny workload, as `#129` already does for CPU. | Zeying, coordinated with Milind since he is porting `milp.rs` | | 2 | sketch-bench | Evaluation table (§6, "Benchmark input"): for the wider config grid at the worst-case parameters of each dataset, export each point's saturated error, `N_sat`, its saturation curve, and costs at `N_sat`, from #130's `study_saturation.py` outputs. For KLL and top-k, add the merged-curve values per `m` from #131. Keep the worst accuracy across inputs. Also keep each curve's value at every checkpoint, which AutoSketch's lookup at `n(S, ℓ)` needs. Record benchmark wall time per point (needed for §7). Lookups follow §6: ASAP uses saturated values with `m = S/x`; AutoSketch uses the curve value at `n(S, ℓ)`. Committed table. Depends on #131 for KLL/top-k. | Zeying | | 3 | sketch-bench | `rqe-optimizer/src/autosketch.rs`: Algorithm 4 ported from ASAPQuery-backend `autosketch_comparison.rs`, generalized from the CMS width/depth grid to each variant's measured parameter axes; one dedicated `Deployment` per RQE. Tests: picks the smallest feasible config on a grid; never shares; its window adapter output is eligible under `candidates::is_eligible`. | Zeying | -| 4 | sketch-bench | `rqe-optimizer/examples/autosketch_vs_asap.rs` (W0/W1 generators, all three methods, JSON output) and `scripts/plot_autosketch_vs_asap.py`; committed results and figures | Zeying | +| 4 | sketch-bench | `rqe-optimizer/examples/autosketch_vs_asap.rs` (`example`/`scaling` generators, all three methods, JSON output) and `scripts/plot_autosketch_vs_asap.py`; committed results and figures | Zeying | | 5 | ASAPQuery | After the MILP lands in `asap-planner-rs`: port PR 1's objective there and rerun PR 4 against it, so the paper reports the planner that ships | Zeying + Milind | PRs 1 and 2 are independent; 3 depends on 2 for a meaningful grid only; 4 @@ -354,7 +354,7 @@ sliding sketch per query. DDSketch, and taken from #131 for KLL and top-k. #131 covers `N ≤ 1e7` and `m ≤ 64`. A deployment needing `m > 64` (e.g. a 1-day lookback over 1-minute windows, `m = 1440`) is outside the measured range. Mark it as extrapolated, - or exclude it for KLL/top-k. Replay W0's chosen plans in sketch-bench once to + or exclude it for KLL/top-k. Replay `example`'s chosen plans in sketch-bench once to confirm the lookups. - Costs and latencies are estimates from per-operation measurements, not end-to-end executions. The execution-based comparison is ASAPQuery-backend From 897dd30c10dfb6d6ee9ca7b43014a8e1267cb3c3 Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 13:43:10 +0000 Subject: [PATCH 10/39] docs(planner): report estimated latency alongside SLA violations Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 713bbe2d..7635c040 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -290,7 +290,9 @@ timings: - The paper's figure shows search + benchmark per method, stacked. - **Total cost** ($/hour) and its breakdown: ingest CPU, query + merge CPU, memory, and which resource binds `n_f`. -- **Latency SLA violations** per method. +- **Estimated query latency and latency SLA violations** per method. + - Estimated latency per RQE: `card(ℓ) · (query_cpu + (S/x − 1) · merge_cpu)` (§4). Report its maximum and median over the RQEs, plus the per-RQE values in the raw output. + - SLA violations: the number of RQEs whose estimated latency exceeds the SLA. Only AutoSketch-Adapted can have any, since the other methods are constrained. - **Estimated accuracy** per RQE (all methods meet it on single-instance measurements by construction). - Active deployments and total sketch instances. From 6bf087d2f9ccacfa5f2e0c310dd84aeddf37786c Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 13:50:15 +0000 Subject: [PATCH 11/39] docs(planner): #131 is merged Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 7635c040..3ca7f9f3 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -32,7 +32,7 @@ interval) and sums the results; the planner is invoked once for the batch. | RQE deployment MILP (HiGHS): candidates `(capability, config, labels, x, y)`, sharing, latency bounds, minimum-CPU objective | sketch-bench `rqe-optimizer/` ([#129](https://github.com/ProjectASAP/sketch-bench/pull/129)) | Merged 2026-10-04. **This is the planner we evaluate for now.** | | Measured per-operation costs (`AtomicCostEntry`: memory/instance, insert/merge/query CPU, accuracy) | sketch-bench `scripts/export_rqe_optimizer_costs.sh` | Merged; 18 rows, 2 configs per sketch variant. | | Saturation study: error vs. `N`, `N_sat`, cost per config and shape | sketch-bench [#130](https://github.com/ProjectASAP/sketch-bench/pull/130) | Merged. Source of the saturated lookup (§6). | -| Accuracy after merging `m` shards (KLL, top-k) | sketch-bench [#131](https://github.com/ProjectASAP/sketch-bench/pull/131) | Open. Needed for KLL/top-k lookups when `m > 1`. | +| Accuracy after merging `m` shards (KLL, top-k) | sketch-bench [#131](https://github.com/ProjectASAP/sketch-bench/pull/131) | Merged 2026-10-05 (`bd644fe`). Source of KLL/top-k lookups when `m > 1`. | | Moving the MILP into ASAPQuery's planner | ASAPQuery `asap-planner-rs/src/optimizer/` (Milind; related: [#776](https://github.com/ProjectASAP/ASAPQuery/pull/776), [#725](https://github.com/ProjectASAP/ASAPQuery/pull/725)) | In progress. The evaluation does not wait for it (§8, PR 5). | No AutoSketch implementation exists in sketch-bench or in this repository. @@ -310,7 +310,7 @@ Figures: | --- | --- | --- | --- | | this | ASAPQuery | This plan (`docs/evaluation/autosketch-vs-planner.md`) | Zeying | | 1 | sketch-bench | `rqe-optimizer`: retained-memory term in `objectives.rs`; `milp::minimize_cost` with per-family price (§4); committed EC2 pricing JSON; dominance pruning also compares retained memory, so it cannot drop a candidate that is cheaper under the new objective. Tests: brute-force agreement on the tiny workload, as `#129` already does for CPU. | Zeying, coordinated with Milind since he is porting `milp.rs` | -| 2 | sketch-bench | Evaluation table (§6, "Benchmark input"): for the wider config grid at the worst-case parameters of each dataset, export each point's saturated error, `N_sat`, its saturation curve, and costs at `N_sat`, from #130's `study_saturation.py` outputs. For KLL and top-k, add the merged-curve values per `m` from #131. Keep the worst accuracy across inputs. Also keep each curve's value at every checkpoint, which AutoSketch's lookup at `n(S, ℓ)` needs. Record benchmark wall time per point (needed for §7). Lookups follow §6: ASAP uses saturated values with `m = S/x`; AutoSketch uses the curve value at `n(S, ℓ)`. Committed table. Depends on #131 for KLL/top-k. | Zeying | +| 2 | sketch-bench | Evaluation table (§6, "Benchmark input"): for the wider config grid at the worst-case parameters of each dataset, export each point's saturated error, `N_sat`, its saturation curve, and costs at `N_sat`, from #130's `study_saturation.py` outputs. For KLL and top-k, add the merged-curve values per `m` from #131. Keep the worst accuracy across inputs. Also keep each curve's value at every checkpoint, which AutoSketch's lookup at `n(S, ℓ)` needs. Record benchmark wall time per point (needed for §7). Lookups follow §6: ASAP uses saturated values with `m = S/x`; AutoSketch uses the curve value at `n(S, ℓ)`. Committed table. | Zeying | | 3 | sketch-bench | `rqe-optimizer/src/autosketch.rs`: Algorithm 4 ported from ASAPQuery-backend `autosketch_comparison.rs`, generalized from the CMS width/depth grid to each variant's measured parameter axes; one dedicated `Deployment` per RQE. Tests: picks the smallest feasible config on a grid; never shares; its window adapter output is eligible under `candidates::is_eligible`. | Zeying | | 4 | sketch-bench | `rqe-optimizer/examples/autosketch_vs_asap.rs` (`example`/`scaling` generators, all three methods, JSON output) and `scripts/plot_autosketch_vs_asap.py`; committed results and figures | Zeying | | 5 | ASAPQuery | After the MILP lands in `asap-planner-rs`: port PR 1's objective there and rerun PR 4 against it, so the paper reports the planner that ships | Zeying + Milind | From 9fd541563ad8006e88e679285a3847e24e1f7087 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Mon, 5 Oct 2026 10:43:54 -0400 Subject: [PATCH 12/39] Update README.md --- docs/README.md | 4 ---- 1 file changed, 4 deletions(-) diff --git a/docs/README.md b/docs/README.md index 14f35ee0..64fa2103 100644 --- a/docs/README.md +++ b/docs/README.md @@ -43,10 +43,6 @@ Task-oriented guides for common operations: - [Deploy to CloudLab](03-how-to-guides/operations/deploy-cloudlab.md) - Deployment guide - [Troubleshooting](03-how-to-guides/operations/troubleshooting.md) - Common issues & solutions -## Evaluation plans - -- [AutoSketch vs. the ASAPQuery planner](evaluation/autosketch-vs-planner.md) - Paper §6.3 plan: methods, cost model, PRs - ## 04. Development Developer practices and infrastructure: From 5e4431af66879444830f6bcac9a1ea853891e318 Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 14:45:30 +0000 Subject: [PATCH 13/39] docs(planner): add the two cost models, absolute SLA and current PR status Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 117 ++++++++++++++++++----- 1 file changed, 94 insertions(+), 23 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 3ca7f9f3..e19876fb 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -101,9 +101,11 @@ Mem = Σ_active D Mem_D In the MILP the `max` is linear: `Mem_D ≥ coef(r, D) · z_{r,D}` for each eligible `r`. -**Price per machine family** — for family `f` with `vCPU_f`, `GiB_f` and +**Price per machine family (steady state, implemented in sketch-bench #137)** — for family `f` with `vCPU_f`, `GiB_f` and on-demand `price_f` ($/hour), the plan needs a fractional instance count -`n_f ≥ CPU / vCPU_f` and `n_f ≥ Mem / GiB_f`; cost is `price_f · n_f`. This +`n_f ≥ CPU / vCPU_f` and `n_f ≥ Mem / GiB_f`; cost is `price_f · n_f`. Here +`CPU` is the steady-state average. The two cost models below replace it in the +reported results. This adds one continuous variable and two constraints, and needs no arbitrary split of an instance's price between CPU and memory. Families: @@ -122,13 +124,65 @@ not change the result. ASAP is solved once per family. AutoSketch-Adapted's plan does not depend on the family; it is scored under each family's cost. +### Two cost models per experiment run + +The average CPU hides that query load is bursty: ingest is continuous, but +query and merge work arrives at each evaluation. Every run is therefore priced +two ways, from one simulated CPU timeline. + +**CPU timeline.** +- Simulate 24 hours in 1-second bins. +- Every RQE first evaluates at `t = 0`, then every `T_r`. This aligned start + is the worst case. +- Each evaluation occupies one core for its estimated latency, starting when + it fires. Work longer than 1 s spills into later bins. +- `CPU(bin)` = ingest rate + busy-core time overlapping the bin. + +From the timeline: +- **total CPU-seconds** = the area under the curve, `Σ_bins CPU(bin) · 1 s`; +- **peak CPU** = `max_bin CPU(bin)`, in vCPU. + +**Model A — usage-based (pay for what is used).** + +```text +$/hour = a · (total CPU-seconds / 24 h) + b · Mem_GiB +``` + +`a` ($/vCPU-hour) and `b` ($/GiB-hour) are a least-squares fit of +`vCPU · a + GiB · b = price` over c7i.xlarge, m7i.xlarge and r7i.xlarge: about +a = 0.0368 and b = 0.00364 with the 2026-10-04 prices. Model A does not depend +on the machine family. + +**Model B — peak-provisioned (buy machines for the peak).** Per family `f`: + +```text +n_f = max(peak CPU / vCPU_f, Mem_GiB / GiB_f) +$/hour = n_f · price_f +``` + +**Optimizing each model.** ASAP is solved separately for model A and for each +family's model B. PerQuery-CostAware and FewestPlans (its cost stage) also use +the model being compared. AutoSketch-Adapted's plan does not depend on cost +and is scored under both. +- Model A is linear: the average-CPU and memory terms weighted by `a` and `b`. +- Model B is linear through the aligned start: the peak is in bin 0, so + `peak ≈ ingest + Σ_r min(latency_{r,D}, 1 s) · z_{r,D}`, where each + (RQE, deployment) latency is a constant. +- After solving, the exact peak is recomputed from the timeline. Runs where it + exceeds the bin-0 value, e.g. an evaluation longer than its interval + overlapping itself, are reported. + +Memory is kept in both models: without it, memory-bound workloads would look +almost free under model A. + **Latency** — per-RQE estimate already in sketch-bench: `card(ℓ) · (query_cpu + (S/x − 1) · merge_cpu)`. ## 5. Constraints -**Accuracy target 95%**, mapped per capability to the metrics the cost table -already records: +**Accuracy target**, swept over {90%, 95%, 99%} in the synthetic workload +(95% elsewhere). A target `p` maps to error ≤ `1 − p` and to precision ≥ `p`. +The 95% case, per capability, using the metrics the cost table already records: | Capability | Metric | Constraint | | --- | --- | --- | @@ -137,11 +191,19 @@ already records: | Cardinality | relative error | ≤ 0.05 | | TopK | precision@k | ≥ 0.95 | -**Latency** — per-RQE limit `L_r = α · min latency over r's eligible -deployments`, swept over `α ∈ {1.5, 2, 5, ∞}`. Using a multiple of the -fastest option keeps every point feasible for ASAP and makes the bound bind. -AutoSketch-Adapted ignores it; its violations are counted and reported, and its -cost is shown for those points but marked as infeasible. +**Latency** — one absolute SLA applies to every RQE, swept over +{0.01, 0.1, 1, 10, 100, 1000} ms and no limit. The synthetic workload extends +the grid as needed. +- An RQE that no method can meet at a given SLA is excluded from every method + at that SLA and reported by ID. Costs at different SLAs therefore cover + different RQE sets; compare methods only at one SLA. +- ASAP, PerQuery-CostAware and FewestPlans must meet the SLA. +- AutoSketch-Adapted ignores it. Its violations are counted, and its cost is + shown for those points but marked as infeasible. + +An earlier version set `L_r = α × the fastest latency of r`. It was dropped: +on `example`, α = 2 forced plans with no merging at 40× the unconstrained +cost. ## 6. Workloads @@ -271,8 +333,8 @@ inputs (§9 Q2). ## 7. Metrics and figures -Reported per (workload, method, machine family, α), median of 10 runs for -timings: +Reported per (workload, method, cost model and machine family, latency SLA), +median of repeated runs for timings: - **Planning time.** Reported in two parts, because the two planners spend their time differently: @@ -288,8 +350,10 @@ timings: pass over the grid, shared by all RQEs and reusable across workloads. It is reported once, next to how many RQEs it served. - The paper's figure shows search + benchmark per method, stacked. -- **Total cost** ($/hour) and its breakdown: ingest CPU, query + merge CPU, - memory, and which resource binds `n_f`. +- **Total cost** ($/hour) under **model A** and under **model B** for each + family, with its inputs: total CPU-seconds, peak CPU, retained GiB, and + which resource binds `n_f` in model B. Baselines are compared under each + model separately, each normalized to ASAP under the same model. - **Estimated query latency and latency SLA violations** per method. - Estimated latency per RQE: `card(ℓ) · (query_cpu + (S/x − 1) · merge_cpu)` (§4). Report its maximum and median over the RQEs, plus the per-RQE values in the raw output. - SLA violations: the number of RQEs whose estimated latency exceeds the SLA. Only AutoSketch-Adapted can have any, since the other methods are constrained. @@ -303,20 +367,27 @@ Figures: 2. Planning time vs. N, log–log (`scaling`). 3. Cost vs. shareability (`scaling`). 4. Cost vs. absolute latency SLA (`example`, `scaling`). +5. Baselines under the two cost models: paired bars per workload, model A + next to model B, each normalized to ASAP. +6. Synthetic workload: cost vs. achieved max estimated latency, one panel per + cost model (main paper figure). ## 8. Who implements what, in which PR -| # | Repo | Change | Owner | +| # | Repo / PR | Scope | Status (2026-10-05) | | --- | --- | --- | --- | -| this | ASAPQuery | This plan (`docs/evaluation/autosketch-vs-planner.md`) | Zeying | -| 1 | sketch-bench | `rqe-optimizer`: retained-memory term in `objectives.rs`; `milp::minimize_cost` with per-family price (§4); committed EC2 pricing JSON; dominance pruning also compares retained memory, so it cannot drop a candidate that is cheaper under the new objective. Tests: brute-force agreement on the tiny workload, as `#129` already does for CPU. | Zeying, coordinated with Milind since he is porting `milp.rs` | -| 2 | sketch-bench | Evaluation table (§6, "Benchmark input"): for the wider config grid at the worst-case parameters of each dataset, export each point's saturated error, `N_sat`, its saturation curve, and costs at `N_sat`, from #130's `study_saturation.py` outputs. For KLL and top-k, add the merged-curve values per `m` from #131. Keep the worst accuracy across inputs. Also keep each curve's value at every checkpoint, which AutoSketch's lookup at `n(S, ℓ)` needs. Record benchmark wall time per point (needed for §7). Lookups follow §6: ASAP uses saturated values with `m = S/x`; AutoSketch uses the curve value at `n(S, ℓ)`. Committed table. | Zeying | -| 3 | sketch-bench | `rqe-optimizer/src/autosketch.rs`: Algorithm 4 ported from ASAPQuery-backend `autosketch_comparison.rs`, generalized from the CMS width/depth grid to each variant's measured parameter axes; one dedicated `Deployment` per RQE. Tests: picks the smallest feasible config on a grid; never shares; its window adapter output is eligible under `candidates::is_eligible`. | Zeying | -| 4 | sketch-bench | `rqe-optimizer/examples/autosketch_vs_asap.rs` (`example`/`scaling` generators, all three methods, JSON output) and `scripts/plot_autosketch_vs_asap.py`; committed results and figures | Zeying | -| 5 | ASAPQuery | After the MILP lands in `asap-planner-rs`: port PR 1's objective there and rerun PR 4 against it, so the paper reports the planner that ships | Zeying + Milind | - -PRs 1 and 2 are independent; 3 depends on 2 for a meaningful grid only; 4 -depends on 1–3. +| this | ASAPQuery #777 | This plan | Draft, updated as decisions change | +| — | sketch-bench #130, #131 | Saturation curves at K ∈ {1e3, 1e5, 1e7}; accuracy after merging `m` shards | Merged | +| 1 | sketch-bench #137 | Retained memory, EC2 pricing, `milp::minimize_cost` (steady-state model), solver scaling | Merged | +| 3 | sketch-bench #135 | AutoSketch-Adapted (Algorithm 4), aligned with the paper's EXAMINE rule and seeding | Merged | +| 2 | sketch-bench #136 | Evaluation table for the trace workloads: per (RQE, config) accuracy for AutoSketch and for ASAP at each `m`, saturation, costs | In review | +| 4 | sketch-bench #138 | Runner, absolute SLA, results for `example`, `scaling`, `traces` | Open; needs a rebase on main and an AutoSketch rerun with #135's final search | +| — | sketch-bench #140 | Saturation curves at K ∈ {1e1, 1e2, 1e4, 1e6} for the synthetic workload | Draft; accuracy done, cost 197 of 240 points | +| — | sketch-bench #139 | Synthetic workload (67 RQEs), FewestPlans, strawmen bound by the SLA | Open; final sweep pending #140 | +| — | sketch-bench, not yet opened | The two cost models (§4): CPU timeline, model A, model B, rerun of every experiment | Not started as a PR | +| 5 | ASAPQuery | After the MILP lands in `asap-planner-rs`: port the objective there and rerun, so the paper reports the planner that ships | Not started; waits for Milind's port | + +Merge order: #136 → rebase and merge #138 → #140 → #139 → two-cost-model PR. ## 9. Decisions From d769d5d6bfb61c3e67947ca17c53f94428abd5e7 Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 14:46:58 +0000 Subject: [PATCH 14/39] docs(planner): #136 is merged Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index e19876fb..e4cc2b6b 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -380,14 +380,14 @@ Figures: | — | sketch-bench #130, #131 | Saturation curves at K ∈ {1e3, 1e5, 1e7}; accuracy after merging `m` shards | Merged | | 1 | sketch-bench #137 | Retained memory, EC2 pricing, `milp::minimize_cost` (steady-state model), solver scaling | Merged | | 3 | sketch-bench #135 | AutoSketch-Adapted (Algorithm 4), aligned with the paper's EXAMINE rule and seeding | Merged | -| 2 | sketch-bench #136 | Evaluation table for the trace workloads: per (RQE, config) accuracy for AutoSketch and for ASAP at each `m`, saturation, costs | In review | +| 2 | sketch-bench #136 | Evaluation table for the trace workloads: per (RQE, config) accuracy for AutoSketch and for ASAP at each `m`, saturation, costs | Merged | | 4 | sketch-bench #138 | Runner, absolute SLA, results for `example`, `scaling`, `traces` | Open; needs a rebase on main and an AutoSketch rerun with #135's final search | | — | sketch-bench #140 | Saturation curves at K ∈ {1e1, 1e2, 1e4, 1e6} for the synthetic workload | Draft; accuracy done, cost 197 of 240 points | | — | sketch-bench #139 | Synthetic workload (67 RQEs), FewestPlans, strawmen bound by the SLA | Open; final sweep pending #140 | | — | sketch-bench, not yet opened | The two cost models (§4): CPU timeline, model A, model B, rerun of every experiment | Not started as a PR | | 5 | ASAPQuery | After the MILP lands in `asap-planner-rs`: port the objective there and rerun, so the paper reports the planner that ships | Not started; waits for Milind's port | -Merge order: #136 → rebase and merge #138 → #140 → #139 → two-cost-model PR. +Merge order: rebase and merge #138 → #140 → #139 → two-cost-model PR. ## 9. Decisions From 71f3785ee2d602409a0f715c615184ea3a27da00 Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 14:50:38 +0000 Subject: [PATCH 15/39] docs(planner): define the synthetic query templates and workload grid Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 124 ++++++++++++++++------- 1 file changed, 86 insertions(+), 38 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index e4cc2b6b..3929cdab 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -216,45 +216,93 @@ cost. ### Synthetic workload -**Data.** -- Series carry `label_0`, with cardinality in {10^1, …, 10^6}, and an - `instance` label with 100 values per `label_0` value. -- Each series is scraped every 10 ms, i.e. 100 samples per second, so - `λ = 100 · 100 · card(label_0)` samples/s. -- The data volume is chosen so that even the smallest windows hold enough - samples for a sketch (1e4 per group for 1 s spatial queries, 6e3 per series - for 1 m temporal ones). Points that still fall below `N_sat` are flagged. -- Key weights for frequency and top-k follow Zipf θ ∈ {0, 0.5, 1.0, 1.5, 2.0}. - Values for quantiles follow Pareto a ∈ {1.1, 2, 3}. -- The saturation curves are extended to K ∈ {1e1, 1e2, 1e4, 1e6} by - measuring, not by interpolation. - -**Queries.** Each spatial query repeats every 1 s, with `S = T = 1 s`, so each -evaluation reads the last second. Each temporal query repeats every 1 m, with -`S = T_range ∈ {1m, 10m, 1h, 6h, 24h}`. - -| # | Query | Capability | Grouping | -|---|---|---|---| -| 1 | `sum by (label_0) (data)` | Freq | `label_0` | -| 2 | `topk by (3, label_0) (data)` | TopK | `label_0` | -| 3 | `quantile by (q, label_0) (data)`, q ∈ {.5, .75, .9, .95, .99} | Quantile | per `label_0` group | -| 4 | `sum_over_time(data[T])` | Freq | per series | -| 5 | `quantile_over_time(q, data[T])`, same five q | Quantile | per series | -| 6 | `rate(data[T])` | Freq over per-series increments | per series | -| 7 | `sum by (label_0) (rate(data[T]))` | Freq over increments | `label_0` | -| 8 | `sum by (label_0) (sum_over_time(data[T]))` | Freq | `label_0` | -| 9 | `topk by (3, label_0) (rate(data[T]))` | TopK over increments | `label_0` | -| 10 | `quantile_over_time(0.9, data[T]) / quantile_over_time(0.5, data[T])` | Two Quantile RQEs | per series | - -Notes on the mapping: +The synthetic workload is the paper's main experiment. It builds many +workloads from a fixed set of PromQL query templates by sweeping the workload +dimensions below, and runs every planning baseline on each workload. + +#### Query templates + +These exercise every summary type the planner supports: frequency, top-k and +quantile, spatial and temporal aggregation, and a binary operator. Each spatial +template repeats every 1 s and reads the last second (`S = T = 1 s`). Each +temporal template repeats every `T` (default 1 m) and reads `S = T_range`, one +RQE per lookback window in the window set `W` (§ workload grid). + +| # | PromQL | Capability | Grouping | RQEs per replica | +|---|---|---|---|---| +| 1 | `sum by (label_0) (data)` | Freq | `label_0` | 1 | +| 2 | `topk by (3, label_0) (data)` | TopK | `label_0` | 1 | +| 3 | `quantile by (q, label_0) (data)`, q ∈ {0.5, 0.75, 0.9, 0.95, 0.99} | Quantile | per `label_0` group | 5 | +| 4 | `sum_over_time(data[T])` | Freq | per series | \|W\| | +| 5 | `quantile_over_time(q, data[T])`, same five q | Quantile | per series | 5·\|W\| | +| 6 | `rate(data[T])` | Freq over per-series increments | per series | \|W\| | +| 7 | `sum by (label_0) (rate(data[T]))` | Freq over increments | `label_0` | \|W\| | +| 8 | `sum by (label_0) (sum_over_time(data[T]))` | Freq | `label_0` | \|W\| | +| 9 | `topk by (3, label_0) (rate(data[T]))` | TopK over increments | `label_0` | \|W\| | +| 10 | `quantile_over_time(0.9, data[T]) / quantile_over_time(0.5, data[T])` | Two Quantile RQEs | per series | 2·\|W\| | + +With all ten templates, one replica has `7 + 12·|W|` RQEs: 67 for the default +five windows. + +Mapping notes: - `rate`/`increase` are modeled as a frequency sum of per-series increments, equivalent to `sum_over_time` over deltas. -- The quantiles of one query, and the two operands of query 10, read the same - stream. ASAP can serve them from one deployment; AutoSketch gets one per RQE. -- One workload instance is the 67 RQEs above (2 spatial + 5 spatial quantiles + 5×5 temporal for queries 4, 6–9 + 5×5 for query 5 + 2×5 for query 10), for one (cardinality, θ or a, - accuracy target, SLA) combination. The figure sweeps the accuracy target over - {90%, 95%, 99%} and an absolute latency SLA grid. For the scalability study, - the RQE set is replicated with distinct `label_0` filters. +- The quantiles of one template, and the two operands of template 10, read the + same stream. ASAP can serve them from one deployment; AutoSketch gets one + deployment per RQE. +- Templates come from the planner's supported query classes: `SpatialAgg` + (`count`/`sum`/`quantile`/`topk by`), `TemporalAgg` + (`count_over_time`/`sum_over_time`/`quantile_over_time`/`increase`/`rate`), + `TemporalAgg SpatialAgg*`, and `AnyAgg AnyAgg`. + +#### Data + +- Series carry `label_0`, plus an `instance` label: `s` series per `label_0` + value. `card(label_0)` and `s` are workload dimensions below. +- Every series is scraped every 10 ms (100 samples/s), so the total rate is + `λ = 100 · s · card(label_0)` samples/s. The volume is chosen so that even + the smallest windows hold enough samples for a sketch; points still below + `N_sat` are flagged. +- Key weights for frequency and top-k follow Zipf θ; quantile values follow + Pareto a. +- Saturation curves cover K ∈ {1e1, …, 1e7} (#130, #140), all measured, never + interpolated. A per-series query has K = `s · card(label_0)` keys. Points with + K above 1e7 are clamped to the largest measured K and flagged as + extrapolated. + +#### Workload grid + +Each dimension has a default (bold). A workload fixes every dimension. + +| Dimension | Values | What it varies | +|---|---|---| +| Query mix (templates) | **all 10**; spatial only {1, 2, 3}; temporal only {4–10}; frequency only {1, 4, 6, 7, 8}; quantile only {3, 5, 10}; top-k only {2, 9} | Summary types, and how much can be shared | +| Number of RQEs: replicas `r` | **1**, 2, 4, 8, 16, 32, 64 | Each replica adds a filter `{label_1="v_i"}` selecting a disjoint subset of series, so it reads its own streams. Total RQEs = `r · (n_spatial + n_temporal·|W|)` | +| Lookback window set `W` | {1h}; {1m, 1h}; {1m, 10m, 1h}; **{1m, 10m, 1h, 6h, 24h}** | Overlapping windows over the same stream: the main sharing opportunity | +| Temporal repeat interval `T` | 10 s, **1 m**, 5 m | Recurrence: query and merge work vs. ingest | +| Group cardinality `card(label_0)` | 1e1, 1e2, **1e3**, 1e4, 1e5, 1e6 | Sketch instances per deployment, keys per sketch | +| Series per group `s` (aggregated series cardinality) | 1, 10, **100**, 1000 | Series aggregated per group: events per group for spatial queries, keys for per-series queries | +| Key skew θ / value tail a | θ ∈ {0, 0.5, **1.0**, 1.5, 2.0}; a ∈ {1.1, **2**, 3} | Sketch size needed for the accuracy target | +| Accuracy target | 90%, **95%**, 99% | §5 | +| Latency SLA | the §5 grid, **no limit** | §5 | + +The full Cartesian product is too large. The sweep is: +1. **Default workload.** Every baseline, every SLA, every cost model. +2. **One dimension at a time.** Vary each dimension over its values with the + others at their defaults. +3. **Two interactions:** + - `r × s`: scale, with planning time against total RQEs; + - `W × card(label_0)`: sharing benefit against state size. + +Every workload runs every baseline (ASAP, AutoSketch-Adapted, +PerQuery-CostAware, FewestPlans) and is priced under model A and model B for +each family (§4). Per (workload, baseline, cost model, SLA), report: +- $/hour, total CPU-seconds, peak CPU, retained GiB; +- max and median estimated latency, and SLA violations; +- active deployments and sketch instances; +- planning time (AutoSketch: search plus charged benchmark time); +- RQEs excluded by the SLA or unservable, and counts of unsaturated or + extrapolated lookups. ### Benchmark input @@ -383,7 +431,7 @@ Figures: | 2 | sketch-bench #136 | Evaluation table for the trace workloads: per (RQE, config) accuracy for AutoSketch and for ASAP at each `m`, saturation, costs | Merged | | 4 | sketch-bench #138 | Runner, absolute SLA, results for `example`, `scaling`, `traces` | Open; needs a rebase on main and an AutoSketch rerun with #135's final search | | — | sketch-bench #140 | Saturation curves at K ∈ {1e1, 1e2, 1e4, 1e6} for the synthetic workload | Draft; accuracy done, cost 197 of 240 points | -| — | sketch-bench #139 | Synthetic workload (67 RQEs), FewestPlans, strawmen bound by the SLA | Open; final sweep pending #140 | +| — | sketch-bench #139 | Synthetic workload: the 10 templates, FewestPlans, strawmen bound by the SLA, and the workload-grid driver (dimensions in §6 "Workload grid": query mix, replicas, window set, repeat interval, `card(label_0)`, series per group, θ/a, accuracy target, SLA), with the sweep script and figures | Open; code for the fixed 67-RQE workload exists. Still to do: the grid driver, then the sweep (after #140 and the two-cost-model PR) | | — | sketch-bench, not yet opened | The two cost models (§4): CPU timeline, model A, model B, rerun of every experiment | Not started as a PR | | 5 | ASAPQuery | After the MILP lands in `asap-planner-rs`: port the objective there and rerun, so the paper reports the planner that ships | Not started; waits for Milind's port | From 76929c41647d4aa1abe7cf934ded8c02819391af Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 15:21:49 +0000 Subject: [PATCH 16/39] docs(planner): spell out the synthetic data model and label cardinalities Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 87 +++++++++++++++++++----- 1 file changed, 71 insertions(+), 16 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 3929cdab..331e253f 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -255,20 +255,75 @@ Mapping notes: (`count_over_time`/`sum_over_time`/`quantile_over_time`/`increase`/`rate`), `TemporalAgg SpatialAgg*`, and `AnyAgg AnyAgg`. -#### Data - -- Series carry `label_0`, plus an `instance` label: `s` series per `label_0` - value. `card(label_0)` and `s` are workload dimensions below. -- Every series is scraped every 10 ms (100 samples/s), so the total rate is - `λ = 100 · s · card(label_0)` samples/s. The volume is chosen so that even - the smallest windows hold enough samples for a sketch; points still below - `N_sat` are flagged. -- Key weights for frequency and top-k follow Zipf θ; quantile values follow - Pareto a. -- Saturation curves cover K ∈ {1e1, …, 1e7} (#130, #140), all measured, never - interpolated. A per-series query has K = `s · card(label_0)` keys. Points with - K above 1e7 are clamped to the largest measured K and flagged as - extrapolated. +#### Data model + +There is one metric, `data`. A **series** is one combination of label values. +Three labels matter: + +| Label | Values | Role | +|---|---|---| +| `label_0` | `C = card(label_0)` values | The grouping label: what `by (label_0)` aggregates by | +| `instance` | `s` values per `label_0` value | Distinguishes the series inside a group. `s` = series per group | +| `label_1` | `r` values, one per replica | Only for replicas (workload grid): replica `i` filters `{label_1="v_i"}` and reads its own disjoint series. With `r = 1` it is absent | + +So a replica has `C · s` series, and the workload has `r · C · s`. Every series +emits one sample every 10 ms (100 samples/s), so one replica's stream carries +`λ = 100 · C · s` samples/s. + +Sample values: +- **Frequency and top-k:** the value is the weight being summed. The total + weight of the keys follows Zipf θ. The key is the `label_0` value for + `by (label_0)` templates and the series for per-series templates. +- **Quantiles:** values are drawn from Pareto a. + +**Only two cardinalities matter.** Queries aggregate by `label_0` or per +series, never by another label. So any other label (a second instance-like +label, a `label_2`, ...) only multiplies the number of series in each group, +and is equivalent to a larger `s`. The model therefore has two independent +cardinality knobs: +- `C`: groups; +- `s`: series per group, the product of the cardinalities of all non-grouping + labels. + +`label_1`'s cardinality equals `r` and is covered by the replica dimension. +Giving every label the same cardinality `c` would tie the knobs together +(`C = c`, `s = c^(L−1)` for `L` labels). It was rejected: `s` explodes (c = 1e3 +with three labels gives 1e9 series, far beyond the measured K ≤ 1e7), and the +effects of more groups and of more series per group could no longer be told +apart. No template groups by `label_1`, keeping the template set as given. + +**How a template becomes sketch input.** A frequency or top-k sketch holds the +groups as keys inside one sketch. A quantile sketch is one sketch per group. + +| Template kind | Sketch instances per deployment | Keys per sketch | Events per sketch per window | +|---|---|---|---| +| `sum`/`topk by (label_0)` (1, 2, 7, 8, 9) | 1 | `C` | `100 · C · s · S` | +| `quantile by (q, label_0)` (3) | `C` | — | `100 · s · S` | +| per-series `sum_over_time`/`rate` (4, 6) | 1 | `C · s` | `100 · C · s · S` | +| per-series `quantile_over_time` (5, 10) | `C · s` | — | `100 · S` | + +**Example.** `C = 3` (`label_0` ∈ {a, b, c}), `s = 2` (`instance` ∈ {i1, i2}), +`r = 1`: six series, `data{label_0="a", instance="i1"}` through +`data{label_0="c", instance="i2"}`, emitting 600 samples/s in total. +- `sum by (label_0) (data)`, `S = 1 s`: one CMS with keys a, b, c; each second + it absorbs 600 weighted updates, 200 per key. +- `quantile by (0.99, label_0) (data)`: three KLL sketches, one per group, each + absorbing 200 values per second. +- `sum_over_time(data[1m])`: one CMS with six keys, one per series, absorbing + 36,000 updates per minute. +- `quantile_over_time(0.99, data[1m])`: six KLL sketches, 6,000 values each per + minute. + +**Modeling choice for spatial templates.** A spatial template evaluates every +1 s over every sample of the last second (`S = T = 1 s`). PromQL's instant +semantics would read only each series' latest sample. Aggregating the whole +second is what a sketch maintained over a 1-second window answers. The +difference is noted wherever spatial results are reported. + +Saturation curves cover K ∈ {1e1, …, 1e7} (#130, #140), all measured and +never interpolated. Per-series templates have K = `C · s`; points above 1e7 are +clamped to the largest measured K and flagged as extrapolated. Points whose +events per sketch fall below `N_sat` are flagged as unsaturated. #### Workload grid @@ -280,8 +335,8 @@ Each dimension has a default (bold). A workload fixes every dimension. | Number of RQEs: replicas `r` | **1**, 2, 4, 8, 16, 32, 64 | Each replica adds a filter `{label_1="v_i"}` selecting a disjoint subset of series, so it reads its own streams. Total RQEs = `r · (n_spatial + n_temporal·|W|)` | | Lookback window set `W` | {1h}; {1m, 1h}; {1m, 10m, 1h}; **{1m, 10m, 1h, 6h, 24h}** | Overlapping windows over the same stream: the main sharing opportunity | | Temporal repeat interval `T` | 10 s, **1 m**, 5 m | Recurrence: query and merge work vs. ingest | -| Group cardinality `card(label_0)` | 1e1, 1e2, **1e3**, 1e4, 1e5, 1e6 | Sketch instances per deployment, keys per sketch | -| Series per group `s` (aggregated series cardinality) | 1, 10, **100**, 1000 | Series aggregated per group: events per group for spatial queries, keys for per-series queries | +| Groups `C = card(label_0)` | 1e1, 1e2, **1e3**, 1e4, 1e5, 1e6 | Keys per frequency sketch; quantile sketches per `by (label_0)` deployment | +| Series per group `s` (product of the non-grouping label cardinalities) | 1, 10, **100**, 1000 | Events per group for spatial templates; keys and sketch instances for per-series templates | | Key skew θ / value tail a | θ ∈ {0, 0.5, **1.0**, 1.5, 2.0}; a ∈ {1.1, **2**, 3} | Sketch size needed for the accuracy target | | Accuracy target | 90%, **95%**, 99% | §5 | | Latency SLA | the §5 grid, **no limit** | §5 | From 7964b5bba1bd0486e8e6c91ade0ff2fd3978b9c7 Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 15:25:16 +0000 Subject: [PATCH 17/39] docs(planner): drop the FewestPlans baseline Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 17 ++++++++--------- 1 file changed, 8 insertions(+), 9 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 331e253f..9169ef8d 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -64,14 +64,13 @@ function (§4). and window choices as ASAP but no batching or sharing. ASAP vs. this ablation isolates the batch/sharing benefit; this ablation vs. AutoSketch-Adapted isolates the objective/window benefit. -4. **FewestPlans** (strawman) — the ASAP MILP minimizing the number of active - deployments first, then cost among plans with that minimum count, under - the same accuracy and latency requirements. - -Only AutoSketch-Adapted ignores latency; the two strawmen must meet the same +Only AutoSketch-Adapted ignores latency; PerQuery-CostAware must meet the same requirements as ASAP, so every point in the cost–latency figure except AutoSketch's is a feasible plan. +A "fewest plans" strawman (minimize the number of deployments, then cost) was +considered and dropped (decided 2026-10-05). + AutoSketch-Adapted is a planner baseline, not a reproduction of the P4 compiler: stage/page/ALU constraints are dropped. Its accuracy probes read sketch-bench measurements instead of running a benchmark inside the search, but @@ -161,7 +160,7 @@ $/hour = n_f · price_f ``` **Optimizing each model.** ASAP is solved separately for model A and for each -family's model B. PerQuery-CostAware and FewestPlans (its cost stage) also use +family's model B. PerQuery-CostAware also uses the model being compared. AutoSketch-Adapted's plan does not depend on cost and is scored under both. - Model A is linear: the average-CPU and memory terms weighted by `a` and `b`. @@ -197,7 +196,7 @@ the grid as needed. - An RQE that no method can meet at a given SLA is excluded from every method at that SLA and reported by ID. Costs at different SLAs therefore cover different RQE sets; compare methods only at one SLA. -- ASAP, PerQuery-CostAware and FewestPlans must meet the SLA. +- ASAP and PerQuery-CostAware must meet the SLA. - AutoSketch-Adapted ignores it. Its violations are counted, and its cost is shown for those points but marked as infeasible. @@ -350,7 +349,7 @@ The full Cartesian product is too large. The sweep is: - `W × card(label_0)`: sharing benefit against state size. Every workload runs every baseline (ASAP, AutoSketch-Adapted, -PerQuery-CostAware, FewestPlans) and is priced under model A and model B for +PerQuery-CostAware) and is priced under model A and model B for each family (§4). Per (workload, baseline, cost model, SLA), report: - $/hour, total CPU-seconds, peak CPU, retained GiB; - max and median estimated latency, and SLA violations; @@ -486,7 +485,7 @@ Figures: | 2 | sketch-bench #136 | Evaluation table for the trace workloads: per (RQE, config) accuracy for AutoSketch and for ASAP at each `m`, saturation, costs | Merged | | 4 | sketch-bench #138 | Runner, absolute SLA, results for `example`, `scaling`, `traces` | Open; needs a rebase on main and an AutoSketch rerun with #135's final search | | — | sketch-bench #140 | Saturation curves at K ∈ {1e1, 1e2, 1e4, 1e6} for the synthetic workload | Draft; accuracy done, cost 197 of 240 points | -| — | sketch-bench #139 | Synthetic workload: the 10 templates, FewestPlans, strawmen bound by the SLA, and the workload-grid driver (dimensions in §6 "Workload grid": query mix, replicas, window set, repeat interval, `card(label_0)`, series per group, θ/a, accuracy target, SLA), with the sweep script and figures | Open; code for the fixed 67-RQE workload exists. Still to do: the grid driver, then the sweep (after #140 and the two-cost-model PR) | +| — | sketch-bench #139 | Synthetic workload: the 10 templates, PerQuery-CostAware bound by the SLA, and the workload-grid driver (dimensions in §6 "Workload grid": query mix, replicas, window set, repeat interval, `card(label_0)`, series per group, θ/a, accuracy target, SLA), with the sweep script and figures | Open; code for the fixed 67-RQE workload exists. Still to do: the grid driver, then the sweep (after #140 and the two-cost-model PR) | | — | sketch-bench, not yet opened | The two cost models (§4): CPU timeline, model A, model B, rerun of every experiment | Not started as a PR | | 5 | ASAPQuery | After the MILP lands in `asap-planner-rs`: port the objective there and rerun, so the paper reports the planner that ships | Not started; waits for Milind's port | From 7d8a17efab69a7749d02bd6e7498e2beb686939e Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 15:27:43 +0000 Subject: [PATCH 18/39] docs(planner): keep only the traces and synthetic workloads Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 28 +++++++++++++----------- 1 file changed, 15 insertions(+), 13 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 9169ef8d..bd3f7c9d 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -201,16 +201,19 @@ the grid as needed. shown for those points but marked as infeasible. An earlier version set `L_r = α × the fastest latency of r`. It was dropped: -on `example`, α = 2 forced plans with no merging at 40× the unconstrained +on the dropped `example` workload, α = 2 forced plans with no merging at 40× the unconstrained cost. ## 6. Workloads +Two workloads, decided 2026-10-05. The earlier `example` (8 RQEs from +`small_problem`) and `scaling` (random RQE batches) workloads were dropped. +Planning-time scaling is now the replica dimension of the synthetic workload +grid. + | ID | Description | Purpose | | --- | --- | --- | -| `example` | `small_problem`'s 8 RQEs (freq, quantile, cardinality, top-k; 1h–1d lookbacks; 60s/300s intervals), tolerances moved to §5 | Readable worked example; one table in the paper | | `synthetic` | Synthetic PromQL workload: the 10 queries in "Synthetic workload" below, over Zipf/Pareto data. Main figure. | Cost–latency trade-off across data and requirements | -| `scaling` | Seeded random batches, `N ∈ {8, 32, 128, 512, 2048}` RQEs. Lookbacks {5m, 15m, 1h, 6h, 1d}, intervals {10s, 60s, 300s}, 4 capabilities, 3 label sets with fixed cardinality/rate. Knob: fraction of RQEs drawn from shared (capability, labels) cohorts, {0, 0.5, 1}. | Planning-time scaling; cost vs. shareability | | `traces` | Real-trace RQEs, one workload per dataset: Alibaba 2022, BOOM and Google 2011. Taken from `asap-tools/dataset-analysis/results/skew_summary.csv` ([#746](https://github.com/ProjectASAP/ASAPQuery/pull/746)): each row's `range_s` is `S` and its `step_s` is `T`. Data parameters and accuracy targets are fit over each whole trace. | Appendix: real-trace results | ### Synthetic workload @@ -377,7 +380,7 @@ and top-k from the lower `θ` bound. Longer samples expose worse cases (sketch-b - **`traces` gives the appendix results.** Its parameters are fit on the full trace. -- **`example`, `scaling` and `synthetic`** use their generators' parameters. +- **`synthetic`** uses its generator's parameters (§6 "Data model"). Each config is benchmarked at these worst-case parameters. AutoSketch §5.2 injects random traffic bursts into synthetic workloads to cover variation over @@ -465,14 +468,13 @@ median of repeated runs for timings: Figures: -1. Total cost by method, grouped by machine family (`example` and `scaling` at N = 128). -2. Planning time vs. N, log–log (`scaling`). -3. Cost vs. shareability (`scaling`). -4. Cost vs. absolute latency SLA (`example`, `scaling`). +1. Synthetic workload, cost vs. achieved max estimated latency, one panel per + cost model (main paper figure). +2. Planning time vs. number of RQEs (synthetic, replica dimension), log–log. +3. Cost vs. each workload-grid dimension (synthetic, one dimension at a time). +4. Cost vs. absolute latency SLA (synthetic default workload, `traces`). 5. Baselines under the two cost models: paired bars per workload, model A next to model B, each normalized to ASAP. -6. Synthetic workload: cost vs. achieved max estimated latency, one panel per - cost model (main paper figure). ## 8. Who implements what, in which PR @@ -483,8 +485,8 @@ Figures: | 1 | sketch-bench #137 | Retained memory, EC2 pricing, `milp::minimize_cost` (steady-state model), solver scaling | Merged | | 3 | sketch-bench #135 | AutoSketch-Adapted (Algorithm 4), aligned with the paper's EXAMINE rule and seeding | Merged | | 2 | sketch-bench #136 | Evaluation table for the trace workloads: per (RQE, config) accuracy for AutoSketch and for ASAP at each `m`, saturation, costs | Merged | -| 4 | sketch-bench #138 | Runner, absolute SLA, results for `example`, `scaling`, `traces` | Open; needs a rebase on main and an AutoSketch rerun with #135's final search | -| — | sketch-bench #140 | Saturation curves at K ∈ {1e1, 1e2, 1e4, 1e6} for the synthetic workload | Draft; accuracy done, cost 197 of 240 points | +| 4 | sketch-bench #138 | Runner, absolute SLA, results for `traces` (the `example`/`scaling` workloads are to be removed in its rebase) | Open; needs a rebase on main and an AutoSketch rerun with #135's final search | +| — | sketch-bench #140 | Saturation curves (accuracy vs. events per sketch, N_sat, costs) at K ∈ {1e1, 1e2, 1e4, 1e6}: the synthetic workload needs these cardinalities and #130 measured only 1e3, 1e5, 1e7 | Draft; accuracy done, cost 197 of 240 points | | — | sketch-bench #139 | Synthetic workload: the 10 templates, PerQuery-CostAware bound by the SLA, and the workload-grid driver (dimensions in §6 "Workload grid": query mix, replicas, window set, repeat interval, `card(label_0)`, series per group, θ/a, accuracy target, SLA), with the sweep script and figures | Open; code for the fixed 67-RQE workload exists. Still to do: the grid driver, then the sweep (after #140 and the two-cost-model PR) | | — | sketch-bench, not yet opened | The two cost models (§4): CPU timeline, model A, model B, rerun of every experiment | Not started as a PR | | 5 | ASAPQuery | After the MILP lands in `asap-planner-rs`: port the objective there and rerun, so the paper reports the planner that ships | Not started; waits for Milind's port | @@ -529,7 +531,7 @@ sliding sketch per query. DDSketch, and taken from #131 for KLL and top-k. #131 covers `N ≤ 1e7` and `m ≤ 64`. A deployment needing `m > 64` (e.g. a 1-day lookback over 1-minute windows, `m = 1440`) is outside the measured range. Mark it as extrapolated, - or exclude it for KLL/top-k. Replay `example`'s chosen plans in sketch-bench once to + or exclude it for KLL/top-k. Replay the synthetic default workload's chosen plans in sketch-bench once to confirm the lookups. - Costs and latencies are estimates from per-operation measurements, not end-to-end executions. The execution-based comparison is ASAPQuery-backend From 13b1584076fe44757ae6c34037c14dbc5eec4daf Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 15:39:41 +0000 Subject: [PATCH 19/39] docs(planner): the evaluation is not rerun on asap-planner-rs Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index bd3f7c9d..06a93673 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -33,7 +33,7 @@ interval) and sums the results; the planner is invoked once for the batch. | Measured per-operation costs (`AtomicCostEntry`: memory/instance, insert/merge/query CPU, accuracy) | sketch-bench `scripts/export_rqe_optimizer_costs.sh` | Merged; 18 rows, 2 configs per sketch variant. | | Saturation study: error vs. `N`, `N_sat`, cost per config and shape | sketch-bench [#130](https://github.com/ProjectASAP/sketch-bench/pull/130) | Merged. Source of the saturated lookup (§6). | | Accuracy after merging `m` shards (KLL, top-k) | sketch-bench [#131](https://github.com/ProjectASAP/sketch-bench/pull/131) | Merged 2026-10-05 (`bd644fe`). Source of KLL/top-k lookups when `m > 1`. | -| Moving the MILP into ASAPQuery's planner | ASAPQuery `asap-planner-rs/src/optimizer/` (Milind; related: [#776](https://github.com/ProjectASAP/ASAPQuery/pull/776), [#725](https://github.com/ProjectASAP/ASAPQuery/pull/725)) | In progress. The evaluation does not wait for it (§8, PR 5). | +| Moving the MILP into ASAPQuery's planner | ASAPQuery `asap-planner-rs/src/optimizer/` (Milind; related: [#776](https://github.com/ProjectASAP/ASAPQuery/pull/776), [#725](https://github.com/ProjectASAP/ASAPQuery/pull/725)) | Out of scope: the evaluation uses sketch-bench `rqe-optimizer` and is not rerun on `asap-planner-rs`. | No AutoSketch implementation exists in sketch-bench or in this repository. @@ -489,7 +489,6 @@ Figures: | — | sketch-bench #140 | Saturation curves (accuracy vs. events per sketch, N_sat, costs) at K ∈ {1e1, 1e2, 1e4, 1e6}: the synthetic workload needs these cardinalities and #130 measured only 1e3, 1e5, 1e7 | Draft; accuracy done, cost 197 of 240 points | | — | sketch-bench #139 | Synthetic workload: the 10 templates, PerQuery-CostAware bound by the SLA, and the workload-grid driver (dimensions in §6 "Workload grid": query mix, replicas, window set, repeat interval, `card(label_0)`, series per group, θ/a, accuracy target, SLA), with the sweep script and figures | Open; code for the fixed 67-RQE workload exists. Still to do: the grid driver, then the sweep (after #140 and the two-cost-model PR) | | — | sketch-bench, not yet opened | The two cost models (§4): CPU timeline, model A, model B, rerun of every experiment | Not started as a PR | -| 5 | ASAPQuery | After the MILP lands in `asap-planner-rs`: port the objective there and rerun, so the paper reports the planner that ships | Not started; waits for Milind's port | Merge order: rebase and merge #138 → #140 → #139 → two-cost-model PR. From fc913c25898b7bf9761a67e1fe4c85758b0ef6e0 Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 16:07:35 +0000 Subject: [PATCH 20/39] docs(planner): model B uses a per-RQE occupancy bound; #141 Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 21 +++++++++++++-------- 1 file changed, 13 insertions(+), 8 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 06a93673..78fd6f93 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -164,12 +164,17 @@ family's model B. PerQuery-CostAware also uses the model being compared. AutoSketch-Adapted's plan does not depend on cost and is scored under both. - Model A is linear: the average-CPU and memory terms weighted by `a` and `b`. -- Model B is linear through the aligned start: the peak is in bin 0, so - `peak ≈ ingest + Σ_r min(latency_{r,D}, 1 s) · z_{r,D}`, where each - (RQE, deployment) latency is a constant. -- After solving, the exact peak is recomputed from the timeline. Runs where it - exceeds the bin-0 value, e.g. an evaluation longer than its interval - overlapping itself, are reported. +- Model B's peak is bounded linearly per RQE. In steady state, an RQE whose + evaluations take `latency` every `T` keeps `ceil(latency / T)` cores busy at + most, counting evaluations that outlast their interval and overlap + themselves. The MILP uses `peak ≤ ingest + Σ_r occupancy_{r,D} · z_{r,D}` + with these constant per-(RQE, deployment) occupancies. + - The bound is exact when no evaluation outlasts its interval. + - An earlier proxy that read only bin 0 undercounted self-overlapping + evaluations, about 40 per synthetic run, and caused sanity violations. + It was replaced (sketch-bench #141). +- After solving, the exact peak is recomputed from the timeline and compared + with the bound. Memory is kept in both models: without it, memory-bound workloads would look almost free under model A. @@ -488,9 +493,9 @@ Figures: | 4 | sketch-bench #138 | Runner, absolute SLA, results for `traces` (the `example`/`scaling` workloads are to be removed in its rebase) | Open; needs a rebase on main and an AutoSketch rerun with #135's final search | | — | sketch-bench #140 | Saturation curves (accuracy vs. events per sketch, N_sat, costs) at K ∈ {1e1, 1e2, 1e4, 1e6}: the synthetic workload needs these cardinalities and #130 measured only 1e3, 1e5, 1e7 | Draft; accuracy done, cost 197 of 240 points | | — | sketch-bench #139 | Synthetic workload: the 10 templates, PerQuery-CostAware bound by the SLA, and the workload-grid driver (dimensions in §6 "Workload grid": query mix, replicas, window set, repeat interval, `card(label_0)`, series per group, θ/a, accuracy target, SLA), with the sweep script and figures | Open; code for the fixed 67-RQE workload exists. Still to do: the grid driver, then the sweep (after #140 and the two-cost-model PR) | -| — | sketch-bench, not yet opened | The two cost models (§4): CPU timeline, model A, model B, rerun of every experiment | Not started as a PR | +| — | sketch-bench #141 (stacked on #139) | The two cost models (§4): CPU timeline, model A, model B with the per-RQE occupancy bound, `minimize_model_cost` with a per-solve time limit; the synthetic workload-grid driver (66 runs over 40 tables) and plots | Open; runs in progress | -Merge order: rebase and merge #138 → #140 → #139 → two-cost-model PR. +Merge order: #138 → #140 → #139 → #141. ## 9. Decisions From 12e98f800903d527227993d4cf89f4c7bd0ca134 Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 16:09:30 +0000 Subject: [PATCH 21/39] docs(planner): describe model B's per-RQE peak occupancy precisely Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 23 ++++++++++++++--------- 1 file changed, 14 insertions(+), 9 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 78fd6f93..0be5339f 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -164,15 +164,20 @@ family's model B. PerQuery-CostAware also uses the model being compared. AutoSketch-Adapted's plan does not depend on cost and is scored under both. - Model A is linear: the average-CPU and memory terms weighted by `a` and `b`. -- Model B's peak is bounded linearly per RQE. In steady state, an RQE whose - evaluations take `latency` every `T` keeps `ceil(latency / T)` cores busy at - most, counting evaluations that outlast their interval and overlap - themselves. The MILP uses `peak ≤ ingest + Σ_r occupancy_{r,D} · z_{r,D}` - with these constant per-(RQE, deployment) occupancies. - - The bound is exact when no evaluation outlasts its interval. - - An earlier proxy that read only bin 0 undercounted self-overlapping - evaluations, about 40 per synthetic run, and caused sanity violations. - It was replaced (sketch-bench #141). +- Model B's peak is bounded linearly per RQE. An RQE's **peak occupancy** is + the most CPU-seconds its evaluations use in any 1-second bin in steady state + (`peak_occupancy` in sketch-bench #141): + - with no self-overlap (`latency ≤ T`), its share of the firing bin, + `min(latency, 1 s)`; e.g. 0.4 s every 60 s gives 0.4; + - otherwise, the maximum over one period of the overlapping evaluations; + e.g. 2.5 s every 1 s gives 0.5 + 1 + 1 = 2.5 cores in bin [2, 3). + + The MILP uses `peak ≤ ingest + Σ_r occupancy_{r,D} · z_{r,D}`, which is linear + because each (RQE, deployment) latency is a constant. It is an upper bound, + since RQEs may peak in different bins. It is exact when no evaluation outlasts + its interval: every RQE then peaks in bin 0, where all of them fire. An + earlier proxy that read only bin 0 undercounted self-overlapping + evaluations, about 40 per synthetic run, and caused sanity violations. - After solving, the exact peak is recomputed from the timeline and compared with the bound. From 1e79267fd3f928dc26979f7ed759afe490ab8e40 Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 16:12:37 +0000 Subject: [PATCH 22/39] docs(planner): report model B under aligned and staggered evaluation starts Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 11 ++++++++++- 1 file changed, 10 insertions(+), 1 deletion(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 0be5339f..367d39ed 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -152,7 +152,16 @@ $/hour = a · (total CPU-seconds / 24 h) + b · Mem_GiB a = 0.0368 and b = 0.00364 with the 2026-10-04 prices. Model A does not depend on the machine family. -**Model B — peak-provisioned (buy machines for the peak).** Per family `f`: +**Model B — peak-provisioned (buy machines for the peak).** Reported for two +evaluation alignments: +- **aligned:** every RQE first fires at `t = 0`, the worst case; +- **staggered:** each RQE first fires at a deterministic pseudo-random offset in + `[0, T_r)`, the typical case (Prometheus spreads rule-group evaluations). + +Plans are optimized for the aligned case only; staggered prices the same plans +by simulation, so the methods' order there is reported but not guaranteed. + +Per family `f`: ```text n_f = max(peak CPU / vCPU_f, Mem_GiB / GiB_f) From f720d32766df4b4629541fc87351bf46a97f9c6f Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 16:14:14 +0000 Subject: [PATCH 23/39] docs(planner): re-solve model B for staggered starts with exact per-bin peaks Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 9 +++++++-- 1 file changed, 7 insertions(+), 2 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 367d39ed..8afbb2e4 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -158,8 +158,13 @@ evaluation alignments: - **staggered:** each RQE first fires at a deterministic pseudo-random offset in `[0, T_r)`, the typical case (Prometheus spreads rule-group evaluations). -Plans are optimized for the aligned case only; staggered prices the same plans -by simulation, so the methods' order there is reported but not guaranteed. +Each alignment is solved separately. Staggered offsets are integers, so the +timeline repeats with period `P = lcm(T_r)` (300 s for the synthetic intervals), +and the staggered peak is exact and linear: +`peak ≥ ingest + Σ_(r,D) occ_{r,D}(b) · z_{r,D}` for every bin `b` in `[0, P)`, +where `occ_{r,D}(b)` is the constant CPU-seconds RQE `r` uses in bin `b` with +deployment `D`'s latency, including self-overlap. If `P` exceeds 86,400 s, the +aligned plan is priced under staggered instead, and flagged. Per family `f`: From e2dfdcb2979588fb5f5fb4f4316253309c07ab44 Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 18:10:48 +0000 Subject: [PATCH 24/39] docs(planner): strictness levels, benchmark and profiling time, refresh stale facts Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 93 +++++++++++++----------- 1 file changed, 49 insertions(+), 44 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 8afbb2e4..6b95fb6a 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -1,6 +1,6 @@ # Evaluation plan: AutoSketch vs. the ASAPQuery planner (paper §6.3) -Status: plan, no results yet. The design decisions are settled in §9. +Status: `traces` results are in sketch-bench #141; the synthetic sweep is running. The design decisions are settled in §9. ## 1. Question @@ -30,12 +30,12 @@ interval) and sums the results; the planner is invoked once for the batch. | Earlier protocol (E1–E3, end-to-end execution) | ASAPQuery-backend `docs/evaluation/autosketch-comparison.md` ([#545](https://github.com/ProjectASAP/ASAPQuery-backend/pull/545)) | Merged. Execution-based; this plan is planner-level and uses estimated costs instead. | | Top-K dashboard comparison | ASAPQuery-backend [#602](https://github.com/ProjectASAP/ASAPQuery-backend/pull/602) | Closed, not merged. | | RQE deployment MILP (HiGHS): candidates `(capability, config, labels, x, y)`, sharing, latency bounds, minimum-CPU objective | sketch-bench `rqe-optimizer/` ([#129](https://github.com/ProjectASAP/sketch-bench/pull/129)) | Merged 2026-10-04. **This is the planner we evaluate for now.** | -| Measured per-operation costs (`AtomicCostEntry`: memory/instance, insert/merge/query CPU, accuracy) | sketch-bench `scripts/export_rqe_optimizer_costs.sh` | Merged; 18 rows, 2 configs per sketch variant. | +| Measured per-operation costs (`AtomicCostEntry`: memory/instance, insert/merge/query CPU, accuracy) | sketch-bench `scripts/export_rqe_optimizer_costs.sh` | Merged; 18 rows. Superseded for this evaluation by the saturation tables (#130, #136, #140). | | Saturation study: error vs. `N`, `N_sat`, cost per config and shape | sketch-bench [#130](https://github.com/ProjectASAP/sketch-bench/pull/130) | Merged. Source of the saturated lookup (§6). | | Accuracy after merging `m` shards (KLL, top-k) | sketch-bench [#131](https://github.com/ProjectASAP/sketch-bench/pull/131) | Merged 2026-10-05 (`bd644fe`). Source of KLL/top-k lookups when `m > 1`. | | Moving the MILP into ASAPQuery's planner | ASAPQuery `asap-planner-rs/src/optimizer/` (Milind; related: [#776](https://github.com/ProjectASAP/ASAPQuery/pull/776), [#725](https://github.com/ProjectASAP/ASAPQuery/pull/725)) | Out of scope: the evaluation uses sketch-bench `rqe-optimizer` and is not rerun on `asap-planner-rs`. | -No AutoSketch implementation exists in sketch-bench or in this repository. +AutoSketch-Adapted is implemented in sketch-bench `rqe-optimizer/src/autosketch.rs` (#135). ## 3. Methods compared @@ -44,8 +44,8 @@ label-set cardinalities and arrival rates, and are scored by the same cost function (§4). 1. **ASAP** — sketch-bench `rqe-optimizer` MILP over the whole batch, with - accuracy and latency constraints, minimizing the §4 cost for one machine - family. + accuracy and latency constraints, solved separately for each cost model + (§4). 2. **AutoSketch-Adapted** — Algorithm 4 run independently per RQE: - search space: the measured configs of the RQE's capability families (`Capability::families()`); @@ -88,8 +88,7 @@ CPU = Σ_active D λ(ℓ_D) · (x_D / y_D) · insert_cpu_D + Σ_r card(ℓ_r) · (query_cpu_D(r) + (S_r / x_D(r) − 1) · merge_cpu_D(r)) / T_r (query + merge) ``` -**Memory** — new. Today's objective tracks peak per-query memory, not retained -state. Retained state of an active deployment `D` holds `x/y` open instances +**Memory** — retained state (sketch-bench #137). Retained state of an active deployment `D` holds `x/y` open instances plus the closed instances needed by the longest lookback it serves: ```text @@ -120,8 +119,6 @@ the AWS Pricing API (us-east-1, Linux, on-demand), committed as edited by hand. Since `n_f` is fractional, instance size within a family does not change the result. -ASAP is solved once per family. AutoSketch-Adapted's plan does not depend on -the family; it is scored under each family's cost. ### Two cost models per experiment run @@ -203,16 +200,20 @@ almost free under model A. ## 5. Constraints -**Accuracy target**, swept over {90%, 95%, 99%} in the synthetic workload -(95% elsewhere). A target `p` maps to error ≤ `1 − p` and to precision ≥ `p`. -The 95% case, per capability, using the metrics the cost table already records: +**Accuracy target**: one of three strictness levels, each with its own target +per capability. The default level matches the `traces` targets from ASAPQuery's +dataset analysis. The synthetic workload sweeps all three; `traces` uses its +fitted targets. -| Capability | Metric | Constraint | -| --- | --- | --- | -| Freq | relative error | ≤ 0.05 | -| Quantile | rank error | ≤ 0.05 | -| Cardinality | relative error | ≤ 0.05 | -| TopK | precision@k | ≥ 0.95 | +| Level | Freq (ARE) | Quantile (rank error) | TopK (precision@k) | Cardinality (relative error) | +| --- | --- | --- | --- | --- | +| loose | ≤ 0.10 | ≤ 0.02 | ≥ 0.90 | ≤ 0.05 | +| **default** | **≤ 0.05** | **≤ 0.01** | **≥ 0.95** | **≤ 0.02** | +| strict | ≤ 0.01 | ≤ 0.005 | ≥ 0.99 | ≤ 0.01 | + +An earlier version mapped one percentage `p` to error ≤ `1 − p` for every +capability. It was dropped (2026-10-05): at 95% it allowed quantile rank error +0.05, so a p99 query could return the p94 value. **Latency** — one absolute SLA applies to every RQE, swept over {0.01, 0.1, 1, 10, 100, 1000} ms and no limit. The synthetic workload extends @@ -364,7 +365,7 @@ Each dimension has a default (bold). A workload fixes every dimension. | Groups `C = card(label_0)` | 1e1, 1e2, **1e3**, 1e4, 1e5, 1e6 | Keys per frequency sketch; quantile sketches per `by (label_0)` deployment | | Series per group `s` (product of the non-grouping label cardinalities) | 1, 10, **100**, 1000 | Events per group for spatial templates; keys and sketch instances for per-series templates | | Key skew θ / value tail a | θ ∈ {0, 0.5, **1.0**, 1.5, 2.0}; a ∈ {1.1, **2**, 3} | Sketch size needed for the accuracy target | -| Accuracy target | 90%, **95%**, 99% | §5 | +| Accuracy target (strictness) | loose, **default**, strict | §5 | | Latency SLA | the §5 grid, **no limit** | §5 | The full Cartesian product is too large. The sweep is: @@ -387,10 +388,10 @@ each family (§4). Per (workload, baseline, cost model, SLA), report: ### Benchmark input -The cost table used for the evaluation needs a wider grid than today's two -configs per variant, otherwise Algorithm 4's neighbor search has nothing to -search: CMS/Count Sketch depth {2..8} × width {256..8192}, KLL k -{50..800}, DD α {0.005..0.05}, HLL precision {10..16}. +Configurations come from the saturation study's grid (#130, #140): CMS, +Count Sketch and CMS-heap top-k with rows ∈ {3, 5} and cols 256–16384 (rows = 5 +costs scaled from rows = 3), KLL k ∈ {50, 200, 800}, DDSketch α 0.005–0.05, HLL +`lg_k` ∈ {10, 12, 14}. #### Data parameters, shared by both methods @@ -470,14 +471,16 @@ median of repeated runs for timings: - *Search time:* AutoSketch is the sum over RQEs of Algorithm 4 wall time, using table lookups. ASAP is candidate generation, dominance pruning and MILP solve. - - *Benchmark time:* AutoSketch benchmarks every probed (config, input size) - per RQE, as in the paper (§5.2, Exp#9: 1–2 minutes per config, about - 6.5 minutes per application). We charge `Σ_r Σ_probes t_bench(config, - n(S_r, ℓ_r))`, where `t_bench` is sketch-bench's measured wall time for - that point over all benchmark inputs; probes already charged for the same - (config, size) are not charged again. ASAP's benchmark time is one profiling - pass over the grid, shared by all RQEs and reusable across workloads. It is - reported once, next to how many RQEs it served. + - *Benchmark time:* AutoSketch benchmarks every probed configuration, as in + the paper (§5.2, Exp#9: 1–2 minutes per config, about 6.5 minutes per + application). Reported two ways, per distinct probed (config, input size): + - a **lower bound**, `n · insert_cpu + query_cpu`, the CPU to insert the + input once and run one query phase; + - a **paper-rate estimate**, 60 s per distinct probe. + - *ASAP's one-time profiling:* the wall time of the sketch-bench saturation + runs that produced its tables (#130, #131, #140). It is shared by all RQEs + and reusable across workloads, so it is reported once, plus amortized per + RQE served. - The paper's figure shows search + benchmark per method, stacked. - **Total cost** ($/hour) under **model A** and under **model B** for each family, with its inputs: total CPU-seconds, peak CPU, retained GiB, and @@ -486,8 +489,8 @@ median of repeated runs for timings: - **Estimated query latency and latency SLA violations** per method. - Estimated latency per RQE: `card(ℓ) · (query_cpu + (S/x − 1) · merge_cpu)` (§4). Report its maximum and median over the RQEs, plus the per-RQE values in the raw output. - SLA violations: the number of RQEs whose estimated latency exceeds the SLA. Only AutoSketch-Adapted can have any, since the other methods are constrained. -- **Estimated accuracy** per RQE (all methods meet it on single-instance - measurements by construction). +- **Estimated accuracy** per RQE: every method meets its target under its own + lookup rule (§6, "Benchmark input"). - Active deployments and total sketch instances. Figures: @@ -506,13 +509,13 @@ Figures: | --- | --- | --- | --- | | this | ASAPQuery #777 | This plan | Draft, updated as decisions change | | — | sketch-bench #130, #131 | Saturation curves at K ∈ {1e3, 1e5, 1e7}; accuracy after merging `m` shards | Merged | -| 1 | sketch-bench #137 | Retained memory, EC2 pricing, `milp::minimize_cost` (steady-state model), solver scaling | Merged | +| 1 | sketch-bench #137 | Retained memory, EC2 pricing, `milp::minimize_cost`, solver scaling | Merged | | 3 | sketch-bench #135 | AutoSketch-Adapted (Algorithm 4), aligned with the paper's EXAMINE rule and seeding | Merged | -| 2 | sketch-bench #136 | Evaluation table for the trace workloads: per (RQE, config) accuracy for AutoSketch and for ASAP at each `m`, saturation, costs | Merged | -| 4 | sketch-bench #138 | Runner, absolute SLA, results for `traces` (the `example`/`scaling` workloads are to be removed in its rebase) | Open; needs a rebase on main and an AutoSketch rerun with #135's final search | -| — | sketch-bench #140 | Saturation curves (accuracy vs. events per sketch, N_sat, costs) at K ∈ {1e1, 1e2, 1e4, 1e6}: the synthetic workload needs these cardinalities and #130 measured only 1e3, 1e5, 1e7 | Draft; accuracy done, cost 197 of 240 points | -| — | sketch-bench #139 | Synthetic workload: the 10 templates, PerQuery-CostAware bound by the SLA, and the workload-grid driver (dimensions in §6 "Workload grid": query mix, replicas, window set, repeat interval, `card(label_0)`, series per group, θ/a, accuracy target, SLA), with the sweep script and figures | Open; code for the fixed 67-RQE workload exists. Still to do: the grid driver, then the sweep (after #140 and the two-cost-model PR) | -| — | sketch-bench #141 (stacked on #139) | The two cost models (§4): CPU timeline, model A, model B with the per-RQE occupancy bound, `minimize_model_cost` with a per-solve time limit; the synthetic workload-grid driver (66 runs over 40 tables) and plots | Open; runs in progress | +| 2 | sketch-bench #136 | Evaluation table for the trace workloads | Merged | +| 4 | sketch-bench #138 | Runner and absolute SLA for `traces` (rebased on main; `example`/`scaling` removed) | Open | +| — | sketch-bench #140 | Saturation curves at K ∈ {1e1, 1e2, 1e4, 1e6} for the synthetic workload | Draft; 62 cost points re-measured serially after the parallel run failed the 10% consistency check | +| — | sketch-bench #139 | Synthetic workload: 10 templates, PerQuery-CostAware bound by the SLA | Open | +| — | sketch-bench #141 | Two cost models (model A; model B aligned and staggered, both solved), strictness levels, workload-grid driver, `traces` results | Open; synthetic sweep and scale study running | Merge order: #138 → #140 → #139 → #141. @@ -553,15 +556,17 @@ sliding sketch per query. the plans.** It is exact by construction for CMS, Count Sketch, HLL and DDSketch, and taken from #131 for KLL and top-k. #131 covers `N ≤ 1e7` and `m ≤ 64`. A deployment needing `m > 64` (e.g. a 1-day lookback over 1-minute - windows, `m = 1440`) is outside the measured range. Mark it as extrapolated, - or exclude it for KLL/top-k. Replay the synthetic default workload's chosen plans in sketch-bench once to + windows, `m = 1440`) is outside the measured range: it is ineligible for KLL + and top-k, and allowed for the exact-merge sketches. Replay the synthetic default workload's chosen plans in sketch-bench once to confirm the lookups. - Costs and latencies are estimates from per-operation measurements, not end-to-end executions. The execution-based comparison is ASAPQuery-backend #545/#547. - AutoSketch-Adapted's deployments (`x = S`, `y = gcd(S, T)`) are in the ASAP candidate set (`candidates.rs` generates every divisor of `S` as a window and - `gcd(x, T)` as a slide). So when its plan meets the latency bounds, the ASAP - MILP can choose the same deployments and pay for shared ones once; ASAP's - cost is never higher. The result to report is the size of the gap and where it comes + `gcd(x, T)` as a slide). So when its plan meets the latency bounds and its + configurations also pass ASAP's lookup rule (the two rules differ on + saturation, §6), the ASAP MILP can choose the same deployments and pay for + shared ones once; ASAP's cost is never higher. The sanity check reports any + case where this does not hold. The result to report is the size of the gap and where it comes from, not that a gap exists. From 6ffe242537eec825a7df98ccbfa0f056a46c6c4d Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 18:10:59 +0000 Subject: [PATCH 25/39] docs(planner): list the measured configuration grid exactly Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 6b95fb6a..fefb07c7 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -389,9 +389,9 @@ each family (§4). Per (workload, baseline, cost model, SLA), report: ### Benchmark input Configurations come from the saturation study's grid (#130, #140): CMS, -Count Sketch and CMS-heap top-k with rows ∈ {3, 5} and cols 256–16384 (rows = 5 -costs scaled from rows = 3), KLL k ∈ {50, 200, 800}, DDSketch α 0.005–0.05, HLL -`lg_k` ∈ {10, 12, 14}. +Count Sketch and CMS-heap top-k with rows ∈ {3, 5} and cols ∈ {256, 1024, +4096, 16384} (rows = 5 costs scaled from rows = 3), KLL k ∈ {50, 200, 800}, +DDSketch α ∈ {0.005, 0.01, 0.02, 0.05}, HLL `lg_k` ∈ {12, 14, 16}. #### Data parameters, shared by both methods From ff8c226937195428d792d671d81861189e04e76c Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 20:40:35 +0000 Subject: [PATCH 26/39] docs(planner): state cost units and the whole-query-phase charge Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 13 +++++++++++++ 1 file changed, 13 insertions(+) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index fefb07c7..29010ca6 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -198,6 +198,19 @@ almost free under model A. **Latency** — per-RQE estimate already in sketch-bench: `card(ℓ) · (query_cpu + (S/x − 1) · merge_cpu)`. +**Units and per-operation costs.** CPU is CPU time (user + system), in +core-seconds, measured by sketch-bench: +- `insert_cpu`: the insert phase's CPU ÷ N, in CPU-seconds per item; +- `merge_cpu`: merging 16 shards ÷ 15, in CPU-seconds per merge; +- `query_cpu`: the benchmark's **whole query phase**, charged per evaluation. + For frequency that is every key seen; for KLL/DDSketch, 101 quantiles; for + cardinality, a fixed repeat count. This overstates a single-quantile query by + up to about 101×, equally for every method, so it inflates absolute estimated + latency but not the comparison (decided 2026-10-05); +- memory: the sketch's self-reported bytes per instance, not process RSS. + +Loads are reported in vCPU (core-seconds per second) and totals in CPU-hours. + ## 5. Constraints **Accuracy target**: one of three strictness levels, each with its own target From b3891ddaeebb2b11c5c89743905a7c41aa90e997 Mon Sep 17 00:00:00 2001 From: zz_y Date: Mon, 5 Oct 2026 21:45:48 +0000 Subject: [PATCH 27/39] docs(planner): drop the staggered-start model B Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 18 ++++-------------- 1 file changed, 4 insertions(+), 14 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 29010ca6..c2fd346d 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -149,19 +149,9 @@ $/hour = a · (total CPU-seconds / 24 h) + b · Mem_GiB a = 0.0368 and b = 0.00364 with the 2026-10-04 prices. Model A does not depend on the machine family. -**Model B — peak-provisioned (buy machines for the peak).** Reported for two -evaluation alignments: -- **aligned:** every RQE first fires at `t = 0`, the worst case; -- **staggered:** each RQE first fires at a deterministic pseudo-random offset in - `[0, T_r)`, the typical case (Prometheus spreads rule-group evaluations). - -Each alignment is solved separately. Staggered offsets are integers, so the -timeline repeats with period `P = lcm(T_r)` (300 s for the synthetic intervals), -and the staggered peak is exact and linear: -`peak ≥ ingest + Σ_(r,D) occ_{r,D}(b) · z_{r,D}` for every bin `b` in `[0, P)`, -where `occ_{r,D}(b)` is the constant CPU-seconds RQE `r` uses in bin `b` with -deployment `D`'s latency, including self-overlap. If `P` exceeds 86,400 s, the -aligned plan is priced under staggered instead, and flagged. +**Model B — peak-provisioned (buy machines for the peak).** Every RQE first +fires at `t = 0` and then every `T_r` (aligned starts, the worst case). A +staggered-start variant was considered and dropped (decided 2026-10-05). Per family `f`: @@ -528,7 +518,7 @@ Figures: | 4 | sketch-bench #138 | Runner and absolute SLA for `traces` (rebased on main; `example`/`scaling` removed) | Open | | — | sketch-bench #140 | Saturation curves at K ∈ {1e1, 1e2, 1e4, 1e6} for the synthetic workload | Draft; 62 cost points re-measured serially after the parallel run failed the 10% consistency check | | — | sketch-bench #139 | Synthetic workload: 10 templates, PerQuery-CostAware bound by the SLA | Open | -| — | sketch-bench #141 | Two cost models (model A; model B aligned and staggered, both solved), strictness levels, workload-grid driver, `traces` results | Open; synthetic sweep and scale study running | +| — | sketch-bench #141 | Two cost models (model A; model B with aligned starts), strictness levels, workload-grid driver, `traces` results | Open; synthetic sweep and scale study running | Merge order: #138 → #140 → #139 → #141. From 997ecf36c82652f82eecd3f34c5d2e33b2adb07f Mon Sep 17 00:00:00 2001 From: zz_y Date: Tue, 6 Oct 2026 00:24:27 +0000 Subject: [PATCH 28/39] docs(planner): charge queries per issued query; fixed-size benchmark bound Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 26 ++++++++++++++++-------- 1 file changed, 17 insertions(+), 9 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index c2fd346d..af82fe21 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -185,18 +185,24 @@ and is scored under both. Memory is kept in both models: without it, memory-bound workloads would look almost free under model A. -**Latency** — per-RQE estimate already in sketch-bench: -`card(ℓ) · (query_cpu + (S/x − 1) · merge_cpu)`. +**Latency** — per-RQE estimate: +`card(ℓ) · (q_r · query_cpu_per_query + (S/x − 1) · merge_cpu)`, the +evaluation's CPU time run serially on one core (an upper bound; parallel +execution across sketch instances would reduce it proportionally). `q_r` is +the number of sketch queries one evaluation issues per instance: +`by (label_0)` frequency: `C` point queries; per-series frequency: `C · s`; +top-k: 1; quantile: 1 per requested quantile; `traces` key queries: the +window's key count, value queries: 1. **Units and per-operation costs.** CPU is CPU time (user + system), in core-seconds, measured by sketch-bench: - `insert_cpu`: the insert phase's CPU ÷ N, in CPU-seconds per item; - `merge_cpu`: merging 16 shards ÷ 15, in CPU-seconds per merge; -- `query_cpu`: the benchmark's **whole query phase**, charged per evaluation. - For frequency that is every key seen; for KLL/DDSketch, 101 quantiles; for - cardinality, a fixed repeat count. This overstates a single-quantile query by - up to about 101×, equally for every method, so it inflates absolute estimated - latency but not the comparison (decided 2026-10-05); +- `query_cpu_per_query`: the benchmark's query-phase CPU ÷ the number of + queries in the phase, charged `q_r` times per instance per evaluation. An + earlier version charged the whole query phase per evaluation; it overstated + single-quantile queries about 101× and made estimated latencies around + 200 s, so it was replaced (decided 2026-10-06); - memory: the sketch's self-reported bytes per instance, not process RSS. Loads are reported in vCPU (core-seconds per second) and totals in CPU-hours. @@ -477,8 +483,10 @@ median of repeated runs for timings: - *Benchmark time:* AutoSketch benchmarks every probed configuration, as in the paper (§5.2, Exp#9: 1–2 minutes per config, about 6.5 minutes per application). Reported two ways, per distinct probed (config, input size): - - a **lower bound**, `n · insert_cpu + query_cpu`, the CPU to insert the - input once and run one query phase; + - a **lower bound**, `N_bench · insert_cpu_per_item + one query phase` + with `N_bench = 1e8`, the size sketch-bench benchmarks at. The paper + also benchmarks a fixed-size representative workload, not the query + window's full data; - a **paper-rate estimate**, 60 s per distinct probe. - *ASAP's one-time profiling:* the wall time of the sketch-bench saturation runs that produced its tables (#130, #131, #140). It is shared by all RQEs From d41f06287218c320a8afd182780c16309739291a Mon Sep 17 00:00:00 2001 From: zz_y Date: Tue, 6 Oct 2026 14:13:14 +0000 Subject: [PATCH 29/39] docs(planner): add the dashboard template set and shared replicas Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 21 +++++++++++++++++++++ 1 file changed, 21 insertions(+) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index af82fe21..87509d01 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -291,6 +291,25 @@ Mapping notes: (`count_over_time`/`sum_over_time`/`quantile_over_time`/`increase`/`rate`), `TemporalAgg SpatialAgg*`, and `AnyAgg AnyAgg`. +#### Dashboard template set + +A second template set models an SLO / monitoring dashboard, latency-quantile +heavy, with Grafana's usual time ranges. Window set +`W_d = {1m, 5m, 15m, 1h, 6h, 24h}`; temporal templates repeat every 1 m (the +dashboard refresh); spatial ones keep `S = T = 1 s`. + +| # | PromQL | Capability | Grouping | RQEs per replica | +|---|---|---|---|---| +| D1 | `quantile_over_time(q, data[w])`, q ∈ {0.5, 0.9, 0.99}, w ∈ `W_d` | Quantile | per series | 18 | +| D2 | `quantile by (q, label_0) (data)`, q ∈ {0.5, 0.9, 0.99} | Quantile | per `label_0` group | 3 | +| D3 | `sum by (label_0) (rate(data[w]))`, w ∈ `W_d` | Freq over increments | `label_0` | 6 | +| D4 | `topk by (3, label_0) (rate(data[w]))`, w ∈ {5m, 1h} | TopK over increments | `label_0` | 2 | +| D5 | `quantile_over_time(0.99, data[w]) / quantile_over_time(0.5, data[w])`, w ∈ {5m, 1h} | Two Quantile RQEs | per series | 4 | + +33 RQEs per replica. D1's quantiles share one stream across overlapping +windows; D5 repeats D1's p99/p50 at 5m and 1h; D3 and D4 share the +`label_0` increment stream. top-k uses k = 3, the k sketch-bench measures. + #### Data model There is one metric, `data`. A **series** is one combination of label values. @@ -369,6 +388,8 @@ Each dimension has a default (bold). A workload fixes every dimension. |---|---|---| | Query mix (templates) | **all 10**; spatial only {1, 2, 3}; temporal only {4–10}; frequency only {1, 4, 6, 7, 8}; quantile only {3, 5, 10}; top-k only {2, 9} | Summary types, and how much can be shared | | Number of RQEs: replicas `r` | **1**, 2, 4, 8, 16, 32, 64 | Each replica adds a filter `{label_1="v_i"}` selecting a disjoint subset of series, so it reads its own streams. Total RQEs = `r · (n_spatial + n_temporal·|W|)` | +| Replica mode | disjoint (each replica filters `{label_1="v_i"}` and reads its own series; planning-time control); **shared** (every replica reads the same stream with a seeded random subset of 3 windows, 3 quantiles from {0.5, 0.75, 0.9, 0.95, 0.99}, and `T` from {10 s, 1 m, 5 m}; identical RQEs are deduplicated): many users or dashboards over the same metrics | How the sharing benefit grows with the number of queries | +| Template set | the 10 templates; **dashboard** (above) | Workload realism | | Lookback window set `W` | {1h}; {1m, 1h}; {1m, 10m, 1h}; **{1m, 10m, 1h, 6h, 24h}** | Overlapping windows over the same stream: the main sharing opportunity | | Temporal repeat interval `T` | 10 s, **1 m**, 5 m | Recurrence: query and merge work vs. ingest | | Groups `C = card(label_0)` | 1e1, 1e2, **1e3**, 1e4, 1e5, 1e6 | Keys per frequency sketch; quantile sketches per `by (label_0)` deployment | From 9c9c8af0945a0423f723c57e42616a4633c63103 Mon Sep 17 00:00:00 2001 From: zz_y Date: Tue, 6 Oct 2026 21:19:10 +0000 Subject: [PATCH 30/39] docs(planner): align the AutoSketch evaluation plan with the per-phase cost model Addresses review on #777: - cost: sketch-bench #145's per-phase CPU and memory with a weighted objective; CPU only first, then Fargate prices. EC2 machine families, the peak-provisioned model and the CPU timeline are dropped. - capabilities: sum and rate/increase map to the exact accumulators (#144), not frequency; strictness applies to quantile and top-k only. - inputs: read the export_rqe_optimizer_costs table (measured_at, merge accuracy, size sweep) instead of the saturation tables; exact accumulators are benchmarked for cost but need no saturation point. - grid: keep template set, shared replicas {1, 8, 64}, C {1e2, 1e3, 1e4}, strictness and SLA {0.01, 0.1, 1, 10} ms; move the rest to planner sensitivity. Disjoint replicas are dropped (spatial filters unsupported). Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 474 ++++++++++------------- 1 file changed, 198 insertions(+), 276 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 87509d01..551172f2 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -1,6 +1,6 @@ # Evaluation plan: AutoSketch vs. the ASAPQuery planner (paper §6.3) -Status: `traces` results are in sketch-bench #141; the synthetic sweep is running. The design decisions are settled in §9. +Status: revised 2026-10-06 to sketch-bench #145's cost model and capabilities (#144). The evaluation code (#138–#141) is reworked onto it before the runs; earlier results are obsolete. The design decisions are settled in §9. ## 1. Question @@ -17,7 +17,7 @@ Algorithm 4) differs from ASAPQuery's planner in four ways that matter here: | Repetition over time | Not modeled | Lookback `S` and repeat interval `T` drive window/slide choice | | Sharing | None: each query gets its own sketch | One deployment may serve several compatible RQEs | | Constraints | Accuracy only | Accuracy and per-RQE query latency | -| Objective | Resource use (memory) | Weighted CPU + memory cost from EC2 prices | +| Objective | Resource use (memory) | Weighted CPU + memory: `w_cpu · CPU + w_mem · memory` | So the comparison invokes AutoSketch **once per RQE** (one QE at one repeat interval) and sums the results; the planner is invoked once for the batch. @@ -29,10 +29,10 @@ interval) and sums the results; the planner is invoked once for the batch. | AutoSketch Algorithm 4 adaptation (LHS seeds, feasibility-directed width/depth neighbor search, pruning) | ASAPQuery-backend `data_plane/examples/autosketch_comparison.rs` ([#547](https://github.com/ProjectASAP/ASAPQuery-backend/pull/547)) | Merged. CMS/Count Sketch/Bloom only; hardcoded CMS grid; executes sketches to measure accuracy. | | Earlier protocol (E1–E3, end-to-end execution) | ASAPQuery-backend `docs/evaluation/autosketch-comparison.md` ([#545](https://github.com/ProjectASAP/ASAPQuery-backend/pull/545)) | Merged. Execution-based; this plan is planner-level and uses estimated costs instead. | | Top-K dashboard comparison | ASAPQuery-backend [#602](https://github.com/ProjectASAP/ASAPQuery-backend/pull/602) | Closed, not merged. | -| RQE deployment MILP (HiGHS): candidates `(capability, config, labels, x, y)`, sharing, latency bounds, minimum-CPU objective | sketch-bench `rqe-optimizer/` ([#129](https://github.com/ProjectASAP/sketch-bench/pull/129)) | Merged 2026-10-04. **This is the planner we evaluate for now.** | -| Measured per-operation costs (`AtomicCostEntry`: memory/instance, insert/merge/query CPU, accuracy) | sketch-bench `scripts/export_rqe_optimizer_costs.sh` | Merged; 18 rows. Superseded for this evaluation by the saturation tables (#130, #136, #140). | -| Saturation study: error vs. `N`, `N_sat`, cost per config and shape | sketch-bench [#130](https://github.com/ProjectASAP/sketch-bench/pull/130) | Merged. Source of the saturated lookup (§6). | -| Accuracy after merging `m` shards (KLL, top-k) | sketch-bench [#131](https://github.com/ProjectASAP/sketch-bench/pull/131) | Merged 2026-10-05 (`bd644fe`). Source of KLL/top-k lookups when `m > 1`. | +| RQE deployment MILP (HiGHS): candidates `(capability, config, labels, x, y)`, sharing, latency bounds | sketch-bench `rqe-optimizer/` ([#129](https://github.com/ProjectASAP/sketch-bench/pull/129)); per-phase cost model and weighted objective [#145](https://github.com/ProjectASAP/sketch-bench/pull/145); exact accumulators and top-k families [#144](https://github.com/ProjectASAP/sketch-bench/pull/144) | Merged. **This is the planner we evaluate.** | +| Measured per-operation costs (`AtomicCostEntry`: memory/instance, insert/merge/query CPU, accuracy) | sketch-bench `scripts/export_rqe_optimizer_costs.sh` | Merged. **The cost table this evaluation reads** (§6). Each row records its measurement conditions (`measured_at`, #155) and accuracy after merging (#154); a sweep over items per instance is #157. | +| Saturation study: error vs. `N`, `N_sat`, cost per config and shape | sketch-bench [#130](https://github.com/ProjectASAP/sketch-bench/pull/130) | Merged. Background for how error and cost depend on `N`; no longer read by the evaluation. | +| Accuracy after merging `m` shards | sketch-bench [#131](https://github.com/ProjectASAP/sketch-bench/pull/131), carried into the cost table by [#154](https://github.com/ProjectASAP/sketch-bench/pull/154) | Merged. Read when a deployment merges `m = S/x > 1` windows. | | Moving the MILP into ASAPQuery's planner | ASAPQuery `asap-planner-rs/src/optimizer/` (Milind; related: [#776](https://github.com/ProjectASAP/ASAPQuery/pull/776), [#725](https://github.com/ProjectASAP/ASAPQuery/pull/725)) | Out of scope: the evaluation uses sketch-bench `rqe-optimizer` and is not rerun on `asap-planner-rs`. | AutoSketch-Adapted is implemented in sketch-bench `rqe-optimizer/src/autosketch.rs` (#135). @@ -44,8 +44,8 @@ label-set cardinalities and arrival rates, and are scored by the same cost function (§4). 1. **ASAP** — sketch-bench `rqe-optimizer` MILP over the whole batch, with - accuracy and latency constraints, solved separately for each cost model - (§4). + accuracy and latency constraints, minimizing the §4 objective at each weight + setting. 2. **AutoSketch-Adapted** — Algorithm 4 run independently per RQE: - search space: the measured configs of the RQE's capability families (`Capability::families()`); @@ -78,132 +78,57 @@ the cost of running those benchmarks is charged to its planning time (§9 Q3). ## 4. Cost model -Disk is excluded (agreed with Milind). Units: CPU in vCPU (CPU-seconds per -second), memory in GiB. +The cost model is sketch-bench `rqe-optimizer`'s (#145), so every method is +scored by the function ASAP optimizes. Disk is excluded. Units: CPU in vCPU +(CPU-seconds per second, the mean over time), memory in GiB. -**CPU** — already in `rqe_optimizer::objectives::score`: +A deployment `D` groups by labels `G` with window `x` and slide `y`; RQE `r` +has lookback `S` and repeat interval `T`. Per-instance costs are measured +(§6): memory `m`, insert `c_ins`, merge `c_mrg`, query `c_qry`. +`λ = card(series) / scrape interval`. -```text -CPU = Σ_active D λ(ℓ_D) · (x_D / y_D) · insert_cpu_D (ingest) - + Σ_r card(ℓ_r) · (query_cpu_D(r) + (S_r / x_D(r) − 1) · merge_cpu_D(r)) / T_r (query + merge) -``` - -**Memory** — retained state (sketch-bench #137). Retained state of an active deployment `D` holds `x/y` open instances -plus the closed instances needed by the longest lookback it serves: - -```text -Mem_D = card(ℓ_D) · mem_bytes_per_instance_D · (x_D + max_{r→D} S_r) / y_D -Mem = Σ_active D Mem_D -``` - -In the MILP the `max` is linear: `Mem_D ≥ coef(r, D) · z_{r,D}` for each -eligible `r`. - -**Price per machine family (steady state, implemented in sketch-bench #137)** — for family `f` with `vCPU_f`, `GiB_f` and -on-demand `price_f` ($/hour), the plan needs a fractional instance count -`n_f ≥ CPU / vCPU_f` and `n_f ≥ Mem / GiB_f`; cost is `price_f · n_f`. Here -`CPU` is the steady-state average. The two cost models below replace it in the -reported results. This -adds one continuous variable and two constraints, and needs no arbitrary split -of an instance's price between CPU and memory. Families: - -| Family | Example instance | Role | +| Phase | CPU (vCPU) | Memory (bytes) | | --- | --- | --- | -| Compute-optimized | c7i.xlarge | cheap CPU, scarce memory | -| General purpose | m7i.xlarge | balanced | -| Memory-optimized | r7i.xlarge | cheap memory, scarce CPU | - -Storage-optimized families are dropped with disk. Prices are fetched once from -the AWS Pricing API (us-east-1, Linux, on-demand), committed as -`rqe-optimizer/data/ec2-pricing-.json` with the query used, and never -edited by hand. Since `n_f` is fractional, instance size within a family does -not change the result. - - -### Two cost models per experiment run - -The average CPU hides that query load is bursty: ingest is continuous, but -query and merge work arrives at each evaluation. Every run is therefore priced -two ways, from one simulated CPU timeline. - -**CPU timeline.** -- Simulate 24 hours in 1-second bins. -- Every RQE first evaluates at `t = 0`, then every `T_r`. This aligned start - is the worst case. -- Each evaluation occupies one core for its estimated latency, starting when - it fires. Work longer than 1 s spills into later bins. -- `CPU(bin)` = ingest rate + busy-core time overlapping the bin. - -From the timeline: -- **total CPU-seconds** = the area under the curve, `Σ_bins CPU(bin) · 1 s`; -- **peak CPU** = `max_bin CPU(bin)`, in vCPU. - -**Model A — usage-based (pay for what is used).** - -```text -$/hour = a · (total CPU-seconds / 24 h) + b · Mem_GiB -``` +| Ingest | `λ · (x/y) · c_ins` | `card(G) · m · x/y` (open windows) | +| Merge | `card(G) · (S/x − 1) · c_mrg / T` | `card(G) · m` (one accumulator per group), 0 when `S = x` | +| Query | `card(G) · c_qry / T` | `card(G) · 8 B`; top-k `card(G) · 32 · 16 B` | +| Storage | 0 | `card(G) · m · ((max S − x)/y + 1)` (closed windows) | -`a` ($/vCPU-hour) and `b` ($/GiB-hour) are a least-squares fit of -`vCPU · a + GiB · b = price` over c7i.xlarge, m7i.xlarge and r7i.xlarge: about -a = 0.0368 and b = 0.00364 with the 2026-10-04 prices. Model A does not depend -on the machine family. +Ingest and storage are paid once per active deployment; merge and query once +per RQE it serves. Memory sums every term, as if every query evaluates at once. -**Model B — peak-provisioned (buy machines for the peak).** Every RQE first -fires at `t = 0` and then every `T_r` (aligned starts, the worst case). A -staggered-start variant was considered and dropped (decided 2026-10-05). +**Objective** — `w_cpu · CPU + w_mem · Memory_GiB`, summed over the plan: -Per family `f`: +1. **CPU only**, `(w_cpu, w_mem) = (1, 0)`: the first run. Memory is still + reported. +2. **Fargate prices**, `w_cpu = 0.0405` $/vCPU-hour and `w_mem = 0.00445` + $/GB-hour (AWS Fargate, us-east-1, Linux/x86, 2026-10-06), so the objective + is in $/hour. CPU costs about 9× memory per unit; serverless pricing charges + exactly these two resources, which is why it fits the model. -```text -n_f = max(peak CPU / vCPU_f, Mem_GiB / GiB_f) -$/hour = n_f · price_f -``` - -**Optimizing each model.** ASAP is solved separately for model A and for each -family's model B. PerQuery-CostAware also uses -the model being compared. AutoSketch-Adapted's plan does not depend on cost -and is scored under both. -- Model A is linear: the average-CPU and memory terms weighted by `a` and `b`. -- Model B's peak is bounded linearly per RQE. An RQE's **peak occupancy** is - the most CPU-seconds its evaluations use in any 1-second bin in steady state - (`peak_occupancy` in sketch-bench #141): - - with no self-overlap (`latency ≤ T`), its share of the firing bin, - `min(latency, 1 s)`; e.g. 0.4 s every 60 s gives 0.4; - - otherwise, the maximum over one period of the overlapping evaluations; - e.g. 2.5 s every 1 s gives 0.5 + 1 + 1 = 2.5 cores in bin [2, 3). - - The MILP uses `peak ≤ ingest + Σ_r occupancy_{r,D} · z_{r,D}`, which is linear - because each (RQE, deployment) latency is a constant. It is an upper bound, - since RQEs may peak in different bins. It is exact when no evaluation outlasts - its interval: every RQE then peaks in bin 0, where all of them fire. An - earlier proxy that read only bin 0 undercounted self-overlapping - evaluations, about 40 per synthetic run, and caused sanity violations. -- After solving, the exact peak is recomputed from the timeline and compared - with the bound. - -Memory is kept in both models: without it, memory-bound workloads would look -almost free under model A. +CPU is the mean, i.e. the area under the CPU-over-time curve: plans are sized +for average load, not bursts. Peak-provisioned pricing (buy machines for the +peak CPU) and per-instance-family EC2 pricing were considered and dropped +(2026-10-06). **Latency** — per-RQE estimate: -`card(ℓ) · (q_r · query_cpu_per_query + (S/x − 1) · merge_cpu)`, the -evaluation's CPU time run serially on one core (an upper bound; parallel -execution across sketch instances would reduce it proportionally). `q_r` is -the number of sketch queries one evaluation issues per instance: -`by (label_0)` frequency: `C` point queries; per-series frequency: `C · s`; -top-k: 1; quantile: 1 per requested quantile; `traces` key queries: the -window's key count, value queries: 1. +`card(G) · (c_qry + (S/x − 1) · c_mrg)`, the evaluation's CPU time run +serially on one core (an upper bound; parallel execution across instances +would reduce it proportionally). `c_qry` is one query of one instance: one +value of a sum or increase accumulator, one top-k list, or one quantile, so a +template asking `n` quantiles is `n` RQEs. **Units and per-operation costs.** CPU is CPU time (user + system), in core-seconds, measured by sketch-bench: -- `insert_cpu`: the insert phase's CPU ÷ N, in CPU-seconds per item; -- `merge_cpu`: merging 16 shards ÷ 15, in CPU-seconds per merge; -- `query_cpu_per_query`: the benchmark's query-phase CPU ÷ the number of - queries in the phase, charged `q_r` times per instance per evaluation. An - earlier version charged the whole query phase per evaluation; it overstated - single-quantile queries about 101× and made estimated latencies around - 200 s, so it was replaced (decided 2026-10-06); -- memory: the sketch's self-reported bytes per instance, not process RSS. +- `c_ins` (`insert_cpu_secs`): the insert phase's CPU ÷ N, in CPU-seconds per item; +- `c_mrg` (`merge_cpu_secs`): merging 16 shards ÷ 15, in CPU-seconds per merge; +- `c_qry` (`query_cpu_secs`): the query phase's CPU ÷ the number of queries in it, per call + (a top-k heap dump is repeated per pass so it is timed above the clock's + floor, #151); +- `m` (`mem_bytes_per_instance`): the self-reported bytes per instance, not process RSS. Top-k counts + its heap and exact accumulators their hash table (#151). Exact accumulators + are priced per group: memory and merge are divided by the group count they + were measured at. Loads are reported in vCPU (core-seconds per second) and totals in CPU-hours. @@ -214,19 +139,26 @@ per capability. The default level matches the `traces` targets from ASAPQuery's dataset analysis. The synthetic workload sweeps all three; `traces` uses its fitted targets. -| Level | Freq (ARE) | Quantile (rank error) | TopK (precision@k) | Cardinality (relative error) | -| --- | --- | --- | --- | --- | -| loose | ≤ 0.10 | ≤ 0.02 | ≥ 0.90 | ≤ 0.05 | -| **default** | **≤ 0.05** | **≤ 0.01** | **≥ 0.95** | **≤ 0.02** | -| strict | ≤ 0.01 | ≤ 0.005 | ≥ 0.99 | ≤ 0.01 | +| Level | Quantile (rank error) | TopK (precision@k) | +| --- | --- | --- | +| loose | ≤ 0.02 | ≥ 0.90 | +| **default** | **≤ 0.01** | **≥ 0.95** | +| strict | ≤ 0.005 | ≥ 0.99 | + +Sum and increase are served by exact accumulators, so their error is 0 and +every level is met. Strictness therefore matters only for quantile and top-k +RQEs. No template asks for cardinality. + +A deployment that merges `m = S/x` windows must meet the target at both +measured merge counts bracketing `m` (count 1 is the single instance), as +`rqe-optimizer` checks it (#154). An earlier version mapped one percentage `p` to error ≤ `1 − p` for every capability. It was dropped (2026-10-05): at 95% it allowed quantile rank error 0.05, so a p99 query could return the p94 value. **Latency** — one absolute SLA applies to every RQE, swept over -{0.01, 0.1, 1, 10, 100, 1000} ms and no limit. The synthetic workload extends -the grid as needed. +{0.01, 0.1, 1, 10} ms and no limit. - An RQE that no method can meet at a given SLA is excluded from every method at that SLA and reported by ID. Costs at different SLAs therefore cover different RQE sets; compare methods only at one SLA. @@ -258,31 +190,36 @@ dimensions below, and runs every planning baseline on each workload. #### Query templates -These exercise every summary type the planner supports: frequency, top-k and -quantile, spatial and temporal aggregation, and a binary operator. Each spatial +These exercise every capability the templates need: sum, rate/increase, top-k +and quantile, spatial and temporal aggregation, and a binary operator. Each spatial template repeats every 1 s and reads the last second (`S = T = 1 s`). Each temporal template repeats every `T` (default 1 m) and reads `S = T_range`, one RQE per lookback window in the window set `W` (§ workload grid). | # | PromQL | Capability | Grouping | RQEs per replica | |---|---|---|---|---| -| 1 | `sum by (label_0) (data)` | Freq | `label_0` | 1 | +| 1 | `sum by (label_0) (data)` | SumOrCount | `label_0` | 1 | | 2 | `topk by (3, label_0) (data)` | TopK | `label_0` | 1 | | 3 | `quantile by (q, label_0) (data)`, q ∈ {0.5, 0.75, 0.9, 0.95, 0.99} | Quantile | per `label_0` group | 5 | -| 4 | `sum_over_time(data[T])` | Freq | per series | \|W\| | +| 4 | `sum_over_time(data[T])` | SumOrCount | per series | \|W\| | | 5 | `quantile_over_time(q, data[T])`, same five q | Quantile | per series | 5·\|W\| | -| 6 | `rate(data[T])` | Freq over per-series increments | per series | \|W\| | -| 7 | `sum by (label_0) (rate(data[T]))` | Freq over increments | `label_0` | \|W\| | -| 8 | `sum by (label_0) (sum_over_time(data[T]))` | Freq | `label_0` | \|W\| | -| 9 | `topk by (3, label_0) (rate(data[T]))` | TopK over increments | `label_0` | \|W\| | +| 6 | `rate(data[T])` | RateOrIncrease | per series | \|W\| | +| 7 | `sum by (label_0) (rate(data[T]))` | RateOrIncrease | `label_0` | \|W\| | +| 8 | `sum by (label_0) (sum_over_time(data[T]))` | SumOrCount | `label_0` | \|W\| | +| 9 | `topk by (3, label_0) (rate(data[T]))` | TopK over per-series increases | `label_0` | \|W\| | | 10 | `quantile_over_time(0.9, data[T]) / quantile_over_time(0.5, data[T])` | Two Quantile RQEs | per series | 2·\|W\| | With all ten templates, one replica has `7 + 12·|W|` RQEs: 67 for the default five windows. Mapping notes: -- `rate`/`increase` are modeled as a frequency sum of per-series increments, - equivalent to `sum_over_time` over deltas. +- Capabilities are `rqe-optimizer`'s (#144). Sum and rate/increase are their + own capabilities, served by exact per-group accumulators (`exact-sum`, + `exact-increase`). Observability queries ask for a group's total, not for the + frequency of arbitrary keys, so no template is a frequency query. +- An exact accumulator has one configuration, so AutoSketch has nothing to + search for sum and increase RQEs. There, ASAP differs from it only by window + choice and sharing. - The quantiles of one template, and the two operands of template 10, read the same stream. ASAP can serve them from one deployment; AutoSketch gets one deployment per RQE. @@ -302,8 +239,8 @@ dashboard refresh); spatial ones keep `S = T = 1 s`. |---|---|---|---|---| | D1 | `quantile_over_time(q, data[w])`, q ∈ {0.5, 0.9, 0.99}, w ∈ `W_d` | Quantile | per series | 18 | | D2 | `quantile by (q, label_0) (data)`, q ∈ {0.5, 0.9, 0.99} | Quantile | per `label_0` group | 3 | -| D3 | `sum by (label_0) (rate(data[w]))`, w ∈ `W_d` | Freq over increments | `label_0` | 6 | -| D4 | `topk by (3, label_0) (rate(data[w]))`, w ∈ {5m, 1h} | TopK over increments | `label_0` | 2 | +| D3 | `sum by (label_0) (rate(data[w]))`, w ∈ `W_d` | RateOrIncrease | `label_0` | 6 | +| D4 | `topk by (3, label_0) (rate(data[w]))`, w ∈ {5m, 1h} | TopK over per-series increases | `label_0` | 2 | | D5 | `quantile_over_time(0.99, data[w]) / quantile_over_time(0.5, data[w])`, w ∈ {5m, 1h} | Two Quantile RQEs | per series | 4 | 33 RQEs per replica. D1's quantiles share one stream across overlapping @@ -319,16 +256,17 @@ Three labels matter: |---|---|---| | `label_0` | `C = card(label_0)` values | The grouping label: what `by (label_0)` aggregates by | | `instance` | `s` values per `label_0` value | Distinguishes the series inside a group. `s` = series per group | -| `label_1` | `r` values, one per replica | Only for replicas (workload grid): replica `i` filters `{label_1="v_i"}` and reads its own disjoint series. With `r = 1` it is absent | -So a replica has `C · s` series, and the workload has `r · C · s`. Every series -emits one sample every 10 ms (100 samples/s), so one replica's stream carries -`λ = 100 · C · s` samples/s. + +There are `C · s` series. Replicas (workload grid) read the same stream, so +they add RQEs, not series. Every series emits one sample every 10 ms +(100 samples/s), so the stream carries `λ = 100 · C · s` samples/s. Sample values: -- **Frequency and top-k:** the value is the weight being summed. The total - weight of the keys follows Zipf θ. The key is the `label_0` value for - `by (label_0)` templates and the series for per-series templates. +- **Sum and increase:** exact, so the value distribution does not affect cost + or accuracy. +- **Top-k:** the value is the weight; the series' total weights within a group + follow Zipf θ. - **Quantiles:** values are drawn from Pareto a. **Only two cardinalities matter.** Queries aggregate by `label_0` or per @@ -340,32 +278,33 @@ cardinality knobs: - `s`: series per group, the product of the cardinalities of all non-grouping labels. -`label_1`'s cardinality equals `r` and is covered by the replica dimension. Giving every label the same cardinality `c` would tie the knobs together (`C = c`, `s = c^(L−1)` for `L` labels). It was rejected: `s` explodes (c = 1e3 with three labels gives 1e9 series, far beyond the measured K ≤ 1e7), and the effects of more groups and of more series per group could no longer be told -apart. No template groups by `label_1`, keeping the template set as given. +apart. -**How a template becomes sketch input.** A frequency or top-k sketch holds the -groups as keys inside one sketch. A quantile sketch is one sketch per group. +**How a template becomes sketch input.** Every deployment keeps one instance +per group of its grouping labels `G` (`card(G)` instances per window), as in +#145's cost model. -| Template kind | Sketch instances per deployment | Keys per sketch | Events per sketch per window | +| Template kind | Instances per window | Kind of instance | Items per instance per window | |---|---|---|---| -| `sum`/`topk by (label_0)` (1, 2, 7, 8, 9) | 1 | `C` | `100 · C · s · S` | -| `quantile by (q, label_0)` (3) | `C` | — | `100 · s · S` | -| per-series `sum_over_time`/`rate` (4, 6) | 1 | `C · s` | `100 · C · s · S` | -| per-series `quantile_over_time` (5, 10) | `C · s` | — | `100 · S` | - -**Example.** `C = 3` (`label_0` ∈ {a, b, c}), `s = 2` (`instance` ∈ {i1, i2}), -`r = 1`: six series, `data{label_0="a", instance="i1"}` through +| `sum`/`rate by (label_0)` (1, 7, 8) | `C` | exact accumulator | `100 · s · S` | +| `topk by (label_0)` (2, 9) | `C` | top-k sketch over the group's `s` series | `100 · s · S` | +| `quantile by (q, label_0)` (3) | `C` | quantile sketch | `100 · s · S` | +| per-series `sum_over_time`/`rate` (4, 6) | `C · s` | exact accumulator | `100 · S` | +| per-series `quantile_over_time` (5, 10) | `C · s` | quantile sketch | `100 · S` | + +**Example.** `C = 3` (`label_0` ∈ {a, b, c}), `s = 2` (`instance` ∈ {i1, i2}): +six series, `data{label_0="a", instance="i1"}` through `data{label_0="c", instance="i2"}`, emitting 600 samples/s in total. -- `sum by (label_0) (data)`, `S = 1 s`: one CMS with keys a, b, c; each second - it absorbs 600 weighted updates, 200 per key. +- `sum by (label_0) (data)`, `S = 1 s`: three exact sums, one per group, each + absorbing 200 values per second. - `quantile by (0.99, label_0) (data)`: three KLL sketches, one per group, each absorbing 200 values per second. -- `sum_over_time(data[1m])`: one CMS with six keys, one per series, absorbing - 36,000 updates per minute. +- `sum_over_time(data[1m])`: six exact sums, one per series, 6,000 values each + per minute. - `quantile_over_time(0.99, data[1m])`: six KLL sketches, 6,000 values each per minute. @@ -375,61 +314,71 @@ semantics would read only each series' latest sample. Aggregating the whole second is what a sketch maintained over a 1-second window answers. The difference is noted wherever spatial results are reported. -Saturation curves cover K ∈ {1e1, …, 1e7} (#130, #140), all measured and -never interpolated. Per-series templates have K = `C · s`; points above 1e7 are -clamped to the largest measured K and flagged as extrapolated. Points whose -events per sketch fall below `N_sat` are flagged as unsaturated. +Items per instance, `100 · s · S` or `100 · S`, generally differ from the +benchmark's. How costs and accuracy are read at that size is §6 +"Benchmark input". #### Workload grid -Each dimension has a default (bold). A workload fixes every dimension. +Each dimension has a default (bold). A workload fixes every dimension. The grid +keeps only what changes the comparison with AutoSketch. | Dimension | Values | What it varies | |---|---|---| -| Query mix (templates) | **all 10**; spatial only {1, 2, 3}; temporal only {4–10}; frequency only {1, 4, 6, 7, 8}; quantile only {3, 5, 10}; top-k only {2, 9} | Summary types, and how much can be shared | -| Number of RQEs: replicas `r` | **1**, 2, 4, 8, 16, 32, 64 | Each replica adds a filter `{label_1="v_i"}` selecting a disjoint subset of series, so it reads its own streams. Total RQEs = `r · (n_spatial + n_temporal·|W|)` | -| Replica mode | disjoint (each replica filters `{label_1="v_i"}` and reads its own series; planning-time control); **shared** (every replica reads the same stream with a seeded random subset of 3 windows, 3 quantiles from {0.5, 0.75, 0.9, 0.95, 0.99}, and `T` from {10 s, 1 m, 5 m}; identical RQEs are deduplicated): many users or dashboards over the same metrics | How the sharing benefit grows with the number of queries | -| Template set | the 10 templates; **dashboard** (above) | Workload realism | -| Lookback window set `W` | {1h}; {1m, 1h}; {1m, 10m, 1h}; **{1m, 10m, 1h, 6h, 24h}** | Overlapping windows over the same stream: the main sharing opportunity | -| Temporal repeat interval `T` | 10 s, **1 m**, 5 m | Recurrence: query and merge work vs. ingest | -| Groups `C = card(label_0)` | 1e1, 1e2, **1e3**, 1e4, 1e5, 1e6 | Keys per frequency sketch; quantile sketches per `by (label_0)` deployment | -| Series per group `s` (product of the non-grouping label cardinalities) | 1, 10, **100**, 1000 | Events per group for spatial templates; keys and sketch instances for per-series templates | -| Key skew θ / value tail a | θ ∈ {0, 0.5, **1.0**, 1.5, 2.0}; a ∈ {1.1, **2**, 3} | Sketch size needed for the accuracy target | -| Accuracy target (strictness) | loose, **default**, strict | §5 | +| Template set | **dashboard**; the 10 templates (with `W` = {1m, 10m, 1h, 6h, 24h}, `T` = 1 m) | Workload realism, and which capabilities appear | +| Replicas `r` | **1**, 8, 64 | Every replica reads the same stream with a seeded random subset of 3 windows, 3 quantiles from {0.5, 0.75, 0.9, 0.95, 0.99} and `T` from {10 s, 1 m, 5 m}; identical RQEs are deduplicated. Many users or dashboards over the same metrics: how the sharing benefit and planning time grow with the number of RQEs | +| Groups `C = card(label_0)` | 1e2, **1e3**, 1e4 | Instances per deployment and items per instance | +| Accuracy target (strictness) | loose, **default**, strict | What AutoSketch optimizes for (§5) | | Latency SLA | the §5 grid, **no limit** | §5 | -The full Cartesian product is too large. The sweep is: -1. **Default workload.** Every baseline, every SLA, every cost model. -2. **One dimension at a time.** Vary each dimension over its values with the - others at their defaults. -3. **Two interactions:** - - `r × s`: scale, with planning time against total RQEs; - - `W × card(label_0)`: sharing benefit against state size. - -Every workload runs every baseline (ASAP, AutoSketch-Adapted, -PerQuery-CostAware) and is priced under model A and model B for -each family (§4). Per (workload, baseline, cost model, SLA), report: -- $/hour, total CPU-seconds, peak CPU, retained GiB; +Fixed: `s = 100`, θ = 1.0, a = 2. + +The sweep is the default workload, then each dimension varied alone with the +others at their defaults. Every workload runs every baseline (ASAP, +AutoSketch-Adapted, PerQuery-CostAware) at both weight settings (§4). Per +(workload, baseline, weights, SLA), report: +- the objective, mean CPU (vCPU) and memory (GiB), and memory per phase; - max and median estimated latency, and SLA violations; -- active deployments and sketch instances; +- active deployments and instances; - planning time (AutoSketch: search plus charged benchmark time); -- RQEs excluded by the SLA or unservable, and counts of unsaturated or - extrapolated lookups. +- RQEs excluded by the SLA or unservable. + +Replicas with disjoint series (each replica filtering `{label_1="v_i"}`) were +dropped: the planner rejects spatial filters, and the shared mode is the case +that shows sharing. + +#### Planner sensitivity (not part of this comparison) + +These dimensions describe the planner, not its gap to AutoSketch. They belong +in the planner's micro-benchmarks: +- query mix (spatial only, temporal only, one capability at a time); +- lookback window set `W` ({1h} to {1m, 10m, 1h, 6h, 24h}); +- repeat interval `T` (10 s, 1 m, 5 m); +- series per group `s` (1 to 1000); +- key skew θ and value tail a; +- interactions `r × s` and `W × C`. ### Benchmark input -Configurations come from the saturation study's grid (#130, #140): CMS, -Count Sketch and CMS-heap top-k with rows ∈ {3, 5} and cols ∈ {256, 1024, -4096, 16384} (rows = 5 costs scaled from rows = 3), KLL k ∈ {50, 200, 800}, -DDSketch α ∈ {0.005, 0.01, 0.02, 0.05}, HLL `lg_k` ∈ {12, 14, 16}. +Every method reads the cost table from sketch-bench +`scripts/export_rqe_optimizer_costs.sh`, the same table the planner reads. +Deployable rows: exact sum, min, max and increase; KLL k ∈ {200, 500}; HLL +`lg_k` ∈ {12, 14}; CMS-heap top-k with rows ∈ {3, 5}, cols = 2048. DDSketch, +CountSketch-heap and UnivMon are measured but not deployable by default. + +**Exact accumulators** have no saturation point: their answer is exact at any +size. They are still benchmarked for cost, on the grouped column specs +(200,000 rows, about 9,900 groups), and priced per group. The sweep over items +per instance (#157) includes them, so their per-group cost is read at the +workload's size like every other row's. #### Data parameters, shared by both methods The benchmark inputs are generated from data parameters fit over the **whole measured dataset**, not from a short sample or the average. These are key skew `θ`, distinct keys `K` per window, value tail index `a`, and per-label-set -`λ` and `card`. Use the worst case across the dataset; for example, size CMS -and top-k from the lower `θ` bound. Longer samples expose worse cases (sketch-bench +`λ` and `card`. Use the worst case across the dataset; for example, size top-k +from the lower `θ` bound. Longer samples expose worse cases (sketch-bench `docs/saturation_conclusions.md`, conclusions 1–5). Fitting follows ASAPQuery #746 and sketch-bench `scripts/recommend_config.py`. @@ -443,57 +392,32 @@ time. We don't need them: worst-case fits over every window of the whole dataset already cover that variation. AutoSketch's benchmark uses the same inputs, which matches the paper: it lets users "use their own trace". -A query with lookback `S` on label set `ℓ` reads about +A window of length `S` on grouping labels `G` holds about ```text -n(S, ℓ) = λ(ℓ) · S / card(ℓ) events per group +n(S, G) = λ · S / card(G) items per instance ``` -#### ASAPQuery: saturated values, because it merges - -ASAP may answer a query by merging `m = S/x` smaller-window sketches into one. -Each of them holds only `n(x, ℓ)` events, which may be below the length at -which its error has settled. So ASAP reads accuracy and cost at saturation, as -measured in sketch-bench's saturation study (#130): - -- Past `N_sat`, error depends on the config alone. -- Per-item insert, merge and query CPU are flat in `N`. -- Memory is fixed by the config; DDSketch and KLL grow only with `ln N`. - -Each (config, dataset parameters) point therefore contributes its error -plateau and its costs at `N_sat`. `N_sat` is taken over the whole measured -dataset: it is the saturation length at the dataset's worst-case parameters -above. A lookup is valid when `n(S, ℓ) ≥ N_sat`. - -- **CMS, Count Sketch, HLL and DDSketch merge exactly.** The merged sketch - equals one sketch over all `n(S, ℓ)` events, so the saturated single-sketch - value applies however small each pane is. -- **KLL and top-k do not.** Merging raises KLL's error, by 1.0–1.1× at - k = 50/200 and 1.12–1.32× at k = 800, and up to 3–4× at small `N`. Top-k - loses up to 40% precision at large `K`. For these sketches, a deployment - with `m > 1` uses the saturated value from the merged curves at `m` shards - (sketch-bench #131). -- If `n(S, ℓ) < N_sat`, the query's window never saturates. ASAP uses the curve's - value at the smallest measured checkpoint `≥ n(S, ℓ)`. -- With uniform keys and large `K`, CMS, Count Sketch and top-k do not saturate - by 1e9 events. Their `N_sat` is a lower bound; flag these points in the - results. - -#### AutoSketch: no saturation requirement - -AutoSketch never merges: it keeps one sketch per query window. In the paper it -benchmarks a config on its workloads and accepts it if the accuracy intent -holds, with no notion of `N_sat`. AutoSketch-Adapted therefore reads the -measured value at its own input size `n(S, ℓ)`, using the dataset-derived -inputs above. - -The windows of a repeating query see different data each time. AutoSketch does -not re-tune for this: it configures once, before deployment, against these -inputs (§9 Q2). +#### Reading the table at the workload's size + +Each row records the size it was measured at (`measured_at`, #155). Until the +size sweep (#157) lands, every row is read at its measured size and the +points whose instance size differs are flagged. With the sweep: + +- **ASAP** reads each deployment's row at its own instance size, + `λ · x / card(G)` items per window. When it merges `m = S/x` windows, the + accuracy check uses the measured merge counts bracketing `m` (#154); merge + error is not monotone in `m`, so both must pass. +- **AutoSketch** never merges (`x = S`). It reads the row at its window's size, + `λ · S / card(G)`, and accepts a config if the target holds there, as the + paper's benchmark-then-accept loop does. It configures once, before + deployment (§9 Q2). +- **Exact accumulators** meet every target at every size; only their cost is + read at the size. ## 7. Metrics and figures -Reported per (workload, method, cost model and machine family, latency SLA), +Reported per (workload, method, weight setting, latency SLA), median of repeated runs for timings: - **Planning time.** Reported in two parts, because the two planners spend @@ -509,47 +433,43 @@ median of repeated runs for timings: also benchmarks a fixed-size representative workload, not the query window's full data; - a **paper-rate estimate**, 60 s per distinct probe. - - *ASAP's one-time profiling:* the wall time of the sketch-bench saturation - runs that produced its tables (#130, #131, #140). It is shared by all RQEs + - *ASAP's one-time profiling:* the wall time of the sketch-bench cost export + that produced its table. It is shared by all RQEs and reusable across workloads, so it is reported once, plus amortized per RQE served. - The paper's figure shows search + benchmark per method, stacked. -- **Total cost** ($/hour) under **model A** and under **model B** for each - family, with its inputs: total CPU-seconds, peak CPU, retained GiB, and - which resource binds `n_f` in model B. Baselines are compared under each - model separately, each normalized to ASAP under the same model. +- **Objective** at each weight setting (§4), with its inputs: mean CPU (vCPU) + and memory (GiB), each split by phase. Baselines are normalized to ASAP at + the same weights. - **Estimated query latency and latency SLA violations** per method. - - Estimated latency per RQE: `card(ℓ) · (query_cpu + (S/x − 1) · merge_cpu)` (§4). Report its maximum and median over the RQEs, plus the per-RQE values in the raw output. + - Estimated latency per RQE: `card(G) · (c_qry + (S/x − 1) · c_mrg)` (§4). Report its maximum and median over the RQEs, plus the per-RQE values in the raw output. - SLA violations: the number of RQEs whose estimated latency exceeds the SLA. Only AutoSketch-Adapted can have any, since the other methods are constrained. - **Estimated accuracy** per RQE: every method meets its target under its own - lookup rule (§6, "Benchmark input"). + reading rule (§6, "Benchmark input"). - Active deployments and total sketch instances. Figures: -1. Synthetic workload, cost vs. achieved max estimated latency, one panel per - cost model (main paper figure). +1. Synthetic workload, objective vs. achieved max estimated latency, one panel + per weight setting (main paper figure). 2. Planning time vs. number of RQEs (synthetic, replica dimension), log–log. -3. Cost vs. each workload-grid dimension (synthetic, one dimension at a time). -4. Cost vs. absolute latency SLA (synthetic default workload, `traces`). -5. Baselines under the two cost models: paired bars per workload, model A - next to model B, each normalized to ASAP. +3. Objective vs. each workload-grid dimension (synthetic, one dimension at a + time). +4. Objective vs. absolute latency SLA (synthetic default workload, `traces`). ## 8. Who implements what, in which PR -| # | Repo / PR | Scope | Status (2026-10-05) | +| # | Repo / PR | Scope | Status (2026-10-06) | | --- | --- | --- | --- | | this | ASAPQuery #777 | This plan | Draft, updated as decisions change | -| — | sketch-bench #130, #131 | Saturation curves at K ∈ {1e3, 1e5, 1e7}; accuracy after merging `m` shards | Merged | -| 1 | sketch-bench #137 | Retained memory, EC2 pricing, `milp::minimize_cost`, solver scaling | Merged | +| — | sketch-bench #144, #145 | Exact accumulators and top-k families; per-phase cost model and weighted objective | Merged | +| — | sketch-bench #151, #152, #154, #155 | Cost-table fixes, value range, accuracy after merging, `measured_at` (#147) | Merged | +| — | sketch-bench #157 | Sweep items per instance; re-export the cost table | Open | +| — | sketch-bench #130, #131 | Saturation curves; accuracy after merging `m` shards | Merged; background | +| 1 | sketch-bench #137 | Retained memory, EC2 pricing, `milp::minimize_cost` | Merged; superseded by #145 | | 3 | sketch-bench #135 | AutoSketch-Adapted (Algorithm 4), aligned with the paper's EXAMINE rule and seeding | Merged | | 2 | sketch-bench #136 | Evaluation table for the trace workloads | Merged | -| 4 | sketch-bench #138 | Runner and absolute SLA for `traces` (rebased on main; `example`/`scaling` removed) | Open | -| — | sketch-bench #140 | Saturation curves at K ∈ {1e1, 1e2, 1e4, 1e6} for the synthetic workload | Draft; 62 cost points re-measured serially after the parallel run failed the 10% consistency check | -| — | sketch-bench #139 | Synthetic workload: 10 templates, PerQuery-CostAware bound by the SLA | Open | -| — | sketch-bench #141 | Two cost models (model A; model B with aligned starts), strictness levels, workload-grid driver, `traces` results | Open; synthetic sweep and scale study running | - -Merge order: #138 → #140 → #139 → #141. +| 4 | sketch-bench #138, #139, #140, #141 | Runner, synthetic workload, saturation at K = 1e1–1e6, two cost models and grid driver | Open; to be reworked onto #145's objective, the cost table and the reduced grid. #140 is no longer needed for the comparison | ## 9. Decisions @@ -578,19 +498,21 @@ lookup-only time would understate AutoSketch's planning cost. **Q4. AutoSketch window adapter.** `x = S`, `y = T` (or `gcd(S, T)`): one sliding sketch per query. -**Q5. Memory model.** Retained state, as in §4. +**Q5. Memory model.** Per phase, as in #145 (§4): open windows, one merge +accumulator per group, query output, and closed windows. -**Q6. Machine-family cost.** The fractional-instance `max` model in §4. +**Q6. Cost.** `w_cpu · CPU + w_mem · memory`: CPU only first, then Fargate's +per-vCPU and per-GB prices (§4). Machine-family and peak-provisioned pricing +were dropped (2026-10-06). ## 10. Known limitations - **Merged accuracy comes from shard-merge measurements, not from replaying - the plans.** It is exact by construction for CMS, Count Sketch, HLL and - DDSketch, and taken from #131 for KLL and top-k. #131 covers `N ≤ 1e7` and - `m ≤ 64`. A deployment needing `m > 64` (e.g. a 1-day lookback over 1-minute - windows, `m = 1440`) is outside the measured range: it is ineligible for KLL - and top-k, and allowed for the exact-merge sketches. Replay the synthetic default workload's chosen plans in sketch-bench once to - confirm the lookups. + the plans.** The cost table measures it at `m` ∈ {4, 16, 64, 256, 1024} over + one benchmark stream (#154). A deployment needing more (e.g. a 1-day + lookback over 1-minute windows, `m = 1440`) reads the 1024 measurement. + Replay the synthetic default workload's chosen plans in sketch-bench once to + confirm. - Costs and latencies are estimates from per-operation measurements, not end-to-end executions. The execution-based comparison is ASAPQuery-backend #545/#547. @@ -598,7 +520,7 @@ sliding sketch per query. candidate set (`candidates.rs` generates every divisor of `S` as a window and `gcd(x, T)` as a slide). So when its plan meets the latency bounds and its configurations also pass ASAP's lookup rule (the two rules differ on - saturation, §6), the ASAP MILP can choose the same deployments and pay for + size, §6), the ASAP MILP can choose the same deployments and pay for shared ones once; ASAP's cost is never higher. The sanity check reports any case where this does not hold. The result to report is the size of the gap and where it comes from, not that a gap exists. From 6d215dbc6ca5ab460e3b22683b7de58758670928 Mon Sep 17 00:00:00 2001 From: zz_y Date: Tue, 6 Oct 2026 21:42:23 +0000 Subject: [PATCH 31/39] docs(planner): use the cost table's Zipf s=1.1 data for the synthetic workload MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Top-k, KLL and DDSketch are measured on Zipf s=1.1 over 100k keys (quantiles on the Zipf ranks). The synthetic workload now uses that distribution instead of Zipf θ=1.0 and Pareto a=2, so the cost table needs no separate run. Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 18 +++++++++++------- 1 file changed, 11 insertions(+), 7 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 551172f2..32c94756 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -179,7 +179,7 @@ grid. | ID | Description | Purpose | | --- | --- | --- | -| `synthetic` | Synthetic PromQL workload: the 10 queries in "Synthetic workload" below, over Zipf/Pareto data. Main figure. | Cost–latency trade-off across data and requirements | +| `synthetic` | Synthetic PromQL workload: the 10 queries in "Synthetic workload" below, over Zipf data. Main figure. | Cost–latency trade-off across data and requirements | | `traces` | Real-trace RQEs, one workload per dataset: Alibaba 2022, BOOM and Google 2011. Taken from `asap-tools/dataset-analysis/results/skew_summary.csv` ([#746](https://github.com/ProjectASAP/ASAPQuery/pull/746)): each row's `range_s` is `S` and its `step_s` is `T`. Data parameters and accuracy targets are fit over each whole trace. | Appendix: real-trace results | ### Synthetic workload @@ -265,9 +265,12 @@ they add RQEs, not series. Every series emits one sample every 10 ms Sample values: - **Sum and increase:** exact, so the value distribution does not affect cost or accuracy. -- **Top-k:** the value is the weight; the series' total weights within a group - follow Zipf θ. -- **Quantiles:** values are drawn from Pareto a. +- **Top-k and quantiles:** the data the cost table is measured on: Zipf + s = 1.1 over a population of 100,000 keys (sketch-bench + `export_rqe_optimizer_costs.sh`). Top-k ranks the keys by weight; quantile + sketches (KLL, DDSketch) are measured on the Zipf ranks as values. The + workload uses the same distribution, so the table needs no separate run for + it. **Only two cardinalities matter.** Queries aggregate by `label_0` or per series, never by another label. So any other label (a second instance-like @@ -331,7 +334,7 @@ keeps only what changes the comparison with AutoSketch. | Accuracy target (strictness) | loose, **default**, strict | What AutoSketch optimizes for (§5) | | Latency SLA | the §5 grid, **no limit** | §5 | -Fixed: `s = 100`, θ = 1.0, a = 2. +Fixed: `s = 100`; data Zipf s = 1.1 over 100,000 keys (§6 "Data model"). The sweep is the default workload, then each dimension varied alone with the others at their defaults. Every workload runs every baseline (ASAP, @@ -355,7 +358,7 @@ in the planner's micro-benchmarks: - lookback window set `W` ({1h} to {1m, 10m, 1h, 6h, 24h}); - repeat interval `T` (10 s, 1 m, 5 m); - series per group `s` (1 to 1000); -- key skew θ and value tail a; +- data distribution (key skew, value tail); - interactions `r × s` and `W × C`. ### Benchmark input @@ -384,7 +387,8 @@ from the lower `θ` bound. Longer samples expose worse cases (sketch-bench - **`traces` gives the appendix results.** Its parameters are fit on the full trace. -- **`synthetic`** uses its generator's parameters (§6 "Data model"). +- **`synthetic`** uses the cost table's own data, Zipf s = 1.1 over 100,000 + keys (§6 "Data model"). Each config is benchmarked at these worst-case parameters. AutoSketch §5.2 injects random traffic bursts into synthetic workloads to cover variation over From 6fafbb2fdad1bf7f913479abb02de969d305e29d Mon Sep 17 00:00:00 2001 From: zz_y Date: Tue, 6 Oct 2026 22:30:36 +0000 Subject: [PATCH 32/39] docs(planner): fix the synthetic data model and list every query - data: C = 1e4 series (label_0, the top-k key), J = 10 jobs, 200 samples/s per series, Zipf s=1.1 over 10,000 keys (the cost table's data at CARDINALITY=10000); fixed, so the comparison does not sweep it. - windows W = {15m, 1h, 6h, 24h}: every per-series summary sees >= 1e5 items; per-summary input sizes listed per grouping. - queries: grouped templates use by (job); top-k is topk(k, sum by (label_0) (...)) with k in {100, 200, 300}, served by one heap-K deployment for k <= K; quantile syntax fixed. Both template sets are enumerated query by query (all checked with promql-parser). - grid: C removed. Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 324 ++++++++++++++--------- 1 file changed, 206 insertions(+), 118 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 32c94756..87854acb 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -91,7 +91,7 @@ has lookback `S` and repeat interval `T`. Per-instance costs are measured | --- | --- | --- | | Ingest | `λ · (x/y) · c_ins` | `card(G) · m · x/y` (open windows) | | Merge | `card(G) · (S/x − 1) · c_mrg / T` | `card(G) · m` (one accumulator per group), 0 when `S = x` | -| Query | `card(G) · c_qry / T` | `card(G) · 8 B`; top-k `card(G) · 32 · 16 B` | +| Query | `card(G) · c_qry / T` | `card(G) · 8 B`; top-k `k · 16 B` | | Storage | 0 | `card(G) · m · ((max S − x)/y + 1)` (closed windows) | Ingest and storage are paid once per active deployment; merge and query once @@ -188,29 +188,70 @@ The synthetic workload is the paper's main experiment. It builds many workloads from a fixed set of PromQL query templates by sweeping the workload dimensions below, and runs every planning baseline on each workload. +#### Data model and scale + +The data model is fixed: the comparison varies the queries and requirements, +not the data. There is one metric, `data`, with two labels: + +| Label | Cardinality | Role | +|---|---|---| +| `label_0` | **C = 1e4** | Identifies the series (one series per value). The top-k key | +| `job` | **J = 10** | Coarse grouping: every series belongs to one job, so each job has `C/J` = 1,000 series | + +| Quantity | Value | +|---|---| +| Series | C = 10,000 | +| Samples per series | 200 per second (one every 5 ms) | +| Stream rate `λ` | 2e6 samples/s; 2e5 per job | +| Values | Zipf s = 1.1 over 10,000 keys: the cost table's data (sketch-bench `export_rqe_optimizer_costs.sh` with `CARDINALITY=10000`). Quantile sketches see the Zipf ranks as values; a top-k key's total over a window follows the same Zipf law | +| Lookback windows `W` | {15m, 1h, 6h, 24h}: at least 15 minutes, so every per-series summary sees at least 1e5 items | +| Repeat interval `T` | 1 m for temporal templates; spatial templates read and repeat every 1 s | + +**Summaries and their input size.** Every deployment keeps one summary per +group of its grouping labels (one per window). The number of items one summary +ingests decides whether a sketch is worth it: + +| Grouping | Summaries per window | Items per summary | Kind | +|---|---|---|---| +| per series, temporal (templates 4, 5, 6, 10; D1, D5) | C = 1e4 | `200 · S`: 1.8e5 (15m) to 1.7e7 (24h) | exact accumulator or quantile sketch | +| `by (job)`, spatial, `S` = 1 s (1, 3; D2) | J = 10 | `2e5 · S` = 2e5 | exact accumulator or quantile sketch | +| `by (job)`, temporal (7, 8; D3) | J = 10 | 1.8e8 to 1.7e10 | exact accumulator | +| top-k over `label_0` (2, 9; D4) | 1 | `λ · S`: 1.8e9 to 1.7e11, over 1e4 keys | top-k sketch (CMS-heap) | + +The cost table is measured at 1e5 to 1e8 items per instance (sketch-bench +#157). Exact accumulators keep constant state and a constant per-item cost, so +their size does not matter. Every quantile sketch falls inside the measured +range. Top-k sketches exceed it and read the 1e8 measurement: CMS-heap's +per-item CPU and memory do not depend on `N` (sketch-bench #130), and Zipf +s = 1.1 keys have saturated by then. + +**Modeling choice for spatial templates.** A spatial template evaluates every +1 s over every sample of the last second (`S = T = 1 s`). PromQL's instant +semantics would read only each series' latest sample. Aggregating the whole +second is what a summary maintained over a 1-second window answers. + #### Query templates These exercise every capability the templates need: sum, rate/increase, top-k -and quantile, spatial and temporal aggregation, and a binary operator. Each spatial -template repeats every 1 s and reads the last second (`S = T = 1 s`). Each -temporal template repeats every `T` (default 1 m) and reads `S = T_range`, one -RQE per lookback window in the window set `W` (§ workload grid). +and quantile, spatial and temporal aggregation, and a binary operator. Each +template expands into one RQE per window `w ∈ W`, per quantile `q` and per +`k ∈ {100, 200, 300}`. -| # | PromQL | Capability | Grouping | RQEs per replica | +| # | PromQL | Capability | Grouping | RQEs | |---|---|---|---|---| -| 1 | `sum by (label_0) (data)` | SumOrCount | `label_0` | 1 | -| 2 | `topk by (3, label_0) (data)` | TopK | `label_0` | 1 | -| 3 | `quantile by (q, label_0) (data)`, q ∈ {0.5, 0.75, 0.9, 0.95, 0.99} | Quantile | per `label_0` group | 5 | -| 4 | `sum_over_time(data[T])` | SumOrCount | per series | \|W\| | -| 5 | `quantile_over_time(q, data[T])`, same five q | Quantile | per series | 5·\|W\| | -| 6 | `rate(data[T])` | RateOrIncrease | per series | \|W\| | -| 7 | `sum by (label_0) (rate(data[T]))` | RateOrIncrease | `label_0` | \|W\| | -| 8 | `sum by (label_0) (sum_over_time(data[T]))` | SumOrCount | `label_0` | \|W\| | -| 9 | `topk by (3, label_0) (rate(data[T]))` | TopK over per-series increases | `label_0` | \|W\| | -| 10 | `quantile_over_time(0.9, data[T]) / quantile_over_time(0.5, data[T])` | Two Quantile RQEs | per series | 2·\|W\| | - -With all ten templates, one replica has `7 + 12·|W|` RQEs: 67 for the default -five windows. +| 1 | `sum by (job) (data)` | SumOrCount | `job` | 1 | +| 2 | `topk(k, sum by (label_0) (sum_over_time(data[w])))` | TopK | none; keys `label_0` | 3·\|W\| = 12 | +| 3 | `quantile by (job) (q, data)`, q ∈ {0.5, 0.75, 0.9, 0.95, 0.99} | Quantile | `job` | 5 | +| 4 | `sum_over_time(data[w])` | SumOrCount | per series | 4 | +| 5 | `quantile_over_time(q, data[w])`, same five q | Quantile | per series | 20 | +| 6 | `rate(data[w])` | RateOrIncrease | per series | 4 | +| 7 | `sum by (job) (rate(data[w]))` | RateOrIncrease | `job` | 4 | +| 8 | `sum by (job) (sum_over_time(data[w]))` | SumOrCount | `job` | 4 | +| 9 | `topk(k, sum by (label_0) (rate(data[w])))` | TopK over per-series increases | none; keys `label_0` | 12 | +| 10 | `quantile_over_time(0.9, data[w]) / quantile_over_time(0.5, data[w])` | Two Quantile RQEs | per series | 8 | + +74 RQEs per replica. Template 10's operands are the same RQEs as template 5's +q = 0.9 and q = 0.5, so 66 are distinct. Mapping notes: - Capabilities are `rqe-optimizer`'s (#144). Sum and rate/increase are their @@ -220,106 +261,149 @@ Mapping notes: - An exact accumulator has one configuration, so AutoSketch has nothing to search for sum and increase RQEs. There, ASAP differs from it only by window choice and sharing. +- Top-k ranks each `label_0` key's total over the window: the heavy hitters + among 1e4 keys. A top-k sketch with heap capacity `K` serves every RQE with + `k ≤ K` on the same stream and window, so ASAP can serve k = 100, 200 and 300 + from one heap-300 deployment; AutoSketch configures each RQE with its own + heap. - The quantiles of one template, and the two operands of template 10, read the same stream. ASAP can serve them from one deployment; AutoSketch gets one deployment per RQE. - Templates come from the planner's supported query classes: `SpatialAgg` - (`count`/`sum`/`quantile`/`topk by`), `TemporalAgg` + (`count`/`sum`/`quantile`/`topk`), `TemporalAgg` (`count_over_time`/`sum_over_time`/`quantile_over_time`/`increase`/`rate`), - `TemporalAgg SpatialAgg*`, and `AnyAgg AnyAgg`. + `TemporalAgg SpatialAgg*`, and `AnyAgg AnyAgg`. Every query below + is checked with `promql-parser`. + +
+Every query of the 10-template set (70 queries, 74 RQEs) + +| # | Template | PromQL | Capability | Grouping | `S` | `T` | RQEs | +|---|---|---|---|---|---|---|---| +| 1 | 1 | `sum by (job) (data)` | SumOrCount | job | 1s | 1s | 1 | +| 2 | 2 | `topk(100, sum by (label_0) (sum_over_time(data[15m])))` | TopK | — (keys: label_0) | 15m | 1m | 1 | +| 3 | 2 | `topk(100, sum by (label_0) (sum_over_time(data[1h])))` | TopK | — (keys: label_0) | 1h | 1m | 1 | +| 4 | 2 | `topk(100, sum by (label_0) (sum_over_time(data[6h])))` | TopK | — (keys: label_0) | 6h | 1m | 1 | +| 5 | 2 | `topk(100, sum by (label_0) (sum_over_time(data[24h])))` | TopK | — (keys: label_0) | 24h | 1m | 1 | +| 6 | 2 | `topk(200, sum by (label_0) (sum_over_time(data[15m])))` | TopK | — (keys: label_0) | 15m | 1m | 1 | +| 7 | 2 | `topk(200, sum by (label_0) (sum_over_time(data[1h])))` | TopK | — (keys: label_0) | 1h | 1m | 1 | +| 8 | 2 | `topk(200, sum by (label_0) (sum_over_time(data[6h])))` | TopK | — (keys: label_0) | 6h | 1m | 1 | +| 9 | 2 | `topk(200, sum by (label_0) (sum_over_time(data[24h])))` | TopK | — (keys: label_0) | 24h | 1m | 1 | +| 10 | 2 | `topk(300, sum by (label_0) (sum_over_time(data[15m])))` | TopK | — (keys: label_0) | 15m | 1m | 1 | +| 11 | 2 | `topk(300, sum by (label_0) (sum_over_time(data[1h])))` | TopK | — (keys: label_0) | 1h | 1m | 1 | +| 12 | 2 | `topk(300, sum by (label_0) (sum_over_time(data[6h])))` | TopK | — (keys: label_0) | 6h | 1m | 1 | +| 13 | 2 | `topk(300, sum by (label_0) (sum_over_time(data[24h])))` | TopK | — (keys: label_0) | 24h | 1m | 1 | +| 14 | 3 | `quantile by (job) (0.5, data)` | Quantile | job | 1s | 1s | 1 | +| 15 | 3 | `quantile by (job) (0.75, data)` | Quantile | job | 1s | 1s | 1 | +| 16 | 3 | `quantile by (job) (0.9, data)` | Quantile | job | 1s | 1s | 1 | +| 17 | 3 | `quantile by (job) (0.95, data)` | Quantile | job | 1s | 1s | 1 | +| 18 | 3 | `quantile by (job) (0.99, data)` | Quantile | job | 1s | 1s | 1 | +| 19 | 4 | `sum_over_time(data[15m])` | SumOrCount | series | 15m | 1m | 1 | +| 20 | 4 | `sum_over_time(data[1h])` | SumOrCount | series | 1h | 1m | 1 | +| 21 | 4 | `sum_over_time(data[6h])` | SumOrCount | series | 6h | 1m | 1 | +| 22 | 4 | `sum_over_time(data[24h])` | SumOrCount | series | 24h | 1m | 1 | +| 23 | 5 | `quantile_over_time(0.5, data[15m])` | Quantile | series | 15m | 1m | 1 | +| 24 | 5 | `quantile_over_time(0.5, data[1h])` | Quantile | series | 1h | 1m | 1 | +| 25 | 5 | `quantile_over_time(0.5, data[6h])` | Quantile | series | 6h | 1m | 1 | +| 26 | 5 | `quantile_over_time(0.5, data[24h])` | Quantile | series | 24h | 1m | 1 | +| 27 | 5 | `quantile_over_time(0.75, data[15m])` | Quantile | series | 15m | 1m | 1 | +| 28 | 5 | `quantile_over_time(0.75, data[1h])` | Quantile | series | 1h | 1m | 1 | +| 29 | 5 | `quantile_over_time(0.75, data[6h])` | Quantile | series | 6h | 1m | 1 | +| 30 | 5 | `quantile_over_time(0.75, data[24h])` | Quantile | series | 24h | 1m | 1 | +| 31 | 5 | `quantile_over_time(0.9, data[15m])` | Quantile | series | 15m | 1m | 1 | +| 32 | 5 | `quantile_over_time(0.9, data[1h])` | Quantile | series | 1h | 1m | 1 | +| 33 | 5 | `quantile_over_time(0.9, data[6h])` | Quantile | series | 6h | 1m | 1 | +| 34 | 5 | `quantile_over_time(0.9, data[24h])` | Quantile | series | 24h | 1m | 1 | +| 35 | 5 | `quantile_over_time(0.95, data[15m])` | Quantile | series | 15m | 1m | 1 | +| 36 | 5 | `quantile_over_time(0.95, data[1h])` | Quantile | series | 1h | 1m | 1 | +| 37 | 5 | `quantile_over_time(0.95, data[6h])` | Quantile | series | 6h | 1m | 1 | +| 38 | 5 | `quantile_over_time(0.95, data[24h])` | Quantile | series | 24h | 1m | 1 | +| 39 | 5 | `quantile_over_time(0.99, data[15m])` | Quantile | series | 15m | 1m | 1 | +| 40 | 5 | `quantile_over_time(0.99, data[1h])` | Quantile | series | 1h | 1m | 1 | +| 41 | 5 | `quantile_over_time(0.99, data[6h])` | Quantile | series | 6h | 1m | 1 | +| 42 | 5 | `quantile_over_time(0.99, data[24h])` | Quantile | series | 24h | 1m | 1 | +| 43 | 6 | `rate(data[15m])` | RateOrIncrease | series | 15m | 1m | 1 | +| 44 | 6 | `rate(data[1h])` | RateOrIncrease | series | 1h | 1m | 1 | +| 45 | 6 | `rate(data[6h])` | RateOrIncrease | series | 6h | 1m | 1 | +| 46 | 6 | `rate(data[24h])` | RateOrIncrease | series | 24h | 1m | 1 | +| 47 | 7 | `sum by (job) (rate(data[15m]))` | RateOrIncrease | job | 15m | 1m | 1 | +| 48 | 7 | `sum by (job) (rate(data[1h]))` | RateOrIncrease | job | 1h | 1m | 1 | +| 49 | 7 | `sum by (job) (rate(data[6h]))` | RateOrIncrease | job | 6h | 1m | 1 | +| 50 | 7 | `sum by (job) (rate(data[24h]))` | RateOrIncrease | job | 24h | 1m | 1 | +| 51 | 8 | `sum by (job) (sum_over_time(data[15m]))` | SumOrCount | job | 15m | 1m | 1 | +| 52 | 8 | `sum by (job) (sum_over_time(data[1h]))` | SumOrCount | job | 1h | 1m | 1 | +| 53 | 8 | `sum by (job) (sum_over_time(data[6h]))` | SumOrCount | job | 6h | 1m | 1 | +| 54 | 8 | `sum by (job) (sum_over_time(data[24h]))` | SumOrCount | job | 24h | 1m | 1 | +| 55 | 9 | `topk(100, sum by (label_0) (rate(data[15m])))` | TopK over increases | — (keys: label_0) | 15m | 1m | 1 | +| 56 | 9 | `topk(100, sum by (label_0) (rate(data[1h])))` | TopK over increases | — (keys: label_0) | 1h | 1m | 1 | +| 57 | 9 | `topk(100, sum by (label_0) (rate(data[6h])))` | TopK over increases | — (keys: label_0) | 6h | 1m | 1 | +| 58 | 9 | `topk(100, sum by (label_0) (rate(data[24h])))` | TopK over increases | — (keys: label_0) | 24h | 1m | 1 | +| 59 | 9 | `topk(200, sum by (label_0) (rate(data[15m])))` | TopK over increases | — (keys: label_0) | 15m | 1m | 1 | +| 60 | 9 | `topk(200, sum by (label_0) (rate(data[1h])))` | TopK over increases | — (keys: label_0) | 1h | 1m | 1 | +| 61 | 9 | `topk(200, sum by (label_0) (rate(data[6h])))` | TopK over increases | — (keys: label_0) | 6h | 1m | 1 | +| 62 | 9 | `topk(200, sum by (label_0) (rate(data[24h])))` | TopK over increases | — (keys: label_0) | 24h | 1m | 1 | +| 63 | 9 | `topk(300, sum by (label_0) (rate(data[15m])))` | TopK over increases | — (keys: label_0) | 15m | 1m | 1 | +| 64 | 9 | `topk(300, sum by (label_0) (rate(data[1h])))` | TopK over increases | — (keys: label_0) | 1h | 1m | 1 | +| 65 | 9 | `topk(300, sum by (label_0) (rate(data[6h])))` | TopK over increases | — (keys: label_0) | 6h | 1m | 1 | +| 66 | 9 | `topk(300, sum by (label_0) (rate(data[24h])))` | TopK over increases | — (keys: label_0) | 24h | 1m | 1 | +| 67 | 10 | `quantile_over_time(0.9, data[15m]) / quantile_over_time(0.5, data[15m])` | 2 × Quantile | series | 15m | 1m | 2 | +| 68 | 10 | `quantile_over_time(0.9, data[1h]) / quantile_over_time(0.5, data[1h])` | 2 × Quantile | series | 1h | 1m | 2 | +| 69 | 10 | `quantile_over_time(0.9, data[6h]) / quantile_over_time(0.5, data[6h])` | 2 × Quantile | series | 6h | 1m | 2 | +| 70 | 10 | `quantile_over_time(0.9, data[24h]) / quantile_over_time(0.5, data[24h])` | 2 × Quantile | series | 24h | 1m | 2 | + +
#### Dashboard template set A second template set models an SLO / monitoring dashboard, latency-quantile -heavy, with Grafana's usual time ranges. Window set -`W_d = {1m, 5m, 15m, 1h, 6h, 24h}`; temporal templates repeat every 1 m (the -dashboard refresh); spatial ones keep `S = T = 1 s`. +heavy, refreshed every 1 m, over the same windows `W`. -| # | PromQL | Capability | Grouping | RQEs per replica | +| # | PromQL | Capability | Grouping | RQEs | |---|---|---|---|---| -| D1 | `quantile_over_time(q, data[w])`, q ∈ {0.5, 0.9, 0.99}, w ∈ `W_d` | Quantile | per series | 18 | -| D2 | `quantile by (q, label_0) (data)`, q ∈ {0.5, 0.9, 0.99} | Quantile | per `label_0` group | 3 | -| D3 | `sum by (label_0) (rate(data[w]))`, w ∈ `W_d` | RateOrIncrease | `label_0` | 6 | -| D4 | `topk by (3, label_0) (rate(data[w]))`, w ∈ {5m, 1h} | TopK over per-series increases | `label_0` | 2 | -| D5 | `quantile_over_time(0.99, data[w]) / quantile_over_time(0.5, data[w])`, w ∈ {5m, 1h} | Two Quantile RQEs | per series | 4 | - -33 RQEs per replica. D1's quantiles share one stream across overlapping -windows; D5 repeats D1's p99/p50 at 5m and 1h; D3 and D4 share the -`label_0` increment stream. top-k uses k = 3, the k sketch-bench measures. - -#### Data model - -There is one metric, `data`. A **series** is one combination of label values. -Three labels matter: - -| Label | Values | Role | -|---|---|---| -| `label_0` | `C = card(label_0)` values | The grouping label: what `by (label_0)` aggregates by | -| `instance` | `s` values per `label_0` value | Distinguishes the series inside a group. `s` = series per group | - - -There are `C · s` series. Replicas (workload grid) read the same stream, so -they add RQEs, not series. Every series emits one sample every 10 ms -(100 samples/s), so the stream carries `λ = 100 · C · s` samples/s. - -Sample values: -- **Sum and increase:** exact, so the value distribution does not affect cost - or accuracy. -- **Top-k and quantiles:** the data the cost table is measured on: Zipf - s = 1.1 over a population of 100,000 keys (sketch-bench - `export_rqe_optimizer_costs.sh`). Top-k ranks the keys by weight; quantile - sketches (KLL, DDSketch) are measured on the Zipf ranks as values. The - workload uses the same distribution, so the table needs no separate run for - it. - -**Only two cardinalities matter.** Queries aggregate by `label_0` or per -series, never by another label. So any other label (a second instance-like -label, a `label_2`, ...) only multiplies the number of series in each group, -and is equivalent to a larger `s`. The model therefore has two independent -cardinality knobs: -- `C`: groups; -- `s`: series per group, the product of the cardinalities of all non-grouping - labels. - -Giving every label the same cardinality `c` would tie the knobs together -(`C = c`, `s = c^(L−1)` for `L` labels). It was rejected: `s` explodes (c = 1e3 -with three labels gives 1e9 series, far beyond the measured K ≤ 1e7), and the -effects of more groups and of more series per group could no longer be told -apart. - -**How a template becomes sketch input.** Every deployment keeps one instance -per group of its grouping labels `G` (`card(G)` instances per window), as in -#145's cost model. - -| Template kind | Instances per window | Kind of instance | Items per instance per window | -|---|---|---|---| -| `sum`/`rate by (label_0)` (1, 7, 8) | `C` | exact accumulator | `100 · s · S` | -| `topk by (label_0)` (2, 9) | `C` | top-k sketch over the group's `s` series | `100 · s · S` | -| `quantile by (q, label_0)` (3) | `C` | quantile sketch | `100 · s · S` | -| per-series `sum_over_time`/`rate` (4, 6) | `C · s` | exact accumulator | `100 · S` | -| per-series `quantile_over_time` (5, 10) | `C · s` | quantile sketch | `100 · S` | - -**Example.** `C = 3` (`label_0` ∈ {a, b, c}), `s = 2` (`instance` ∈ {i1, i2}): -six series, `data{label_0="a", instance="i1"}` through -`data{label_0="c", instance="i2"}`, emitting 600 samples/s in total. -- `sum by (label_0) (data)`, `S = 1 s`: three exact sums, one per group, each - absorbing 200 values per second. -- `quantile by (0.99, label_0) (data)`: three KLL sketches, one per group, each - absorbing 200 values per second. -- `sum_over_time(data[1m])`: six exact sums, one per series, 6,000 values each - per minute. -- `quantile_over_time(0.99, data[1m])`: six KLL sketches, 6,000 values each per - minute. - -**Modeling choice for spatial templates.** A spatial template evaluates every -1 s over every sample of the last second (`S = T = 1 s`). PromQL's instant -semantics would read only each series' latest sample. Aggregating the whole -second is what a sketch maintained over a 1-second window answers. The -difference is noted wherever spatial results are reported. - -Items per instance, `100 · s · S` or `100 · S`, generally differ from the -benchmark's. How costs and accuracy are read at that size is §6 -"Benchmark input". +| D1 | `quantile_over_time(q, data[w])`, q ∈ {0.5, 0.9, 0.99} | Quantile | per series | 12 | +| D2 | `quantile by (job) (q, data)`, q ∈ {0.5, 0.9, 0.99} | Quantile | `job` | 3 | +| D3 | `sum by (job) (rate(data[w]))` | RateOrIncrease | `job` | 4 | +| D4 | `topk(k, sum by (label_0) (rate(data[w])))`, w ∈ {15m, 1h} | TopK over per-series increases | none; keys `label_0` | 6 | +| D5 | `quantile_over_time(0.99, data[w]) / quantile_over_time(0.5, data[w])`, w ∈ {15m, 1h} | Two Quantile RQEs | per series | 4 | + +29 RQEs per replica; D5's operands repeat D1's, so 25 are distinct. D1's +quantiles share one stream across overlapping windows; D3 and D4 share the +increase stream. + +
+Every query of the dashboard set (27 queries, 29 RQEs) + +| # | Template | PromQL | Capability | Grouping | `S` | `T` | RQEs | +|---|---|---|---|---|---|---|---| +| 1 | D1 | `quantile_over_time(0.5, data[15m])` | Quantile | series | 15m | 1m | 1 | +| 2 | D1 | `quantile_over_time(0.5, data[1h])` | Quantile | series | 1h | 1m | 1 | +| 3 | D1 | `quantile_over_time(0.5, data[6h])` | Quantile | series | 6h | 1m | 1 | +| 4 | D1 | `quantile_over_time(0.5, data[24h])` | Quantile | series | 24h | 1m | 1 | +| 5 | D1 | `quantile_over_time(0.9, data[15m])` | Quantile | series | 15m | 1m | 1 | +| 6 | D1 | `quantile_over_time(0.9, data[1h])` | Quantile | series | 1h | 1m | 1 | +| 7 | D1 | `quantile_over_time(0.9, data[6h])` | Quantile | series | 6h | 1m | 1 | +| 8 | D1 | `quantile_over_time(0.9, data[24h])` | Quantile | series | 24h | 1m | 1 | +| 9 | D1 | `quantile_over_time(0.99, data[15m])` | Quantile | series | 15m | 1m | 1 | +| 10 | D1 | `quantile_over_time(0.99, data[1h])` | Quantile | series | 1h | 1m | 1 | +| 11 | D1 | `quantile_over_time(0.99, data[6h])` | Quantile | series | 6h | 1m | 1 | +| 12 | D1 | `quantile_over_time(0.99, data[24h])` | Quantile | series | 24h | 1m | 1 | +| 13 | D2 | `quantile by (job) (0.5, data)` | Quantile | job | 1s | 1s | 1 | +| 14 | D2 | `quantile by (job) (0.9, data)` | Quantile | job | 1s | 1s | 1 | +| 15 | D2 | `quantile by (job) (0.99, data)` | Quantile | job | 1s | 1s | 1 | +| 16 | D3 | `sum by (job) (rate(data[15m]))` | RateOrIncrease | job | 15m | 1m | 1 | +| 17 | D3 | `sum by (job) (rate(data[1h]))` | RateOrIncrease | job | 1h | 1m | 1 | +| 18 | D3 | `sum by (job) (rate(data[6h]))` | RateOrIncrease | job | 6h | 1m | 1 | +| 19 | D3 | `sum by (job) (rate(data[24h]))` | RateOrIncrease | job | 24h | 1m | 1 | +| 20 | D4 | `topk(100, sum by (label_0) (rate(data[15m])))` | TopK over increases | — (keys: label_0) | 15m | 1m | 1 | +| 21 | D4 | `topk(100, sum by (label_0) (rate(data[1h])))` | TopK over increases | — (keys: label_0) | 1h | 1m | 1 | +| 22 | D4 | `topk(200, sum by (label_0) (rate(data[15m])))` | TopK over increases | — (keys: label_0) | 15m | 1m | 1 | +| 23 | D4 | `topk(200, sum by (label_0) (rate(data[1h])))` | TopK over increases | — (keys: label_0) | 1h | 1m | 1 | +| 24 | D4 | `topk(300, sum by (label_0) (rate(data[15m])))` | TopK over increases | — (keys: label_0) | 15m | 1m | 1 | +| 25 | D4 | `topk(300, sum by (label_0) (rate(data[1h])))` | TopK over increases | — (keys: label_0) | 1h | 1m | 1 | +| 26 | D5 | `quantile_over_time(0.99, data[15m]) / quantile_over_time(0.5, data[15m])` | 2 × Quantile | series | 15m | 1m | 2 | +| 27 | D5 | `quantile_over_time(0.99, data[1h]) / quantile_over_time(0.5, data[1h])` | 2 × Quantile | series | 1h | 1m | 2 | + +
#### Workload grid @@ -328,13 +412,12 @@ keeps only what changes the comparison with AutoSketch. | Dimension | Values | What it varies | |---|---|---| -| Template set | **dashboard**; the 10 templates (with `W` = {1m, 10m, 1h, 6h, 24h}, `T` = 1 m) | Workload realism, and which capabilities appear | -| Replicas `r` | **1**, 8, 64 | Every replica reads the same stream with a seeded random subset of 3 windows, 3 quantiles from {0.5, 0.75, 0.9, 0.95, 0.99} and `T` from {10 s, 1 m, 5 m}; identical RQEs are deduplicated. Many users or dashboards over the same metrics: how the sharing benefit and planning time grow with the number of RQEs | -| Groups `C = card(label_0)` | 1e2, **1e3**, 1e4 | Instances per deployment and items per instance | +| Template set | **dashboard**; the 10 templates | Workload realism, and which capabilities appear | +| Replicas `r` | **1**, 8, 64 | Every replica reads the same stream with a seeded random subset of 3 windows from `W`, 3 quantiles from {0.5, 0.75, 0.9, 0.95, 0.99}, a `k` from {100, 200, 300} and `T` from {10 s, 1 m, 5 m}; identical RQEs are deduplicated. Many users or dashboards over the same metrics: how the sharing benefit and planning time grow with the number of RQEs | | Accuracy target (strictness) | loose, **default**, strict | What AutoSketch optimizes for (§5) | | Latency SLA | the §5 grid, **no limit** | §5 | -Fixed: `s = 100`; data Zipf s = 1.1 over 100,000 keys (§6 "Data model"). +The data model is fixed (§6 "Data model and scale"); no grid dimension changes the data. The sweep is the default workload, then each dimension varied alone with the others at their defaults. Every workload runs every baseline (ASAP, @@ -355,20 +438,24 @@ that shows sharing. These dimensions describe the planner, not its gap to AutoSketch. They belong in the planner's micro-benchmarks: - query mix (spatial only, temporal only, one capability at a time); -- lookback window set `W` ({1h} to {1m, 10m, 1h, 6h, 24h}); +- lookback window set `W`, including windows under 15 minutes; - repeat interval `T` (10 s, 1 m, 5 m); -- series per group `s` (1 to 1000); +- data scale: `C`, `J` and the sample rate; - data distribution (key skew, value tail); -- interactions `r × s` and `W × C`. +- interactions such as `r × C` and `W × C`. ### Benchmark input Every method reads the cost table from sketch-bench `scripts/export_rqe_optimizer_costs.sh`, the same table the planner reads. Deployable rows: exact sum, min, max and increase; KLL k ∈ {200, 500}; HLL -`lg_k` ∈ {12, 14}; CMS-heap top-k with rows ∈ {3, 5}, cols = 2048. DDSketch, +`lg_k` ∈ {12, 14}; CMS-heap top-k with rows ∈ {3, 5}, cols = 2048 and heap +∈ {100, 200, 300}, each scored at every `k` up to its heap. DDSketch, CountSketch-heap and UnivMon are measured but not deployable by default. +The table is exported with `CARDINALITY=10000`, the workload's key count, and +measured at 1e5 to 1e8 items per instance (#157). + **Exact accumulators** have no saturation point: their answer is exact at any size. They are still benchmarked for cost, on the grouped column specs (200,000 rows, about 9,900 groups), and priced per group. The sweep over items @@ -387,8 +474,8 @@ from the lower `θ` bound. Longer samples expose worse cases (sketch-bench - **`traces` gives the appendix results.** Its parameters are fit on the full trace. -- **`synthetic`** uses the cost table's own data, Zipf s = 1.1 over 100,000 - keys (§6 "Data model"). +- **`synthetic`** uses the cost table's own data, Zipf s = 1.1 over 10,000 + keys (§6 "Data model and scale"). Each config is benchmarked at these worst-case parameters. AutoSketch §5.2 injects random traffic bursts into synthetic workloads to cover variation over @@ -469,6 +556,7 @@ Figures: | — | sketch-bench #144, #145 | Exact accumulators and top-k families; per-phase cost model and weighted objective | Merged | | — | sketch-bench #151, #152, #154, #155 | Cost-table fixes, value range, accuracy after merging, `measured_at` (#147) | Merged | | — | sketch-bench #157 | Sweep items per instance; re-export the cost table | Open | +| — | sketch-bench (to open) | Top-k for this workload: heap size as a parameter, precision scored at k ∈ {100, 200, 300}; `rqe-optimizer` RQEs carry `k`, and a heap-`K` deployment serves `k ≤ K`; export at `CARDINALITY=10000`, 1e5–1e8 items | Not started | | — | sketch-bench #130, #131 | Saturation curves; accuracy after merging `m` shards | Merged; background | | 1 | sketch-bench #137 | Retained memory, EC2 pricing, `milp::minimize_cost` | Merged; superseded by #145 | | 3 | sketch-bench #135 | AutoSketch-Adapted (Algorithm 4), aligned with the paper's EXAMINE rule and seeding | Merged | From 4409716edef7a0e086737f1cc9457c37c861a691 Mon Sep 17 00:00:00 2001 From: zz_y Date: Tue, 6 Oct 2026 23:05:55 +0000 Subject: [PATCH 33/39] docs(planner): read sketch accuracy from the saturation curves (sketch-bench #156) Keep the fixed synthetic data (Zipf s=1.1 over 10,000 keys, quantiles on the Zipf ranks) and top-k with k in {100, 200, 300} served by heap-K deployments. Everything else follows sketch-bench #156/#162: - sketch accuracy from the saturation curves at n(S, G); CPU, memory and exact rows from the cost table; grid configs only. - merging: every pane saturated, then read at n(S, G); KLL/top-k merge penalty is #158. AutoSketch uses the same lookup with m = 1. - DDSketch scored in value relative error (#162), with its own targets. - a targeted saturation run at the workload's data shape (theta 1.1, K 1e4, Zipf-rank quantiles, CMS-heap heaps 100/200/300) so lookups land on measured points. Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 138 +++++++++++++---------- 1 file changed, 81 insertions(+), 57 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 87854acb..487460d8 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -30,9 +30,9 @@ interval) and sums the results; the planner is invoked once for the batch. | Earlier protocol (E1–E3, end-to-end execution) | ASAPQuery-backend `docs/evaluation/autosketch-comparison.md` ([#545](https://github.com/ProjectASAP/ASAPQuery-backend/pull/545)) | Merged. Execution-based; this plan is planner-level and uses estimated costs instead. | | Top-K dashboard comparison | ASAPQuery-backend [#602](https://github.com/ProjectASAP/ASAPQuery-backend/pull/602) | Closed, not merged. | | RQE deployment MILP (HiGHS): candidates `(capability, config, labels, x, y)`, sharing, latency bounds | sketch-bench `rqe-optimizer/` ([#129](https://github.com/ProjectASAP/sketch-bench/pull/129)); per-phase cost model and weighted objective [#145](https://github.com/ProjectASAP/sketch-bench/pull/145); exact accumulators and top-k families [#144](https://github.com/ProjectASAP/sketch-bench/pull/144) | Merged. **This is the planner we evaluate.** | -| Measured per-operation costs (`AtomicCostEntry`: memory/instance, insert/merge/query CPU, accuracy) | sketch-bench `scripts/export_rqe_optimizer_costs.sh` | Merged. **The cost table this evaluation reads** (§6). Each row records its measurement conditions (`measured_at`, #155) and accuracy after merging (#154); a sweep over items per instance is #157. | -| Saturation study: error vs. `N`, `N_sat`, cost per config and shape | sketch-bench [#130](https://github.com/ProjectASAP/sketch-bench/pull/130) | Merged. Background for how error and cost depend on `N`; no longer read by the evaluation. | -| Accuracy after merging `m` shards | sketch-bench [#131](https://github.com/ProjectASAP/sketch-bench/pull/131), carried into the cost table by [#154](https://github.com/ProjectASAP/sketch-bench/pull/154) | Merged. Read when a deployment merges `m = S/x > 1` windows. | +| Measured per-operation costs (`AtomicCostEntry`: memory/instance, insert/merge/query CPU, accuracy) | sketch-bench `scripts/export_rqe_optimizer_costs.sh` | Merged. **Source of CPU and memory** (§6), and of exact accumulators' rows. Each row records its measurement conditions (`measured_at`, #155). Sketch rows' accuracy is not read from it (#156 Q9). | +| Saturation study: error vs. `N`, `N_sat`, cost per config and shape | sketch-bench [#130](https://github.com/ProjectASAP/sketch-bench/pull/130) | Merged. **Source of sketch accuracy**: the planner reads the error curve at each instance's item count (#156). | +| Accuracy after merging `m` shards | sketch-bench [#131](https://github.com/ProjectASAP/sketch-bench/pull/131), carried into the cost table by [#154](https://github.com/ProjectASAP/sketch-bench/pull/154) | Merged. Merge penalties for KLL and top-k are a follow-up ([#158](https://github.com/ProjectASAP/sketch-bench/issues/158)); until then #156 requires every merged pane to be saturated. | | Moving the MILP into ASAPQuery's planner | ASAPQuery `asap-planner-rs/src/optimizer/` (Milind; related: [#776](https://github.com/ProjectASAP/ASAPQuery/pull/776), [#725](https://github.com/ProjectASAP/ASAPQuery/pull/725)) | Out of scope: the evaluation uses sketch-bench `rqe-optimizer` and is not rerun on `asap-planner-rs`. | AutoSketch-Adapted is implemented in sketch-bench `rqe-optimizer/src/autosketch.rs` (#135). @@ -139,19 +139,23 @@ per capability. The default level matches the `traces` targets from ASAPQuery's dataset analysis. The synthetic workload sweeps all three; `traces` uses its fitted targets. -| Level | Quantile (rank error) | TopK (precision@k) | -| --- | --- | --- | -| loose | ≤ 0.02 | ≥ 0.90 | -| **default** | **≤ 0.01** | **≥ 0.95** | -| strict | ≤ 0.005 | ≥ 0.99 | +| Level | Quantile, KLL (rank error) | Quantile, DDSketch (value relative error) | TopK (precision@k) | +| --- | --- | --- | --- | +| loose | ≤ 0.02 | ≤ 0.02 | ≥ 0.90 | +| **default** | **≤ 0.01** | **≤ 0.01** | **≥ 0.95** | +| strict | ≤ 0.005 | ≤ 0.005 | ≥ 0.99 | + +Each sketch is scored in the metric its guarantee is stated in (#156 Q9): +KLL in rank error, DDSketch in value relative error `|x̃ − x_q| / |x_q|` +(sketch-bench #162). DDSketch is not deployable by default, so its column only +applies with `--allow-undeployable-families`. Sum and increase are served by exact accumulators, so their error is 0 and every level is met. Strictness therefore matters only for quantile and top-k RQEs. No template asks for cardinality. -A deployment that merges `m = S/x` windows must meet the target at both -measured merge counts bracketing `m` (count 1 is the single instance), as -`rqe-optimizer` checks it (#154). +How the accuracy of a deployment is read, including when it merges windows, +is §6 "Reading accuracy and cost". An earlier version mapped one percentage `p` to error ≤ `1 − p` for every capability. It was dropped (2026-10-05): at 95% it allowed quantile rank error @@ -218,12 +222,13 @@ ingests decides whether a sketch is worth it: | `by (job)`, temporal (7, 8; D3) | J = 10 | 1.8e8 to 1.7e10 | exact accumulator | | top-k over `label_0` (2, 9; D4) | 1 | `λ · S`: 1.8e9 to 1.7e11, over 1e4 keys | top-k sketch (CMS-heap) | -The cost table is measured at 1e5 to 1e8 items per instance (sketch-bench -#157). Exact accumulators keep constant state and a constant per-item cost, so -their size does not matter. Every quantile sketch falls inside the measured -range. Top-k sketches exceed it and read the 1e8 measurement: CMS-heap's -per-item CPU and memory do not depend on `N` (sketch-bench #130), and Zipf -s = 1.1 keys have saturated by then. +Exact accumulators keep constant state and a constant per-item cost, so their +size does not matter. Sketch accuracy is read from the saturation curve at the +instance's item count (§6 "Reading accuracy and cost"). Every quantile sketch +falls inside the curves' range (1e3 to 1e8 items). Top-k sketches exceed it; +#156 reads the error at the largest measured N when the point has saturated +there, which Zipf s = 1.1 keys do. CMS-heap's per-item CPU and memory do not +depend on `N` (sketch-bench #130). **Modeling choice for spatial templates.** A spatial template evaluates every 1 s over every sample of the last second (`S = T = 1 s`). PromQL's instant @@ -446,21 +451,34 @@ in the planner's micro-benchmarks: ### Benchmark input -Every method reads the cost table from sketch-bench -`scripts/export_rqe_optimizer_costs.sh`, the same table the planner reads. -Deployable rows: exact sum, min, max and increase; KLL k ∈ {200, 500}; HLL -`lg_k` ∈ {12, 14}; CMS-heap top-k with rows ∈ {3, 5}, cols = 2048 and heap -∈ {100, 200, 300}, each scored at every `k` up to its heap. DDSketch, -CountSketch-heap and UnivMon are measured but not deployable by default. - -The table is exported with `CARDINALITY=10000`, the workload's key count, and -measured at 1e5 to 1e8 items per instance (#157). +Two sketch-bench measurements feed every method, per sketch-bench #156: + +- **Accuracy of sketch rows:** the saturation curves (`saturation.csv`, + `saturation_curve.csv`), error vs. items `N` per (config, data shape). The + planner reads them through `--saturation-dir`. +- **CPU and memory, and exact rows:** the cost table from + `scripts/export_rqe_optimizer_costs.sh`, exported with `CARDINALITY=10000`, + the workload's key count. + +Sketch configs are the saturation grid's (#156 Q4): KLL k ∈ {50, 200, 800}, +HLL `lg_k` ∈ {12, 14, 16}, CMS-heap top-k with rows ∈ {3, 5}, cols ∈ {256, +1024, 4096, 16384}. A sketch config with no curve is not eligible. +DDSketch, CountSketch-heap and UnivMon are not deployable by default. + +**A saturation run at the synthetic workload's data shape.** The grid has no +point at θ = 1.1 and K = 1e4, measures quantiles on Pareto data, and scores +top-k only at k = 32. The synthetic workload therefore needs one targeted run +of `scripts/study_saturation.py` at its own data, so every lookup lands on a +measured point instead of the worst of the bracketing grid points (#156 Q8): +- Zipf s = 1.1 over 10,000 keys, the same data as the cost table; +- KLL and DDSketch on the Zipf ranks, DDSketch in value relative error (#162); +- CMS-heap with heap ∈ {100, 200, 300}, each scored at every `k` ∈ {100, 200, + 300} up to its heap; +- `N` up to 1e9 for top-k, 1e8 for the quantile sketches. **Exact accumulators** have no saturation point: their answer is exact at any size. They are still benchmarked for cost, on the grouped column specs -(200,000 rows, about 9,900 groups), and priced per group. The sweep over items -per instance (#157) includes them, so their per-group cost is read at the -workload's size like every other row's. +(200,000 rows, about 9,900 groups), and priced per group. #### Data parameters, shared by both methods @@ -474,8 +492,9 @@ from the lower `θ` bound. Longer samples expose worse cases (sketch-bench - **`traces` gives the appendix results.** Its parameters are fit on the full trace. -- **`synthetic`** uses the cost table's own data, Zipf s = 1.1 over 10,000 - keys (§6 "Data model and scale"). +- **`synthetic`** uses its own fixed data, Zipf s = 1.1 over 10,000 keys + (§6 "Data model and scale"): `data_shape` = (θ = 1.1, K = 1e4), quantiles on + the Zipf ranks. Each config is benchmarked at these worst-case parameters. AutoSketch §5.2 injects random traffic bursts into synthetic workloads to cover variation over @@ -489,22 +508,28 @@ A window of length `S` on grouping labels `G` holds about n(S, G) = λ · S / card(G) items per instance ``` -#### Reading the table at the workload's size - -Each row records the size it was measured at (`measured_at`, #155). Until the -size sweep (#157) lands, every row is read at its measured size and the -points whose instance size differs are flagged. With the sweep: - -- **ASAP** reads each deployment's row at its own instance size, - `λ · x / card(G)` items per window. When it merges `m = S/x` windows, the - accuracy check uses the measured merge counts bracketing `m` (#154); merge - error is not monotone in `m`, so both must pass. -- **AutoSketch** never merges (`x = S`). It reads the row at its window's size, - `λ · S / card(G)`, and accepts a config if the target holds there, as the - paper's benchmark-then-accept loop does. It configures once, before - deployment (§9 Q2). -- **Exact accumulators** meet every target at every size; only their cost is - read at the size. +#### Reading accuracy and cost + +Accuracy follows sketch-bench #156. With `n(S, G)` the items one instance holds +and `m = S/x` the windows a query merges: + +- **No merge (`m = 1`):** the error is the curve's value at `n(S, G)`. + Between checkpoints, the worse of the two neighbours; below 1,000 items, not + eligible; above the largest measured `N`, that `N`'s value if the point has + saturated, else not eligible (#156 Q6). +- **Merging (`m > 1`):** every pane, `n(x, G)` items, must reach `N_sat`; the + error is then read at `n(S, G)` as above (#156 Q11). CMS, CountSketch, HLL and + DDSketch merge exactly, so this is their single-sketch error. KLL and top-k + lose accuracy when merged; that penalty is the follow-up #158. +- **Top-k:** precision is read at the RQE's `k` from the curve of the + deployment's heap size, so a heap-300 deployment serves k = 100, 200 and 300. +- **AutoSketch** never merges (`x = S`): the same lookup with `m = 1` at its + window's size, `n(S, G)` (#156 Q12). It configures once, before deployment + (§9 Q2). +- **Exact accumulators** have zero error at every size. + +CPU and memory come from the cost table at the configs' measured size; +per-item costs are flat in `N` (sketch-bench #130). ## 7. Metrics and figures @@ -554,14 +579,15 @@ Figures: | --- | --- | --- | --- | | this | ASAPQuery #777 | This plan | Draft, updated as decisions change | | — | sketch-bench #144, #145 | Exact accumulators and top-k families; per-phase cost model and weighted objective | Merged | -| — | sketch-bench #151, #152, #154, #155 | Cost-table fixes, value range, accuracy after merging, `measured_at` (#147) | Merged | -| — | sketch-bench #157 | Sweep items per instance; re-export the cost table | Open | -| — | sketch-bench (to open) | Top-k for this workload: heap size as a parameter, precision scored at k ∈ {100, 200, 300}; `rqe-optimizer` RQEs carry `k`, and a heap-`K` deployment serves `k ≤ K`; export at `CARDINALITY=10000`, 1e5–1e8 items | Not started | +| — | sketch-bench #151, #152, #154, #155 | Cost-table fixes, value range, accuracy after merging, `measured_at` (#147) | Merged; #154's merge rule is replaced by #156 | +| — | sketch-bench #156, #162, #158 | Accuracy from the saturation curves at n(T) with `data_shape`; DDSketch in value relative error; merge penalties | Open; #156 blocked on #162 | +| — | sketch-bench #157 | Re-export the cost table (grid configs, `CARDINALITY=10000`); per-item CPU and memory vs. size | Open; its accuracy part is covered by #156 | +| — | sketch-bench (to open) | Top-k for this workload: heap size as a parameter, precision scored at k ∈ {100, 200, 300}; `rqe-optimizer` RQEs carry `k`, and a heap-`K` deployment serves `k ≤ K`; the targeted saturation run at θ = 1.1, K = 1e4 (§6) | Not started | | — | sketch-bench #130, #131 | Saturation curves; accuracy after merging `m` shards | Merged; background | | 1 | sketch-bench #137 | Retained memory, EC2 pricing, `milp::minimize_cost` | Merged; superseded by #145 | | 3 | sketch-bench #135 | AutoSketch-Adapted (Algorithm 4), aligned with the paper's EXAMINE rule and seeding | Merged | | 2 | sketch-bench #136 | Evaluation table for the trace workloads | Merged | -| 4 | sketch-bench #138, #139, #140, #141 | Runner, synthetic workload, saturation at K = 1e1–1e6, two cost models and grid driver | Open; to be reworked onto #145's objective, the cost table and the reduced grid. #140 is no longer needed for the comparison | +| 4 | sketch-bench #138, #139, #140, #141 | Runner, synthetic workload, saturation at K = 1e1–1e6, two cost models and grid driver | Open; to be reworked onto #145's objective, #156's lookup and the reduced grid. #140's K = 1e4 points are superseded for the synthetic workload by the targeted run (§6) | ## 9. Decisions @@ -599,12 +625,10 @@ were dropped (2026-10-06). ## 10. Known limitations -- **Merged accuracy comes from shard-merge measurements, not from replaying - the plans.** The cost table measures it at `m` ∈ {4, 16, 64, 256, 1024} over - one benchmark stream (#154). A deployment needing more (e.g. a 1-day - lookback over 1-minute windows, `m = 1440`) reads the 1024 measurement. - Replay the synthetic default workload's chosen plans in sketch-bench once to - confirm. +- **Merged accuracy is read, not replayed.** Exact for CMS, CountSketch, HLL + and DDSketch; for KLL and top-k the merge penalty is ignored until #158, with + every pane required to be saturated. Replay the synthetic default workload's + chosen plans in sketch-bench once to confirm. - Costs and latencies are estimates from per-operation measurements, not end-to-end executions. The execution-based comparison is ASAPQuery-backend #545/#547. From f9ead5e3830ed22cb78055c4d66f8c3e7d6b4a3c Mon Sep 17 00:00:00 2001 From: zz_y Date: Tue, 6 Oct 2026 23:21:06 +0000 Subject: [PATCH 34/39] docs(planner): fix top-k queries at k = 32 k = 32 is the heap size sketch-bench measures top-k at (CMS_HEAP_TOP_K), so top-k needs no heap parameter, no k on the RQE and no extra saturation points. 58 RQEs per replica for the 10 templates (50 distinct), 25 for the dashboard (21 distinct); query lists regenerated and parser-checked. Co-Authored-By: Claude Opus 5.5 --- docs/evaluation/autosketch-vs-planner.md | 177 ++++++++++------------- 1 file changed, 77 insertions(+), 100 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 487460d8..f2d4cee9 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -91,7 +91,7 @@ has lookback `S` and repeat interval `T`. Per-instance costs are measured | --- | --- | --- | | Ingest | `λ · (x/y) · c_ins` | `card(G) · m · x/y` (open windows) | | Merge | `card(G) · (S/x − 1) · c_mrg / T` | `card(G) · m` (one accumulator per group), 0 when `S = x` | -| Query | `card(G) · c_qry / T` | `card(G) · 8 B`; top-k `k · 16 B` | +| Query | `card(G) · c_qry / T` | `card(G) · 8 B`; top-k `32 · 16 B` | | Storage | 0 | `card(G) · m · ((max S − x)/y + 1)` (closed windows) | Ingest and storage are paid once per active deployment; merge and query once @@ -239,24 +239,24 @@ second is what a summary maintained over a 1-second window answers. These exercise every capability the templates need: sum, rate/increase, top-k and quantile, spatial and temporal aggregation, and a binary operator. Each -template expands into one RQE per window `w ∈ W`, per quantile `q` and per -`k ∈ {100, 200, 300}`. +template expands into one RQE per window `w ∈ W` and per quantile `q`. Top-k +uses k = 32. | # | PromQL | Capability | Grouping | RQEs | |---|---|---|---|---| | 1 | `sum by (job) (data)` | SumOrCount | `job` | 1 | -| 2 | `topk(k, sum by (label_0) (sum_over_time(data[w])))` | TopK | none; keys `label_0` | 3·\|W\| = 12 | +| 2 | `topk(32, sum by (label_0) (sum_over_time(data[w])))` | TopK | none; keys `label_0` | 4 | | 3 | `quantile by (job) (q, data)`, q ∈ {0.5, 0.75, 0.9, 0.95, 0.99} | Quantile | `job` | 5 | | 4 | `sum_over_time(data[w])` | SumOrCount | per series | 4 | | 5 | `quantile_over_time(q, data[w])`, same five q | Quantile | per series | 20 | | 6 | `rate(data[w])` | RateOrIncrease | per series | 4 | | 7 | `sum by (job) (rate(data[w]))` | RateOrIncrease | `job` | 4 | | 8 | `sum by (job) (sum_over_time(data[w]))` | SumOrCount | `job` | 4 | -| 9 | `topk(k, sum by (label_0) (rate(data[w])))` | TopK over per-series increases | none; keys `label_0` | 12 | +| 9 | `topk(32, sum by (label_0) (rate(data[w])))` | TopK over per-series increases | none; keys `label_0` | 4 | | 10 | `quantile_over_time(0.9, data[w]) / quantile_over_time(0.5, data[w])` | Two Quantile RQEs | per series | 8 | -74 RQEs per replica. Template 10's operands are the same RQEs as template 5's -q = 0.9 and q = 0.5, so 66 are distinct. +58 RQEs per replica. Template 10's operands are the same RQEs as template 5's +q = 0.9 and q = 0.5, so 50 are distinct. Mapping notes: - Capabilities are `rqe-optimizer`'s (#144). Sum and rate/increase are their @@ -267,10 +267,8 @@ Mapping notes: search for sum and increase RQEs. There, ASAP differs from it only by window choice and sharing. - Top-k ranks each `label_0` key's total over the window: the heavy hitters - among 1e4 keys. A top-k sketch with heap capacity `K` serves every RQE with - `k ≤ K` on the same stream and window, so ASAP can serve k = 100, 200 and 300 - from one heap-300 deployment; AutoSketch configures each RQE with its own - heap. + among 1e4 keys. k = 32 is the heap size sketch-bench measures top-k at + (`CMS_HEAP_TOP_K`), so every top-k RQE reads a measured curve. - The quantiles of one template, and the two operands of template 10, read the same stream. ASAP can serve them from one deployment; AutoSketch gets one deployment per RQE. @@ -281,80 +279,64 @@ Mapping notes: is checked with `promql-parser`.
-Every query of the 10-template set (70 queries, 74 RQEs) +Every query of the 10-template set (54 queries, 58 RQEs) | # | Template | PromQL | Capability | Grouping | `S` | `T` | RQEs | |---|---|---|---|---|---|---|---| | 1 | 1 | `sum by (job) (data)` | SumOrCount | job | 1s | 1s | 1 | -| 2 | 2 | `topk(100, sum by (label_0) (sum_over_time(data[15m])))` | TopK | — (keys: label_0) | 15m | 1m | 1 | -| 3 | 2 | `topk(100, sum by (label_0) (sum_over_time(data[1h])))` | TopK | — (keys: label_0) | 1h | 1m | 1 | -| 4 | 2 | `topk(100, sum by (label_0) (sum_over_time(data[6h])))` | TopK | — (keys: label_0) | 6h | 1m | 1 | -| 5 | 2 | `topk(100, sum by (label_0) (sum_over_time(data[24h])))` | TopK | — (keys: label_0) | 24h | 1m | 1 | -| 6 | 2 | `topk(200, sum by (label_0) (sum_over_time(data[15m])))` | TopK | — (keys: label_0) | 15m | 1m | 1 | -| 7 | 2 | `topk(200, sum by (label_0) (sum_over_time(data[1h])))` | TopK | — (keys: label_0) | 1h | 1m | 1 | -| 8 | 2 | `topk(200, sum by (label_0) (sum_over_time(data[6h])))` | TopK | — (keys: label_0) | 6h | 1m | 1 | -| 9 | 2 | `topk(200, sum by (label_0) (sum_over_time(data[24h])))` | TopK | — (keys: label_0) | 24h | 1m | 1 | -| 10 | 2 | `topk(300, sum by (label_0) (sum_over_time(data[15m])))` | TopK | — (keys: label_0) | 15m | 1m | 1 | -| 11 | 2 | `topk(300, sum by (label_0) (sum_over_time(data[1h])))` | TopK | — (keys: label_0) | 1h | 1m | 1 | -| 12 | 2 | `topk(300, sum by (label_0) (sum_over_time(data[6h])))` | TopK | — (keys: label_0) | 6h | 1m | 1 | -| 13 | 2 | `topk(300, sum by (label_0) (sum_over_time(data[24h])))` | TopK | — (keys: label_0) | 24h | 1m | 1 | -| 14 | 3 | `quantile by (job) (0.5, data)` | Quantile | job | 1s | 1s | 1 | -| 15 | 3 | `quantile by (job) (0.75, data)` | Quantile | job | 1s | 1s | 1 | -| 16 | 3 | `quantile by (job) (0.9, data)` | Quantile | job | 1s | 1s | 1 | -| 17 | 3 | `quantile by (job) (0.95, data)` | Quantile | job | 1s | 1s | 1 | -| 18 | 3 | `quantile by (job) (0.99, data)` | Quantile | job | 1s | 1s | 1 | -| 19 | 4 | `sum_over_time(data[15m])` | SumOrCount | series | 15m | 1m | 1 | -| 20 | 4 | `sum_over_time(data[1h])` | SumOrCount | series | 1h | 1m | 1 | -| 21 | 4 | `sum_over_time(data[6h])` | SumOrCount | series | 6h | 1m | 1 | -| 22 | 4 | `sum_over_time(data[24h])` | SumOrCount | series | 24h | 1m | 1 | -| 23 | 5 | `quantile_over_time(0.5, data[15m])` | Quantile | series | 15m | 1m | 1 | -| 24 | 5 | `quantile_over_time(0.5, data[1h])` | Quantile | series | 1h | 1m | 1 | -| 25 | 5 | `quantile_over_time(0.5, data[6h])` | Quantile | series | 6h | 1m | 1 | -| 26 | 5 | `quantile_over_time(0.5, data[24h])` | Quantile | series | 24h | 1m | 1 | -| 27 | 5 | `quantile_over_time(0.75, data[15m])` | Quantile | series | 15m | 1m | 1 | -| 28 | 5 | `quantile_over_time(0.75, data[1h])` | Quantile | series | 1h | 1m | 1 | -| 29 | 5 | `quantile_over_time(0.75, data[6h])` | Quantile | series | 6h | 1m | 1 | -| 30 | 5 | `quantile_over_time(0.75, data[24h])` | Quantile | series | 24h | 1m | 1 | -| 31 | 5 | `quantile_over_time(0.9, data[15m])` | Quantile | series | 15m | 1m | 1 | -| 32 | 5 | `quantile_over_time(0.9, data[1h])` | Quantile | series | 1h | 1m | 1 | -| 33 | 5 | `quantile_over_time(0.9, data[6h])` | Quantile | series | 6h | 1m | 1 | -| 34 | 5 | `quantile_over_time(0.9, data[24h])` | Quantile | series | 24h | 1m | 1 | -| 35 | 5 | `quantile_over_time(0.95, data[15m])` | Quantile | series | 15m | 1m | 1 | -| 36 | 5 | `quantile_over_time(0.95, data[1h])` | Quantile | series | 1h | 1m | 1 | -| 37 | 5 | `quantile_over_time(0.95, data[6h])` | Quantile | series | 6h | 1m | 1 | -| 38 | 5 | `quantile_over_time(0.95, data[24h])` | Quantile | series | 24h | 1m | 1 | -| 39 | 5 | `quantile_over_time(0.99, data[15m])` | Quantile | series | 15m | 1m | 1 | -| 40 | 5 | `quantile_over_time(0.99, data[1h])` | Quantile | series | 1h | 1m | 1 | -| 41 | 5 | `quantile_over_time(0.99, data[6h])` | Quantile | series | 6h | 1m | 1 | -| 42 | 5 | `quantile_over_time(0.99, data[24h])` | Quantile | series | 24h | 1m | 1 | -| 43 | 6 | `rate(data[15m])` | RateOrIncrease | series | 15m | 1m | 1 | -| 44 | 6 | `rate(data[1h])` | RateOrIncrease | series | 1h | 1m | 1 | -| 45 | 6 | `rate(data[6h])` | RateOrIncrease | series | 6h | 1m | 1 | -| 46 | 6 | `rate(data[24h])` | RateOrIncrease | series | 24h | 1m | 1 | -| 47 | 7 | `sum by (job) (rate(data[15m]))` | RateOrIncrease | job | 15m | 1m | 1 | -| 48 | 7 | `sum by (job) (rate(data[1h]))` | RateOrIncrease | job | 1h | 1m | 1 | -| 49 | 7 | `sum by (job) (rate(data[6h]))` | RateOrIncrease | job | 6h | 1m | 1 | -| 50 | 7 | `sum by (job) (rate(data[24h]))` | RateOrIncrease | job | 24h | 1m | 1 | -| 51 | 8 | `sum by (job) (sum_over_time(data[15m]))` | SumOrCount | job | 15m | 1m | 1 | -| 52 | 8 | `sum by (job) (sum_over_time(data[1h]))` | SumOrCount | job | 1h | 1m | 1 | -| 53 | 8 | `sum by (job) (sum_over_time(data[6h]))` | SumOrCount | job | 6h | 1m | 1 | -| 54 | 8 | `sum by (job) (sum_over_time(data[24h]))` | SumOrCount | job | 24h | 1m | 1 | -| 55 | 9 | `topk(100, sum by (label_0) (rate(data[15m])))` | TopK over increases | — (keys: label_0) | 15m | 1m | 1 | -| 56 | 9 | `topk(100, sum by (label_0) (rate(data[1h])))` | TopK over increases | — (keys: label_0) | 1h | 1m | 1 | -| 57 | 9 | `topk(100, sum by (label_0) (rate(data[6h])))` | TopK over increases | — (keys: label_0) | 6h | 1m | 1 | -| 58 | 9 | `topk(100, sum by (label_0) (rate(data[24h])))` | TopK over increases | — (keys: label_0) | 24h | 1m | 1 | -| 59 | 9 | `topk(200, sum by (label_0) (rate(data[15m])))` | TopK over increases | — (keys: label_0) | 15m | 1m | 1 | -| 60 | 9 | `topk(200, sum by (label_0) (rate(data[1h])))` | TopK over increases | — (keys: label_0) | 1h | 1m | 1 | -| 61 | 9 | `topk(200, sum by (label_0) (rate(data[6h])))` | TopK over increases | — (keys: label_0) | 6h | 1m | 1 | -| 62 | 9 | `topk(200, sum by (label_0) (rate(data[24h])))` | TopK over increases | — (keys: label_0) | 24h | 1m | 1 | -| 63 | 9 | `topk(300, sum by (label_0) (rate(data[15m])))` | TopK over increases | — (keys: label_0) | 15m | 1m | 1 | -| 64 | 9 | `topk(300, sum by (label_0) (rate(data[1h])))` | TopK over increases | — (keys: label_0) | 1h | 1m | 1 | -| 65 | 9 | `topk(300, sum by (label_0) (rate(data[6h])))` | TopK over increases | — (keys: label_0) | 6h | 1m | 1 | -| 66 | 9 | `topk(300, sum by (label_0) (rate(data[24h])))` | TopK over increases | — (keys: label_0) | 24h | 1m | 1 | -| 67 | 10 | `quantile_over_time(0.9, data[15m]) / quantile_over_time(0.5, data[15m])` | 2 × Quantile | series | 15m | 1m | 2 | -| 68 | 10 | `quantile_over_time(0.9, data[1h]) / quantile_over_time(0.5, data[1h])` | 2 × Quantile | series | 1h | 1m | 2 | -| 69 | 10 | `quantile_over_time(0.9, data[6h]) / quantile_over_time(0.5, data[6h])` | 2 × Quantile | series | 6h | 1m | 2 | -| 70 | 10 | `quantile_over_time(0.9, data[24h]) / quantile_over_time(0.5, data[24h])` | 2 × Quantile | series | 24h | 1m | 2 | +| 2 | 2 | `topk(32, sum by (label_0) (sum_over_time(data[15m])))` | TopK | — (keys: label_0) | 15m | 1m | 1 | +| 3 | 2 | `topk(32, sum by (label_0) (sum_over_time(data[1h])))` | TopK | — (keys: label_0) | 1h | 1m | 1 | +| 4 | 2 | `topk(32, sum by (label_0) (sum_over_time(data[6h])))` | TopK | — (keys: label_0) | 6h | 1m | 1 | +| 5 | 2 | `topk(32, sum by (label_0) (sum_over_time(data[24h])))` | TopK | — (keys: label_0) | 24h | 1m | 1 | +| 6 | 3 | `quantile by (job) (0.5, data)` | Quantile | job | 1s | 1s | 1 | +| 7 | 3 | `quantile by (job) (0.75, data)` | Quantile | job | 1s | 1s | 1 | +| 8 | 3 | `quantile by (job) (0.9, data)` | Quantile | job | 1s | 1s | 1 | +| 9 | 3 | `quantile by (job) (0.95, data)` | Quantile | job | 1s | 1s | 1 | +| 10 | 3 | `quantile by (job) (0.99, data)` | Quantile | job | 1s | 1s | 1 | +| 11 | 4 | `sum_over_time(data[15m])` | SumOrCount | series | 15m | 1m | 1 | +| 12 | 4 | `sum_over_time(data[1h])` | SumOrCount | series | 1h | 1m | 1 | +| 13 | 4 | `sum_over_time(data[6h])` | SumOrCount | series | 6h | 1m | 1 | +| 14 | 4 | `sum_over_time(data[24h])` | SumOrCount | series | 24h | 1m | 1 | +| 15 | 5 | `quantile_over_time(0.5, data[15m])` | Quantile | series | 15m | 1m | 1 | +| 16 | 5 | `quantile_over_time(0.5, data[1h])` | Quantile | series | 1h | 1m | 1 | +| 17 | 5 | `quantile_over_time(0.5, data[6h])` | Quantile | series | 6h | 1m | 1 | +| 18 | 5 | `quantile_over_time(0.5, data[24h])` | Quantile | series | 24h | 1m | 1 | +| 19 | 5 | `quantile_over_time(0.75, data[15m])` | Quantile | series | 15m | 1m | 1 | +| 20 | 5 | `quantile_over_time(0.75, data[1h])` | Quantile | series | 1h | 1m | 1 | +| 21 | 5 | `quantile_over_time(0.75, data[6h])` | Quantile | series | 6h | 1m | 1 | +| 22 | 5 | `quantile_over_time(0.75, data[24h])` | Quantile | series | 24h | 1m | 1 | +| 23 | 5 | `quantile_over_time(0.9, data[15m])` | Quantile | series | 15m | 1m | 1 | +| 24 | 5 | `quantile_over_time(0.9, data[1h])` | Quantile | series | 1h | 1m | 1 | +| 25 | 5 | `quantile_over_time(0.9, data[6h])` | Quantile | series | 6h | 1m | 1 | +| 26 | 5 | `quantile_over_time(0.9, data[24h])` | Quantile | series | 24h | 1m | 1 | +| 27 | 5 | `quantile_over_time(0.95, data[15m])` | Quantile | series | 15m | 1m | 1 | +| 28 | 5 | `quantile_over_time(0.95, data[1h])` | Quantile | series | 1h | 1m | 1 | +| 29 | 5 | `quantile_over_time(0.95, data[6h])` | Quantile | series | 6h | 1m | 1 | +| 30 | 5 | `quantile_over_time(0.95, data[24h])` | Quantile | series | 24h | 1m | 1 | +| 31 | 5 | `quantile_over_time(0.99, data[15m])` | Quantile | series | 15m | 1m | 1 | +| 32 | 5 | `quantile_over_time(0.99, data[1h])` | Quantile | series | 1h | 1m | 1 | +| 33 | 5 | `quantile_over_time(0.99, data[6h])` | Quantile | series | 6h | 1m | 1 | +| 34 | 5 | `quantile_over_time(0.99, data[24h])` | Quantile | series | 24h | 1m | 1 | +| 35 | 6 | `rate(data[15m])` | RateOrIncrease | series | 15m | 1m | 1 | +| 36 | 6 | `rate(data[1h])` | RateOrIncrease | series | 1h | 1m | 1 | +| 37 | 6 | `rate(data[6h])` | RateOrIncrease | series | 6h | 1m | 1 | +| 38 | 6 | `rate(data[24h])` | RateOrIncrease | series | 24h | 1m | 1 | +| 39 | 7 | `sum by (job) (rate(data[15m]))` | RateOrIncrease | job | 15m | 1m | 1 | +| 40 | 7 | `sum by (job) (rate(data[1h]))` | RateOrIncrease | job | 1h | 1m | 1 | +| 41 | 7 | `sum by (job) (rate(data[6h]))` | RateOrIncrease | job | 6h | 1m | 1 | +| 42 | 7 | `sum by (job) (rate(data[24h]))` | RateOrIncrease | job | 24h | 1m | 1 | +| 43 | 8 | `sum by (job) (sum_over_time(data[15m]))` | SumOrCount | job | 15m | 1m | 1 | +| 44 | 8 | `sum by (job) (sum_over_time(data[1h]))` | SumOrCount | job | 1h | 1m | 1 | +| 45 | 8 | `sum by (job) (sum_over_time(data[6h]))` | SumOrCount | job | 6h | 1m | 1 | +| 46 | 8 | `sum by (job) (sum_over_time(data[24h]))` | SumOrCount | job | 24h | 1m | 1 | +| 47 | 9 | `topk(32, sum by (label_0) (rate(data[15m])))` | TopK over increases | — (keys: label_0) | 15m | 1m | 1 | +| 48 | 9 | `topk(32, sum by (label_0) (rate(data[1h])))` | TopK over increases | — (keys: label_0) | 1h | 1m | 1 | +| 49 | 9 | `topk(32, sum by (label_0) (rate(data[6h])))` | TopK over increases | — (keys: label_0) | 6h | 1m | 1 | +| 50 | 9 | `topk(32, sum by (label_0) (rate(data[24h])))` | TopK over increases | — (keys: label_0) | 24h | 1m | 1 | +| 51 | 10 | `quantile_over_time(0.9, data[15m]) / quantile_over_time(0.5, data[15m])` | 2 × Quantile | series | 15m | 1m | 2 | +| 52 | 10 | `quantile_over_time(0.9, data[1h]) / quantile_over_time(0.5, data[1h])` | 2 × Quantile | series | 1h | 1m | 2 | +| 53 | 10 | `quantile_over_time(0.9, data[6h]) / quantile_over_time(0.5, data[6h])` | 2 × Quantile | series | 6h | 1m | 2 | +| 54 | 10 | `quantile_over_time(0.9, data[24h]) / quantile_over_time(0.5, data[24h])` | 2 × Quantile | series | 24h | 1m | 2 |
@@ -368,15 +350,15 @@ heavy, refreshed every 1 m, over the same windows `W`. | D1 | `quantile_over_time(q, data[w])`, q ∈ {0.5, 0.9, 0.99} | Quantile | per series | 12 | | D2 | `quantile by (job) (q, data)`, q ∈ {0.5, 0.9, 0.99} | Quantile | `job` | 3 | | D3 | `sum by (job) (rate(data[w]))` | RateOrIncrease | `job` | 4 | -| D4 | `topk(k, sum by (label_0) (rate(data[w])))`, w ∈ {15m, 1h} | TopK over per-series increases | none; keys `label_0` | 6 | +| D4 | `topk(32, sum by (label_0) (rate(data[w])))`, w ∈ {15m, 1h} | TopK over per-series increases | none; keys `label_0` | 2 | | D5 | `quantile_over_time(0.99, data[w]) / quantile_over_time(0.5, data[w])`, w ∈ {15m, 1h} | Two Quantile RQEs | per series | 4 | -29 RQEs per replica; D5's operands repeat D1's, so 25 are distinct. D1's +25 RQEs per replica; D5's operands repeat D1's, so 21 are distinct. D1's quantiles share one stream across overlapping windows; D3 and D4 share the increase stream.
-Every query of the dashboard set (27 queries, 29 RQEs) +Every query of the dashboard set (23 queries, 25 RQEs) | # | Template | PromQL | Capability | Grouping | `S` | `T` | RQEs | |---|---|---|---|---|---|---|---| @@ -399,14 +381,10 @@ increase stream. | 17 | D3 | `sum by (job) (rate(data[1h]))` | RateOrIncrease | job | 1h | 1m | 1 | | 18 | D3 | `sum by (job) (rate(data[6h]))` | RateOrIncrease | job | 6h | 1m | 1 | | 19 | D3 | `sum by (job) (rate(data[24h]))` | RateOrIncrease | job | 24h | 1m | 1 | -| 20 | D4 | `topk(100, sum by (label_0) (rate(data[15m])))` | TopK over increases | — (keys: label_0) | 15m | 1m | 1 | -| 21 | D4 | `topk(100, sum by (label_0) (rate(data[1h])))` | TopK over increases | — (keys: label_0) | 1h | 1m | 1 | -| 22 | D4 | `topk(200, sum by (label_0) (rate(data[15m])))` | TopK over increases | — (keys: label_0) | 15m | 1m | 1 | -| 23 | D4 | `topk(200, sum by (label_0) (rate(data[1h])))` | TopK over increases | — (keys: label_0) | 1h | 1m | 1 | -| 24 | D4 | `topk(300, sum by (label_0) (rate(data[15m])))` | TopK over increases | — (keys: label_0) | 15m | 1m | 1 | -| 25 | D4 | `topk(300, sum by (label_0) (rate(data[1h])))` | TopK over increases | — (keys: label_0) | 1h | 1m | 1 | -| 26 | D5 | `quantile_over_time(0.99, data[15m]) / quantile_over_time(0.5, data[15m])` | 2 × Quantile | series | 15m | 1m | 2 | -| 27 | D5 | `quantile_over_time(0.99, data[1h]) / quantile_over_time(0.5, data[1h])` | 2 × Quantile | series | 1h | 1m | 2 | +| 20 | D4 | `topk(32, sum by (label_0) (rate(data[15m])))` | TopK over increases | — (keys: label_0) | 15m | 1m | 1 | +| 21 | D4 | `topk(32, sum by (label_0) (rate(data[1h])))` | TopK over increases | — (keys: label_0) | 1h | 1m | 1 | +| 22 | D5 | `quantile_over_time(0.99, data[15m]) / quantile_over_time(0.5, data[15m])` | 2 × Quantile | series | 15m | 1m | 2 | +| 23 | D5 | `quantile_over_time(0.99, data[1h]) / quantile_over_time(0.5, data[1h])` | 2 × Quantile | series | 1h | 1m | 2 |
@@ -418,7 +396,7 @@ keeps only what changes the comparison with AutoSketch. | Dimension | Values | What it varies | |---|---|---| | Template set | **dashboard**; the 10 templates | Workload realism, and which capabilities appear | -| Replicas `r` | **1**, 8, 64 | Every replica reads the same stream with a seeded random subset of 3 windows from `W`, 3 quantiles from {0.5, 0.75, 0.9, 0.95, 0.99}, a `k` from {100, 200, 300} and `T` from {10 s, 1 m, 5 m}; identical RQEs are deduplicated. Many users or dashboards over the same metrics: how the sharing benefit and planning time grow with the number of RQEs | +| Replicas `r` | **1**, 8, 64 | Every replica reads the same stream with a seeded random subset of 3 windows from `W`, 3 quantiles from {0.5, 0.75, 0.9, 0.95, 0.99} and `T` from {10 s, 1 m, 5 m}; identical RQEs are deduplicated. Many users or dashboards over the same metrics: how the sharing benefit and planning time grow with the number of RQEs | | Accuracy target (strictness) | loose, **default**, strict | What AutoSketch optimizes for (§5) | | Latency SLA | the §5 grid, **no limit** | §5 | @@ -466,14 +444,13 @@ HLL `lg_k` ∈ {12, 14, 16}, CMS-heap top-k with rows ∈ {3, 5}, cols ∈ {256, DDSketch, CountSketch-heap and UnivMon are not deployable by default. **A saturation run at the synthetic workload's data shape.** The grid has no -point at θ = 1.1 and K = 1e4, measures quantiles on Pareto data, and scores -top-k only at k = 32. The synthetic workload therefore needs one targeted run +point at θ = 1.1 and K = 1e4, and measures quantiles on Pareto data. The synthetic +workload therefore needs one targeted run of `scripts/study_saturation.py` at its own data, so every lookup lands on a measured point instead of the worst of the bracketing grid points (#156 Q8): - Zipf s = 1.1 over 10,000 keys, the same data as the cost table; - KLL and DDSketch on the Zipf ranks, DDSketch in value relative error (#162); -- CMS-heap with heap ∈ {100, 200, 300}, each scored at every `k` ∈ {100, 200, - 300} up to its heap; +- CMS-heap at its measured heap, k = 32; - `N` up to 1e9 for top-k, 1e8 for the quantile sketches. **Exact accumulators** have no saturation point: their answer is exact at any @@ -521,8 +498,8 @@ and `m = S/x` the windows a query merges: error is then read at `n(S, G)` as above (#156 Q11). CMS, CountSketch, HLL and DDSketch merge exactly, so this is their single-sketch error. KLL and top-k lose accuracy when merged; that penalty is the follow-up #158. -- **Top-k:** precision is read at the RQE's `k` from the curve of the - deployment's heap size, so a heap-300 deployment serves k = 100, 200 and 300. +- **Top-k:** `precision_at_k` at k = 32, the heap size every top-k config is + measured at. - **AutoSketch** never merges (`x = S`): the same lookup with `m = 1` at its window's size, `n(S, G)` (#156 Q12). It configures once, before deployment (§9 Q2). @@ -582,7 +559,7 @@ Figures: | — | sketch-bench #151, #152, #154, #155 | Cost-table fixes, value range, accuracy after merging, `measured_at` (#147) | Merged; #154's merge rule is replaced by #156 | | — | sketch-bench #156, #162, #158 | Accuracy from the saturation curves at n(T) with `data_shape`; DDSketch in value relative error; merge penalties | Open; #156 blocked on #162 | | — | sketch-bench #157 | Re-export the cost table (grid configs, `CARDINALITY=10000`); per-item CPU and memory vs. size | Open; its accuracy part is covered by #156 | -| — | sketch-bench (to open) | Top-k for this workload: heap size as a parameter, precision scored at k ∈ {100, 200, 300}; `rqe-optimizer` RQEs carry `k`, and a heap-`K` deployment serves `k ≤ K`; the targeted saturation run at θ = 1.1, K = 1e4 (§6) | Not started | +| — | sketch-bench (to open) | The targeted saturation run at θ = 1.1, K = 1e4 with quantiles on the Zipf ranks (§6) | Not started | | — | sketch-bench #130, #131 | Saturation curves; accuracy after merging `m` shards | Merged; background | | 1 | sketch-bench #137 | Retained memory, EC2 pricing, `milp::minimize_cost` | Merged; superseded by #145 | | 3 | sketch-bench #135 | AutoSketch-Adapted (Algorithm 4), aligned with the paper's EXAMINE rule and seeding | Merged | From 5f57cdfc971f61dc85dc03bdf96968e23059f74d Mon Sep 17 00:00:00 2001 From: zz_y Date: Wed, 7 Oct 2026 23:27:05 +0000 Subject: [PATCH 35/39] docs(planner): one p95 accuracy level; inputs from one sketch-bench study MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - §5: a single p95 level in each family's metric (error ≤ 0.05, top-k precision ≥ 0.95) for synthetic and traces; the three strictness levels and fitted trace targets are dropped, and so is the grid's strictness dimension. - §2/§6: costs and accuracy come from one study on asap_sketchlib 0.3.0: the cost table from --phase optimizer-cost (#174), a full-cross grid with the cost shape (#186), KLL merge curves (#179), top-k heap m·k (#182), per-query k with curves at 10/32/100 (#185) and the theoretical fallback (#180). Quantiles see Pareto a = 2, the cost shape. - §6: traces are Alibaba 2022 and Google 2011; BOOM is left out for now. - §8/§10: statuses as of 2026-10-07; the top-k m·k read is an assumption. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis --- docs/evaluation/autosketch-vs-planner.md | 174 ++++++++++++----------- 1 file changed, 90 insertions(+), 84 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index f2d4cee9..96df2162 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -1,6 +1,6 @@ # Evaluation plan: AutoSketch vs. the ASAPQuery planner (paper §6.3) -Status: revised 2026-10-06 to sketch-bench #145's cost model and capabilities (#144). The evaluation code (#138–#141) is reworked onto it before the runs; earlier results are obsolete. The design decisions are settled in §9. +Status: revised 2026-10-07. One accuracy level, p95 (§5). Costs and accuracy come from one sketch-bench study on asap_sketchlib 0.3.0 (#174, #178–#186, §6). The evaluation code is sketch-bench #138, which #139 and #141 were folded into. The design decisions are settled in §9. ## 1. Question @@ -30,9 +30,11 @@ interval) and sums the results; the planner is invoked once for the batch. | Earlier protocol (E1–E3, end-to-end execution) | ASAPQuery-backend `docs/evaluation/autosketch-comparison.md` ([#545](https://github.com/ProjectASAP/ASAPQuery-backend/pull/545)) | Merged. Execution-based; this plan is planner-level and uses estimated costs instead. | | Top-K dashboard comparison | ASAPQuery-backend [#602](https://github.com/ProjectASAP/ASAPQuery-backend/pull/602) | Closed, not merged. | | RQE deployment MILP (HiGHS): candidates `(capability, config, labels, x, y)`, sharing, latency bounds | sketch-bench `rqe-optimizer/` ([#129](https://github.com/ProjectASAP/sketch-bench/pull/129)); per-phase cost model and weighted objective [#145](https://github.com/ProjectASAP/sketch-bench/pull/145); exact accumulators and top-k families [#144](https://github.com/ProjectASAP/sketch-bench/pull/144) | Merged. **This is the planner we evaluate.** | -| Measured per-operation costs (`AtomicCostEntry`: memory/instance, insert/merge/query CPU, accuracy) | sketch-bench `scripts/export_rqe_optimizer_costs.sh` | Merged. **Source of CPU and memory** (§6), and of exact accumulators' rows. Each row records its measurement conditions (`measured_at`, #155). Sketch rows' accuracy is not read from it (#156 Q9). | -| Saturation study: error vs. `N`, `N_sat`, cost per config and shape | sketch-bench [#130](https://github.com/ProjectASAP/sketch-bench/pull/130) | Merged. **Source of sketch accuracy**: the planner reads the error curve at each instance's item count (#156). | -| Accuracy after merging `m` shards | sketch-bench [#131](https://github.com/ProjectASAP/sketch-bench/pull/131), carried into the cost table by [#154](https://github.com/ProjectASAP/sketch-bench/pull/154) | Merged. Merge penalties for KLL and top-k are a follow-up ([#158](https://github.com/ProjectASAP/sketch-bench/issues/158)); until then #156 requires every merged pane to be saturated. | +| Measured per-operation costs (`AtomicCostEntry`: memory/instance, insert/merge/query CPU, accuracy) | sketch-bench `scripts/study_saturation.py --phase optimizer-cost` → `rqe_atomic_costs.json` ([#174](https://github.com/ProjectASAP/sketch-bench/issues/174), [#178](https://github.com/ProjectASAP/sketch-bench/pull/178)) | Merged. **Source of CPU and memory** (§6), and of exact accumulators' rows. Measured serially, alone on one machine, at the synthetic data's shape. Each row records its measurement conditions (`measured_at`) and names its `accuracy_metric`; its accuracy is the seed mean, and is cross-checked against the curves when the planner loads. | +| Saturation study: error vs. `N` per config and data shape | sketch-bench `scripts/study_saturation.py --phase accuracy` ([#130](https://github.com/ProjectASAP/sketch-bench/pull/130), [#186](https://github.com/ProjectASAP/sketch-bench/pull/186)) | Merged. **Source of sketch accuracy**: the planner reads the error curve at each instance's item count. The grid is a full cross of data shapes and includes the cost table's shape as a row and column; the planner refuses a grid with holes. | +| Accuracy after merging `m` windows | sketch-bench [#131](https://github.com/ProjectASAP/sketch-bench/pull/131); merge curves read by the planner [#179](https://github.com/ProjectASAP/sketch-bench/pull/179); top-k heap sized `m · k` [#182](https://github.com/ProjectASAP/sketch-bench/pull/182) | Merged. KLL reads measured merge curves. A top-k heap of `m · k` is read as one sketch (§6). | +| Per-query top-k `k` | sketch-bench [#185](https://github.com/ProjectASAP/sketch-bench/pull/185) | Merged. Curves at k ∈ {10, 32, 100}; the query's own `k` sizes and prices the heap. | +| Theoretical fallback | sketch-bench [#180](https://github.com/ProjectASAP/sketch-bench/pull/180) | Merged. Where no curve covers a point, the sketch's guarantee at 95% confidence, never better than what was measured (§6). | | Moving the MILP into ASAPQuery's planner | ASAPQuery `asap-planner-rs/src/optimizer/` (Milind; related: [#776](https://github.com/ProjectASAP/ASAPQuery/pull/776), [#725](https://github.com/ProjectASAP/ASAPQuery/pull/725)) | Out of scope: the evaluation uses sketch-bench `rqe-optimizer` and is not rerun on `asap-planner-rs`. | AutoSketch-Adapted is implemented in sketch-bench `rqe-optimizer/src/autosketch.rs` (#135). @@ -91,7 +93,7 @@ has lookback `S` and repeat interval `T`. Per-instance costs are measured | --- | --- | --- | | Ingest | `λ · (x/y) · c_ins` | `card(G) · m · x/y` (open windows) | | Merge | `card(G) · (S/x − 1) · c_mrg / T` | `card(G) · m` (one accumulator per group), 0 when `S = x` | -| Query | `card(G) · c_qry / T` | `card(G) · 8 B`; top-k `32 · 16 B` | +| Query | `card(G) · c_qry / T` | `card(G) · 8 B`; top-k `k · 16 B` | | Storage | 0 | `card(G) · m · ((max S − x)/y + 1)` (closed windows) | Ingest and storage are paid once per active deployment; merge and query once @@ -134,32 +136,28 @@ Loads are reported in vCPU (core-seconds per second) and totals in CPU-hours. ## 5. Constraints -**Accuracy target**: one of three strictness levels, each with its own target -per capability. The default level matches the `traces` targets from ASAPQuery's -dataset analysis. The synthetic workload sweeps all three; `traces` uses its -fitted targets. +**Accuracy target**: one level, p95: 95% accuracy in each family's own +metric. It applies to every workload, `synthetic` and `traces` alike. -| Level | Quantile, KLL (rank error) | Quantile, DDSketch (value relative error) | TopK (precision@k) | +| Quantile, KLL (rank error) | Quantile, DDSketch (value relative error) | Cardinality, HLL (relative error) | TopK (precision@k) | | --- | --- | --- | --- | -| loose | ≤ 0.02 | ≤ 0.02 | ≥ 0.90 | -| **default** | **≤ 0.01** | **≤ 0.01** | **≥ 0.95** | -| strict | ≤ 0.005 | ≤ 0.005 | ≥ 0.99 | +| ≤ 0.05 | ≤ 0.05 | ≤ 0.05 | ≥ 0.95 | -Each sketch is scored in the metric its guarantee is stated in (#156 Q9): +Each sketch is scored in the metric its guarantee is stated in: KLL in rank error, DDSketch in value relative error `|x̃ − x_q| / |x_q|` -(sketch-bench #162). DDSketch is not deployable by default, so its column only -applies with `--allow-undeployable-families`. +(sketch-bench #162). Sum and increase are served by exact accumulators, so their error is 0 and -every level is met. Strictness therefore matters only for quantile and top-k -RQEs. No template asks for cardinality. +the target is always met. The target therefore matters only for quantile and +top-k RQEs. No synthetic template asks for cardinality. How the accuracy of a deployment is read, including when it merges windows, is §6 "Reading accuracy and cost". -An earlier version mapped one percentage `p` to error ≤ `1 − p` for every -capability. It was dropped (2026-10-05): at 95% it allowed quantile rank error -0.05, so a p99 query could return the p94 value. +Earlier versions used three strictness levels (loose, default, strict, with +default rank error ≤ 0.01) and fitted per-trace targets. Both were replaced by +the single p95 level (2026-10-07). At 95%, a quantile's rank error may reach +0.05, so a p99 query may return a value between the p94 and the p99. **Latency** — one absolute SLA applies to every RQE, swept over {0.01, 0.1, 1, 10} ms and no limit. @@ -184,7 +182,7 @@ grid. | ID | Description | Purpose | | --- | --- | --- | | `synthetic` | Synthetic PromQL workload: the 10 queries in "Synthetic workload" below, over Zipf data. Main figure. | Cost–latency trade-off across data and requirements | -| `traces` | Real-trace RQEs, one workload per dataset: Alibaba 2022, BOOM and Google 2011. Taken from `asap-tools/dataset-analysis/results/skew_summary.csv` ([#746](https://github.com/ProjectASAP/ASAPQuery/pull/746)): each row's `range_s` is `S` and its `step_s` is `T`. Data parameters and accuracy targets are fit over each whole trace. | Appendix: real-trace results | +| `traces` | Real-trace RQEs, one workload per dataset: Alibaba 2022 and Google 2011 (BOOM is left out for now). Taken from `asap-tools/dataset-analysis/results/skew_summary.csv` ([#746](https://github.com/ProjectASAP/ASAPQuery/pull/746)): each row's `range_s` is `S` and its `step_s` is `T`. Data parameters are fit over each whole trace; the accuracy target is the p95 level (§5). | Appendix: real-trace results | ### Synthetic workload @@ -207,7 +205,7 @@ not the data. There is one metric, `data`, with two labels: | Series | C = 10,000 | | Samples per series | 200 per second (one every 5 ms) | | Stream rate `λ` | 2e6 samples/s; 2e5 per job | -| Values | Zipf s = 1.1 over 10,000 keys: the cost table's data (sketch-bench `export_rqe_optimizer_costs.sh` with `CARDINALITY=10000`). Quantile sketches see the Zipf ranks as values; a top-k key's total over a window follows the same Zipf law | +| Values | The cost table's data shape (sketch-bench `study_saturation.py`, `COST_*`): keys are Zipf s = 1.1 over 10,000 keys, so a top-k key's total over a window follows that Zipf law; quantile sketches see Pareto a = 2 values | | Lookback windows `W` | {15m, 1h, 6h, 24h}: at least 15 minutes, so every per-series summary sees at least 1e5 items | | Repeat interval `T` | 1 m for temporal templates; spatial templates read and repeat every 1 s | @@ -224,11 +222,11 @@ ingests decides whether a sketch is worth it: Exact accumulators keep constant state and a constant per-item cost, so their size does not matter. Sketch accuracy is read from the saturation curve at the -instance's item count (§6 "Reading accuracy and cost"). Every quantile sketch -falls inside the curves' range (1e3 to 1e8 items). Top-k sketches exceed it; -#156 reads the error at the largest measured N when the point has saturated -there, which Zipf s = 1.1 keys do. CMS-heap's per-item CPU and memory do not -depend on `N` (sketch-bench #130). +instance's item count (§6 "Reading accuracy and cost"). The grid's curves run +to 1e7 items; the points the workloads read past that are measured to 1e9 +(§6 "Benchmark input"). Past the last measured `N`, a saturated curve's last +value is read, which covers the larger top-k windows. CMS-heap's per-item CPU +and memory do not depend on `N` (sketch-bench #130). **Modeling choice for spatial templates.** A spatial template evaluates every 1 s over every sample of the last second (`S = T = 1 s`). PromQL's instant @@ -267,8 +265,9 @@ Mapping notes: search for sum and increase RQEs. There, ASAP differs from it only by window choice and sharing. - Top-k ranks each `label_0` key's total over the window: the heavy hitters - among 1e4 keys. k = 32 is the heap size sketch-bench measures top-k at - (`CMS_HEAP_TOP_K`), so every top-k RQE reads a measured curve. + among 1e4 keys. Every top-k template uses k = 32. `k` is a per-query + variable (sketch-bench #185); the study measures curves at k ∈ {10, 32, 100}, + so every top-k RQE here reads a measured curve. - The quantiles of one template, and the two operands of template 10, read the same stream. ASAP can serve them from one deployment; AutoSketch gets one deployment per RQE. @@ -397,7 +396,6 @@ keeps only what changes the comparison with AutoSketch. |---|---|---| | Template set | **dashboard**; the 10 templates | Workload realism, and which capabilities appear | | Replicas `r` | **1**, 8, 64 | Every replica reads the same stream with a seeded random subset of 3 windows from `W`, 3 quantiles from {0.5, 0.75, 0.9, 0.95, 0.99} and `T` from {10 s, 1 m, 5 m}; identical RQEs are deduplicated. Many users or dashboards over the same metrics: how the sharing benefit and planning time grow with the number of RQEs | -| Accuracy target (strictness) | loose, **default**, strict | What AutoSketch optimizes for (§5) | | Latency SLA | the §5 grid, **no limit** | §5 | The data model is fixed (§6 "Data model and scale"); no grid dimension changes the data. @@ -429,29 +427,28 @@ in the planner's micro-benchmarks: ### Benchmark input -Two sketch-bench measurements feed every method, per sketch-bench #156: - -- **Accuracy of sketch rows:** the saturation curves (`saturation.csv`, - `saturation_curve.csv`), error vs. items `N` per (config, data shape). The - planner reads them through `--saturation-dir`. -- **CPU and memory, and exact rows:** the cost table from - `scripts/export_rqe_optimizer_costs.sh`, exported with `CARDINALITY=10000`, - the workload's key count. - -Sketch configs are the saturation grid's (#156 Q4): KLL k ∈ {50, 200, 800}, -HLL `lg_k` ∈ {12, 14, 16}, CMS-heap top-k with rows ∈ {3, 5}, cols ∈ {256, -1024, 4096, 16384}. A sketch config with no curve is not eligible. -DDSketch, CountSketch-heap and UnivMon are not deployable by default. - -**A saturation run at the synthetic workload's data shape.** The grid has no -point at θ = 1.1 and K = 1e4, and measures quantiles on Pareto data. The synthetic -workload therefore needs one targeted run -of `scripts/study_saturation.py` at its own data, so every lookup lands on a -measured point instead of the worst of the bracketing grid points (#156 Q8): -- Zipf s = 1.1 over 10,000 keys, the same data as the cost table; -- KLL and DDSketch on the Zipf ranks, DDSketch in value relative error (#162); -- CMS-heap at its measured heap, k = 32; -- `N` up to 1e9 for top-k, 1e8 for the quantile sketches. +One sketch-bench study, on asap_sketchlib 0.3.0, feeds every method. The +planner reads it through `--saturation-dir DIR`: + +- **Accuracy of sketch rows:** the saturation curves, error vs. items `N` per + (config, data shape), seed mean of 3 seeds: + - `DIR/out_grid_1e7_cost/`: the grid to `N` = 1e7. The cost table's shape + (Zipf 1.1 over 1e4 keys; Pareto a = 2) is a full row and column of it, so + the grid stays a full cross (#186). KLL also has merge curves at the + shard counts the workloads' deployments reach; + - `DIR/out_1e9/`: the points the workloads read past 1e7: top-k at the cost + shape, and quantiles at Pareto a = 2 and 3 (alibaba's quantiles), to 1e9 + and 3.16e8 items. +- **CPU and memory, and exact rows:** the cost table, + `DIR/optimizer_cost/rqe_atomic_costs.json`, from `--phase optimizer-cost`, + measured serially, alone on one machine, at the cost shape. + +Sketch configs are the grid's: KLL k ∈ {50, 200, 800}, DDSketch α ∈ {0.005, +0.01, 0.02, 0.05}, HLL `lg_k` ∈ {12, 14, 16}, CMS-heap top-k with rows ∈ +{3, 5}, cols ∈ {256, 1024, 4096, 16384} and k ∈ {10, 32, 100}. A sketch +config with no curve is not eligible. Only the families ASAPQuery deploys +are candidates (`DEPLOYABLE_FAMILIES`): the exact accumulators, KLL, +DDSketch, HLL and CMS-heap top-k. **Exact accumulators** have no saturation point: their answer is exact at any size. They are still benchmarked for cost, on the grouped column specs @@ -469,9 +466,9 @@ from the lower `θ` bound. Longer samples expose worse cases (sketch-bench - **`traces` gives the appendix results.** Its parameters are fit on the full trace. -- **`synthetic`** uses its own fixed data, Zipf s = 1.1 over 10,000 keys - (§6 "Data model and scale"): `data_shape` = (θ = 1.1, K = 1e4), quantiles on - the Zipf ranks. +- **`synthetic`** uses its own fixed data, the cost table's shape (§6 "Data + model and scale"): `data_shape` = (θ = 1.1, K = 1e4) for keys, a = 2 for + quantile values. Each config is benchmarked at these worst-case parameters. AutoSketch §5.2 injects random traffic bursts into synthetic workloads to cover variation over @@ -487,22 +484,30 @@ n(S, G) = λ · S / card(G) items per instance #### Reading accuracy and cost -Accuracy follows sketch-bench #156. With `n(S, G)` the items one instance holds -and `m = S/x` the windows a query merges: - -- **No merge (`m = 1`):** the error is the curve's value at `n(S, G)`. - Between checkpoints, the worse of the two neighbours; below 1,000 items, not - eligible; above the largest measured `N`, that `N`'s value if the point has - saturated, else not eligible (#156 Q6). -- **Merging (`m > 1`):** every pane, `n(x, G)` items, must reach `N_sat`; the - error is then read at `n(S, G)` as above (#156 Q11). CMS, CountSketch, HLL and - DDSketch merge exactly, so this is their single-sketch error. KLL and top-k - lose accuracy when merged; that penalty is the follow-up #158. -- **Top-k:** `precision_at_k` at k = 32, the heap size every top-k config is - measured at. +With `n(S, G)` the items one instance holds and `m = S/x` the windows a +query merges, every method reads the same curves: + +- **Data shape:** the worse of the grid points bracketing the workload's + shape (θ and `K`, or a). The cost shape is on the grid, so the synthetic + workload reads its own points. +- **No merge (`m = 1`):** the curve's value at `n(S, G)`. Between + checkpoints, the worse of the two neighbours; below the first checkpoint, + not eligible; past the last, its value if the curve has saturated by then. +- **Merging (`m > 1`):** HLL and DDSketch merge exactly, so they read the + plain curve. KLL reads its merge curve at `m` shards, the worse of the two + measured shard counts around `m` (#179). +- **Top-k:** precision@k at the query's own `k`, or at the next larger + measured `k` (harder) (#185). The planner sizes each heap `m · k`; a heap + of at least `m · k` is read as one unmerged sketch (#182). This is an + assumption, not a guarantee: the merged heaps are taken to still hold the + true top `k`. +- **No measurement:** where no curve covers a point (past an unsaturated + curve, past the largest measured `k` or shard count), the sketch's + theoretical guarantee at 95% confidence, never better than the last + measured value (#180). Every RQE in this evaluation reads a measurement + (`accuracy_source` in the raw output). - **AutoSketch** never merges (`x = S`): the same lookup with `m = 1` at its - window's size, `n(S, G)` (#156 Q12). It configures once, before deployment - (§9 Q2). + window's size, `n(S, G)`. It configures once, before deployment (§9 Q2). - **Exact accumulators** have zero error at every size. CPU and memory come from the cost table at the configs' measured size; @@ -526,8 +531,8 @@ median of repeated runs for timings: also benchmarks a fixed-size representative workload, not the query window's full data; - a **paper-rate estimate**, 60 s per distinct probe. - - *ASAP's one-time profiling:* the wall time of the sketch-bench cost export - that produced its table. It is shared by all RQEs + - *ASAP's one-time profiling:* the wall time of the sketch-bench study runs + that produced its curves and cost table. It is shared by all RQEs and reusable across workloads, so it is reported once, plus amortized per RQE served. - The paper's figure shows search + benchmark per method, stacked. @@ -552,19 +557,20 @@ Figures: ## 8. Who implements what, in which PR -| # | Repo / PR | Scope | Status (2026-10-06) | +| # | Repo / PR | Scope | Status (2026-10-07) | | --- | --- | --- | --- | | this | ASAPQuery #777 | This plan | Draft, updated as decisions change | | — | sketch-bench #144, #145 | Exact accumulators and top-k families; per-phase cost model and weighted objective | Merged | -| — | sketch-bench #151, #152, #154, #155 | Cost-table fixes, value range, accuracy after merging, `measured_at` (#147) | Merged; #154's merge rule is replaced by #156 | -| — | sketch-bench #156, #162, #158 | Accuracy from the saturation curves at n(T) with `data_shape`; DDSketch in value relative error; merge penalties | Open; #156 blocked on #162 | -| — | sketch-bench #157 | Re-export the cost table (grid configs, `CARDINALITY=10000`); per-item CPU and memory vs. size | Open; its accuracy part is covered by #156 | -| — | sketch-bench (to open) | The targeted saturation run at θ = 1.1, K = 1e4 with quantiles on the Zipf ranks (§6) | Not started | +| — | sketch-bench #151, #152, #155 | Cost-table fixes, value range, `measured_at` (#147) | Merged | +| — | sketch-bench #174 (#178) | One study as the only source: the cost table from `--phase optimizer-cost`, `accuracy_metric` per row, seed-mean accuracy | Merged | +| — | sketch-bench #179, #182, #185 | KLL merge curves; top-k heap `m · k`; per-query top-k `k` | Merged | +| — | sketch-bench #180, #186 | Theoretical fallback; the cost shape as a full grid row and column | Merged | +| — | asap_sketchlib 0.3.0 | Sketch runtime the study and the planner use | Pinned; the study was rerun on it | | — | sketch-bench #130, #131 | Saturation curves; accuracy after merging `m` shards | Merged; background | | 1 | sketch-bench #137 | Retained memory, EC2 pricing, `milp::minimize_cost` | Merged; superseded by #145 | | 3 | sketch-bench #135 | AutoSketch-Adapted (Algorithm 4), aligned with the paper's EXAMINE rule and seeding | Merged | | 2 | sketch-bench #136 | Evaluation table for the trace workloads | Merged | -| 4 | sketch-bench #138, #139, #140, #141 | Runner, synthetic workload, saturation at K = 1e1–1e6, two cost models and grid driver | Open; to be reworked onto #145's objective, #156's lookup and the reduced grid. #140's K = 1e4 points are superseded for the synthetic workload by the targeted run (§6) | +| 4 | sketch-bench #138 | Runner, synthetic workload and grid, figures (#139 and #141 folded in; #140 closed) | Open; rebased on main, p95 runs done | ## 9. Decisions @@ -602,10 +608,10 @@ were dropped (2026-10-06). ## 10. Known limitations -- **Merged accuracy is read, not replayed.** Exact for CMS, CountSketch, HLL - and DDSketch; for KLL and top-k the merge penalty is ignored until #158, with - every pane required to be saturated. Replay the synthetic default workload's - chosen plans in sketch-bench once to confirm. +- **Merged accuracy is read, not replayed.** Exact for HLL and DDSketch; + KLL reads measured merge curves. A top-k heap of `m · k` is assumed to keep + the true top `k` after merging, with no theoretical basis (§6). Replay the + synthetic default workload's chosen plans in sketch-bench once to confirm. - Costs and latencies are estimates from per-operation measurements, not end-to-end executions. The execution-based comparison is ASAPQuery-backend #545/#547. From 6be88a607133db5562f880d7ed66541ef861e5da Mon Sep 17 00:00:00 2001 From: zz_y Date: Thu, 8 Oct 2026 14:07:29 +0000 Subject: [PATCH 36/39] =?UTF-8?q?docs(planner):=20metrics=20dimension=20fo?= =?UTF-8?q?r=20planning-time=20scaling;=20shared=20r=20=E2=88=88=20{1,=208?= =?UTF-8?q?}?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Shared replicas on one metric stop adding RQEs past r = 8, so r = 64 is dropped. m ∈ {1, 8, 16} copies of the template set, each on its own metric, grow RQEs as 21 · m and drive the planning-time figure. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis --- docs/evaluation/autosketch-vs-planner.md | 11 ++++++----- 1 file changed, 6 insertions(+), 5 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 96df2162..f744404d 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -176,7 +176,7 @@ cost. Two workloads, decided 2026-10-05. The earlier `example` (8 RQEs from `small_problem`) and `scaling` (random RQE batches) workloads were dropped. -Planning-time scaling is now the replica dimension of the synthetic workload +Planning-time scaling is now the metrics dimension of the synthetic workload grid. | ID | Description | Purpose | @@ -395,7 +395,8 @@ keeps only what changes the comparison with AutoSketch. | Dimension | Values | What it varies | |---|---|---| | Template set | **dashboard**; the 10 templates | Workload realism, and which capabilities appear | -| Replicas `r` | **1**, 8, 64 | Every replica reads the same stream with a seeded random subset of 3 windows from `W`, 3 quantiles from {0.5, 0.75, 0.9, 0.95, 0.99} and `T` from {10 s, 1 m, 5 m}; identical RQEs are deduplicated. Many users or dashboards over the same metrics: how the sharing benefit and planning time grow with the number of RQEs | +| Shared replicas `r` | **1**, 8 | Every replica reads the same stream with a seeded random subset of 3 windows from `W`, 3 quantiles from {0.5, 0.75, 0.9, 0.95, 0.99} and `T` from {10 s, 1 m, 5 m}; identical RQEs are deduplicated. Many users or dashboards over the same metrics: the sharing benefit. The dashboard has 12 such combinations, so RQEs stop growing past r = 8 (52 at r = 8, 57 at r = 64) | +| Metrics `m` | **1**, 8, 16 | `m` copies of the template set, copy `i` on its own metric `data_i` with the same data model; nothing is shared across copies, so RQEs grow as 21 · `m` (168, 336). Planning time vs. the number of RQEs | | Latency SLA | the §5 grid, **no limit** | §5 | The data model is fixed (§6 "Data model and scale"); no grid dimension changes the data. @@ -411,8 +412,8 @@ AutoSketch-Adapted, PerQuery-CostAware) at both weight settings (§4). Per - RQEs excluded by the SLA or unservable. Replicas with disjoint series (each replica filtering `{label_1="v_i"}`) were -dropped: the planner rejects spatial filters, and the shared mode is the case -that shows sharing. +dropped: the planner rejects spatial filters. The metrics dimension gives +disjoint copies without filters, each copy on its own metric (2026-10-08). #### Planner sensitivity (not part of this comparison) @@ -550,7 +551,7 @@ Figures: 1. Synthetic workload, objective vs. achieved max estimated latency, one panel per weight setting (main paper figure). -2. Planning time vs. number of RQEs (synthetic, replica dimension), log–log. +2. Planning time vs. number of RQEs (synthetic, metrics dimension), log–log. 3. Objective vs. each workload-grid dimension (synthetic, one dimension at a time). 4. Objective vs. absolute latency SLA (synthetic default workload, `traces`). From 3be9859486e619ffaa7f026471721c94772f4a78 Mon Sep 17 00:00:00 2001 From: zz_y Date: Thu, 8 Oct 2026 14:24:01 +0000 Subject: [PATCH 37/39] docs(planner): AutoSketch benchmarks each probed config once per metric Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis --- docs/evaluation/autosketch-vs-planner.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index f744404d..22f2fa05 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -526,7 +526,8 @@ median of repeated runs for timings: MILP solve. - *Benchmark time:* AutoSketch benchmarks every probed configuration, as in the paper (§5.2, Exp#9: 1–2 minutes per config, about 6.5 minutes per - application). Reported two ways, per distinct probed (config, input size): + application). Each probed config is benchmarked once per metric, on that + metric's data, so the time grows with the metrics dimension. Reported two ways: - a **lower bound**, `N_bench · insert_cpu_per_item + one query phase` with `N_bench = 1e8`, the size sketch-bench benchmarks at. The paper also benchmarks a fixed-size representative workload, not the query From 1490317f21423443989c96d01e558d475e549ae4 Mon Sep 17 00:00:00 2001 From: zz_y Date: Thu, 8 Oct 2026 15:11:05 +0000 Subject: [PATCH 38/39] docs(planner): finer SLA grid; AutoSketch meets every SLA it is held to Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis --- docs/evaluation/autosketch-vs-planner.md | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 22f2fa05..7d53da18 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -160,13 +160,15 @@ the single p95 level (2026-10-07). At 95%, a quantile's rank error may reach 0.05, so a p99 query may return a value between the p94 and the p99. **Latency** — one absolute SLA applies to every RQE, swept over -{0.01, 0.1, 1, 10} ms and no limit. +{0.003, 0.01, 0.03, 0.1, 0.3, 1, 3, 10} ms and no limit. - An RQE that no method can meet at a given SLA is excluded from every method at that SLA and reported by ID. Costs at different SLAs therefore cover different RQE sets; compare methods only at one SLA. - ASAP and PerQuery-CostAware must meet the SLA. - AutoSketch-Adapted ignores it. Its violations are counted, and its cost is - shown for those points but marked as infeasible. + shown for those points but marked as infeasible. In practice it meets every + SLA: it never merges, so each RQE's latency is within 1.4% of the lowest any + deployment reaches, and RQEs no method can meet are excluded for all. An earlier version set `L_r = α × the fastest latency of r`. It was dropped: on the dropped `example` workload, α = 2 forced plans with no merging at 40× the unconstrained From adcd6a26ef9b86583e227613d85a102dff893763 Mon Sep 17 00:00:00 2001 From: zz_y Date: Thu, 8 Oct 2026 20:43:13 +0000 Subject: [PATCH 39/39] docs(planner): cost by use, two latency versions, the mixed set, results MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - §4: plans priced by use, w_cpu * AUC(CPU) + w_mem * AUC(memory), CPU elastic: ingest on ceil(rho) workers split by sample, compaction at window close, query jobs; memory as ingest, storage, compaction and query, each counted once; latency is a batch's longest chain (model in sketch-bench #188). A peak-billed cost model was dropped. - §5: latency is reported, in two versions: no SLA, with each method's cost-latency frontier, and a batch latency SLA over {100 ms .. 10 s}. - §3: PerQuery is the same MILP over each RQE's own candidates (no sharing). - §6: the evaluated set is the mixed template set, with shared r in {1, 8} and metrics m in {1, 8, 16}; the dashboard set is not evaluated. - §7: AutoSketch's planning time is search plus its measured benchmark (approxbench accuracy runs at 1e8 items); figures are the frontier, cost vs. SLA and planning time; absolute costs only. - §8-§10: PRs (#188, ASAPQuery #812, #138 stacked on #188), decisions Q3, Q5-Q7, limitations (elastic CPU, compaction merge cost). - §11: results on the synthetic mixed set. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01FhcWJExZcmVqS6r6rEtjis --- docs/evaluation/autosketch-vs-planner.md | 352 ++++++++++++++--------- 1 file changed, 213 insertions(+), 139 deletions(-) diff --git a/docs/evaluation/autosketch-vs-planner.md b/docs/evaluation/autosketch-vs-planner.md index 7d53da18..72a44be6 100644 --- a/docs/evaluation/autosketch-vs-planner.md +++ b/docs/evaluation/autosketch-vs-planner.md @@ -1,6 +1,6 @@ # Evaluation plan: AutoSketch vs. the ASAPQuery planner (paper §6.3) -Status: revised 2026-10-07. One accuracy level, p95 (§5). Costs and accuracy come from one sketch-bench study on asap_sketchlib 0.3.0 (#174, #178–#186, §6). The evaluation code is sketch-bench #138, which #139 and #141 were folded into. The design decisions are settled in §9. +Status: revised 2026-10-08. One accuracy level, p95 (§5). Costs and accuracy come from one sketch-bench study on asap_sketchlib 0.3.0 (#174, #178–#186, §6). Plans are priced **by use** and their latency is reported, in two versions: no SLA, with each method's cost–latency frontier, and a batch latency SLA (§4, §5; model in sketch-bench #188, `docs/rqe_sketch_deployment_v1.md`, "Cost by use and batch latency"). The synthetic evaluation runs on the mixed template set (§6). The evaluation code is sketch-bench #138, stacked on #188. Results so far are in §11. The design decisions are settled in §9. ## 1. Question @@ -29,7 +29,7 @@ interval) and sums the results; the planner is invoked once for the batch. | AutoSketch Algorithm 4 adaptation (LHS seeds, feasibility-directed width/depth neighbor search, pruning) | ASAPQuery-backend `data_plane/examples/autosketch_comparison.rs` ([#547](https://github.com/ProjectASAP/ASAPQuery-backend/pull/547)) | Merged. CMS/Count Sketch/Bloom only; hardcoded CMS grid; executes sketches to measure accuracy. | | Earlier protocol (E1–E3, end-to-end execution) | ASAPQuery-backend `docs/evaluation/autosketch-comparison.md` ([#545](https://github.com/ProjectASAP/ASAPQuery-backend/pull/545)) | Merged. Execution-based; this plan is planner-level and uses estimated costs instead. | | Top-K dashboard comparison | ASAPQuery-backend [#602](https://github.com/ProjectASAP/ASAPQuery-backend/pull/602) | Closed, not merged. | -| RQE deployment MILP (HiGHS): candidates `(capability, config, labels, x, y)`, sharing, latency bounds | sketch-bench `rqe-optimizer/` ([#129](https://github.com/ProjectASAP/sketch-bench/pull/129)); per-phase cost model and weighted objective [#145](https://github.com/ProjectASAP/sketch-bench/pull/145); exact accumulators and top-k families [#144](https://github.com/ProjectASAP/sketch-bench/pull/144) | Merged. **This is the planner we evaluate.** | +| RQE deployment MILP (HiGHS): candidates `(capability, config, labels, x, y)`, sharing, latency bounds | sketch-bench `rqe-optimizer/` ([#129](https://github.com/ProjectASAP/sketch-bench/pull/129)); per-phase cost model and weighted objective [#145](https://github.com/ProjectASAP/sketch-bench/pull/145); exact accumulators and top-k families [#144](https://github.com/ProjectASAP/sketch-bench/pull/144) | Merged. **This is the planner we evaluate**, with the cost by use and batch latency of sketch-bench #188 (`milp::minimize_usage_cost`). | | Measured per-operation costs (`AtomicCostEntry`: memory/instance, insert/merge/query CPU, accuracy) | sketch-bench `scripts/study_saturation.py --phase optimizer-cost` → `rqe_atomic_costs.json` ([#174](https://github.com/ProjectASAP/sketch-bench/issues/174), [#178](https://github.com/ProjectASAP/sketch-bench/pull/178)) | Merged. **Source of CPU and memory** (§6), and of exact accumulators' rows. Measured serially, alone on one machine, at the synthetic data's shape. Each row records its measurement conditions (`measured_at`) and names its `accuracy_metric`; its accuracy is the seed mean, and is cross-checked against the curves when the planner loads. | | Saturation study: error vs. `N` per config and data shape | sketch-bench `scripts/study_saturation.py --phase accuracy` ([#130](https://github.com/ProjectASAP/sketch-bench/pull/130), [#186](https://github.com/ProjectASAP/sketch-bench/pull/186)) | Merged. **Source of sketch accuracy**: the planner reads the error curve at each instance's item count. The grid is a full cross of data shapes and includes the cost table's shape as a row and column; the planner refuses a grid with holes. | | Accuracy after merging `m` windows | sketch-bench [#131](https://github.com/ProjectASAP/sketch-bench/pull/131); merge curves read by the planner [#179](https://github.com/ProjectASAP/sketch-bench/pull/179); top-k heap sized `m · k` [#182](https://github.com/ProjectASAP/sketch-bench/pull/182) | Merged. KLL reads measured merge curves. A top-k heap of `m · k` is read as one sketch (§6). | @@ -42,12 +42,14 @@ AutoSketch-Adapted is implemented in sketch-bench `rqe-optimizer/src/autosketch. ## 3. Methods compared All methods read the same `AtomicCostTable`, the same RQEs and the same -label-set cardinalities and arrival rates, and are scored by the same cost +label-set cardinalities and arrival rates, and are priced by the same cost function (§4). -1. **ASAP** — sketch-bench `rqe-optimizer` MILP over the whole batch, with - accuracy and latency constraints, minimizing the §4 objective at each weight - setting. +1. **ASAP** — sketch-bench `rqe-optimizer`'s `milp::minimize_usage_cost` over + the whole batch, with sharing: the cheapest plan billed by use (§4) that + meets every RQE's accuracy target, at each weight setting. Version 1 also + solves it under a sweep of latency bounds to trace its cost–latency + frontier; version 2 under each SLA of a grid (§5). 2. **AutoSketch-Adapted** — Algorithm 4 run independently per RQE: - search space: the measured configs of the RQE's capability families (`Capability::families()`); @@ -60,79 +62,94 @@ function (§4). merges nothing; - no sharing: every RQE gets its own deployment, even when two RQEs pick an identical one, so ingest and memory are paid per RQE; - - latency is ignored during search, then checked after. -3. **PerQuery-CostAware** (strawman) — the ASAP MILP solved on each RQE alone, - with the same accuracy and latency requirements, and the results summed. It uses the same objective - and window choices as ASAP but no batching or sharing. ASAP vs. this ablation - isolates the batch/sharing benefit; this ablation vs. AutoSketch-Adapted - isolates the objective/window benefit. -Only AutoSketch-Adapted ignores latency; PerQuery-CostAware must meet the same -requirements as ASAP, so every point in the cost–latency figure except -AutoSketch's is a feasible plan. + - latency is ignored; its plan is priced and its latency computed like the + others'. It is one point, not a frontier. +3. **PerQuery-CostAware** (ablation) — the same MILP with sharing ruled out: + each RQE may use only its own candidates, kept as separate copies, so + every RQE pays its own ingest (`minimize_usage_cost`'s `allowed`). Same + objective, window choices, latency bounds and SLAs as ASAP. ASAP vs. this + ablation isolates the batch/sharing benefit; this ablation vs. + AutoSketch-Adapted isolates the objective/window benefit. A "fewest plans" strawman (minimize the number of deployments, then cost) was considered and dropped (decided 2026-10-05). AutoSketch-Adapted is a planner baseline, not a reproduction of the P4 compiler: stage/page/ALU constraints are dropped. Its accuracy probes read -sketch-bench measurements instead of running a benchmark inside the search, but -the cost of running those benchmarks is charged to its planning time (§9 Q3). +sketch-bench measurements instead of running a benchmark inside the search; +the measured time of running those benchmarks is charged to its planning time +(§7, §9 Q3). ## 4. Cost model -The cost model is sketch-bench `rqe-optimizer`'s (#145), so every method is -scored by the function ASAP optimizes. Disk is excluded. Units: CPU in vCPU -(CPU-seconds per second, the mean over time), memory in GiB. +Plans are priced **by use**: `cost = w_cpu · AUC(CPU) + w_mem · AUC(memory)`, +the mean vCPUs and GiB over time, with CPU elastic (a job gets a core +whenever it is ready). This is billing as on fine-grained autoscaling +platforms (Cloud Run with request-based billing, Dataflow, Flink with an +autoscaler), idealized. Billing for provisioned capacity (VMs, containers +billed by allocation) is not modeled; a peak-billed cost model was +considered and dropped (2026-10-08). The full model, with every formula and +a plain-English explanation, is sketch-bench #188, +`docs/rqe_sketch_deployment_v1.md`, "Cost by use and batch latency"; this +section summarizes it. A deployment `D` groups by labels `G` with window `x` and slide `y`; RQE `r` -has lookback `S` and repeat interval `T`. Per-instance costs are measured -(§6): memory `m`, insert `c_ins`, merge `c_mrg`, query `c_qry`. -`λ = card(series) / scrape interval`. +has lookback `S` and repeat interval `T`, and merges `n = S/x` windows. +Per-instance costs are measured (§6): memory `m`, insert `c_ins`, merge +`c_mrg`, query `c_qry`. `inst` is the instances per window (`card(G)`, or 1 +for a sketch shared by all groups); `w = inst · m` is one window's memory. -| Phase | CPU (vCPU) | Memory (bytes) | +**CPU** (mean vCPUs): + +| Part | When | CPU | | --- | --- | --- | -| Ingest | `λ · (x/y) · c_ins` | `card(G) · m · x/y` (open windows) | -| Merge | `card(G) · (S/x − 1) · c_mrg / T` | `card(G) · m` (one accumulator per group), 0 when `S = x` | -| Query | `card(G) · c_qry / T` | `card(G) · 8 B`; top-k `k · 16 B` | -| Storage | 0 | `card(G) · m · ((max S − x)/y + 1)` (closed windows) | +| Ingest (the precompute) | continuously | `ρ = λ · (x/y) · c_ins` per active deployment | +| Compaction | each window close | `(k − 1) · inst · c_mrg / y` per active deployment | +| Query | each firing | `ℓ / T` per RQE, `ℓ = card(G) · c_qry + inst · (n − 1) · c_mrg` | + +Ingest runs on `k = ⌈ρ⌉` parallel workers split by sample (each keeps its +own open windows); at each window close a compaction job merges the `k` +partial copies into one stored instance. A query merges the stored windows +of its lookback, then estimates. Every job uses at most one core: +sketch-bench's operations are single-threaded (CPU time equals wall time). + +**Memory** (mean GiB), each part counted once: -Ingest and storage are paid once per active deployment; merge and query once -per RQE it serves. Memory sums every term, as if every query evaluates at once. +| Part | Bytes | Held | +| --- | --- | --- | +| Ingest | `w · (x/y) · k` (open windows, per worker) | always | +| Storage | `w · ((max S − x)/y + 1)` (compacted closed windows) | always | +| Compaction | `k · w` (the closed window's partial copies) | while it runs | +| Query | `w · [n > 1]` (merge accumulators) + `card(G)` × output bytes | while it runs | -**Objective** — `w_cpu · CPU + w_mem · Memory_GiB`, summed over the plan: +**Weights:** -1. **CPU only**, `(w_cpu, w_mem) = (1, 0)`: the first run. Memory is still +1. **CPU only**, `(w_cpu, w_mem) = (1, 0)`, in vCPU. Memory is still reported. 2. **Fargate prices**, `w_cpu = 0.0405` $/vCPU-hour and `w_mem = 0.00445` - $/GB-hour (AWS Fargate, us-east-1, Linux/x86, 2026-10-06), so the objective - is in $/hour. CPU costs about 9× memory per unit; serverless pricing charges - exactly these two resources, which is why it fits the model. - -CPU is the mean, i.e. the area under the CPU-over-time curve: plans are sized -for average load, not bursts. Peak-provisioned pricing (buy machines for the -peak CPU) and per-instance-family EC2 pricing were considered and dropped -(2026-10-06). - -**Latency** — per-RQE estimate: -`card(G) · (c_qry + (S/x − 1) · c_mrg)`, the evaluation's CPU time run -serially on one core (an upper bound; parallel execution across instances -would reduce it proportionally). `c_qry` is one query of one instance: one -value of a sum or increase accumulator, one top-k list, or one quantile, so a -template asking `n` quantiles is `n` RQEs. + $/GB-hour (AWS Fargate, us-east-1, Linux/x86, 2026-10-06), so the cost is + in $/hour. + +**Latency.** A batch (all queries fired at one instant) finishes when its +last query does. With elastic CPU no job waits, so a batch's latency is its +longest **chain**: the newest window's compaction, then the query, each on +one core, `(k − 1) · inst · c_mrg + ℓ`. A plan's latency is the longest +chain over its RQEs (the job-placement algorithm in #188's doc, §6, has this +as its closed form). Median and p90 over the RQEs are reported too. **Units and per-operation costs.** CPU is CPU time (user + system), in core-seconds, measured by sketch-bench: -- `c_ins` (`insert_cpu_secs`): the insert phase's CPU ÷ N, in CPU-seconds per item; -- `c_mrg` (`merge_cpu_secs`): merging 16 shards ÷ 15, in CPU-seconds per merge; -- `c_qry` (`query_cpu_secs`): the query phase's CPU ÷ the number of queries in it, per call - (a top-k heap dump is repeated per pass so it is timed above the clock's - floor, #151); -- `m` (`mem_bytes_per_instance`): the self-reported bytes per instance, not process RSS. Top-k counts - its heap and exact accumulators their hash table (#151). Exact accumulators - are priced per group: memory and merge are divided by the group count they - were measured at. - -Loads are reported in vCPU (core-seconds per second) and totals in CPU-hours. +- `c_ins` (`insert_cpu_secs`): the insert phase's CPU ÷ N, per item; +- `c_mrg` (`merge_cpu_secs`): merging 16 shards ÷ 15, per merge (compaction + reuses it for the workers' smaller partial copies, an approximation for + KLL and DDSketch); +- `c_qry` (`query_cpu_secs`): the query phase's CPU ÷ the number of queries + in it, per call (a top-k heap dump is repeated per pass so it is timed + above the clock's floor, #151); +- `m` (`mem_bytes_per_instance`): the self-reported bytes per instance, not + process RSS. Top-k counts its heap and exact accumulators their hash table + (#151). Exact accumulators are priced per group: memory and merge are + divided by the group count they were measured at. ## 5. Constraints @@ -159,20 +176,27 @@ default rank error ≤ 0.01) and fitted per-trace targets. Both were replaced by the single p95 level (2026-10-07). At 95%, a quantile's rank error may reach 0.05, so a p99 query may return a value between the p94 and the p99. -**Latency** — one absolute SLA applies to every RQE, swept over -{0.003, 0.01, 0.03, 0.1, 0.3, 1, 3, 10} ms and no limit. -- An RQE that no method can meet at a given SLA is excluded from every method - at that SLA and reported by ID. Costs at different SLAs therefore cover - different RQE sets; compare methods only at one SLA. -- ASAP and PerQuery-CostAware must meet the SLA. -- AutoSketch-Adapted ignores it. Its violations are counted, and its cost is - shown for those points but marked as infeasible. In practice it meets every - SLA: it never merges, so each RQE's latency is within 1.4% of the lowest any - deployment reaches, and RQEs no method can meet are excluded for all. - -An earlier version set `L_r = α × the fastest latency of r`. It was dropped: -on the dropped `example` workload, α = 2 forced plans with no merging at 40× the unconstrained -cost. +**Latency** — the evaluation's message is lower cost **and** lower latency, +so latency is a reported result, in two versions: + +- **Version 1, no SLA.** Each method's cheapest plan, with its latency + reported. ASAP and PerQuery are also solved under latency bounds `L` + (12 log-spaced from the tightest feasible bound, the largest over RQEs of + each RQE's fastest chain, up to the unbounded plan's latency, plus + AutoSketch's latency): each gives the cheapest plan at most `L` slow, and + together they are the method's cost–latency frontier. AutoSketch is one + point. Main message: at AutoSketch's latency, what each method costs. +- **Version 2, a batch latency SLA.** ASAP and PerQuery must keep the batch + latency at most `L`, for `L` ∈ {100, 300, 1000, 3000, 10000} ms. AutoSketch + ignores the SLA; whether its latency meets each one is reported. + +With elastic CPU, a bound `L` is exactly a filter on pairs (rule out every +deployment whose chain exceeds `L`), so both versions stay exact MILPs. + +Earlier versions bounded each RQE's own latency over a µs-scale SLA grid, +excluding RQEs no method could meet (2026-10-07), and before that set +`L_r = α ×` the fastest latency of `r`. Both were replaced by the batch +latency of a plan priced by use (2026-10-08). ## 6. Workloads @@ -183,7 +207,7 @@ grid. | ID | Description | Purpose | | --- | --- | --- | -| `synthetic` | Synthetic PromQL workload: the 10 queries in "Synthetic workload" below, over Zipf data. Main figure. | Cost–latency trade-off across data and requirements | +| `synthetic` | Synthetic PromQL workload: the 10-template mixed set in "Synthetic workload" below, over Zipf data. Main figure. | Cost–latency trade-off across scales | | `traces` | Real-trace RQEs, one workload per dataset: Alibaba 2022 and Google 2011 (BOOM is left out for now). Taken from `asap-tools/dataset-analysis/results/skew_summary.csv` ([#746](https://github.com/ProjectASAP/ASAPQuery/pull/746)): each row's `range_s` is `S` and its `step_s` is `T`. Data parameters are fit over each whole trace; the accuracy target is the p95 level (§5). | Appendix: real-trace results | ### Synthetic workload @@ -391,27 +415,19 @@ increase stream. #### Workload grid -Each dimension has a default (bold). A workload fixes every dimension. The grid -keeps only what changes the comparison with AutoSketch. +The evaluated template set is the **mixed** set (the 10 templates, 50 +distinct RQEs). The dashboard set above is kept as a description of a +dashboard but is not evaluated (2026-10-08). A workload fixes every +dimension; the sweep is the default, then each dimension alone. | Dimension | Values | What it varies | |---|---|---| -| Template set | **dashboard**; the 10 templates | Workload realism, and which capabilities appear | -| Shared replicas `r` | **1**, 8 | Every replica reads the same stream with a seeded random subset of 3 windows from `W`, 3 quantiles from {0.5, 0.75, 0.9, 0.95, 0.99} and `T` from {10 s, 1 m, 5 m}; identical RQEs are deduplicated. Many users or dashboards over the same metrics: the sharing benefit. The dashboard has 12 such combinations, so RQEs stop growing past r = 8 (52 at r = 8, 57 at r = 64) | -| Metrics `m` | **1**, 8, 16 | `m` copies of the template set, copy `i` on its own metric `data_i` with the same data model; nothing is shared across copies, so RQEs grow as 21 · `m` (168, 336). Planning time vs. the number of RQEs | -| Latency SLA | the §5 grid, **no limit** | §5 | - -The data model is fixed (§6 "Data model and scale"); no grid dimension changes the data. - -The sweep is the default workload, then each dimension varied alone with the -others at their defaults. Every workload runs every baseline (ASAP, -AutoSketch-Adapted, PerQuery-CostAware) at both weight settings (§4). Per -(workload, baseline, weights, SLA), report: -- the objective, mean CPU (vCPU) and memory (GiB), and memory per phase; -- max and median estimated latency, and SLA violations; -- active deployments and instances; -- planning time (AutoSketch: search plus charged benchmark time); -- RQEs excluded by the SLA or unservable. +| Shared replicas `r` | **1**, 8 | Every replica reads the same stream with a seeded random subset of 3 windows from `W`, 3 quantiles from {0.5, 0.75, 0.9, 0.95, 0.99} and `T` from {10 s, 1 m, 5 m}; identical RQEs are deduplicated (50 and 92 RQEs). Many users over the same metrics: the sharing benefit | +| Metrics `m` | **1**, 8, 16 | `m` copies of the template set, copy `i` on its own metric `data_i` with the same data model; nothing is shared across copies, so RQEs grow as 50 · `m` (400, 800). Planning time vs. the number of RQEs | + +Every workload runs every method at both weight settings, in both latency +versions (§5). The data model is fixed (§6 "Data model and scale"); no grid +dimension changes the data. Replicas with disjoint series (each replica filtering `{label_1="v_i"}`) were dropped: the planner rejects spatial filters. The metrics dimension gives @@ -518,50 +534,47 @@ per-item costs are flat in `N` (sketch-bench #130). ## 7. Metrics and figures -Reported per (workload, method, weight setting, latency SLA), -median of repeated runs for timings: - -- **Planning time.** Reported in two parts, because the two planners spend - their time differently: - - *Search time:* AutoSketch is the sum over RQEs of Algorithm 4 wall time, - using table lookups. ASAP is candidate generation, dominance pruning and - MILP solve. - - *Benchmark time:* AutoSketch benchmarks every probed configuration, as in - the paper (§5.2, Exp#9: 1–2 minutes per config, about 6.5 minutes per - application). Each probed config is benchmarked once per metric, on that - metric's data, so the time grows with the metrics dimension. Reported two ways: - - a **lower bound**, `N_bench · insert_cpu_per_item + one query phase` - with `N_bench = 1e8`, the size sketch-bench benchmarks at. The paper - also benchmarks a fixed-size representative workload, not the query - window's full data; - - a **paper-rate estimate**, 60 s per distinct probe. +Reported per (workload, method, weight setting, and bound or SLA), median of +repeated runs for timings: + +- **Cost** by use (§4), with its mean CPU (vCPU, by part: ingest, + compaction, query) and memory (GiB, by part: ingest, storage, compaction, + query). Absolute costs, never ratios to ASAP. +- **Latency:** the plan's batch latency (its longest chain), and the median + and p90 over its RQEs; per-RQE values in the raw output. Version 2 also + reports whether AutoSketch meets each SLA. +- **Planning time:** + - *ASAP and PerQuery:* candidate generation plus MILP solve. + - *AutoSketch:* its search (the sum over RQEs of Algorithm 4, using table + lookups) **plus its measured benchmark time.** In the paper every probe + is a benchmark run (§5.2, Exp#9: 1–2 minutes per config). Here each + distinct probed (config, data shape) runs approxbench's accuracy + benchmark at 1e8 items, generating the data, computing the exact + baseline and scoring, and its wall time is measured + (`scripts/autosketch_benchmark_time.py`, serially on an idle machine). + A config is charged once per metric, on that metric's data, so the time + grows with the metrics dimension. 60 s per probe (the paper's rate) is + shown only as a reference. - *ASAP's one-time profiling:* the wall time of the sketch-bench study runs - that produced its curves and cost table. It is shared by all RQEs - and reusable across workloads, so it is reported once, plus amortized per - RQE served. - - The paper's figure shows search + benchmark per method, stacked. -- **Objective** at each weight setting (§4), with its inputs: mean CPU (vCPU) - and memory (GiB), each split by phase. Baselines are normalized to ASAP at - the same weights. -- **Estimated query latency and latency SLA violations** per method. - - Estimated latency per RQE: `card(G) · (c_qry + (S/x − 1) · c_mrg)` (§4). Report its maximum and median over the RQEs, plus the per-RQE values in the raw output. - - SLA violations: the number of RQEs whose estimated latency exceeds the SLA. Only AutoSketch-Adapted can have any, since the other methods are constrained. + that produced its curves and cost table, reported once. - **Estimated accuracy** per RQE: every method meets its target under its own reading rule (§6, "Benchmark input"). -- Active deployments and total sketch instances. +- Active deployments and ingest workers. -Figures: +Figures (synthetic mixed set): -1. Synthetic workload, objective vs. achieved max estimated latency, one panel - per weight setting (main paper figure). -2. Planning time vs. number of RQEs (synthetic, metrics dimension), log–log. -3. Objective vs. each workload-grid dimension (synthetic, one dimension at a - time). -4. Objective vs. absolute latency SLA (synthetic default workload, `traces`). +1. Version 1: each method's cost–latency frontier, one panel per workload × + weight setting (ASAP and PerQuery as lines, AutoSketch as a point). Main + paper figure. +2. Version 2: cost vs. batch latency SLA, one panel per workload × weight + setting; AutoSketch's cost as a reference, marked where it meets or + misses each SLA. +3. Planning time vs. number of RQEs (metrics dimension), with AutoSketch's + search alone and with its measured benchmark. ## 8. Who implements what, in which PR -| # | Repo / PR | Scope | Status (2026-10-07) | +| # | Repo / PR | Scope | Status (2026-10-08) | | --- | --- | --- | --- | | this | ASAPQuery #777 | This plan | Draft, updated as decisions change | | — | sketch-bench #144, #145 | Exact accumulators and top-k families; per-phase cost model and weighted objective | Merged | @@ -574,7 +587,9 @@ Figures: | 1 | sketch-bench #137 | Retained memory, EC2 pricing, `milp::minimize_cost` | Merged; superseded by #145 | | 3 | sketch-bench #135 | AutoSketch-Adapted (Algorithm 4), aligned with the paper's EXAMINE rule and seeding | Merged | | 2 | sketch-bench #136 | Evaluation table for the trace workloads | Merged | -| 4 | sketch-bench #138 | Runner, synthetic workload and grid, figures (#139 and #141 folded in; #140 closed) | Open; rebased on main, p95 runs done | +| — | sketch-bench #188 | Cost by use (ingest workers split by sample, compaction, memory parts), batch latency, `milp::minimize_usage_cost` with a latency bound and `allowed` candidates; design in `docs/rqe_sketch_deployment_v1.md` | Open | +| — | ASAPQuery #812 | Dataset analysis with 6h and 24h windows over a day of Alibaba (trace workloads) | Open | +| 4 | sketch-bench #138 | Runner, synthetic workload and grid, both latency versions, AutoSketch's measured benchmark, figures (#139 and #141 folded in; #140 closed) | Open; stacked on #188; synthetic mixed runs done (§11); traces to rerun | ## 9. Decisions @@ -597,18 +612,27 @@ charged.** In the paper every probe is a benchmark run (§5.2, Algorithm 4 line 5: "Evaluate T by c"), and benchmarking dominates search time (Exp#9). Running sketch-bench inside the search would give the same accuracy answers as reading the same measurements, so the plan is unchanged. Planning time adds the -measured benchmark time of each distinct probed (config, size) point (§7). A -lookup-only time would understate AutoSketch's planning cost. +**measured** benchmark time of each distinct probed (config, data shape), +charged once per metric (§7); a lookup-only time would understate +AutoSketch's planning cost. An earlier lower bound (`1e8 · c_ins` + one query +phase) was replaced by the measurement (2026-10-08). **Q4. AutoSketch window adapter.** `x = S`, `y = T` (or `gcd(S, T)`): one sliding sketch per query. -**Q5. Memory model.** Per phase, as in #145 (§4): open windows, one merge -accumulator per group, query output, and closed windows. +**Q5. Memory model.** By use (§4): ingest (open windows, one copy per +worker), storage (compacted closed windows), compaction (partial copies while +it runs) and query (merge accumulators and output while it runs), each +counted once. -**Q6. Cost.** `w_cpu · CPU + w_mem · memory`: CPU only first, then Fargate's -per-vCPU and per-GB prices (§4). Machine-family and peak-provisioned pricing -were dropped (2026-10-06). +**Q6. Cost.** By use, `w_cpu · AUC(CPU) + w_mem · AUC(memory)`: CPU only +first, then Fargate's per-vCPU and per-GB prices (§4). Machine-family pricing +(2026-10-06) and a peak-billed cost model (2026-10-08) were dropped. + +**Q7. Latency.** Reported, not a requirement to violate: the plan's batch +latency, its longest chain under elastic CPU. Version 1 has no SLA and +traces each method's cost–latency frontier; version 2 requires a batch +latency SLA (§5). Decided 2026-10-08. ## 10. Known limitations @@ -619,11 +643,61 @@ were dropped (2026-10-06). - Costs and latencies are estimates from per-operation measurements, not end-to-end executions. The execution-based comparison is ASAPQuery-backend #545/#547. +- CPU is idealized as elastic (capacity follows the load at once). Real + autoscalers add scale-up delay, warm minimum instances billed while idle + and billing granularity, so the cost by use is a lower bound for them. +- Compaction prices merging the workers' smaller partial copies at the + measured per-merge cost, exact for fixed-size sketches and an + approximation for KLL and DDSketch. - AutoSketch-Adapted's deployments (`x = S`, `y = gcd(S, T)`) are in the ASAP candidate set (`candidates.rs` generates every divisor of `S` as a window and - `gcd(x, T)` as a slide). So when its plan meets the latency bounds and its - configurations also pass ASAP's lookup rule (the two rules differ on - size, §6), the ASAP MILP can choose the same deployments and pay for - shared ones once; ASAP's cost is never higher. The sanity check reports any + `gcd(x, T)` as a slide). So when its configurations also pass ASAP's lookup + rule (the two rules differ on size, §6), the ASAP MILP can choose the same + deployments and pay for shared ones once; ASAP's unbounded cost is never + higher. The sanity check reports any case where this does not hold. The result to report is the size of the gap and where it comes from, not that a gap exists. + +## 11. Results (synthetic mixed set, 2026-10-08) + +Absolute costs by use; latency is the plan's batch latency. AutoSketch's +latency is 1011 ms in every workload. Results, figures and summaries are in +sketch-bench #138, `rqe-optimizer/results/autosketch-vs-asap-synthetic/` +(version 1) and `…-synthetic-sla/` (version 2). + +**Version 1** (CPU only, vCPU; Fargate in $/hour): + +| Workload | RQEs | ASAP cheapest (latency) | PerQuery cheapest (latency) | AutoSketch | ASAP at 1011 ms | PerQuery at 1011 ms | +|---|---|---|---|---|---|---| +| mixed | 50 | 3.32 (12.1 s) | 9.55 (3.84 s) | 3410 | 5.12 | 12.0 | +| mixed, Fargate | 50 | 0.220 $/h (2.94 s) | 0.720 $/h (2.97 s) | 163 $/h | 0.284 $/h | 0.813 $/h | +| mixed, r = 8 | 92 | 3.39 (2.94 s) | 16.7 (3.84 s) | 4089 | 5.15 | 19.3 | +| mixed, m = 8 | 400 | 26.5 (12.1 s) | 76.4 (3.84 s) | 27,280 | 41.0 | 96.2 | +| mixed, m = 16 | 800 | 53.1 (12.1 s) | 153 (3.84 s) | 54,560 | 81.9 | 192 | + +The tightest feasible bound is 92 ms in every workload; there ASAP costs +41.6 vCPU and PerQuery 66.6 (mixed). ASAP's frontier is below PerQuery's at +every latency, and both are two to three orders of magnitude below +AutoSketch. + +**Version 2** (CPU only, vCPU; AutoSketch meets only the 3 s and 10 s SLAs): + +| Workload | SLA | ASAP | PerQuery | AutoSketch | +|---|---|---|---|---| +| mixed | 100 ms | 34.5 | 59.2 | 3410 (misses) | +| mixed | 1 s | 5.12 | 12.0 | 3410 (misses) | +| mixed | 3 s | 3.32 | 9.74 | 3410 | +| mixed, m = 16 | 100 ms | 553 | 947 | 54,560 (misses) | +| mixed, m = 16 | 3 s | 53.2 | 156 | 54,560 | + +**Planning time** (AutoSketch: search + measured benchmark, 8 distinct probes +per metric of 20–37 s each): + +| RQEs | ASAP | PerQuery | AutoSketch | +|---|---|---|---| +| 50 | 0.76 s | 0.65 s | 0.002 s + 212 s | +| 400 | 6.8 s | 5.6 s | 0.015 s + 1697 s | +| 800 | 14.7 s | 11.7 s | 0.031 s + 3394 s | + +The trace workloads (Alibaba, Google, with 6h and 24h windows, ASAPQuery +#812) are still to be rerun on this model.