Skip to content

Stage 3 per-second cost model, memory pricing and calibration - #594

Draft
zzylol wants to merge 2 commits into
stack/509-e2-keyed-additive-topkfrom
stack/509-w1-stage3-cost
Draft

zzylol wants to merge 2 commits into
stack/509-e2-keyed-additive-topkfrom
stack/509-w1-stage3-cost

Conversation

@zzylol

@zzylol zzylol commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Rebased on main d4869a7 (DF 54).

Wave 1 chain: #594 → #593 → #592 → #591 → #595 → #596 → #597 → #598 → #599 (on #589)

Why

Stage 3 priced each candidate per workload evaluation, which can't compare ingestion-time work (it runs as data arrives) with query-time work (it runs on each evaluation), or one-off queries with repeating ones. Stage 2 materialization (#509, step 5) needs a common basis first. The user approved decisions S1, S2 and S5. Design doc: docs/design_docs/proposals/stage3-cost-model.md.

What

  • Per-second cost (S1). An ingestion-time node is priced over one second of ingested rows, λ. λ is the declared ingestion_rate, else input_cardinality / data_ingestion_interval, else the default. A query-time node costs its per-evaluation cost × r, where r = evaluation_rate_of over the distinct repeat intervals of the roots reaching it. Equal intervals count once. One-off and Unknown recurrence is amortized as invocations / H, with H = 1 h (S5).
  • Memory term (S2). w_mem × retained bytes, with w_mem = 1.25e-7 cost/(byte·s) (1 GB ≈ 1/8 vCPU). It applies only to ingestion-time state that query time reads. Retained bytes = 2 × groups × state bytes.
  • Stage3Calibration on PlanningModels (with_calibration) replaces the hard-coded calibration. The unit is cost_per_second (COST_UNIT → COST_PER_SECOND), and the cost source names the calibration version.
  • RootDemand { accuracy, recurrence, predictability } (in asap_types::workload) replaces &[Option<AccuracyTarget>] in plan_stages and the selection API. The facade and stage_pipeline fill it from the workload entries.
  • Scan pricing fix (found by the Example 3 acceptance work, test: #509 Examples 2–4 acceptance specs and end-to-end tests #590). A scan's rows now follow the longest range plus offset that reads it, instead of 1 min. A time range passes only its own span.
  • Docs. The new design doc is linked from the proposals README and from docs: propose workload-wide planning, summary sharing, and materialization #509's Stage 3 section. The Example 1 acceptance doc now gives costs per second.

How

price in crates/plan-selection/src/lib.rs reads each node's timing from the physical DAG and finds which roots reach the node. Costs still add up node by node, so the tree DP's additivity assumption holds, and a shared node is still charged once.

Before / After

Before After
Unit cpu_ms_per_workload_evaluation cost_per_second
Example 1 selected P58 at 52.201 P58 at 5.220 (52.201 × 0.1; both panels every 10 s)
Example 1 runner-up P42 at 53.201 P42 at 5.320
Example 1 ranking — unchanged (every candidate × 0.1)
Example 3 Pattern A, raw plan, shared vs separate scans 61,904 vs 25,229 (sharing looks costlier) shared ≤ separate (5-year scan, ranges pass 1–3 years)

Tests

New tests:

  • an ingestion-time node is priced at λ;
  • the memory term equals w × retained bytes;
  • a node reached by two roots with equal intervals is charged once;
  • one-off recurrence is amortized over H;
  • the shared long-range scan regression. It fails on the old scan logic: shared 61,904 vs separate 25,229.
  • Example 1's ranking is unchanged and P58 costs 5.220.

cargo test --workspace: 1,523 passed, 10 ignored. #589 had 1,517 passed and 10 ignored; the difference is the 6 new tests. fmt and clippy (-D warnings) pass. Viewer tests: 29 run, OK, with 6 skipped because py_mini_racer is missing. The Example 1 fixture is regenerated and validated by crates/devtools/tests/stage_pipeline.rs.

🤖 Generated with Claude Code

zzylol added a commit that referenced this pull request Oct 4, 2026
Example 1 declares http_requests_total a counter, so Stage 3 selects P60 at
46.201 × 0.1 = 4.620 cost per second. Update the per-second ranking test and
the Stage 3 cost model's worked example to P60.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 4, 2026
Pass 1 takes the declared metric types since #593, and #594's year-long scan
test selects its variant by the Sharing enum instead of a shared flag.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 4, 2026
plan_stages takes per-root demand since #594, and DataWorkload declares metric
types since #593; Example 2's pipeline test runs each query once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol and others added 2 commits October 5, 2026 04:48
…ibration

Stage 3 prices every candidate in cost per second of wall time:
- an ingestion-time node over one second of ingested rows (λ: declared
  ingestion rate, else cardinality / ingestion interval, else the default);
- a query-time node per evaluation times the evaluation rate of the roots
  reaching it: distinct repeat intervals add, equal intervals count once,
  one-off and unknown recurrence is amortized over the horizon H (1 h);
- ingestion-time state read at query time costs w_mem × retained bytes
  (2 windows × groups × state bytes), w_mem = 1.25e-7 cost/(byte·s).

The coefficients and H move into Stage3Calibration on PlanningModels
(with_calibration). plan_stages and the selection API take a per-root
RootDemand (accuracy, recurrence, predictability), filled from workload
entries by the facade and the stage_pipeline devtool. COST_UNIT becomes
COST_PER_SECOND ("cost_per_second") and the cost source names the
calibration version.

Also fixes scan pricing found by the Example 3 acceptance work: a scan's
rows follow the longest range plus offset reading it (not 1 min), and a
time range passes only its own span, so a shared 5-year scan is no longer
costlier than separate scans.

Example 1 still selects P58: 52.201 per evaluation becomes 5.220 per
second; every candidate scales by 0.1, so the ranking is unchanged. The
Example 1 viewer fixture is regenerated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scope of the Stage 3 planning model for designers: what is priced, the
per-second basis for ingestion-time and query-time nodes, recurrence and
horizon, the memory term (1.25e-7 cost/(byte·s), 1 GB ≈ 1/8 vCPU) and its
configuration, illustrative default statistics, interaction with the tree
DP and shared nodes (one node, one computation; "not materialized" is
Stage 2 duplicating the node), a worked Example 1 table, and what is out
of scope. Linked from the proposals README and #509 Stage 3. The Example 1
acceptance doc's costs are restated per second.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the stack/509-e2-keyed-additive-topk branch from 6575cfe to 4bcdab4 Compare October 5, 2026 06:21
@zzylol
zzylol force-pushed the stack/509-w1-stage3-cost branch from 8764a6b to 3ee0baf Compare October 5, 2026 06:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant