Skip to content

feat(asap-tools): add paced feeder for the data-plane eval - #793

Draft
zzylol wants to merge 1 commit into
mainfrom
feat/paced-feeder
Draft

zzylol wants to merge 1 commit into
mainfrom
feat/paced-feeder

Conversation

@zzylol

@zzylol zzylol commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds asap-tools/data-sources/paced-feeder, the data feeder for the data-plane vs ClickHouse evaluation (#789). It generates the synthetic workload on the fly and sends it in real time, one second of event time per wall-clock second:

  • ASAPQuery: Prometheus remote write to /api/v1/write (snappy-compressed protobuf).
  • ClickHouse: INSERT … FORMAT RowBinary into (ts DateTime64(3), label_0 String, instance String, value Float64).

The data has C groups (label_0) × s series per group (instance), --samples-per-sec samples per series, and Pareto(shape, scale) values. Each value is a stateless hash of (seed, series, sample index), so two feeders with the same flags send identical data and no dataset file is needed. At 1e6 series × 900 s, a file would be about 9e8 rows (~50 GB) per workload.

Pacing. Each tick is split into batches sent concurrently, and the next tick waits for all of them, so a series' samples never arrive out of order. --stats-out writes per-tick rows, bytes, start lag and elapsed time, plus achieved_rows_per_sec, overrun_ticks and errors. Any failed request makes the feeder exit non-zero.

Batch sizes. The default batch size depends on the sink. Remote write uses 50,000 rows, which compresses to about 0.76 MB per request, under axum's 2 MB default body limit in ASAPQuery's ingest. ClickHouse uses 100,000 rows: at 50,000 rows, 1e6 rows/s fell behind (ticks took about 1.15 s, since each INSERT creates a part).

It is a standalone crate (empty [workspace], own Cargo.lock), like otlp_exporter and fake_exporter.

Testing

  • Unit tests (6):
    • every series gets a sample in each sub-second slot of a tick;
    • values depend only on seed, series and sample index;
    • series map to the right group and instance labels;
    • values follow the Pareto distribution (empirical vs theoretical p50/p90/p99 within 5%);
    • remote write decodes, with structs mirroring prometheus_remote_write.rs, to the right labels (sorted) and samples;
    • RowBinary round-trips every column.
  • Live runs on a CloudLab c6320, against ClickHouse 26.10 (single binary) and the precompute_engine binary from main:
Run Result
ClickHouse, 1e5 rows/s, 10 s 1,000,000 rows; 100,000 distinct series; 10 distinct timestamps; all values ≥ 1; no overruns
ClickHouse, 1e6 rows/s, 10 s, feeder pinned to 2 cores 1e6 rows/s, 0 overruns, feeder 0.15 cores
ASAPQuery remote write, 1e6 rows/s, 10 s, feeder pinned to 2 cores 1e6 rows/s, 0 overruns, 0 failed requests, feeder 0.34 cores. The engine logged the parsed series data{instance="i0",label_0="g0"} routed to the configured aggregation

cargo clippy --all-targets -- -D warnings and cargo fmt --check are clean.

🤖 Generated with Claude Code

Generate a seeded synthetic workload (C groups x s series, Pareto values)
on the fly and send it in real time to ASAPQuery over Prometheus remote
write or to ClickHouse with RowBinary INSERTs. Values depend only on the
seed, series and sample index, so two feeders send identical data. Per-tick
timing is written as JSON so a run can be checked for keeping up.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant