Repository navigation
Conversation
Generate a seeded synthetic workload (C groups x s series, Pareto values) on the fly and send it in real time to ASAPQuery over Prometheus remote write or to ClickHouse with RowBinary INSERTs. Values depend only on the seed, series and sample index, so two feeders send identical data. Per-tick timing is written as JSON so a run can be checked for keeping up. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
asap-tools/data-sources/paced-feeder, the data feeder for the data-plane vs ClickHouse evaluation (#789). It generates the synthetic workload on the fly and sends it in real time, one second of event time per wall-clock second:/api/v1/write(snappy-compressed protobuf).INSERT … FORMAT RowBinaryinto(ts DateTime64(3), label_0 String, instance String, value Float64).The data has
Cgroups (label_0) ×sseries per group (instance),--samples-per-secsamples per series, and Pareto(shape, scale) values. Each value is a stateless hash of(seed, series, sample index), so two feeders with the same flags send identical data and no dataset file is needed. At 1e6 series × 900 s, a file would be about 9e8 rows (~50 GB) per workload.Pacing. Each tick is split into batches sent concurrently, and the next tick waits for all of them, so a series' samples never arrive out of order.
--stats-outwrites per-tick rows, bytes, start lag and elapsed time, plusachieved_rows_per_sec,overrun_ticksand errors. Any failed request makes the feeder exit non-zero.Batch sizes. The default batch size depends on the sink. Remote write uses 50,000 rows, which compresses to about 0.76 MB per request, under axum's 2 MB default body limit in ASAPQuery's ingest. ClickHouse uses 100,000 rows: at 50,000 rows, 1e6 rows/s fell behind (ticks took about 1.15 s, since each INSERT creates a part).
It is a standalone crate (empty
[workspace], ownCargo.lock), likeotlp_exporterandfake_exporter.Testing
prometheus_remote_write.rs, to the right labels (sorted) and samples;precompute_enginebinary frommain:data{instance="i0",label_0="g0"}routed to the configured aggregationcargo clippy --all-targets -- -D warningsandcargo fmt --checkare clean.🤖 Generated with Claude Code