You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[Perf] Reduce per-sample ingest cost on the remote-write path #817
A small data-plane comparison of ASAPQuery against ClickHouse found that ASAP's per-sample ingest cost decides whether it is cheaper overall. Query latency is already much lower.
Where the numbers come from. The run used ASAPQuery's query_engine_rust, not this backend. It is the setup from the data-plane eval plan (ProjectASAP/ASAPQuery#789), with the paced feeder from ProjectASAP/ASAPQuery#793.
Setup:
1,000 groups × 100 series, 1 sample/s per series = 1e5 samples/s, Pareto(1.5) values.
Per-group p50/p90/p99 over the last closed 1-minute window, queried every 10 s.
DDSketch(α = 0.01) on both sides.
Each system pinned to 16 cores on a CloudLab c6320, 5 minutes per arm, one trial.
ClickHouse, Null raw table + quantilesDD MV
ASAPQuery, DDSketch
Query latency p50 / p95
41 / 90 ms
10.5 / 16 ms
CPU, average (ingest-dominated)
0.13 cores (~1.3 µs/sample)
0.24 cores (~2.4 µs/sample)
RSS, average
0.75 GiB
0.15 GiB
Usage-based cost (CPU + memory)
$0.0075/h
$0.0095/h
Peak-provisioned cost (m7i)
$0.018/h
$0.015/h
Accuracy was the same: every answer within α.
So ASAP answers queries about 4× faster, but its steady ingest costs about 1.9× more CPU per sample, which makes it about 27% more expensive under usage-based pricing. The two ingest paths differ in shape:
ClickHouse: columnar batches (RowBinary, 100k rows per INSERT). The MV aggregates each block in one vectorized pass, so per-row work is amortized.
ASAP: per-sample work. Protobuf + snappy decode, label parsing and string building, series-key hashing, routing each sample, then a sketch update per sample.
The same pattern in this backend
The backend was not measured. Its remote-write path has the same per-sample shape and does more work per series. In data_plane/src/drivers/ingest/prometheus_remote_write.rs:
canonicalize_request, for every TimeSeries of every request:
builds a label HashMap by cloning each name and value;
format!s the series key;
computes the attribute fingerprint;
builds the population key via a BTreeMap and serde_json::to_string.
route_messages, for every sample:
extracts the group key for each matching config (extract_group_key_for / extract_group_key_from_labels);
allocates series_key.to_string() per routed sample.
Prometheus remote write usually sends one or a few samples per series per request, so this per-series work is effectively per-sample work. The same series are re-canonicalized on every request.
Proposed directions
Series cache. Map raw label bytes, or a hash of them, to the canonical series, its group keys and its routing targets (sid / materialization). Do this once per series, not once per request, and invalidate on physical-plan generation change. Hot-path samples would then skip label HashMaps, key formatting, JSON population keys and per-sample to_string().
Batch routing. Group a request's samples by destination once, and hand each worker or materialization a contiguous batch, instead of routing sample by sample.
Batched sketch updates. Let accumulators take a slice of values per (series or group, window), e.g. a DDSketch bulk add, instead of one call per sample.
Ingest format. Remote write's protobuf + snappy is heavier than ClickHouse's RowBinary. Measure decode cost separately. If it is significant, consider a lighter batch format for high-rate producers, or decode without materializing owned Strings.
Metric: report CPU per sample and RSS for the data-plane process, and keep a ClickHouse arm on the same stream for reference.
Profiling:perf record -g on the ingest threads should show how the ~µs per sample splits across decode, canonicalization, routing and sketch update. That split should decide the order of 1–4.
Suggested target: per-sample ingest CPU at or below ClickHouse's ~1.3 µs/sample for the same plan, without changing query-side behavior.
Context
A small data-plane comparison of ASAPQuery against ClickHouse found that ASAP's per-sample ingest cost decides whether it is cheaper overall. Query latency is already much lower.
Where the numbers come from. The run used ASAPQuery's
query_engine_rust, not this backend. It is the setup from the data-plane eval plan (ProjectASAP/ASAPQuery#789), with the paced feeder from ProjectASAP/ASAPQuery#793.Setup:
quantilesDDMVAccuracy was the same: every answer within α.
So ASAP answers queries about 4× faster, but its steady ingest costs about 1.9× more CPU per sample, which makes it about 27% more expensive under usage-based pricing. The two ingest paths differ in shape:
RowBinary, 100k rows per INSERT). The MV aggregates each block in one vectorized pass, so per-row work is amortized.The same pattern in this backend
The backend was not measured. Its remote-write path has the same per-sample shape and does more work per series. In
data_plane/src/drivers/ingest/prometheus_remote_write.rs:canonicalize_request, for everyTimeSeriesof every request:HashMapby cloning each name and value;format!s the series key;BTreeMapandserde_json::to_string.route_messages, for every sample:extract_group_key_for/extract_group_key_from_labels);series_key.to_string()per routed sample.Prometheus remote write usually sends one or a few samples per series per request, so this per-series work is effectively per-sample work. The same series are re-canonicalized on every request.
Proposed directions
HashMaps, key formatting, JSON population keys and per-sampleto_string().RowBinary. Measure decode cost separately. If it is significant, consider a lighter batch format for high-rate producers, or decode without materializing ownedStrings.How to measure
--groups 1000 --series-per-group 100for 1e5 samples/s.perf record -gon the ingest threads should show how the ~µs per sample splits across decode, canonicalization, routing and sketch update. That split should decide the order of 1–4.Suggested target: per-sample ingest CPU at or below ClickHouse's ~1.3 µs/sample for the same plan, without changing query-side behavior.