Skip to content

[Perf] Reduce per-sample ingest cost on the remote-write path #817

Description

@zzylol

Context

A small data-plane comparison of ASAPQuery against ClickHouse found that ASAP's per-sample ingest cost decides whether it is cheaper overall. Query latency is already much lower.

Where the numbers come from. The run used ASAPQuery's query_engine_rust, not this backend. It is the setup from the data-plane eval plan (ProjectASAP/ASAPQuery#789), with the paced feeder from ProjectASAP/ASAPQuery#793.

Setup:

  • 1,000 groups × 100 series, 1 sample/s per series = 1e5 samples/s, Pareto(1.5) values.
  • Per-group p50/p90/p99 over the last closed 1-minute window, queried every 10 s.
  • DDSketch(α = 0.01) on both sides.
  • Each system pinned to 16 cores on a CloudLab c6320, 5 minutes per arm, one trial.
ClickHouse, Null raw table + quantilesDD MV ASAPQuery, DDSketch
Query latency p50 / p95 41 / 90 ms 10.5 / 16 ms
CPU, average (ingest-dominated) 0.13 cores (~1.3 µs/sample) 0.24 cores (~2.4 µs/sample)
RSS, average 0.75 GiB 0.15 GiB
Usage-based cost (CPU + memory) $0.0075/h $0.0095/h
Peak-provisioned cost (m7i) $0.018/h $0.015/h

Accuracy was the same: every answer within α.

So ASAP answers queries about 4× faster, but its steady ingest costs about 1.9× more CPU per sample, which makes it about 27% more expensive under usage-based pricing. The two ingest paths differ in shape:

  • ClickHouse: columnar batches (RowBinary, 100k rows per INSERT). The MV aggregates each block in one vectorized pass, so per-row work is amortized.
  • ASAP: per-sample work. Protobuf + snappy decode, label parsing and string building, series-key hashing, routing each sample, then a sketch update per sample.

The same pattern in this backend

The backend was not measured. Its remote-write path has the same per-sample shape and does more work per series. In data_plane/src/drivers/ingest/prometheus_remote_write.rs:

  • canonicalize_request, for every TimeSeries of every request:
    • builds a label HashMap by cloning each name and value;
    • format!s the series key;
    • computes the attribute fingerprint;
    • builds the population key via a BTreeMap and serde_json::to_string.
  • route_messages, for every sample:
    • extracts the group key for each matching config (extract_group_key_for / extract_group_key_from_labels);
    • allocates series_key.to_string() per routed sample.

Prometheus remote write usually sends one or a few samples per series per request, so this per-series work is effectively per-sample work. The same series are re-canonicalized on every request.

Proposed directions

  1. Series cache. Map raw label bytes, or a hash of them, to the canonical series, its group keys and its routing targets (sid / materialization). Do this once per series, not once per request, and invalidate on physical-plan generation change. Hot-path samples would then skip label HashMaps, key formatting, JSON population keys and per-sample to_string().
  2. Batch routing. Group a request's samples by destination once, and hand each worker or materialization a contiguous batch, instead of routing sample by sample.
  3. Batched sketch updates. Let accumulators take a slice of values per (series or group, window), e.g. a DDSketch bulk add, instead of one call per sample.
  4. Ingest format. Remote write's protobuf + snappy is heavier than ClickHouse's RowBinary. Measure decode cost separately. If it is significant, consider a lighter batch format for high-rate producers, or decode without materializing owned Strings.

How to measure

  • Load generator: the ASAPQuery paced feeder (feat(asap-tools): add paced feeder for the data-plane eval ASAPQuery#793). It sends a seeded workload at a fixed rate over remote write, e.g. --groups 1000 --series-per-group 100 for 1e5 samples/s.
  • Metric: report CPU per sample and RSS for the data-plane process, and keep a ClickHouse arm on the same stream for reference.
  • Profiling: perf record -g on the ingest threads should show how the ~µs per sample splits across decode, canonicalization, routing and sketch update. That split should decide the order of 1–4.

Suggested target: per-sample ingest CPU at or below ClickHouse's ~1.3 µs/sample for the same plan, without changing query-side behavior.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions