Skip to content

feat(asap-tools): label query sets and skew fits for cluster traces - #746

Draft
zzylol wants to merge 7 commits into
mainfrom
feat/trace-label-skew
Draft

zzylol wants to merge 7 commits into
mainfrom
feat/trace-label-skew

Conversation

@zzylol

@zzylol zzylol commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds asap-tools/dataset-analysis/, which covers two tasks:

  1. Label query sets: PromQL queries, with their exact group-by labels, for three traces:
    • Google ClusterData 2011
    • Alibaba microservices v2022
    • Datadog BOOM
  2. Skew fits: for each query, a lower bound, a maximum-likelihood fit and an upper bound of the skew parameter.
    • Queries that aggregate over keys (the group-by labels) get a discrete Zipf θ, weighted by row count and by value sum.
    • Queries that aggregate over values get a power-law tail α. It is fitted with powerlaw, the method of Clauset, Shalizi & Newman, and compared against lognormal and exponential fits.
    • lower/upper are the min/max of per-window MLEs; mle is the fit on all the data.
    • Windows are 5 min for Google, 1 min for Alibaba, and 20 equal chunks per series for BOOM.

results/skew_summary.csv (59 rows) is the handoff to the sketch-bench saturation study.

Data

fetch_data.sh downloads a fixed sample:

  • Google: task_usage and task_events part-00000, plus all of job_events.
  • Alibaba v2022: NodeMetrics_0, MSMetrics_0, and 10 consecutive CallGraph and MCRRT shards (30 min). These are read straight from tar.gz.
  • BOOM: 20 sampled series plus the taxonomy.

A full run takes about 4.5 min with 48 workers.

Key θ (lower / mle / upper), count-weighted; value-weighted in brackets

Query θ
Google sum by (user) (cpu_rate) 1.22 / 1.24 / 1.28 [1.25 / 1.30 / 1.42]
Google sum by (user, priority) 1.23 / 1.25 / 1.28 [1.27 / 1.32 / 1.44]
Google sum by (priority) 1.03 / 1.10 / 1.17 [1.44 / 1.61 / 1.75]
Google sum by (job_id) 1.11 / 1.10 / 1.14
Google sum by (scheduling_class) 0.41 / 0.58 / 0.76
Google sum by (machine_id) (uniform baseline) 0.23 / 0.21 / 0.25
Alibaba MSRTMCR by (msname) 1.34 / 1.54 / 1.90 [1.03 / 1.11 / 1.31]
Alibaba MSRTMCR by (nodeid) 1.07 / 1.13 / 1.23
CallGraph count by (rpctype) 1.26 / 1.32 / 1.39
CallGraph count by (um) / by (dm) 1.11 / 1.15 / 1.13, 1.10 / 1.13 / 1.11
CallGraph count by (service, um, dm) 0.91 / 0.98 / 0.93
MSMetrics sum by (msname) 0.83 (the same in every window)
NodeMetrics by (nodeid) (uniform baseline) ≈0

Known issues (why this is a draft)

  • mle can fall outside [lower, upper]. For most CallGraph queries, the fit on all 30 minutes is steeper than the fit on any single minute: pooling adds rare keys to the tail and more mass to the heavy keys. So mle is not bracketed by the window bounds. The definition of the bounds needs a decision.
  • Value α is not ready to use as-is.
    • Fits use up to 5,000 sampled values, and the tail must keep at least 100 values, so α describes roughly the top ≥2% of values.
    • The Google fits lose to lognormal.
    • For BOOM, power_law_ok is a majority vote over variates while R and p are medians. As a result ds-671-10S and ds-1840-D show ok=True with absurd α.
  • powerlaw is pinned to 1.5. Version 2.0 caps α at 3 by default and falls back to a slow numerical fit.
  • BOOM: the tags were removed and each variate was z-scored. So there is no key θ for BOOM, and α is fitted per variate on x - min(x).
  • The um/dm labels include the placeholder values UNKNOWN and USER, which are counted as ordinary keys.

Also changed

.github/workflows/python.yml gets a test-dataset-analysis job (black, isort, flake8, mypy, unittest), path-filtered to this directory.

Validation

  • black, isort, flake8, mypy and shellcheck pass.
  • 18 unit tests pass. They cover Zipf θ recovery within ±0.05, Pareto α recovery within ±5%, the bounds logic, and failure paths.

Citations

  • Google ClusterData 2011 traces (2025)
  • Luo et al., "The Power of Prediction: Microservice Auto Scaling via Workload Learning", SoCC 2022
  • Cohen et al., "This Time is Different: An Observability Perspective on Time Series Foundation Models", arXiv 2505.14766

🤖 Generated with Claude Code

zzylol and others added 7 commits September 29, 2026 18:21
Add asap-tools/dataset-analysis: query sets over Google 2011, Alibaba
v2022 and BOOM traces, a fetch script, and fit_skew.py, which fits a
discrete Zipf theta per key query and a power-law alpha per value query,
with lower/upper bounds from per-window fits. Add a CI job running its
lint, type check and unit tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Output of fit_skew.py over the data fetched by fetch_data.sh.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- lower/upper now span the per-window fits and the pooled fit, so
  lower <= mle <= upper.
- Bounds are computed at several window lengths (window_lengths_s per
  dataset) by merging finest-window aggregates; results gain window_len_s.
- Value fits use up to 100k samples and pick xmin over a 50-quantile grid
  (p50..p99.9, at least 100 tail values); losing fits report best_alt.
- BOOM series report ok_frac and alpha over passing variates only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
power_law_ok now needs the power law to significantly beat both the
lognormal and the exponential. Otherwise best_alt names the significantly
better alternative or "inconclusive". BOOM ok_frac uses the same rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace power_law_ok and best_alt with tail_class (light, power_law,
lognormal, heavy_inconclusive). BOOM series report the majority class,
ok_frac as the share of non-light variates, and alpha and R/p medians
over the non-light variates.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
K and rows were whole-sample totals repeated on every window length.
Rename them to K_total/rows_total and add K_win_{min,median,max} and
rows_win_{min,median,max} from the existing per-window aggregates
(value queries get rows_win_* only), so the saturation study can read N
and K at the sketch window and the query lookback. Write the summary
with enough digits to keep row counts exact.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fetch and analyze a longer sample: Google task_usage/task_events parts
0-119 (about 7 days), Alibaba CallGraph/MCRRTUpdate shards 0-119 (6 h),
MSMetricsUpdate 0-47 and NodeMetricsUpdate 0-1 (1 day). Drop the
30-minute clip on Alibaba tables and let a table override the dataset's
window_lengths_s, so window lengths reach 6 h (CallGraph, MSRTMCR), 1 day
(MSMetrics, NodeMetrics) and 7 days (Google). Joins are loaded once per
table instead of once per file, so the task_events join covers all 120
parts. Regenerate results/skew_summary.csv.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant