Conversation
Add asap-tools/dataset-analysis: query sets over Google 2011, Alibaba v2022 and BOOM traces, a fetch script, and fit_skew.py, which fits a discrete Zipf theta per key query and a power-law alpha per value query, with lower/upper bounds from per-window fits. Add a CI job running its lint, type check and unit tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Output of fit_skew.py over the data fetched by fetch_data.sh. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- lower/upper now span the per-window fits and the pooled fit, so lower <= mle <= upper. - Bounds are computed at several window lengths (window_lengths_s per dataset) by merging finest-window aggregates; results gain window_len_s. - Value fits use up to 100k samples and pick xmin over a 50-quantile grid (p50..p99.9, at least 100 tail values); losing fits report best_alt. - BOOM series report ok_frac and alpha over passing variates only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
power_law_ok now needs the power law to significantly beat both the lognormal and the exponential. Otherwise best_alt names the significantly better alternative or "inconclusive". BOOM ok_frac uses the same rule. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace power_law_ok and best_alt with tail_class (light, power_law, lognormal, heavy_inconclusive). BOOM series report the majority class, ok_frac as the share of non-light variates, and alpha and R/p medians over the non-light variates. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
K and rows were whole-sample totals repeated on every window length.
Rename them to K_total/rows_total and add K_win_{min,median,max} and
rows_win_{min,median,max} from the existing per-window aggregates
(value queries get rows_win_* only), so the saturation study can read N
and K at the sketch window and the query lookback. Write the summary
with enough digits to keep row counts exact.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fetch and analyze a longer sample: Google task_usage/task_events parts 0-119 (about 7 days), Alibaba CallGraph/MCRRTUpdate shards 0-119 (6 h), MSMetricsUpdate 0-47 and NodeMetricsUpdate 0-1 (1 day). Drop the 30-minute clip on Alibaba tables and let a table override the dataset's window_lengths_s, so window lengths reach 6 h (CallGraph, MSRTMCR), 1 day (MSMetrics, NodeMetrics) and 7 days (Google). Joins are loaded once per table instead of once per file, so the task_events join covers all 120 parts. Regenerate results/skew_summary.csv. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
asap-tools/dataset-analysis/, which covers two tasks:powerlaw, the method of Clauset, Shalizi & Newman, and compared against lognormal and exponential fits.lower/upperare the min/max of per-window MLEs;mleis the fit on all the data.results/skew_summary.csv(59 rows) is the handoff to the sketch-bench saturation study.Data
fetch_data.shdownloads a fixed sample:task_usageandtask_eventspart-00000, plus all ofjob_events.A full run takes about 4.5 min with 48 workers.
Key θ (lower / mle / upper), count-weighted; value-weighted in brackets
sum by (user) (cpu_rate)sum by (user, priority)sum by (priority)sum by (job_id)sum by (scheduling_class)sum by (machine_id)(uniform baseline)by (msname)by (nodeid)count by (rpctype)count by (um)/by (dm)count by (service, um, dm)sum by (msname)by (nodeid)(uniform baseline)Known issues (why this is a draft)
mlecan fall outside[lower, upper]. For most CallGraph queries, the fit on all 30 minutes is steeper than the fit on any single minute: pooling adds rare keys to the tail and more mass to the heavy keys. Somleis not bracketed by the window bounds. The definition of the bounds needs a decision.power_law_okis a majority vote over variates while R and p are medians. As a result ds-671-10S and ds-1840-D showok=Truewith absurd α.powerlawis pinned to 1.5. Version 2.0 caps α at 3 by default and falls back to a slow numerical fit.x - min(x).um/dmlabels include the placeholder valuesUNKNOWNandUSER, which are counted as ordinary keys.Also changed
.github/workflows/python.ymlgets atest-dataset-analysisjob (black, isort, flake8, mypy, unittest), path-filtered to this directory.Validation
Citations
🤖 Generated with Claude Code